vmware-knight
Provides tools for managing VMware vCenter Server and standalone ESXi hosts, including VM lifecycle, deployment, guest operations, cluster management, datastore browsing, and alarm operations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@vmware-knightlist all VMs on lab-vcenter"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
VMware Knight
VMware Knight is an MCP-powered VMware operations toolkit for managing vCenter Server and standalone ESXi hosts through AI-assisted and command-line workflows.
It is designed for environments where operators need to work across multiple VMware targets without constantly remembering long FQDNs, IP addresses, or editing configuration files by hand.
Two VMware Knight features are especially important:
Friendly target tags — assign a simple alias such as
lab-vcenter,prod-vcenter, oresxi-lab-01to each VMware target and use that alias directly from Claude, Codex, the MCP server, or the CLI.Interactive management wizard — add, edit, remove, tag, test, and configure VMware targets through a guided terminal interface.
VMware Knight also includes VMware lifecycle, deployment, guest operations, cluster management, datastore browsing, alarm operations, investigation workflows, scheduled scanning, plan/apply workflows, and an MCP server for AI-assisted operations.
VMware Knight is a community project and is not an official VMware product.
Platforms: macOS, Linux and Windows 10/11.
macOS users: follow the macOS user guide. Its PATH commands (Steps 3 and 5) are required.
Windows users: follow the Windows user guide in Option C — Windows. Its PATH commands for
uvand Git (Steps 3 and 6) are required; the install fails without them. For Codex, also see docs/windows-codex.md.
About the Author
Technology Lead | Principal Solutions Architect | Strategic Advisory | Enterprise Networking | Cybersecurity | Data Center | Agentic AI Enthusiast | MBA | CCDE #20210029 | 4× CCIE #28551 | CISSP | 2× NSE 7 | OpenShift Certified | Cisco Champion | Cisco Fire Jumper Elite
Ehsan Emad is a technology leader and principal solutions architect with extensive experience across enterprise networking, cybersecurity, data center architecture, infrastructure automation, and technical advisory.
He is the creator of VMware Knight, Cisco Knight Multi-Device MCP, Networking With Ehsan, and the TechLoungeCast podcast.
VMware Knight was created to make VMware infrastructure easier to operate through modern AI and MCP workflows while keeping target selection explicit, understandable, and manageable across environments containing multiple vCenter Servers and standalone ESXi hosts.
GitHub: CyberKnightLabs
LinkedIn: Ehsan Emad
Networking With Ehsan: networkingwithehsan.com
TechLoungeCast: techloungecast.com
Podcast: Listen on Apple Podcasts
1. Friendly Target Tags
Managing several vCenter Servers and standalone ESXi hosts can become difficult when every operation depends on remembering IP addresses, FQDNs, or long infrastructure names.
VMware Knight adds an optional friendly tag/alias to every VMware target.
Instead of asking an AI assistant to operate against:
vc.networkingwithehsan.localyou can assign:
lab-vcenterand then simply ask:
Using VMware Knight, list all VMs on lab-vcenter.The same idea works for standalone ESXi hosts:
esxi-lab-01
esxi-lab-02
esxi-edge-01Related MCP server: VMware-AIops
How target resolution works
A configured VMware target can be referenced by:
Friendly tag
Configured target name
Hostname or IP address
Example:
targets:
- name: vcenter
tag: lab-vcenter
host: vc.networkingwithehsan.local
type: vcenter
username: administrator@vsphere.local
port: 443
verify_ssl: falseAll of the following resolve to the same target:
lab-vcenter
vcenter
vc.networkingwithehsan.localThe friendly tag is normally the easiest identifier to use with Claude or Codex.
Example prompts
Show me the VMs on lab-vcenter.Check cluster health on prod-vcenter.List the VMs running on esxi-lab-01.Show active alarms on dr-vcenter.Tag safety
VMware Knight validates identifiers to prevent ambiguous target selection.
A tag cannot conflict with the name, hostname/IP, or tag of another configured target.
Target matching is case-insensitive, so these refer to the same configured identifier:
lab-vcenter
LAB-VCENTER
Lab-vCenterTags are optional. Existing target names and hostname/IP references remain valid.
For multi-vCenter and multi-ESXi environments, descriptive tags are recommended.
2. Interactive Management Wizard
VMware Knight includes a guided terminal wizard so you do not need to manually edit the VMware target YAML configuration for normal day-to-day management.
Start it with:
vmware-knight wizardThe wizard displays:
VMware Knight Management Wizard
1. List VMware targets
2. Add VMware target
3. Edit VMware target
4. Remove VMware target
5. Set/Change target tag
6. Test connections
7. Configure MCP clients
8. Exit2.1 List VMware targets
Choose:
1. List VMware targetsThe wizard displays the configured targets with information such as:
target name
friendly tag
target type
hostname or IP address
username
port
This makes it easy to confirm the exact identifiers available to the MCP server.
2.2 Add a VMware target
Choose:
2. Add VMware targetThe wizard guides you through adding either:
a vCenter Server
a standalone ESXi host
You are prompted for the target information, including:
target name
friendly tag/alias
hostname or IP address
target type
username
HTTPS port
TLS certificate verification preference
password
The configuration is validated before it is accepted.
Example:
Target name: vcenter
Friendly tag/alias: lab-vcenter
vCenter/ESXi host: vc.networkingwithehsan.local
Target type:
1. vCenter
2. ESXi
Select target type: 1
Username: administrator@vsphere.local
Port: 443
Verify TLS certificate? No
Password: ********After enrollment, the target can be referenced by the friendly tag:
lab-vcenter2.3 Edit a VMware target
Choose:
3. Edit VMware targetThe wizard shows the configured targets and asks which one you want to modify.
You can update:
name
tag
hostname/IP
vCenter or ESXi target type
username
port
TLS verification
password
If you rename a target, VMware Knight migrates the stored credential reference to the new target name.
The configuration is validated before the change is accepted.
2.4 Remove a VMware target
Choose:
4. Remove VMware targetThe wizard asks you to select the target and confirm removal.
Removing a target removes:
the target from VMware Knight configuration
its stored password entry
The wizard requires confirmation before completing the removal.
2.5 Set or change a target tag
Choose:
5. Set/Change target tagThis option provides a fast way to change only the friendly alias without editing the rest of the target configuration.
Example:
vcenter -> lab-vcenteror:
esxi-01 -> esxi-lab-01The new tag is validated against the identifiers of all other targets before it is saved.
2.6 Test VMware connections
Choose:
6. Test connectionsVMware Knight runs its connection and environment checks against the configured targets.
Use this after:
adding a new vCenter
adding an ESXi host
changing a password
changing a hostname/IP
modifying TLS settings
troubleshooting MCP connectivity
You can also run the diagnostic command directly:
vmware-knight doctor2.7 Configure MCP clients
Choose:
7. Configure MCP clientsVMware Knight currently focuses its wizard-driven MCP configuration on:
Claude Desktop
OpenAI Codex
The wizard asks for the VMware Knight executable path and writes or updates the selected MCP client configuration.
The default executable location is typically:
~/.local/bin/vmware-knight # macOS / Linux
%USERPROFILE%\.local\bin\vmware-knight.exe # WindowsBoth options work on Windows: the wizard uses the Windows config locations, adds .exe to the path, and passes your home folder to the server (see Claude Desktop on Windows).
The generated MCP server name is:
vmware-knight3. Installation
macOS user guide
macOS user guide — installation. Use this on a Mac that has never run VMware Knight. Run the commands in Terminal, one step at a time, and don't move on until each check passes. No administrator rights are needed, except for the Apple installer in Step 1.
Required on macOS: the PATH commands in Steps 3 and 5. You must run them so Terminal can find
uvandvmware-knight. Programs installed into~/.local/binare not found until that folder is on your PATH. This is especially common whenuvcame from Homebrew, which does not add~/.local/binfor you. That causeszsh: command not found: vmware-knight.
Step 1 — Install Git. A new Mac has no Git until Apple's Command Line Tools are installed. VMware Knight is installed straight from GitHub, so uv needs Git. Check first:
git --versionIf it prints a version, go to Step 2. If a window asks you to install the command line developer tools, click Install and wait for it to finish. If nothing appears, start the installer yourself, then run git --version again:
xcode-select --installStep 2 — Install uv. uv downloads a suitable Python automatically, so you don't need to install Python separately:
curl -LsSf https://astral.sh/uv/install.sh | shIf you use Homebrew, brew install uv also works.
Step 3 — Add ~/.local/bin to PATH (required). You must run this command. It makes the current Terminal window see programs in ~/.local/bin, where the uv installer and VMware Knight put their commands:
export PATH="$HOME/.local/bin:$PATH"Then check both tools. Both commands must print a version:
uv --version
git --versionStep 4 — Install VMware Knight.
uv tool install git+https://github.com/CyberKnightLabs/vmware-knight-mcp.gitTo install a specific release instead of the latest code, add the tag, for example ...vmware-knight-mcp.git@v1.12.10.
Step 5 — Keep VMware Knight on PATH (required). You must run both commands. uv tool update-shell adds ~/.local/bin to your shell profile (~/.zshrc) so every new Terminal window finds vmware-knight. The export line applies it to this window too:
uv tool update-shell
export PATH="$HOME/.local/bin:$PATH"Step 6 — Check the install.
which vmware-knight
vmware-knight --helpwhich should print a path like:
/Users/USERNAME/.local/bin/vmware-knightStep 7 — Run the wizard to add your targets and tags, test the connections, and connect Claude Desktop or Codex (option 7):
vmware-knight wizardThe wizard saves your targets in ~/.vmware-knight/ and writes the full path of VMware Knight into your MCP client's config, so Claude Desktop and Codex find it without relying on PATH. Fully quit and reopen the client after option 7, then start a new chat.
On Linux, the steps are the same, except for Step 1: install Git with your package manager, for example sudo apt install git or sudo dnf install git.
Option A — Clone from GitHub
This is the recommended installation method. On a new Mac, follow the macOS user guide first. It covers installing uv and Git, and the required PATH steps.
git clone https://github.com/CyberKnightLabs/vmware-knight-mcp.git
cd vmware-knight-mcpInstall VMware Knight with uv:
uv tool install .You can also install directly from GitHub without cloning:
uv tool install git+https://github.com/CyberKnightLabs/vmware-knight-mcp.gitOr with pip:
pip install git+https://github.com/CyberKnightLabs/vmware-knight-mcp.gitVerify the CLI installation:
vmware-knight --helpA successful installation provides these executables:
vmware-knight
vmware-knight-mcp
vmware-knight-mcpis a stdio MCP server entry point. It is normally started by an MCP client and may appear to wait silently if launched directly from a terminal. Usevmware-knight --helpto verify the installation.
Option C — Windows
Windows user guide — installation. VMware Knight runs natively on Windows 10 and 11. Run these commands in PowerShell (no administrator rights needed), one step at a time, and don't move on until each check passes.
Required on Windows: the PATH commands in Steps 3 and 6. You must run them so PowerShell can find
uv, Git andvmware-knight. Do not skip them, even if you opened a new window.Why: a PowerShell window that is already open does not see programs installed after it started. This includes a new tab in the same Windows Terminal, which inherits the old PATH. That is what causes errors like
uv : The term 'uv' is not recognizedorGit executable not found. Step 3 below fixes it without closing anything.
Step 1 — Install uv. uv downloads a suitable Python automatically, so you don't need to install Python separately.
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"Step 2 — Install Git. VMware Knight is installed straight from GitHub, so uv needs Git. If winget asks you to accept its source agreement, answer Y.
winget install --id Git.Git -e --source wingetIf winget is not available (some older Windows 10 installs), install Git from git-scm.com/download/win and keep the default "Git from the command line" option.
Step 3 — Add uv and Git to PATH (required). You must run this command. It reads the PATH that the two installers just saved, so the current window can find uv and git:
$env:Path = [Environment]::GetEnvironmentVariable("Path", "Machine") + ";" + [Environment]::GetEnvironmentVariable("Path", "User")Step 4 — Check both tools. Both commands must print a version:
uv --version
git --versionIf git --version still says not recognized, check that Git is on disk:
Test-Path "C:\Program Files\Git\cmd\git.exe"If that prints True, Git is installed but its installer did not add it to PATH. You must add it permanently for your user, then repeat Steps 3 and 4:
$userPath = [string][Environment]::GetEnvironmentVariable("Path", "User")
[Environment]::SetEnvironmentVariable("Path", ($userPath.TrimEnd(";") + ";C:\Program Files\Git\cmd").TrimStart(";"), "User")If it prints False, Git did not install: run Step 2 again, or use the installer from git-scm.com.
Step 5 — Install VMware Knight.
uv tool install git+https://github.com/CyberKnightLabs/vmware-knight-mcp.gitTo install a specific release instead of the latest code, add the tag, for example ...vmware-knight-mcp.git@v1.12.10.
Step 6 — Add VMware Knight to PATH (required). You must run both commands. uv tool update-shell adds %USERPROFILE%\.local\bin to your PATH. Step 3's command makes the current window see it:
uv tool update-shell
$env:Path = [Environment]::GetEnvironmentVariable("Path", "Machine") + ";" + [Environment]::GetEnvironmentVariable("Path", "User")Step 7 — Check the install.
where.exe vmware-knight
vmware-knight --helpwhere.exe should print a path like:
C:\Users\USERNAME\.local\bin\vmware-knight.exeThen run the wizard to add your targets, exactly as on macOS:
vmware-knight wizardWindows-specific notes:
Topic | Windows |
Config folder |
|
Executable |
|
Codex config |
|
Claude Desktop config |
|
|
|
Find the executable |
|
Option B — Development Installation
Clone the repository:
git clone https://github.com/CyberKnightLabs/vmware-knight-mcp.git
cd vmware-knight-mcpCreate the project environment:
uv syncRun VMware Knight directly from the source tree:
uv run vmware-knight --helpStart the management wizard:
uv run vmware-knight wizardStart the MCP server:
uv run vmware-knight mcp4. First-Time Setup
After installing VMware Knight, start the interactive management wizard:
vmware-knight wizardThe wizard can:
1. List VMware targets
2. Add VMware target
3. Edit VMware target
4. Remove VMware target
5. Set/Change target tag
6. Test connections
7. Configure MCP clients
8. ExitA recommended first-time sequence is:
1. Add your vCenter or standalone ESXi target.
2. Assign a friendly tag such as lab-vcenter or production-vcenter.
3. Repeat for any additional VMware targets.
4. Run Test connections.
5. Configure your MCP client.
6. Start a fresh MCP client session.
7. Ask the AI client to use VMware Knight against one of your friendly tags.Example environment:
Name Tag Type Host
-------------- --------------- --------- -------------------------------
vcenter lab-vcenter vcenter vc.example.local
esxi-01 esxi-lab-01 esxi 10.10.10.51
esxi-02 esxi-lab-02 esxi 10.10.10.52After setup, the AI assistant can use the friendly tag directly:
Using VMware Knight, show me all VMs on lab-vcenter.Configure Codex
VMware Knight can install its MCP configuration directly into Codex.
Windows users: the wizard (
vmware-knight wizard→ 7 → 2) and the command below both work on Windows. Usewhere.exe vmware-knightinstead ofwhich, and pass the.exepath, for example--path "$env:USERPROFILE\.local\bin\vmware-knight.exe". There is also a standalone installer script that finds the executable for you; see docs/windows-codex.md.
First locate the installed VMware Knight executable:
which vmware-knightTypical uv installations expose it through:
~/.local/bin/vmware-knightInstall VMware Knight into Codex:
vmware-knight mcp-config install \
--agent codex \
--path ~/.local/bin/vmware-knight \
--yesVMware Knight resolves the real executable path and writes the MCP entry into:
~/.codex/config.tomlThe resulting Codex configuration is similar to:
[mcp_servers.vmware-knight]
command = "/Users/USERNAME/.local/share/uv/tools/vmware-knight/bin/vmware-knight"
args = ["mcp"]
enabled = true
startup_timeout_sec = 120After installation, start a fresh Codex session so the MCP configuration is loaded.
If the Codex CLI is available, verify the server with:
codex mcp list
codex mcp get vmware-knightOn macOS installations where Codex is bundled inside the ChatGPT application but is not on your shell PATH, you can use:
/Applications/ChatGPT.app/Contents/Resources/codex mcp listand:
/Applications/ChatGPT.app/Contents/Resources/codex mcp get vmware-knightA healthy configuration should report:
vmware-knight
enabled: true
transport: stdio
args: mcpThen test the live MCP connection from a fresh Codex session with a request such as:
Use the VMware Knight MCP tool cluster_health_summary against target lab-vcenter with top_n=2.A successful Codex session should show a tool call similar to:
Called vmware-knight.cluster_health_summary(...)This confirms Codex is using VMware Knight through MCP rather than reading local files.
5. Configuration Files
VMware Knight stores its configuration under:
~/.vmware-knight/ # macOS / Linux
%USERPROFILE%\.vmware-knight\ # WindowsPrimary files include:
~/.vmware-knight/config.yaml
~/.vmware-knight/.envThe main configuration can also be overridden with:
VMWARE_KNIGHT_CONFIGExample:
VMWARE_KNIGHT_CONFIG=/path/to/config.yaml vmware-knight doctorOn Windows (PowerShell):
$env:VMWARE_KNIGHT_CONFIG = "C:\path\to\config.yaml"; vmware-knight doctorExample target configuration
targets:
- name: vcenter
tag: lab-vcenter
host: vc.networkingwithehsan.local
type: vcenter
username: administrator@vsphere.local
port: 443
verify_ssl: false
- name: esxi-01
tag: esxi-lab-01
host: 10.10.10.51
type: esxi
username: root
port: 443
verify_ssl: falseFor normal operation, use the wizard instead of manually editing this file.
6. MCP Client Setup
Claude Desktop
VMware Knight can configure Claude Desktop automatically through the management wizard.
Start the wizard from any terminal after VMware Knight is installed:
vmware-knight wizardThen select:
7. Configure MCP clients
1. Claude DesktopWhen prompted for the VMware Knight executable path, the default is typically:
~/.local/bin/vmware-knightPress Enter to accept the default path, then confirm the installation with Y.
VMware Knight resolves the real executable path automatically and writes the MCP configuration to:
~/Library/Application Support/Claude/claude_desktop_config.json # macOS
%APPDATA%\Claude\claude_desktop_config.json # WindowsA typical generated entry looks like:
{
"mcpServers": {
"vmware-knight": {
"command": "/Users/USERNAME/.local/share/uv/tools/vmware-knight/bin/vmware-knight",
"args": ["mcp"]
}
}
}After installation, fully quit Claude Desktop and reopen it so the new MCP server is loaded. Start a fresh chat before testing.
A simple validation prompt is:
Using VMware Knight, list the configured VMware targets.For an end-to-end live MCP test, use a request such as:
Use the VMware Knight MCP tool cluster_health_summary against target lab-vcenter with top_n=2.
Do not inspect local files. Return the two live top issues from vCenter.If Claude returns live vCenter data through VMware Knight, the MCP integration is working correctly.
If vmware-knight is not found before starting the wizard, verify the installation with:
which vmware-knight
vmware-knight --help
vmware-knight-mcpis the stdio MCP server entry point used by MCP clients. You normally do not need to run it manually.
Claude Desktop on Windows
The wizard works on Windows too: vmware-knight wizard → 7 → 1, accept the default path, and confirm. It writes to the right place:
Claude Desktop install | Config file |
Installer from claude.ai |
|
Microsoft Store |
|
On Windows, the generated entry also includes the .exe path and an env section with your home folder, so the server finds %USERPROFILE%\.vmware-knight:
{
"mcpServers": {
"vmware-knight": {
"command": "C:\\Users\\USERNAME\\.local\\bin\\vmware-knight.exe",
"args": ["mcp"],
"env": {
"USERPROFILE": "C:\\Users\\USERNAME",
"SYSTEMROOT": "C:\\WINDOWS",
"PYTHONUTF8": "1"
}
}
}
}Then fully quit Claude Desktop (right-click its tray icon → Quit), reopen it, and start a new chat.
Manual setup, if you prefer to edit the file yourself or the wizard picked the wrong file:
In Claude Desktop, open Settings → Developer → Edit Config. This opens
claude_desktop_config.jsonin the right place for your install (usually%APPDATA%\Claude\claude_desktop_config.json).Find your executable path with
where.exe vmware-knight.Add VMware Knight under
mcpServers. In JSON, every backslash must be doubled:
{
"mcpServers": {
"vmware-knight": {
"command": "C:\\Users\\USERNAME\\.local\\bin\\vmware-knight.exe",
"args": ["mcp"]
}
}
}If the file already has other servers, add "vmware-knight": { ... } next to them inside the existing mcpServers object.
Fully quit Claude Desktop (right-click its tray icon → Quit), reopen it, and start a new chat.
OpenAI Codex
The wizard can also configure Codex:
vmware-knight wizardThen select:
7. Configure MCP clients
2. CodexA typical Codex configuration looks like:
[mcp_servers.vmware-knight]
command = "/Users/yourname/.local/bin/vmware-knight"
args = ["mcp"]
enabled = true
startup_timeout_sec = 120On Windows, the wizard also adds .exe to the path and passes your home folder to the server:
[mcp_servers.vmware-knight]
command = "C:\\Users\\yourname\\.local\\bin\\vmware-knight.exe"
args = ["mcp"]
enabled = true
startup_timeout_sec = 120
[mcp_servers.vmware-knight.env]
USERPROFILE = "C:\\Users\\yourname"
SYSTEMROOT = "C:\\WINDOWS"
PYTHONUTF8 = "1"
# ...plus APPDATA, LOCALAPPDATA, TEMP and TMPCodex does not need Full Access to use VMware Knight: MCP tools run outside the Codex shell sandbox. If Codex only works with Full Access, the MCP entry is not loading; see docs/windows-codex.md.
Restart Codex after changing its MCP configuration.
Example:
Using VMware Knight, show me the VMs on esxi-lab-01.7. How VMware Knight Works
The basic workflow is:
User
|
v
Claude / Codex
|
v
VMware Knight MCP Server
|
+--> resolve target by tag / name / host
|
v
pyVmomi / vSphere APIs
|
+--> vCenter Server
| |
| +--> Clusters
| +--> ESXi Hosts
| +--> Virtual Machines
| +--> Datastores
|
+--> Standalone ESXi HostWhen the AI sends a target such as:
lab-vcenterVMware Knight resolves that identifier to the configured VMware target before establishing the connection.
8. Capabilities Overview
VMware Knight includes operations across the following areas:
Area | Capabilities |
VM Lifecycle | Power operations, VM creation, deletion, reconfiguration, snapshots, clone, migration |
Deployment | OVA deployment, templates, linked clones, ISO attachment, batch deployment |
Guest Operations | Command execution, file upload, file download |
Plan / Apply | Multi-step execution plans with review and rollback support |
Cluster Management | Cluster information, creation, deletion, HA/DRS configuration, host membership |
Datastore | Browse datastore content and discover deployment images |
Networking | Distributed port group and host VMkernel management workflows |
Alarm Management | List, acknowledge, and reset triggered alarms |
Investigation | VM, host, datastore, cluster, and cross-vCenter investigation workflows |
Scanning | Scheduled alarm, event, and ESXi host-log scanning |
Notifications | JSONL logging and webhook-based notification workflows |
MCP | Structured VMware operations through Claude and Codex |
CLI | Direct command-line operation without an AI client |
9. Quick Investigation Workflows
VMware Knight includes read-oriented investigation commands that help identify problems before making changes.
Examples:
vmware-knight attentionShow issues requiring attention across configured vCenter targets.
vmware-knight summaryShow a cluster-oriented health summary.
vmware-knight investigate vm web-01 --hours 72Correlate information around a VM.
vmware-knight investigate host esxi-01Correlate information around a host.
vmware-knight investigate datastore datastore1Correlate information around a datastore.
Offline HTML output is also available for supported investigation workflows:
vmware-knight investigate vm web-01 --htmlSome investigation workflows delegate read-only aggregation to the companion vmware-monitor library.
10. VM Lifecycle
VMware Knight supports common virtual-machine lifecycle operations.
Operation | CLI Example | vCenter | ESXi |
Power On |
| Yes | Yes |
Graceful Shutdown |
| Yes | Yes |
Force Power Off |
| Yes | Yes |
Create VM |
| Yes | Yes |
Delete VM |
| Yes | Yes |
Reconfigure VM |
| Yes | Yes |
Create Snapshot |
| Yes | Yes |
List Snapshots |
| Yes | Yes |
Revert Snapshot |
| Yes | Yes |
Delete Snapshot |
| Yes | Yes |
Task Status |
| Yes | Yes |
Clone VM |
| Yes | Yes |
vMotion |
| Yes | No |
Set TTL |
| Yes | Yes |
Cancel TTL |
| Yes | Yes |
List TTLs |
| Yes | Yes |
Clean Slate |
| Yes | Yes |
Example using a target tag:
vmware-knight vm power-on web-01 --target lab-vcenter11. Guest Operations
Guest operations require VMware Tools inside the guest operating system.
Examples:
vmware-knight vm guest-exec web-01 \
--cmd /bin/bash \
--args "-c 'whoami'" \
--user rootUpload a file:
vmware-knight vm guest-upload web-01 \
--local ./script.sh \
--guest /tmp/script.sh \
--user rootDownload a file:
vmware-knight vm guest-download web-01 \
--guest /var/log/syslog \
--local ./syslog.txt \
--user rootGuest usernames are explicit. VMware Knight does not silently assume root.
12. Plan / Apply Workflow
For multi-step operations, VMware Knight supports a plan-oriented workflow.
Typical sequence:
Create Plan
|
v
Validate Operations
|
v
Review Affected Objects
|
v
Confirm
|
v
Apply Sequentially
|
+--> Success
|
+--> Failure --> Rollback available operationsCore MCP operations include:
vm_create_plan
vm_apply_plan
vm_rollback_planPlans are stored under:
~/.vmware-knight/plans/Successful plans are automatically removed, and stale plans are cleaned up.
13. VM Deployment and Provisioning
VMware Knight includes several deployment workflows.
Operation | Example |
Deploy OVA |
|
Deploy Template |
|
Linked Clone |
|
Attach ISO |
|
Mark Template |
|
Batch Clone |
|
Batch Deploy |
|
Example:
vmware-knight deploy ova ./ubuntu.ova \
--name ubuntu-lab-01 \
--datastore datastore1 \
--target lab-vcenter14. Cluster Management
Cluster operations include:
Operation | Command |
Cluster Info |
|
Create Cluster |
|
Delete Cluster |
|
Add Host |
|
Remove Host |
|
Configure HA/DRS |
|
Example:
vmware-knight cluster info LAB-CLUSTER --target lab-vcenterRemoving a host from a cluster requires the appropriate VMware state, including maintenance-mode requirements where applicable.
15. Alarm Management
VMware Knight supports VMware alarm workflows.
List triggered alarms:
vmware-knight alarm list --target lab-vcenterAcknowledge an alarm:
vmware-knight alarm acknowledge esxi-01 "Host memory usage"Reset alarms:
vmware-knight alarm reset esxi-01 "Host memory usage"Reset operations should be reviewed carefully because VMware alarm-clear APIs may affect multiple alarms within the selected scope.
16. Datastore Operations
Browse datastore content:
vmware-knight datastore browse datastore1 --path "iso/"Scan for deployment images:
vmware-knight datastore scan-images --target lab-vcenterImage discovery can locate content such as:
ISO
OVA
OVF
VMDK17. Scheduled Scanning
VMware Knight includes scheduled scanning for operational visibility.
Start the daemon:
vmware-knight daemon startCheck status:
vmware-knight daemon statusStop it:
vmware-knight daemon stopRun a one-time scan:
vmware-knight scan nowThe scanner can work across multiple configured targets and process operational information such as:
triggered alarms
vCenter events
ESXi host logs
scan findings
incomplete/unreachable target conditions
Structured scan output is written under the VMware Knight configuration directory.
18. Safety Model
Infrastructure automation requires explicit safety controls.
VMware Knight includes safeguards around write and destructive operations.
Depending on the operation and interface, these include:
dry-run or preview behavior
explicit confirmation
additional confirmation for destructive operations
plan-before-apply workflows
audit logging
rollback information where supported
explicit target selection
identifier collision validation
no silent target switching
The friendly-tag system is also part of the safety model because it lets operators use understandable infrastructure identifiers while VMware Knight still resolves the target against a validated inventory.
Always review destructive operations before approving them.
19. Common Workflows
Deploy a lab VM
vmware-knight datastore browse datastore1 \
--pattern "*.ova" \
--target lab-vcentervmware-knight deploy ova ./ubuntu.ova \
--name lab-vm \
--datastore datastore1 \
--target lab-vcentervmware-knight vm snapshot-create lab-vm \
--name baseline \
--target lab-vcentervmware-knight vm set-ttl lab-vm \
--minutes 480 \
--target lab-vcenterWork with multiple VMware environments
Configure tags such as:
prod-vcenter
dr-vcenter
lab-vcenter
esxi-lab-01
esxi-lab-02Then ask Claude or Codex:
Using VMware Knight, show me all VMs on lab-vcenter.Using VMware Knight, check active alarms on prod-vcenter.Using VMware Knight, list VMs on esxi-lab-01.This is easier and safer than repeatedly typing raw IP addresses.
20. CLI Quick Reference
Diagnostics:
vmware-knight doctor
vmware-knight doctor --skip-authWizard:
vmware-knight wizardMCP server:
vmware-knight mcpVM operations:
vmware-knight vm power-on my-vm --target lab-vcenter
vmware-knight vm power-off my-vm --target lab-vcenter
vmware-knight vm create new-vm --cpu 4 --memory 8192 --disk 100 --target lab-vcenter
vmware-knight vm delete my-vm --target lab-vcenter
vmware-knight vm snapshot-create my-vm --name before-upgrade --target lab-vcenter
vmware-knight vm snapshot-list my-vm --target lab-vcenter
vmware-knight vm snapshot-revert my-vm --name before-upgrade --target lab-vcenter
vmware-knight vm clone my-vm --new-name my-vm-clone --target lab-vcenter
vmware-knight vm migrate my-vm --to-host esxi-02 --target lab-vcenter
vmware-knight vm set-ttl my-vm --minutes 60 --target lab-vcenter
vmware-knight vm list-ttl --target lab-vcenterDeployment:
vmware-knight deploy ova ./ubuntu.ova --name my-vm --datastore ds1 --target lab-vcenter
vmware-knight deploy template golden-ubuntu --name new-vm --target lab-vcenter
vmware-knight deploy linked-clone --source base-vm --snapshot clean --name test-vm --target lab-vcenter
vmware-knight deploy iso my-vm --iso "[datastore1] iso/ubuntu.iso" --target lab-vcenterCluster:
vmware-knight cluster info LAB-CLUSTER --target lab-vcenter
vmware-knight cluster create LAB-CLUSTER --ha --drs --target lab-vcenter
vmware-knight cluster configure LAB-CLUSTER --ha --drs --target lab-vcenterAlarms:
vmware-knight alarm list --target lab-vcenterDatastore:
vmware-knight datastore browse datastore1 --path "iso/" --target lab-vcenter
vmware-knight datastore scan-images --target lab-vcenterScanning:
vmware-knight scan now
vmware-knight daemon start
vmware-knight daemon status
vmware-knight daemon stop21. Project Structure
vmware-knight-mcp/
├── vmware_knight/
│ ├── config.py
│ ├── connection.py
│ ├── doctor.py
│ ├── init_wizard.py
│ ├── setup_wizard.py
│ ├── cli/
│ ├── ops/
│ ├── scanner/
│ ├── notify/
│ └── mcp_server/
├── skills/
│ └── vmware-knight/
├── examples/
├── tests/
├── config.example.yaml
├── pyproject.toml
├── RELEASE_NOTES.md
└── README.md22. Technology
VMware Knight is built primarily with:
Python
pyVmomi
Typer
FastMCP / Model Context Protocol
YAML-based target configuration
local credential/environment storage
automated pytest coverage
23. Testing
Run the project tests with:
uv run pytest -qThe project includes tests for areas such as:
configuration loading
target resolution
friendly tag resolution
identifier collision handling
wizard operations
credential migration
MCP configuration
lifecycle safety gates
plan/apply behavior
cluster operations
guest operations
alarms
deployment
regression coverage
24. Updating VMware Knight
If installed directly from GitHub with uv, reinstall from the repository:
uv tool install --force \
git+https://github.com/CyberKnightLabs/vmware-knight-mcp.gitThe same command works in PowerShell on Windows, written on one line:
uv tool install --force git+https://github.com/CyberKnightLabs/vmware-knight-mcp.gitFor a cloned repository:
cd vmware-knight-mcp
git pull origin main
uv sync25. Troubleshooting
Check the installation
vmware-knight --helpOn macOS, if Terminal says zsh: command not found: vmware-knight (or uv), ~/.local/bin is not on your PATH. Run Step 5 of the macOS user guide, then check again:
uv tool update-shell
export PATH="$HOME/.local/bin:$PATH"
which vmware-knightIf git asks you to install the command line developer tools, see Step 1 of the macOS user guide.
On Windows, if PowerShell says uv, git or vmware-knight is not recognized, the program is usually installed but the window has an old PATH. Reload it and check again:
$env:Path = [Environment]::GetEnvironmentVariable("Path", "Machine") + ";" + [Environment]::GetEnvironmentVariable("Path", "User")
where.exe uv git vmware-knightIf vmware-knight is still missing, run uv tool update-shell and reload again. If git is still missing, see Step 4 of Option C — Windows. The same applies to uv tool install failing with Git executable not found.
Check VMware configuration
vmware-knight doctorCheck targets
vmware-knight wizardChoose:
1. List VMware targetsTest connectivity
From the wizard choose:
6. Test connectionsMCP client does not show VMware Knight
Confirm that:
the
vmware-knightexecutable path existsthe MCP configuration uses the correct executable
the MCP server name is
vmware-knightthe client has been restarted after changing its MCP configuration
For Codex, a typical block is:
[mcp_servers.vmware-knight]
command = "/Users/yourname/.local/bin/vmware-knight"
args = ["mcp"]
enabled = true
startup_timeout_sec = 120On Windows, also check:
the
commandpath ends in.exein Codex's TOML, the path is either in single quotes (
'C:\Users\...') or has doubled backslashes ("C:\\Users\\..."); single backslashes inside double quotes break the whole filecodex mcp get vmware-knightshows the entry; if it reports a parse error, re-run the wizard or the Windows installer script
TLS errors in a lab
If your lab uses a self-signed certificate, the wizard allows TLS certificate verification to be disabled for that target.
Do this only when you understand and accept the security implications.
26. Repository
GitHub:
https://github.com/CyberKnightLabs/vmware-knight-mcpClone:
git clone https://github.com/CyberKnightLabs/vmware-knight-mcp.git27. License
MIT
See LICENSE for the complete license text.
Available Tools
60 toolsacknowledge_vcenter_alarmA
[WRITE] Acknowledge a triggered vCenter alarm — marks it as seen WITHOUT clearing it.
This only marks the alarm; it does not clear it. The alarm stays in the active list with acknowledged=true until its condition clears or it is reset. To remove it entirely after fixing the root cause use reset_vcenter_alarm instead. Get exact entity_name and alarm_name from list_vcenter_alarms first; an unknown pair returns a not-found error.
Returns: Dict: entity_name, alarm_name, action ("acknowledged"), acknowledged (true).
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| alarm_name | Yes | Exact alarm definition name, e.g. "Virtual machine CPU usage". | |
| entity_name | Yes | VM, ESXi host, or cluster the alarm fired on (from list_vcenter_alarms output). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and destructiveHint=false, and the description adds meaningful state behavior beyond that: alarm remains active, acknowledged=true persists until condition clears, and an unknown pair returns a not-found error. It also explicitly labels the action as [WRITE], reinforcing the mutation without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Each sentence earns its place: purpose, non-effect, persistence behavior, alternative tool, prerequisite/error, and return format. The [WRITE] marker front-loads the key action, and the structure is scannable without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema, the description covers the action, the state change, the sibling alternative, the required data source, error behavior, and the return fields. An agent has enough to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters. The description adds value by instructing the agent to source entity_name and alarm_name from list_vcenter_alarms output and warning that an unknown pair causes a not-found error, which is practical guidance beyond the schema field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('acknowledge') and resource ('vCenter alarm'), and immediately distinguishes the action by saying it marks as seen WITHOUT clearing. It explicitly names reset_vcenter_alarm as the different alternative, so an agent can tell this tool apart from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use ('marks it as seen'), when-not-to-use (does not clear it), when to choose the sibling instead (remove entirely after fixing cause, use reset_vcenter_alarm), and a concrete prerequisite (get exact entity_name and alarm_name from list_vcenter_alarms first). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_host_vmkA
[WRITE] Add a static-IP VMkernel adapter on a DVS portgroup - preview/confirm gated.
Deliberately minimal shape for throwaway test vmks on L2-only segments (e.g. a TEP VLAN): static IPv4, NO gateway, NO services enabled. The DVS port allocation is handled internally - pass the distributed portgroup name. confirm=False validates (host + portgroup exist, IP/netmask/MTU legal, IP not already on the host) and returns the exact spec without writing; confirm=True creates and returns the assigned device name. Verify with list_host_vmks; remove with remove_host_vmk. Audited.
Returns: Preview dict (action="preview") or result dict (action="created", device=e.g. "vmk2"). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| ip | Yes | Static IPv4 address for the vmk. | |
| mtu | No | MTU for the vmk (default 1500; 9000 for jumbo tests). | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True creates it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| netmask | Yes | Subnet mask (e.g. 255.255.255.0). | |
| host_name | Yes | ESXi host to add the vmk on. | |
| portgroup | Yes | Distributed portgroup name to connect to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses the preview/confirm gate, exactly what confirm=False does (validates and returns spec without writing), what confirm=True does (creates and returns device name), and the error return format. It also notes the operation is audited. This is rich behavioral context that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the write marker and core purpose, followed by concise usage context, parameter behavior, and a returns section. There is no filler; every sentence adds operational insight. The structure makes it easy for an agent to parse the key facts quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no output schema, the description fully covers what an agent needs: the operation, the safety gate, validation behavior, return shapes including errors, and pointers to verification and removal. It also names relevant sibling tools. The context is complete enough to call this tool correctly without additional lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage, baseline is 3. The description adds value by explaining the confirm parameter's two-phase behavior, clarifying that the DVS port allocation is handled internally so only the portgroup name is needed, and stating that no gateway or services are part of the intended shape. This meaningfully supplements the schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Add a static-IP VMkernel adapter on a DVS portgroup'. It also highlights the unique preview/confirm gating and minimal shape, which clearly differentiates it from sibling tools like list_host_vmks, remove_host_vmk, and set_vmk_service. An agent can tell exactly what this tool does and how it differs without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context: 'Deliberately minimal shape for throwaway test vmks on L2-only segments' and states the constraints (static IPv4, no gateway, no services). It also directs verification and removal to list_host_vmks and remove_host_vmk. It lacks an explicit when-not-to-use alternative for more complex adapter configurations, but the minimal-shape framing implies the limitation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_iso_to_vmA
[WRITE] Mount a datastore ISO into a VM's virtual CD-ROM drive.
Reconfigures the existing CD-ROM (replacing any mounted ISO) or adds one on the VM's IDE controller; fails with a clear message if the VM has no IDE controller. Works whether the VM is powered on or off. Find ISO paths first with browse_datastore using pattern "*.iso". Returns a status string confirming attachment, or a VM-not-found / no-IDE-controller error.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target from config.yaml; omit for the default target. | |
| vm_name | Yes | Exact VM name as shown in vCenter inventory. | |
| iso_ds_path | Yes | Datastore path in bracket format, e.g. "[datastore1] iso/ubuntu-22.04.iso". |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses key behavioral traits: it replaces any mounted ISO, reconfigures the existing CD-ROM or adds one on the IDE controller, works whether the VM is powered on or off, and returns a status string or specific errors. This adds substantial value beyond readOnlyHint/idempotentHint/destructiveHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tight and well-structured: a one-line action statement followed by concise behavioral details, failure modes, and return value. Every sentence adds necessary information without fluff, and the main purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a complete input schema, an output schema, and relevant annotations, the description covers purpose, usage steps, edge cases (no IDE controller), power state, and return behavior. Nothing an agent needs to correctly invoke it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the parameters (vm_name, iso_ds_path, target) are already fully documented in the schema. The description adds a useful hint about finding ISO paths with browse_datastore, but does not provide additional semantic detail for the parameters themselves. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Mount a datastore ISO into a VM's virtual CD-ROM drive.' It clearly defines the operation and distinguishes it from siblings like browse_datastore (finding ISOs) and vm_reconfigure (general VM reconfiguration). The [WRITE] prefix also signals mutation, aligning with the write intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context: it tells agents to find ISO paths first with browse_datastore using pattern '*.iso', and explains when the operation fails (no IDE controller). It does not spell out 'when not to use' alternatives, but the prerequisite and failure condition provide clear usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_clone_vmsA
[WRITE] Batch clone multiple VMs from a source VM (gold image).
Each clone: full copy → optional reconfigure → optional snapshot → optional power on. Returns one dict per VM with its status. Clones run sequentially, so a long vm_names list may take a while — prefer batch_linked_clone_vms for disposable test copies.
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | Override CPU count for all clones (optional). | |
| target | No | Optional vCenter/ESXi target name from config. | |
| power_on | No | Power on each clone after creation. | |
| vm_names | Yes | Names for the new VMs. | |
| memory_mb | No | Override memory for all clones (optional). | |
| snapshot_name | No | Snapshot each clone with this name (optional). | |
| source_vm_name | Yes | Source VM to clone from. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The [WRITE] prefix, sequential execution warning, full-copy pipeline description, and return-value note ('one dict per VM') all add behavior beyond the annotations. Since readOnlyHint is false and idempotentHint is false, these details are genuinely useful and consistent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose, workflow, return shape, performance caveat, and alternative are all covered in four concise sentences. No filler or redundant restatement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a batch operation with a full input schema canvassing all 7 parametersable, the description covers the key operational expectations: sequential execution, optional post-clone steps, return format, and when to choose the linked-clone sibling. Nothing essential for selecting and invoking it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds useful semantic context by calling the source a 'gold image' and mapping the pipeline stages to the optional reconfigure/snapshot/power-on parameters, which is beyond the raw schema field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Batch clone multiple VMs from a source VM'. It also clarifies the workflow (full copy, optional reconfigure, snapshot, power on) and differentiates from the linked-clone sibling by naming batch_linked_clone_vms explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: use this for full independent clones, and 'prefer batch_linked_clone_vms for disposable test copies'. It also sets expectations for long vm_names lists by noting clones run sequentially.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_deploy_from_specA
[WRITE] Deploy multiple VMs in one call from a declarative YAML spec file.
Use for fleet provisioning (several VMs, shared defaults); for a single VM prefer deploy_vm_from_template, vm_clone, deploy_vm_from_ova, or deploy_linked_clone. The channel is chosen by spec keys: "source" (full clone), "template", "linked_clone: {source, snapshot}", per-VM "ova", else empty-VM creation (optionally "iso"). A "defaults" block sets cpu/memory_mb/disk_gb/network/ datastore/snapshot/power_on, overridable per VM. VMs deploy sequentially and one VM's failure does not stop the rest.
Returns: One dict per VM: name, status ("ok" or "error"), and messages with per-step results.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config.yaml; omit to use the default target. | |
| spec_path | Yes | Local filesystem path to the deploy.yaml specification file. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses non-obvious execution behavior: VMs deploy sequentially, one VM's failure does not stop the rest, and a per-VM status dict is returned. This goes beyond the annotations, which only indicate write/non-idempotent behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized: purpose first, usage guidance second, spec semantics and failure behavior next, and Returns clearly delimited. Every sentence earns its place for a complex batch operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only two parameters, the description fully covers target selection, spec file meaning, YAML channel semantics, defaults/overrides, execution order, failure isolation, and return format. Nothing essential is missing for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds essential meaning to spec_path by explaining the YAML contract: source/template/linked_clone/ova/iso keys, defaults block, and per-VM overrides. This is critical semantic content absent from the schema's one-line path description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+method: 'Deploy multiple VMs in one call from a declarative YAML spec file.' It clearly separates itself from single-VM siblings by naming deploy_vm_from_template, vm_clone, deploy_vm_from_ova, and deploy_linked_clone as alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States when to use it ('fleet provisioning (several VMs, shared defaults)') and when not to ('for a single VM prefer ...'). It also explains channel selection via spec keys, so the agent understands the tool covers multiple deployment modes from one spec.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
batch_linked_clone_vmsA
[WRITE] Batch create linked clones from a VM snapshot (fastest batch provisioning).
Clones share the source disk via copy-on-write, so the source must stay intact. Returns one dict per clone. Prefer batch_clone_vms for independent copies.
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | Override CPU count (optional). | |
| target | No | Optional vCenter/ESXi target from config. | |
| power_on | No | Power on each clone. | |
| vm_names | Yes | Names for the new linked clones. | |
| memory_mb | No | Override memory (optional). | |
| snapshot_name | Yes | Clone base (from vm_list_snapshots). | |
| source_vm_name | Yes | Source VM to clone from. | |
| baseline_snapshot | No | Snapshot each clone with this name (optional). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnlyHint=false and destructiveHint=false; the description adds meaningful behavioral context: clones share the source disk via copy-on-write, the source must remain intact, and each clone returns one dict. This goes beyond what the structured annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core action, and each sentence earns its place: purpose, behavioral constraint, return shape, and sibling routing. No filler or redundant schema repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters and a rich output schema, the description covers the critical non-obvious aspects: copy-on-write semantics, source immutability, return format, and the relevant alternative. It does not need to restate parameter details since the schema fully documents them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description does not add parameter-level meaning beyond the schema, though it reinforces the cloning mechanism tied to source_vm_name and snapshot_name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Batch create linked clones from a VM snapshot'. It clearly distinguishes itself from batch_clone_vms by noting the alternative is for independent copies, so an agent can tell them apart immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool over the sibling: 'Prefer batch_clone_vms for independent copies.' It also communicates a hard precondition—'the source must stay intact'—which informs safe usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browse_datastoreARead-onlyIdempotent
[READ] Browse files in a vSphere datastore directory.
Use this to find OVA/ISO/VMDK paths before calling deploy_vm_from_ova or attach_iso_to_vm; for an estate-wide image sweep use scan_datastore_images.
Returns the list envelope: 'items' is one row per file, and 'returned'/'total'/'truncated' state completeness. Every match in the searched folders is returned, so truncated is always false.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Subdirectory path (empty string for root). | |
| target | No | Optional vCenter/ESXi target name from config. | |
| pattern | No | Glob pattern to filter files (e.g. "*.ova", "*.iso", "*"). | * |
| datastore_name | Yes | Name of the datastore to browse. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds valuable behavioral context beyond that: it explains the list envelope ('items' one row per file, 'returned'/'total'/'truncated' completeness) and guarantees that 'Every match in the searched folders is returned, so truncated is always false.' This is genuinely useful behavioral disclosure beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences, each earning its place: core action, usage guidance with alternatives, and return-envelope behavior. The '[READ]' prefix is slightly redundant with annotations but aids quick scanning and does not inflate the length. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by explaining the response structure and the truncation guarantee. It also provides workflow context and names the alternative tool, making the definition self-sufficient for correct selection and invocation. Minor details like recursion behavior are not essential for this directory-browsing use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (path, target, pattern, datastore_name) are already documented in the input schema. The description does not add parameter-specific meaning beyond contextual hints about common file types (OVA/ISO/VMDK), so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Browse files in a vSphere datastore directory.' It explicitly differentiates from related siblings by stating it is for finding OVA/ISO/VMDK paths before deploy_vm_from_ova or attach_iso_to_vm, and that scan_datastore_images is the alternative for estate-wide sweeps. This makes the tool's purpose unambiguous relative to its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Use this to find OVA/ISO/VMDK paths before calling deploy_vm_from_ova or attach_iso_to_vm; for an estate-wide image sweep use scan_datastore_images.' It names the specific alternatives and the conditions that select them, leaving no ambiguity about when to invoke this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_add_hostA
[WRITE] Move an ESXi host that vCenter already manages into a cluster.
The host must already be in vCenter inventory (standalone or in another cluster) — this does NOT register brand-new hosts and takes no host credentials; use the vCenter UI for first-time registration. Idempotent: a host already in the cluster returns success without change. Maintenance mode is not required to join (it IS required by cluster_remove_host). Check membership first with cluster_info. Returns a status string: moved, already-in-cluster, or a not-found error.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter target from config.yaml; omit for the default target. | |
| host_name | Yes | Host name as shown in vCenter inventory, usually the FQDN, e.g. "esxi-01.lab.local". | |
| cluster_name | Yes | Destination cluster (create with cluster_create). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly states 'Idempotent: a host already in the cluster returns success without change,' but the annotations declare idempotentHint=false. This is a direct contradiction, so despite adding otherwise useful behavioral context, this dimension must score the lowest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with '[WRITE]' and a clear one-line purpose. Each subsequent sentence adds distinct value, though it is longer than strictly necessary; the structure is logical and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given output schema exists and parameter coverage is complete, the description covers prerequisites, idempotency, maintenance-mode expectations, a recommended pre-check, and the return status values. An agent has enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds helpful context like 'takes no host credentials' and that cluster_name references a cluster created via cluster_create, but it does not substantially expand parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Move an ESXi host that vCenter already manages into a cluster.' It clearly scopes the operation to already-managed hosts and distinguishes this tool from registration workflows and from cluster_remove_host.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use and when-not-to-use guidance: no brand-new host registration, no host credentials, use the vCenter UI for first-time registration. It also references cluster_info as a precursor and contrasts maintenance-mode requirements with cluster_remove_host.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_configureA
[WRITE] Reconfigure cluster HA/DRS settings.
Returns a status string naming what changed. Pass only the fields to change; None leaves a setting untouched — then verify with cluster_info. drs_behavior applies only when DRS is enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| ha | No | Enable (True) or disable (False) HA, or None to leave unchanged. | |
| drs | No | Enable (True) or disable (False) DRS, or None to leave unchanged. | |
| name | Yes | Cluster name. | |
| target | No | Optional vCenter target name from config. | |
| drs_behavior | No | DRS behavior: "fullyAutomated", "partiallyAutomated", or "manual". |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, destructiveHint=false, idempotentHint=false) establish the write/safety profile, so the description carries a lower burden but still adds real value: it discloses the return contract ('status string naming what changed'), the partial-update behavior (untouched settings remain), and the conditional dependency of drs_behavior on DRS being enabled. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero filler. Purpose is front-loaded in sentence one, the return contract and delta semantics follow, and the conditional caveat closes. Every sentence earns its place, and the [WRITE] tag is a compact one-token signal. No redundancy with schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with 100% schema coverage and an output schema, the description covers the essential decision-making surface: what it changes, how partial updates work, what it returns, and how to verify the result. Notable omissions are prerequisites (e.g., cluster must exist) and any caveat about side effects of disabling HA/DRS on running workloads, but these are minor given the rich schema and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 given every parameter (ha, drs, name, target, drs_behavior) already has a schema description. The description adds genuine meaning beyond the schema by stating the delta semantics ('None leaves a setting untouched', reinforcing the nullable defaults) and the conditional rule that drs_behavior only applies when DRS is enabled — information absent from the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Reconfigure') and a concrete resource ('cluster HA/DRS settings'), which clearly separates this mutation tool from cluster_create, cluster_delete, and cluster_info. The [WRITE] tag reinforces the mutation intent. It doesn't explicitly name a sibling it is not, but the verb+resource combo leaves little ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage semantics: 'Pass only the fields to change; None leaves a setting untouched' defines the delta update pattern, and 'verify with cluster_info' routes the agent to the correct read-side sibling. The precondition 'drs_behavior applies only when DRS is enabled' is a genuine guardrail. It doesn't explicitly state when to use cluster_create instead, but 'reconfigure' makes that boundary implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_createA
[WRITE] Create a new empty cluster in a datacenter, optionally enabling HA and DRS.
Fails with a clear error (no partial state) if the name already exists or drs_behavior is invalid. Then add hosts with cluster_add_host, change HA/DRS later with cluster_configure, and verify with cluster_info. Returns a status string naming the features enabled.
| Name | Required | Description | Default |
|---|---|---|---|
| ha | No | True enables vSphere HA (default False). | |
| drs | No | True enables DRS (default False). | |
| name | Yes | Name for the new cluster; must be unique in the datacenter. | |
| target | No | vCenter target from config.yaml; omit for the default target. | |
| datacenter | No | Datacenter name; omit for the first one on the target. | |
| drs_behavior | No | "fullyAutomated" (default), "partiallyAutomated", or "manual". Only takes effect when drs=True. | fullyAutomated |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a write operation and not idempotent, so the description adds value by disclosing failure behavior: it fails with a clear error and no partial state if the name exists or drs_behavior is invalid. It also discloses the return value as a status string naming enabled features, going beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and every sentence contributes: purpose, failure/atomicity behavior, related workflow, and return value. There is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters and a full schema, the description supplies the critical non-schema context: failure conditions, no partial state, follow-up sibling operations, and the shape of the returned status. Combined with the rich schema and output schema, an agent has everything it needs to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents each parameter and its defaults. The description adds a bit of semantic context by warning that invalid drs_behavior causes failure and that HA/DRS are optional, but it does not significantly expand on the parameter meanings already provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Create a new empty cluster in a datacenter, optionally enabling HA and DRS.' It clearly distinguishes itself from sibling tools by stating the cluster is empty and by explicitly connecting to cluster_add_host, cluster_configure, and cluster_info for subsequent actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage context: use this to create an empty cluster, then add hosts with cluster_add_host, modify HA/DRS with cluster_configure, and verify with cluster_info. It names the alternatives and indicates the sequencing, which leaves little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_deleteADestructive
[WRITE] Delete an empty cluster (no hosts must remain).
Without confirm=True this only previews: it returns blast_radius (cluster name and id, host/VM/datastore counts and names, blockers) and deletes nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Refused: a cluster that still has hosts or VMs (evacuate members with cluster_remove_host first), or one whose members could not be read. Check cluster_info first. Returns a dict (action, blast_radius).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name of the cluster to delete. | |
| target | No | Optional vCenter target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint=true, but the description meaningfully extends this by explaining the preview/apply dual mode, the blast_radius return value, and the refusal cases. It also warns against self-confirming without user consent, which is behavioral context far beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then organizes safety behavior, refusal conditions, and return type in clear sections. Every sentence carries operational meaning; there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructive nature, the description fully equips an agent: prerequisites (cluster_info), associated sibling action (cluster_remove_host), preview behavior, return shape, and explicit refusal conditions. No output schema exists, so the stated return dict (action, blast_radius) is essential and present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter is already documented. The description reinforces confirm's preview-vs-apply semantics and the returned dict, but it does not substantially add new parameter-level meaning beyond what the schema provides; it instead adds usage and safety context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Delete an empty cluster', and immediately scopes it with '(no hosts must remain)'. It differentiates itself from related operations by naming cluster_remove_host as the evacuation step, so an agent can distinguish deletion from member removal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: only empty clusters, check cluster_info first, use cluster_remove_host for evacuation, and never set confirm=True before showing the preview to the user. It also states refusal conditions, making the decision boundary clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_health_summaryARead-onlyIdempotent
[READ] One-glance health rollup for every cluster — "is anything on fire?".
Aggregates hosts, VM power state, live CPU/memory pressure, triggered alarms and
datastores thin-provisioned past 100% of capacity per cluster, assigns each a
status ("ok"/"warn"/"critical"), and flattens the anomalies into a ranked
top_issues focus list — the fast triage view for "what's wrong right now?".
Read-only (delegates to the vmware-monitor library).
Use it FIRST for a cross-cluster glance, then drill in with
vm_investigation_bundle or host_investigation_bundle.
Returns {totals, top_issues, issues_total, clusters, snapshot,
customization_hint}. Lead with top_issues (worst first, each with a
drill-down hint), clusters as context, customization_hint last.
Point-in-time — no trending.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | Cap the top_issues focus list (default 10; 0 hides it). | |
| target | No | Optional vCenter/ESXi target name from config. Uses default if omitted. | |
| include_vms | No | Roll up VM power counts (default True). False skips the VM pass on very large fleets. | |
| cluster_filter | No | Case-insensitive substring to show only matching clusters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and idempotent behavior, and the description adds value beyond them: it delegates to vmware-monitor, returns a snapshot without trending, ranks top_issues worst-first, and reveals the response shape. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core question ('is anything on fire?'), then organizes behavior, usage, and return format in tight paragraphs. Every sentence adds information, and the return-field ordering guidance earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates by documenting the return object and which fields to lead with. It also names sibling drill-down tools, explains status semantics, and clarifies the point-in-time nature — leaving no critical gap for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all four parameters well. The description adds useful context about top_issues and return ordering but doesn't materially enrich parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description names a specific verb and resource ('health rollup for every cluster') and defines the artifact precisely: statuses per cluster plus a ranked top_issues list. It also distinguishes itself from deeper investigation tools by explicitly positioning this as the fast cross-cluster triage view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs the agent to 'Use it FIRST for a cross-cluster glance, then drill in with vm_investigation_bundle or host_investigation_bundle.' It also sets expectations with 'Point-in-time — no trending,' so the agent knows not to use it for temporal analysis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_infoARead-onlyIdempotent
[READ] Get detailed cluster information: member hosts, HA/DRS config, resource capacity.
Read-only, no side effects. Use before cluster_add_host / cluster_remove_host (shows membership and per-host maintenance mode) and to verify cluster_configure changes.
Returns: Dict with name, host_count, hosts (each: name, connection_state, power_state, maintenance_mode), ha_enabled, ha_admission_control, drs_enabled, drs_behavior, total/effective CPU (MHz) and memory (GB). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Exact cluster name. | |
| target | No | vCenter target name from config.yaml; omit to use the default target. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already declaring readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, the description's 'Read-only, no side effects' is largely redundant. It earns credit by adding the error-return contract ('Errors return a dict with error + hint') and by enumerating the full return schema, which annotations cannot convey. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured and front-loaded: [READ] tag, one-sentence purpose, usage guidance, then a compact return-format block. Every sentence earns its place; the only mild redundancy ('Read-only, no side effects' vs annotations) is brief and reinforces safety.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with no output schema, an agent has everything needed: the exact return dict (all fields listed down to per-host maintenance_mode and CPU/MHz memory/GB units), the error format, usage timing, and a read-only safety profile from annotations. For a 2-parameter read tool, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — 'name' is documented as 'Exact cluster name' and 'target' as the config.yaml vCenter target with default behavior. The description adds no per-parameter meaning beyond the schema, so the baseline 3 applies; the description's return-field list is useful but concerns outputs, not parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get detailed cluster information') plus the exact contents (member hosts, HA/DRS config, resource capacity). The [READ] prefix and the named write-siblings (cluster_add_host, cluster_remove_host, cluster_configure) make the read-vs-mutate distinction unambiguous among the many cluster_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use context: 'Use before cluster_add_host / cluster_remove_host' and 'to verify cluster_configure changes.' This is clear and actionable, but it names no when-not-to-use case or stated alternative (e.g., cluster_health_summary for health-focused checks), so it stops short of the full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cluster_remove_hostADestructive
[WRITE] Remove a host from a cluster (host must be in maintenance mode).
Without confirm=True this only previews: it returns blast_radius (host name and id, maintenance mode, VM count, powered-on VM count, blockers) and moves nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Refused: a host not in maintenance mode, a host with powered-on VMs, and a host whose state could not be read. Run cluster_info first for the exact member host names. The host is not deleted — it stays in vCenter inventory standalone; use cluster_add_host to move it back. Returns a dict.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| host_name | Yes | ESXi host name to remove (from cluster_info output). | |
| cluster_name | Yes | Cluster to remove the host from. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses preview-vs-apply behavior, refusal conditions, maintenance-mode requirement, powered-on VM blocking, and the fact that the host remains in vCenter inventory. This adds material context beyond the destructiveHint annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place, covering the preview flow, user-consent constraint, refusal cases, inventory behavior, and sibling reference. It is front-loaded with the [WRITE] tag and the safety-critical preview warning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, human-in-the-loop tool with no output schema, this description is complete: it identifies the returned dict, the blast_radius fields, the failure modes, and the prerequisite cluster_info call. An agent can invoke it safely and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers all parameters, but the description enriches them significantly: confirm=false means preview/return blast radius, confirm=true applies, and host_name should be the exact name from cluster_info output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a specific verb ('Remove'), resource ('host from a cluster'), and a hard precondition ('must be in maintenance mode'). It also distinguishes itself from cluster_delete by clarifying that the host is not deleted, and from cluster_add_host as a way to move it back.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs the agent to run cluster_info first for exact host names, to preview with confirm=False, to show the blast radius to the user, and not to set confirm=True on its own. It names cluster_add_host as the alternative for reversing the operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_vm_to_templateA
[WRITE] Convert a powered-off VM to a vSphere template.
Returns a status string. Use this to freeze a golden image: afterwards the VM cannot be powered on and serves only as a clone source for deploy_vm_from_template.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| vm_name | Yes | Name of the VM to convert (must be powered off). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses important behavioral effects: the VM becomes a template, cannot be powered on, and serves only as a clone source. It also states the return value is a status string. Since the annotations only indicate readOnly=false and destructive=false, the description adds meaningful state-change context without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no filler. It front-loads the operation verb and resource, then adds the most important consequence and downstream usage. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-changing tool with two parameters and an output schema, the description covers the prerequisite, the effect, and the follow-up tool. It does not mention failure modes or whether the conversion is reversible, but these are secondary given the schema and output schema already provide structural coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents vm_name and target fully. The description reinforces that the VM must be powered off but adds no new parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate because the schema carries the load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: "Convert a powered-off VM to a vSphere template." It clearly differentiates from siblings by explaining the resulting role as a clone source for deploy_vm_from_template, leaving no ambiguity about what this tool accomplishes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear use case: "freeze a golden image" and explains the consequence that the VM cannot be powered on afterward. It does not explicitly name alternatives or when-not-to-use conditions, but the downstream workflow with deploy_vm_from_template provides enough context for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_drs_ruleA
[WRITE] Create a VM-VM DRS rule (affinity or anti-affinity) - preview/confirm gated.
rule_type "affinity" keeps the listed VMs together; "antiAffinity" keeps them apart (e.g. redundant appliance pairs on separate hosts). Requires >=2 distinct VMs, all members of the cluster. VM-Host rules hang off cluster VM/host groups and are not created here. Verify with list_drs_rules. Audited.
Returns: Preview dict (action="preview", would_create) or result dict (action="created", created incl. the assigned key). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| cluster | Yes | Exact cluster name. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True creates it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| enabled | No | Create the rule enabled (default) or disabled. | |
| vm_names | Yes | VM names the rule governs (>=2, all in the cluster). | |
| rule_name | Yes | Name for the new rule; must be unique on the cluster. | |
| rule_type | Yes | "affinity" or "antiAffinity". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description significantly augments the annotations by disclosing the preview/confirm gating, the exact behavior of confirm=false vs true, the return dict shapes, error shape, auditability, and membership requirements. Annotations only say readOnlyHint=false, so this added context is genuinely valuable and non-contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and gating flag, then gives rule semantics, constraints, exclusions, verification path, and return format in compact bullets. Every sentence earns its place and there is no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description properly explains the return shapes for preview, created, and error cases. Combined with 100% schema parameter coverage, the tool is fully transparent about what happens at each confirm state, what constraints apply, and how to verify the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds extra meaning for rule_type by explaining the behavioral difference between affinity and antiAffinity, and it reinforces the VM count/cluster membership precondition, going slightly beyond the schema's field-level descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: "Create a VM-VM DRS rule (affinity or anti-affinity)". It distinguishes this tool from sibling DRS tools by explicitly stating VM-Host rules are not created here and pointing to list_drs_rules for verification.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage context: it is a preview/confirm gated creation flow, requires >=2 distinct cluster VMs, and excludes VM-Host rules. It references list_drs_rules for verification, though it does not enumerate every sibling alternative (e.g., set_drs_rule_enabled vs create), so it stops short of a full when-not map.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_dvs_portgroupA
[WRITE] Create a VLAN-tagged portgroup on a dvSwitch - preview/confirm gated.
confirm=False (default) validates everything (switch exists, name free, binding and VLAN legal) and returns the exact spec that WOULD be created without writing anything. confirm=True creates the portgroup and waits for the task. Verify afterwards with list_dvs_portgroups. Audited.
binding="ephemeral" creates a portgroup with no pre-created port pool, attachable from the ESXi host client even when vCenter is down - use for a self-hosted VCSA's own management portgroup. num_ports is ignored for ephemeral. lateBinding is deprecated by vSphere and not offered.
Returns: Preview dict (action="preview", would_create) or result dict (action="created", created). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Name for the new portgroup; must be unique on the switch. | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| binding | No | "earlyBinding" (default) or "ephemeral". | earlyBinding |
| confirm | No | False (default) returns the blast radius and changes nothing. True creates it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| vlan_id | Yes | VLAN ID to tag (0-4094; 0 = none). | |
| dvs_name | Yes | Name of the distributed virtual switch to create it on. | |
| num_ports | No | Port count for earlyBinding portgroups (default 8). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses important behaviors not in the annotations: preview/confirm gating, that confirm=False writes nothing, that confirm=True waits for the task, that the operation is audited, and the exact return dict shapes for success and error cases. Annotations only indicate readOnly=false and destructive=false, so the description carries the full behavioral burden and does so thoroughly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat long but well-structured: a one-line purpose header, a gating/validation paragraph, an ephemeral binding paragraph, and a returns section. Every section carries essential information, and the purpose is front-loaded. Minor redundancy with parameter schema descriptions keeps it from being perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and a complex 7-parameter tool with open-world hints, the description covers what an agent needs: preview/confirm semantics, return value structure (including error dicts), validation behavior, and edge-case binding behavior. It leaves no critical gap for selecting or invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema by explaining ephemeral binding implications (attachable when vCenter is down, num_ports ignored) and clarifying the output of confirm=False versus confirm=True. It does not deeply explain target or dvs_name, but the schema already covers them adequately, so this is solidly above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Create a VLAN-tagged portgroup on a dvSwitch.' It also immediately signals the preview/confirm gating, which clearly distinguishes this from read-only or other network tools. The scope (VLAN-tagged, dvSwitch) makes it unambiguous among siblings like list_dvs_portgroups or VM tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit workflow guidance: confirm=False to validate and preview, confirm=True to actually create, and verify afterwards with list_dvs_portgroups. It also names a concrete scenario for ephemeral binding (self-hosted VCSA management portgroup) and explains when num_ports is ignored. This is strong when-to-use and alternative-routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cross_vcenter_attentionARead-onlyIdempotent
[READ] "What needs attention now?" across EVERY configured vCenter — one list.
Runs cluster_health_summary against every configured target and returns one
globally ranked top_issues list (worst first, each tagged with its
vcenter) plus a per-target rollup — "where do I look first, anywhere in the
estate?". Degrades gracefully: an unreachable target is listed under
unreachable and the rest still aggregate. Delegates to the vmware-monitor
library (read-only). Lead with top_issues, then drill in with
vm_investigation_bundle. Point-in-time.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | Cap the merged top_issues focus list (default 10). | |
| cluster_filter | No | Case-insensitive cluster substring applied to every target. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint/idempotentHint annotations, the description discloses important runtime behavior: it degrades gracefully when a target is unreachable, lists unreachable targets separately, aggregates remaining results, and is point-in-time. It also names the underlying library and confirms read-only behavior, giving an agent confidence in side-effect safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, front-loading the core value proposition and then providing behavior, fallback semantics, and follow-up guidance. Every sentence contributes useful information; there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description sufficiently describes the return shape: a globally ranked top_issues list tagged by vcenter, a per-target rollup, and an unreachable list. It also covers failure behavior, read-only safety, point-in-time semantics, and a suggested drill-in path, making it complete for a read-only aggregation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters (top_n and cluster_filter) are already clearly documented in the schema. The tool description adds context about ranking and aggregation but does not need to repeat parameter details. Baseline 3 is appropriate because the schema carries the parameter documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific purpose: aggregate attention-worthy issues across every configured vCenter into one ranked list. It explicitly names the underlying operation (runs cluster_health_summary against every target) and distinguishes itself from single-target tools by emphasizing the cross-vCenter scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: as the first estate-wide triage step, and it explicitly suggests following up with vm_investigation_bundle for deeper drill-in. It does not explicitly state when not to use it or compare it directly to cluster_health_summary as an alternative, but the 'across EVERY configured vCenter' scope makes the usage context unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
datastore_investigation_bundleARead-onlyIdempotent
[READ] "What is happening around this datastore?" — one correlated drill-down.
Correlates a datastore's capacity/free space/accessibility, the hosts that mount it, a rollup of the VMs it backs, alarms across datastore/host, and a merged event timeline. Delegates to the vmware-monitor library (read-only); do not dump the result raw. (Per-datastore latency is a separate perf report.)
Use this AFTER cluster_health_summary flags storage pressure. Point-in-time.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | Event-timeline look-back window in hours (default 24). | |
| target | No | Optional vCenter/ESXi target name from config (default if omitted). | |
| datastore_name | Yes | Exact datastore name. Unknown names return a teaching error. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description reinforces this with '[READ]' and 'read-only'. The description adds valuable behavioral guidance beyond annotations: it delegates to the vmware-monitor library, should not be dumped raw, and is point-in-time. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured: a short headline, a clear bullet-like list of correlated data, a behavioral note, and a usage trigger. Every sentence earns its place, and the key purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, the description compensates by enumerating what the bundle contains: capacity, free space, accessibility, hosts, VM rollups, alarms, and an event timeline. Combined with the usage trigger and read-only caveat, an agent has enough context to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema already documents datastore_name with exact-match semantics, hours as an event-timeline look-back window, and target as an optional vCenter/ESXi selection. The description adds only a high-level 'point-in-time' hint rather than parameter-specific details, so it meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb/resource combination: it 'correlates' a datastore's capacity, free space, accessibility, mounting hosts, backed VMs, alarms, and event timeline. This clearly distinguishes it from sibling investigation bundles (vm_investigation_bundle, host_investigation_bundle) and individual datastore browsing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it: 'AFTER cluster_health_summary flags storage pressure.' It also gives a when-not signal by noting that per-datastore latency is a separate perf report. It does not name a specific sibling tool for that alternative, so it stops short of a full routing matrix.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_drs_ruleADestructive
[WRITE] Delete a VM-VM DRS rule - confirm-gated, guarded.
REFUSES non-VM-VM rules: VM-Host rules can carry licensing/compliance placement constraints (must-run-on licensed hosts) and hang off cluster groups - manage those in the vSphere UI. The preview and result both record the full rule definition so a mistaken delete can be recreated from the audit trail. Audited.
Returns: Preview dict (action="preview", would_delete) or result dict (action="deleted", deleted). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| cluster | Yes | Exact cluster name. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True deletes it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| rule_name | Yes | Exact rule name (see list_drs_rules). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint annotation by explaining the confirm-gated preview/delete flow, that both preview and result record the full rule definition for audit recovery, that the operation is audited, and that errors return a dict with 'error' and 'hint'. This gives the agent a clear model of side effects and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured and dense: the WRITE marker and purpose are first, then the key exclusion, then behavioral/audit context, then return shape. Every sentence earns its place and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive confirm-gated tool with no output schema and no enum constraints, this description covers the essential operational context: what it deletes, what it refuses, the preview/delete contract, auditability, and error shape. An agent has enough to invoke it safely and interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already provides strong descriptions for confirm, cluster, rule_name, and target. The description adds useful behavioral context about the preview/result flow, but it does not substantially add parameter-level meaning beyond what the schema already says.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Starts with a specific verb and resource: 'Delete a VM-VM DRS rule'. It also narrows the scope to VM-VM rules only and clearly distinguishes this from sibling tools like create_drs_rule, set_drs_rule_enabled, and list_drs_rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when not to use the tool: non-VM-VM rules are refused and should be managed in the vSphere UI. It also reinforces the confirm-gated workflow by warning that the caller should not blindly set confirm=true before a preview has been seen.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deploy_linked_cloneA
[WRITE] Create a linked clone from a VM snapshot — near-instant, minimal disk usage.
The clone writes to a copy-on-write delta over the source's base disk, so it depends on the source staying intact. Fastest provisioning for test/dev fleets; use vm_clone for fully independent copies. Requires the named snapshot — run vm_list_snapshots first. Returns a status string with the new clone name.
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | Override vCPU count; omit to keep the source's value. | |
| target | No | vCenter/ESXi target from config.yaml; omit for the default. | |
| new_name | Yes | Name for the new linked clone; must not already exist. | |
| power_on | No | Power the clone on after creation (default False). | |
| memory_mb | No | Override memory in MB; omit to keep the source's value. | |
| snapshot_name | Yes | Clone base (from vm_list_snapshots). | |
| source_vm_name | Yes | Source VM name (must have at least one snapshot). | |
| baseline_snapshot | No | If set, snapshots the new clone with this name. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only state low-level hints (readOnly=false, destructive=false), so the description carries the behavioral burden. It discloses the copy-on-write mechanism, the source-dependency risk ('depends on the source staying intact'), and the return type ('status string with the new clone name'). This adds meaningful behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose first, then behavior, use case, alternative, prerequisite, and return value. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 8 parameters, the output schema and 100% parameter coverage handle the details. The description adds exactly the missing context an agent needs: what the tool does, when to choose it, the critical source dependency, and the required snapshot workflow. Nothing essential is omitted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are fully described in the schema. The description adds context for snapshot_name/source_vm_name via the prerequisite and source-dependency note, but most parameter meaning is already present in the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action ('Create a linked clone from a VM snapshot'), the core characteristics ('near-instant, minimal disk usage'), and explicitly differentiates from vm_clone ('use vm_clone for fully independent copies'). This makes the tool's role unmistakable among many VM-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use guidance ('Fastest provisioning for test/dev fleets'), names the alternative for a different need (vm_clone), and states a hard prerequisite ('Requires the named snapshot — run vm_list_snapshots first'). It also implies a when-not via the dependency on the source staying intact.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deploy_vm_from_ovaA
[WRITE] Create a new VM by importing a local .ova file (OVF parse + VMDK upload).
Use for local OVA files; for vSphere templates use deploy_vm_from_template, to copy an existing VM vm_clone. Returns a status string naming the new VM. Upload time scales with OVA size; fails before creating anything if the datastore is not found.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target from config.yaml; omit for the default. | |
| vm_name | Yes | Name for the new VM; must not already exist. | |
| ova_path | Yes | Local path to the .ova file (must be readable by this server). | |
| power_on | No | Power the VM on after import (default False). | |
| folder_path | No | vCenter folder path; omit for the datacenter root. | |
| network_name | No | Port group for the NICs (default "VM Network"). | VM Network |
| snapshot_name | No | If set, snapshots the new VM with this name. | |
| datastore_name | Yes | Target datastore (see browse_datastore). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral details beyond the annotations: it returns a status string naming the new VM, upload time scales with OVA size, and it 'fails before creating anything if the datastore is not found.' This gives the agent a clear picture of side effects, performance, and a failure atomicity guarantee that annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences with no filler. It front-loads the core actioneb7a and [WRITE] marker, then the usage differentiation, return value, and behavior caveats. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter creation tool, the description plus fully described schema covers what the agent needs: the operation, usage conditions, return value, failure behavior, and performance expectation. The output schema exists and the annotations cover mutability, so no critical decision-making context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 8 parameters including defaults and semantics. The description adds no parameter-specific detail beyond mentioning the datastore failure condition, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Create a new VM by importing a local .ova file' and even names the underlying mechanism (OVF parse + VMDK upload). It also explicitly distinguishes itself from deploy_vm_from_template and vm_clone, making its unique purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for local OVA files; for vSphere templates use deploy_vm_from_template, to copy an existing VM vm_clone' gives explicit when-to-use guidance and names the relevant alternatives. It also adds a practical caveat about upload time and a failure condition (datastore not found), helping an agent decide when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deploy_vm_from_templateA
[WRITE] Deploy a new VM by cloning from a vSphere template.
Returns a status string. Use for VMs marked as templates; for a running source VM use vm_clone. Requires an existing template — convert_vm_to_template makes one.
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | Override CPU count (optional). | |
| target | No | Optional vCenter/ESXi target from config. | |
| new_name | Yes | Name for the new VM. | |
| power_on | No | Power on after deployment. | |
| memory_mb | No | Override memory in MB (optional). | |
| snapshot_name | No | Snapshot the new VM with this name. | |
| template_name | Yes | Source vSphere template name. | |
| datastore_name | No | Target datastore (template's own if omitted). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The [WRITE] prefix reinforces that this is a mutating operation, consistent with readOnlyHint=false. The description adds value by stating that a return status string is produced and by disclosing the prerequisite that an existing template is required. It does not dwell on safety since destructiveHint=false, so a 4 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence in the description earns its place: purpose, return value, correct use case, and prerequisite. It is compact, front-loaded, and avoids restating schema fields or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, annotations are present, and all parameters are already described in the schema, the description covers the remaining contextual needs: when to use it, what it returns, and what must exist beforehand. Nothing essential is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all eight parameters and their meanings. The description adds no parameter-specific detail beyond the prerequisite relationship to template_name, so the baseline score of 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Deploy a new VM by cloning from a vSphere template.' It also distinguishes itself from vm_clone by explicitly saying it is for VMs marked as templates, not running source VMs. This is enough for an agent to separate it from the relevant siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage conditions: use for templates, use vm_clone for a running source VM, and require an existing template, with convert_vm_to_template named as the way to create one. This is clear, actionable routing guidance that leaves little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
host_investigation_bundleARead-onlyIdempotent
[READ] "What is happening around this ESXi host?" — one correlated drill-down.
Correlates a host's state (connection, CPU/memory, ESXi version, uptime), its cluster context, a rollup of the VMs it runs, the datastores it mounts, alarms across host/cluster/datastore, live performance, and a merged event timeline. Delegates to the vmware-monitor library (read-only); do not dump the result raw.
Use this AFTER cluster_health_summary flags a host. Point-in-time.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | Event-timeline look-back window in hours (default 24). | |
| target | No | Optional vCenter/ESXi target name from config (default if omitted). | |
| host_name | Yes | Host name as shown in vCenter/ESXi, in any case, or its short form, or the host's management IP. On a standalone ESXi target, the target's name or tag also works. Unknown names list the hosts that exist. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavioral context beyond those: it delegates to the vmware-monitor library, instructs the agent not to dump the result raw, and notes the result is point-in-time. This helps set expectations about how the tool behaves and how its output should be consumed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear question and purpose, followed by a compact list of correlated data, then workflow guidance. It is a little list-heavy, but every sentence contributes to either purpose, usage, or safe handling. No filler or tautology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only investigation bundle with no output schema, the description covers the main data categories returned, the trigger condition, read-only safety, and the point-in-time nature. It does not describe return formatting or potential error behavior, which would make it fully complete, but it gives sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema coverage is 100%, with each parameter already described in detail (host_name aliasing, hours look-back, optional target). The description does not add parameter-level semantics beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, specific question and states the tool correlates a host's state, cluster context, VMs, datastores, alarms, live performance, and a merged event timeline. This clearly identifies the resource (host) and the action (correlated drill-down), and differentiates it from sibling investigation bundles by explicitly centering on the host.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit usage trigger: 'Use this AFTER cluster_health_summary flags a host.' This names a sibling tool and the condition under which this tool becomes relevant. It does not spell out when-not-to-use alternatives like vm_investigation_bundle or datastore_investigation_bundle, but the workflow guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_drs_rulesARead-onlyIdempotent
[READ] List a cluster's DRS rules: VM-VM affinity/anti-affinity and VM-Host.
Per rule: key, name, type (affinity / antiAffinity / vmHost), enabled, mandatory, and the member VM names (VM-VM) or group names (VM-Host). The verify pair for create/delete/set_drs_rule_enabled.
Returns: Dict with cluster, count, and rules sorted by name. Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| cluster | Yes | Exact cluster name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The [READ] prefix and the return-value contract (a dict with cluster, count, and rules sorted by name; errors as dict with error + hint) add meaningful behavior beyond the readOnlyHint and idempotentHint annotations. It also discloses per-rule fields, giving an agent a concrete expectation of what will be returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary action, then follows with per-rule details, usage context, and return format. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description explains the return structure and error format, and the input side is fully covered by the schema. The tool's purpose, scope, and edge-case behavior are all described, leaving no critical gap for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers both parameters (target and cluster) with descriptions and 100% coverage, so the description adds no parameter-specific meaning beyond what's structured. The mention of 'cluster' in the description is redundant with the schema's 'Exact cluster name.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'List a cluster's DRS rules' and enumerates the exact rule types (VM-VM affinity/anti-affinity and VM-Host), giving a specific verb and resource. It also distinguishes itself from sibling mutation tools by explicitly stating it is the verify pair for create/delete/set_drs_rule_enabled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states 'The verify pair for create/delete/set_drs_rule_enabled,' explicitly placing the tool as the read-back verification after those mutations. This is a clear, actionable guideline for when an agent should invoke it, and it distinguishes it from the three mutation siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dvs_portgroupsARead-onlyIdempotent
[READ] List distributed virtual portgroups, optionally scoped to one dvSwitch.
Per portgroup: name, parent dvSwitch, binding type (earlyBinding / ephemeral), VLAN setting (id, trunk ranges, or pvlan), configured port count, and whether it is the switch's uplink portgroup. Use to verify create_dvs_portgroup results or to survey network config before changes.
Returns:
The family list envelope {items, returned, limit, total, truncated,
hint}; each item is one portgroup. portgroups is kept as a deprecated
alias for items. Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max portgroups to return (default 200). | |
| offset | No | Skip this many portgroups first (paging). | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| dvs_name | No | dvSwitch name to scope to; omit to list across all switches. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, and the description adds substantial behavioral context: the exact fields returned per portgroup, the family list envelope, the deprecated alias for items, and the error dict shape. This goes well beyond what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action, followed by useful use-case and output details. Every sentence adds value: scoping, per-item fields, when to use it, and the response shape and error format. It is informative without being bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the absence of an output schema, the description fully explains the return envelope, item contents, deprecated alias, and error behavior. Combined with the well-documented input schema and safety annotations, an agent has enough to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters and their defaults. The description only reinforces that dvs_name scopes the listing, which adds context but does not materially expand on the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('List distributed virtual portgroups') and adds the optional dvSwitch scoping. It distinguishes itself from siblings like create_dvs_portgroup by clearly being the read-side counterpart, and no sibling covers the same resource and action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it: 'verify create_dvs_portgroup results or to survey network config before changes.' It does not spell out when not to use it or name alternative tools, so it stops one step short of the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_host_vmksARead-onlyIdempotent
[READ] List VMkernel adapters, optionally scoped to one ESXi host.
Per vmk: device, IP/netmask/dhcp, MTU, MAC, portgroup (standard or DVS), netstack, and which host services it is selected for (management, vmotion, vsan, ...). The verify pair for add_host_vmk/remove_host_vmk.
Gotcha - this list is not automatically a complete estate inventory.
vCenter answers property reads for a host it has lost contact with out of
its own cache, so hosts it never reached appear here as rows with
reachable: false, device: null and a note naming the connectionState,
rather than being dropped. Check hosts_unreachable (and the
unreachable_note present only when it is non-zero) before reporting the
result as the full picture, and do not read a null field on such a row as a
measurement - it means nobody looked. Adapters shown for an unreachable
host are vCenter's last cached view.
Returns:
The family list envelope {items, returned, limit, total, truncated,
hint} plus hosts_unreachable (int) and, when that is non-zero,
unreachable_note (str). Each item is one VMkernel adapter, or one
unread host when reachable is false; services is null when a host's
service map could not be read. truncated stays a paging fact - an
incomplete estate is reported by hosts_unreachable, not by it.
vmks is kept as a deprecated alias for items. Errors return
"error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max vmks to return (default 100). | |
| offset | No | Skip this many vmks first (paging). | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| host_name | No | ESXi host name; omit to list across all hosts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral context: vCenter serves cached data for unreachable hosts, rows appear with reachable:false and null device, hosts_unreachable and unreachable_note fields, and truncated is a paging fact rather than an estate-completeness signal. This is exactly the kind of non-obvious behavior an agent needs to interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place: the gotcha about unreachable hosts is critical and the return-format section is dense but necessary. The '[READ]' prefix and front-loaded purpose sentence make the core action immediately clear. It could be slightly tightened, but the length is justified by the behavioral complexity it discloses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description fully compensates: it explains the envelope fields, the unreachable-host edge case, the deprecated alias, and error shape. The annotations cover safety, the schema covers parameters, and the description covers interpretation. Nothing an agent needs to call this correctly and interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters (limit, offset, target, host_name). The description adds the semantic nuance that host_name scopes to one ESXi host and omitting it lists across all hosts, which is already in the schema. It doesn't add new parameter-level detail beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'List VMkernel adapters, optionally scoped to one ESXi host.' It then enumerates the per-vmk fields returned, which distinguishes it from sibling tools like add_host_vmk/remove_host_vmk and set_vmk_service. The 'verify pair' phrasing further anchors its role relative to mutation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it (list/verify VMkernel adapters) and names its sibling relationship: 'The verify pair for add_host_vmk/remove_host_vmk.' It also gives a strong when-not-to-trust-it warning: check hosts_unreachable before reporting the result as the full picture. This is explicit usage guidance beyond what the schema or annotations provide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_vcenter_alarmsARead-onlyIdempotent
[READ] List active/triggered alarms across the vCenter inventory.
Start here for alarm work: it supplies the exact entity_name/alarm_name pair that acknowledge_vcenter_alarm and reset_vcenter_alarm require.
Returns the list envelope: 'items' holds severity (critical/warning/info), entity name and type, alarm name, acknowledged flag, and trigger time; 'returned'/'limit'/'total'/'truncated'/'hint' state completeness, so a limited page is never mistaken for the whole picture. 'total' is the real active-alarm count — every alarm is collected before the limit is applied.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max number of alarms to return (None = all). Use when many alarms are active. | |
| target | No | Optional vCenter target name from config. Uses default if omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, and non-destructive. The description adds meaningful behavioral detail beyond those hints: it explains the pagination envelope, that 'total' is the real active-alarm count, and that every alarm is collected before the limit is applied. This prevents misinterpretation of truncated results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and usage guidance, then follows with the necessary return-envelope details. Each sentence adds relevant information, and there is no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates by fully specifying the return envelope fields and their meaning. The two optional parameters are already documented in the schema, and the safety profile is covered by annotations, so the description is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both parameters. The description adds value by clarifying pagination semantics around 'limit' and 'total'—specifically that 'total' is the true count and that limiting is applied after collection. This goes beyond the schema's simple 'max number to return' phrasing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'List active/triggered alarms across the vCenter inventory.' It also distinguishes this tool from alarm-related siblings by positioning it as the starting point that supplies the entity_name/alarm_name pair needed by acknowledge_vcenter_alarm and reset_vcenter_alarm.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Start here for alarm work' and explains that the output provides the exact pair that acknowledge_vcenter_alarm and reset_vcenter_alarm require. This clearly instructs the agent when to use this tool relative to its named siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_host_vmkADestructive
[WRITE] Remove a VMkernel adapter - confirm-gated, guarded, fail-closed.
REFUSES (unless force_unprotected=True) when the vmk is selected for any host service (management/vmotion/vsan/...), lives on a non-default netstack (NSX TEPs on vxlan, dedicated vmotion/provisioning stacks - never visible in the service map), carries a default gateway route, or when any of that CANNOT be verified - unverifiable is treated as unsafe, never as clear. Test vmks created by add_host_vmk trip none of these and remove cleanly without force.
ABSOLUTE, no override: the host's only management-enabled vmk is never removable - this call rides the interface it would delete.
force_unprotected=True (together with confirm=True) overrides the non-absolute protections for deliberate teardown; the override and every bypassed protection are recorded in the result (and the audit trail).
Returns: Preview dict (action="preview") or result dict (action="removed", plus forced/protections_bypassed when overridden). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| vmk | Yes | Device name to remove (e.g. "vmk2"). | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True removes it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| host_name | Yes | ESXi host the vmk lives on. | |
| force_unprotected | No | True bypasses the non-absolute protections above. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations' destructiveHint, the description explains the confirm-gated preview flow, fail-closed semantics for unverifiable state, the absolute protection on the management vmk, and the audit trail. It also details return shapes for preview, success, and error cases. This is rich behavioral context that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place for a destructive, guarded operation. It front-loads the core action and safety posture, then covers guardrails, overrides, and return behavior without wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a dangerous mutation with no output schema, the description fully explains preview vs. removal results, error shape, override behavior, and audit outcomes. Combined with fully documented schema parameters secret, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds extra meaning to confirm and force_unprotected (e.g., "confirm=False returns the blast radius", "force_unprotected=True together with confirm=True overrides"), which clarifies the interplay between parameters beyond their individual schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with "[WRITE] Remove a VMkernel adapter," giving a specific verb and resource. It clearly distinguishes itself from sibling tools like add_host_vmk and list_host_vmks by scoping exactly what it does and its destructive nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-not conditions: it refuses for management-enabled vmks, service-bound vmks, non-default netstacks, and default gateway routes unless forced. It also identifies a safe use case ("Test vmks created by add_host_vmk trip none of these"). It doesn't explicitly route to sibling alternatives for inspecting or reconfiguring vmks, but the exclusion conditions are strong enough for an agent to gauge suitability.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_vcenter_alarmA
[WRITE] Clear triggered vCenter alarms back to normal state.
Use this after resolving the underlying issue; to merely mark an alarm seen use acknowledge_vcenter_alarm instead. Get entity_name and alarm_name from list_vcenter_alarms first.
Gotcha: vSphere has no per-alarm clear — this clears ALL triggered alarms matching the named alarm's entity type (host/VM/all) and current status (red/yellow), so confirm the blast radius with the user first. Returns a dict whose 'scope' field states exactly what was cleared.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter target name from config. | |
| alarm_name | Yes | Exact alarm definition name from list_vcenter_alarms output. | |
| entity_name | Yes | Name of the entity with the alarm (VM name, host name, or cluster name). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, destructiveHint=false), the description discloses the crucial gotcha: vSphere has no per-alarm clear, so this tool clears ALL triggered alarms matching the entity type and status. It also discloses the return value's shape (a dict with a 'scope' field). This is significant behavioral context that the annotations do not provide, making the 'blast radius' explicit and reducing the risk of unintended side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear [WRITE] tag and a one-sentence purpose. Every subsequent sentence adds essential usage or behavioral information: when to use, the alternative, the parameter source, the critical gotcha, and the return value. There is no fluff or redundancy; each line earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description explains what the tool returns ('a dict whose 'scope' field states exactly what was cleared'). It covers the core workflow (fetch parameters from list_vcenter_alarms), the safety caveat (blast radius), and the conceptual difference from acknowledge_vcenter_alarm. For a tool with modest complexity and rich annotations, this description is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% parameter description coverage, including where alarm_name comes from (list_vcenter_alarms output) and what entity_name means. The description's instruction to 'Get entity_name and alarm_name from list_vcenter_alarms first' restates what the schema already says for alarm_name, adding no new semantic meaning to the parameters. It is useful as a usage ordering cue, but not as parameter semantics, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Clear triggered vCenter alarms back to normal state') and differentiates from the sibling tool by explicitly naming acknowledge_vcenter_alarm as the alternative for merely marking an alarm seen. This makes the tool's purpose unambiguous and clearly distinguishes it from closely related operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance ('Use this after resolving the underlying issue'), identifies the alternative (acknowledge_vcenter_alarm), and instructs the agent to obtain parameters from list_vcenter_alarms first. It also flags a critical prerequisite by advising to confirm the blast radius with the user. This is exactly the practical routing information an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_datastore_imagesARead-onlyIdempotent
[READ] Scan all accessible datastores for deployable images (OVA/ISO/OVF/VMDK).
Returns the images found and refreshes the cache at ~/.vmware-knight/image_registry.json. Use this when you do not know which datastore holds an image; prefer browse_datastore once you do, because this walks every datastore and may take minutes on a large estate.
Returns:
The family list envelope {items, returned, limit, total, truncated,
hint} plus last_scan; each item is one image. images is kept as a
deprecated alias for items.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the safety profile is covered. The description adds valuable context: it refreshes a local cache file, may be slow on large estates, and describes the return envelope (including a deprecated alias). This goes beyond what annotations convey and is consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized: a one-line purpose, a usage note, and a returns section. It front-loads the core action and packs the return format into a compact paragraph. While not as terse as the best examples, every sentence earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a scan tool with one optional parameter and no output schema, the description explains the return envelope, the cache side effect, and the performance caveat. It gives the agent everything needed to decide when to call it and what to expect, making it complete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'target' is fully described in the schema (coverage 100%) as 'Optional vCenter/ESXi target name from config.' The description does not add any additional meaning or syntax details about this parameter, so the baseline of 3 applies since the schema handles it completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair, 'Scan all accessible datastores for deployable images (OVA/ISO/OVF/VMDK)', and immediately contrasts itself with browse_datastore, so an agent can distinguish it without opening the schema. It names the exact image types it looks for, which is concrete and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the triggering condition ('Use this when you do not know which datastore holds an image') and the alternative ('prefer browse_datastore once you do'), including a rationale about performance ('walks every datastore and may take minutes on a large estate'). This is textbook when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_drs_rule_enabledAIdempotent
[WRITE] Enable or disable an existing DRS rule - preview/confirm gated.
The day-2 toggle: anti-affinity rules often must be disabled while a cluster is temporarily too small to satisfy them, then re-enabled when hosts return. Idempotent - matching state returns a noop, no write. Names are matched exactly; ambiguous names refuse. Audited.
Returns: Preview dict (action="preview"), noop dict (action="noop"), or result dict (action="set", rule_now). Errors return "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| cluster | Yes | Exact cluster name. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| enabled | Yes | True enables the rule; False disables it. | |
| rule_name | Yes | Exact rule name (see list_drs_rules). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations say write, idempotent, non-destructive; the description adds the preview/confirm gate, noop on matching state, exact-match/ambiguous-name behavior, and audit trail. This is exactly the behavioral context an agent needs beyond the flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Starts with a purpose + write marker, then a short motivating scenario, then bullets for idempotency/name matching/audit and return shapes. No filler; all sentences carry information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies return variants (preview/noop/set/error) and key fields, plus safety and matching behavior. An agent can invoke it correctly without needing external lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and parameter descriptions are already strong, so the baseline applies. The description reinforces exact rule_name matching and preview/confirm behavior but does not need to explain fields the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states specific verb 'enable or disable' on 'existing DRS rule' and immediately distinguishes the tool from the create/delete/list DRS siblings. The preview/confirm gating and day-2 toggle context make its niche unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete scenario: anti-affinity rules disabled when cluster temporarily too small and re-enabled when hosts return. It does not explicitly name sibling alternatives, but 'existing rule' and 'day-2 toggle' make the boundary to create/delete obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_vmk_serviceAIdempotent
[WRITE] Enable/disable a host service on an existing vmk - preview/confirm gated.
Completes the add_host_vmk story: adapters are created serviceless by design, then tagged here (e.g. enable "vmotion" on a new vMotion vmk). Idempotent - re-applying the current state returns a no-write noop. Verify with list_host_vmks (the services field). Audited.
Valid services are the vSphere nicType names: management, vmotion, vsan, vSphereProvisioning, faultToleranceLogging, vSphereReplication, vSphereReplicationNFC, vSphereBackupNFC, ptp, and the nvme/vsan variants. Note "vSphereProvisioning", not "provisioning".
FAIL CLOSED: refuses both directions when the host's service map cannot be read. ABSOLUTE, no override: disabling management on the host's only management-enabled vmk - the call rides the interface it would untag.
Returns: Preview dict (action="preview"), noop dict (action="noop") when the state already matches, or result dict (action="set", services_now). Errors return a dict with "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| vmk | Yes | Device name to change (e.g. "vmk3"). | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. Do not set True because the user asked earlier; they have not seen the preview yet. | |
| enabled | Yes | True selects the vmk for the service; False deselects. | |
| service | Yes | Service/nicType name to enable or disable. | |
| host_name | Yes | ESXi host the vmk lives on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses preview/confirm gating, idempotent no-write noop behavior, audit logging, fail-closed behavior when the service map cannot be read, and the absolute safeguard against disabling management on the only management-enabled vmk. This is substantial non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every sentence earns its place: it front-loads the operation and gating model, organizes safety behaviors clearly, and ends with a concise return-value breakdown. There is no filler or repeated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of an output schema, the description is unusually complete. It covers valid inputs, preview/confirm behavior, idempotency, failure modes, safety guards, and the exact shapes of all return dicts. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description adds real value for parameters: it enumerates valid nicType service names, warns about the exact 'vSphereProvisioning' spelling, and clarifies confirm semantics, including the instruction not to set confirm=true before the user has seen the preview. This goes well beyond the schema's basic descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with '[WRITE] Enable/disable a host service on an existing vmk - preview/confirm gated,' naming a specific verb, resource, and operation mode. It also distinguishes itself from add_host_vmk by explaining that adapters are created serviceless and then tagged by this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it completes the add_host_vmk story and tells the agent to verify results with list_host_vmks. It does not explicitly list when-not-to-use conditions or contrast with every sibling, but the context is sufficiently clear for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_apply_planADestructive
[WRITE] Execute a previously created plan step by step.
Without confirm=True this only previews: it returns blast_radius listing every step (index, action, target object), the count and indices of the destructive ones (delete, power off, revert, guest commands...), and blockers — and runs nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Each step is shown with its full parameters (passwords redacted). Each destructive step is measured as its own tool would measure it (vm_delete, vm_power_off, vm_revert_snapshot, vm_guest_exec, cluster_delete ...), and that tool's blockers refuse the plan. A step on something an earlier step creates or changes is marked check "deferred": it is measured immediately before it runs, and the plan stops there if it fails.
Refused: a target other than the one the plan was created against (no target is a target of its own), a step its tool would refuse, anything it could not read, a delete_vm step without its acknowledge_blast_radius (from a vm_delete preview), and any iscsi_* or storage_rescan step — those are gated in vmware-storage; run storage_iscsi_* / storage_rescan there.
With confirm=True steps run sequentially. On failure: stops immediately, keeps the plan file with per-step results, and returns rollback_available. On success: deletes the plan file. If a step fails and rollback_available is true, ask the user whether to rollback, then call vm_rollback_plan.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | The vCenter/ESXi target the plan was created against. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| plan_id | Yes | The plan ID returned by vm_create_plan. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds substantial behavioral detail: confirms are preview-only, destructive steps inherit the measuring and blocking semantics of their underlying tools, deferred steps are measured immediately before execution, the plan file is deleted on success and retained on failure, and rollback_available is returned after a failure. This is well beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured into preview behavior, measurement semantics, refusal cases, and execution outcomes. It is front-loaded with the central purpose and every sentence carries safety-relevant information, though a few clauses are dense enough that they could be tightened without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain return behavior, and it does: it lists what the blast_radius contains, how destructive steps are measured, what happens on failure (stops, keeps plan file, returns rollback_available) and success (deletes plan file). Combined with the cross-tool routing and rollback guidance, nothing essential is missing for an agent to invoke this tool safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the schema already documents plan_id, target, and confirm. The description adds extra semantic nuance around confirm=False/True behavior, the requirement that target match the plan's original target, and refusal cases such as delete_vm requiring acknowledge_blast_radius. That pushes it above baseline but not to 5, since the schema already carries most parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Execute a previously created plan step by step.' It immediately separates this tool from plan-creation and rollback siblings by naming the apply workflow and explicitly referencing vm_rollback_plan and vm_create_plan contexts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is explicit about when to use the tool and what the agent must do first: show the preview and get explicit user decision, never set confirm=True on its own, and ask before rolling back. It also lists excluded step types (iscsi_*, storage_rescan) and routes the agent to the correct sibling tools, which is strong, actionable guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_cancel_ttlAIdempotent
[WRITE] Cancel an existing TTL for a VM (prevents auto-deletion).
Returns a status string. Use vm_list_ttl first for the exact vm_name. This only removes the schedule and never touches the VM itself.
| Name | Required | Description | Default |
|---|---|---|---|
| vm_name | Yes | Name of the VM whose TTL should be cancelled. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a write operation and non-destructive behavior; the description adds specific context: 'only removes the schedule and never touches the VM itself.' This clarifies the scope beyond the raw annotation flags without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler. The [WRITE] marker and action are front-loaded, followed by return type, prerequisite, and safety clarification—each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter operation with an output schema and annotations, the description covers purpose, return behavior, prerequisite, and non-destructive scope. Nothing necessary for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter, so the baseline is 3. The description adds the practical requirement to obtain the 'exact vm_name' via vm_list_ttl, which goes beyond the schema's generic parameter description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Cancel an existing TTL for a VM' with the effect 'prevents auto-deletion.' The mention of vm_list_ttl for the exact vm_name disambiguates it from the related TTL tools in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs the agent to 'Use vm_list_ttl first for the exact vm_name,' establishing the correct precondition. It does not explicitly state when not to use this tool, but the workflow context is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_clean_slateADestructive
[WRITE] Revert a VM to its baseline snapshot (Clean Slate).
Without confirm=True this only previews: it returns blast_radius (VM identity, power state and whether it is powered off first, the snapshot and when it was taken, snapshot count, blockers) and changes nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
With confirm=True: powers off the VM first if it is running, then reverts to the named snapshot. Use this to reset a lab/dev VM to a clean starting state after a task completes. Irreversible — everything written since the snapshot is lost. Refused: no snapshot of that name, more than one, or a VM whose state cannot be read. Returns a dict (action, blast_radius).
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| vm_name | Yes | Name of the VM to revert. | |
| snapshot_name | No | Snapshot name to revert to (default: "baseline"). | baseline |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by explaining the preview-only mode, blast_radius contents, automatic power-off behavior, irreversibility, refusal conditions, and the return dict. An agent knows exactly what will happen in each confirm mode and what can go wrong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, dense paragraphs front-load the action and then walk through preview, apply, irreversibility, and refusal cases. Every sentence earns its place; the length is justified by the destructive and guarded nature of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema, the description supplies everything needed: mode semantics, side effects, failure conditions, and return shape. The remaining parameter details are already covered by the input schema, so there is no essential gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions already cover all parameters, so the baseline is 3. The description adds meaningful value by explaining the safety-critical semantic difference between confirm=False and confirm=True, including the power-off side effect and the named snapshot default. This reinforces rather than merely repeats the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: 'Revert a VM to its baseline snapshot (Clean Slate).' This clearly scopes the tool to baseline reset rather than generic snapshot operations, and the [WRITE] marker matches the mutating nature. The purpose is immediately distinguishable from siblings like vm_revert_snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear guidance on when to use it ('reset a lab/dev VM... after a task completes') and strong guardrail behavior ('Do not set confirm=True on your own... Show that to the user'). It does not explicitly name the alternative tool for arbitrary snapshot reverts, so it misses the highest bar, but the usage context is otherwise explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_cloneA
[WRITE] Clone a VM. Without to_host/to_datastore the clone lands on the source's host+datastore.
Returns a status string naming the clone. Full independent copy — slow and full disk cost; prefer deploy_linked_clone for near-instant test copies and batch_clone_vms for many at once. Cloning a running VM may capture a crash-consistent disk.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config. | |
| to_host | No | Target ESXi host name (default: source's host). | |
| vm_name | Yes | Source VM (or template) name. | |
| new_name | Yes | Name for the new clone. | |
| power_on | No | Power on the clone after creation. | |
| to_datastore | No | Target datastore name (default: source's datastore). |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnly=false and destructive=false. The description adds valuable behavioral context: the clone is a full independent copy, it is slow and consumes full disk space, and cloning a running VM may capture a crash-consistent disk. It also discloses the return value format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded paragraphs with no filler. The '[WRITE]' marker, core operation, default behavior, return type, cost/performance caveat, alternatives, and consistency warning all earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with six parameters and an output schema, the description fully covers the practical calling context: placement defaults, return value, cost profile, alternatives, and runtime consistency caveat. Nothing needed for correct selection or invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all six parameters with 100% coverage, so the baseline is 3. The description synthesizes the to_host/to_datastore defaults into one phrase ('source's host+datastore'), but this largely restates what the schema already says rather than adding new parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Clone a VM') and makes the output explicit ('Returns a status string naming the clone'). It also distinguishes itself from siblings by naming deploy_linked_clone and batch_clone_vms as different tools for different use cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to prefer alternatives: 'prefer deploy_linked_clone for near-instant test copies and batch_clone_vms for many at once.' It also provides a practical caveat about cloning a running VM, which helps the agent decide whether this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_createA
[WRITE] Create a new empty VM with the given hardware sizing.
Creates a powered-off VM with one disk and one NIC. To populate it, attach an ISO (attach_iso_to_vm) and power it on, or use deploy_vm_from_ova or vm_clone for a ready-to-run guest. Fails before creating anything if the datastore is not found. Returns a status string with the new VM name.
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | vCPU count (default 2). | |
| target | No | vCenter/ESXi target from config.yaml; omit for the default. | |
| disk_gb | No | Primary disk size in GB (default 40). | |
| vm_name | Yes | Name for the new VM; must not already exist. | |
| memory_mb | No | Memory in MB (default 4096). | |
| folder_path | No | vCenter folder path; omit for the datacenter root. | |
| network_name | No | Port group for the NIC (default "VM Network"). | VM Network |
| datastore_name | No | Target datastore; omit for the first accessible one. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a non-read-only, non-idempotent, non-destructive write operation. The description adds useful behavioral context beyond that: the VM starts powered off, includes one disk and one NIC, fails atomically if the datastore is not found, and returns a status string with the new VM name. It does not cover permission requirements or cleanup behavior, but it is solid given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, then provides workflow guidance and failure behavior in just four sentences. Every sentence adds value, and no space is wasted repeating schema field details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich input schema, output schema, and annotations, the description supplies the missing behavioral and workflow context needed for correct use: what kind of VM is created, how to populate it, and what happens on datastore failure. It is complete enough for an agent to select and invoke the tool appropriately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and every parameter already has a clear description with defaults. The description adds little parameter-specific meaning beyond tying the hardware sizing parameters together ('one disk and one NIC'), so the schema carries the heavy lifting. This matches the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a new empty VM') and clarifies the result: a powered-off VM with one disk and one NIC. It distinguishes itself from siblings like deploy_vm_from_ova and vm_clone by emphasizing that this creates an empty VM rather than a ready-to-run guest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when to use this tool versus alternatives: attach an ISO and power on, or use deploy_vm_from_ova or vm_clone for a ready-to-run guest. This gives clear decision guidance and names the sibling tools, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_create_planA
[WRITE] Create an execution plan for multi-step VM operations.
Use for 2+ steps or 2+ VMs. Validates actions, checks the targets exist in vSphere, and generates a plan with rollback info per step.
Each operation is a dict with "action" key plus action-specific params. Allowed actions: power_on, power_off, reset, suspend, create_vm, delete_vm, reconfigure, create_snapshot, delete_snapshot, revert_snapshot, clone, migrate, deploy_ova, deploy_template, linked_clone, attach_iso, convert_to_template.
Returns plan dict with plan_id, steps, summary (vms_affected, irreversible_steps, rollback_available). Show to user for confirmation before calling vm_apply_plan.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| operations | Yes | List of operation dicts, each with "action" + params. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses that the tool validates actions, checks that targets exist in vSphere, and generates rollback information per step. It also states the important workflow guardrail that the plan must be shown for confirmation before applying. It does not detail plan persistence or lifecycle, but the annotations already cover the basic safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized: purpose, usage condition, behavior, parameter format, allowed actions, and return contract are each presented clearly. It is longer than minimal, but the allowed action list is necessary because the schema does not provide enum constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex planning tool with no output schema, the description covers the full call contract: operation format, allowed actions, validation behavior, return structure, and the follow-up step with vm_apply_plan. The workflow implication that this tool does not execute the VM operations is clear from the explicit 'before calling vm_apply_plan' instruction.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description adds real value by specifying that each operation is a dict with an 'action' key and enumerating all allowed action values. This compensates for the schema's generic items object and lack of enums. The target parameter is already adequately documented in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Create an execution plan for multi-step VM operations.' It clearly distinguishes this from direct VM-operation siblings by framing it as the planning step, and it differentiates from vm_apply_plan by calling out the handoff workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use for 2+ steps or 2+ VMs,' giving a concrete criterion for when this tool should be preferred. It also clarifies that the resulting plan should be shown to the user before vm_apply_plan is called, though it does not explicitly name the single-operation siblings as alternatives to use instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_create_snapshotA
[WRITE] Create a snapshot of a VM.
Returns a status string. Use this before a risky change so vm_revert_snapshot can undo it, then reclaim the space with vm_delete_snapshot — snapshots left for days grow delta disks and must not be treated as backups.
| Name | Required | Description | Default |
|---|---|---|---|
| memory | No | Include memory state (heavier, allows resume). | |
| target | No | vCenter/ESXi target name from config. | |
| quiesce | No | Quiesce guest filesystem (requires running VMware Tools). | |
| vm_name | Yes | VM to snapshot. | |
| description | No | Optional description. | |
| snapshot_name | Yes | Snapshot name. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which are minimal), the description discloses that the tool returns a status string and warns about delta-disk growth. It does not elaborate on async behavior or failure modes, but the annotations and output schema cover the core write/safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the action and return type, and every sentence serves a purpose. The lifecycle warning and sibling references add value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create-snapshot tool with an output schema and a 100%-described input schema, the description provides the missing operational context: when to snapshot, how to undo, how to reclaim space, and why snapshots are not backups. Nothing essential is left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents every parameter, including defaults and descriptions, so the description carries little additional parameter-level semantics. It adds global context about snapshot lifecycle but does not deepen understanding of individual parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Create a snapshot of a VM') and immediately orients the agent by referencing the related revert and delete siblings. This makes the tool's purpose unambiguous and distinguishes it from nearby snapshot lifecycle tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool ('before a risky change') and when to follow up with vm_revert_snapshot and vm_delete_snapshot. It also warns against treating snapshots as backups, leaving no ambiguity about appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_deleteADestructive
[WRITE] Delete a VM and its disks and snapshots (irreversible).
Without confirm=True this only previews: it returns blast_radius (identity, host, disks, total size, snapshot count, blockers) and destroys nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user said "delete" earlier: they have not seen what it destroys yet.
To delete, call again with confirm=True and acknowledge_blast_radius set to the preview's acknowledge_with object, unchanged. The VM is re-measured first; if it changed (another snapshot, a different VM under the same name), nothing is deleted and you must preview again.
Refused outright: a powered-on or suspended VM (power it off with vm_power_off first), a VM whose disks or identity cannot be read, and a name that matches more than one VM. Use vm_set_ttl instead when the VM should only expire later.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config. | |
| confirm | No | False (default) previews; True deletes, with the acknowledgement. | |
| vm_name | Yes | Exact name of the VM to delete. | |
| acknowledge_blast_radius | No | The preview's acknowledge_with object. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, and the description goes well beyond them. It explains the two-phase preview/delete behavior, the re-measurement safety check before deletion, the irreversible destruction of disks and snapshots, and the exact refusal conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence carries operational weight, and the most critical facts (irreversibility, two-phase behavior, user consent) are front-loaded. The structure moves from purpose to preview requirement to deletion protocol to refusals and alternatives without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description explains what the preview returns (blast_radius with identity, host, disks, size, snapshot count, blockers, and acknowledge_with) and covers all refusals and safety checks. An agent has everything needed to call this tool safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds significant meaning beyond the schema: confirm=False previews while confirm=True deletes, acknowledge_blast_radius must be the preview's acknowledge_with object unchanged, and the toolbox is re-measured before deletion. This makes the two-phase contract unambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: delete a VM, including its disks and snapshots, and notes irreversibility. It also distinguishes this tool from related relatives by directing the agent to vm_set_ttl for expiration-only cases and vm_power_off for powered-on VMs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is explicitly gated: preview first with confirm=False, show the blast radius to the user, then delete with confirm=True and the unchanged acknowledge_blast_radius. It gives clear when-not-to-use conditions (powered-on/suspended VMs, unreadable disks, ambiguous names) and names the alternative tool for each relevant case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_delete_snapshotADestructive
[WRITE] Permanently delete a named snapshot, consolidating its delta disk into the parent.
Without confirm=True this only previews: it returns blast_radius (VM and instance UUID, the snapshot's name, id and creation time, remove_children, how many child snapshots sit below it and how many snapshots would be removed) and deletes nothing. Show it to the user and get their decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Frees disk space and does NOT change the VM's current state (unlike vm_revert_snapshot). Works while the VM is powered on. Refused: a snapshot name that is not found, a name that matches more than one snapshot on the VM (vSphere allows duplicates; rename one first), and a snapshot tree that cannot be read. Run vm_list_snapshots first for exact names.
Consolidation is slow for old/large deltas (often minutes). By default (wait=False) this returns a task id immediately so it does not block your context — poll it with vm_task_status. Set wait=True only for small snapshots (blocks up to 30 min).
Returns: Dict with action (preview, snapshot_delete_started, snapshot_deleted), blast_radius, and result (carries the task id to poll via vm_task_status).
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | False (default) = async, return task id at once; True = block. | |
| target | No | vCenter/ESXi target from config.yaml; omit for the default target. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| vm_name | Yes | Exact name of the VM owning the snapshot. | |
| snapshot_name | Yes | Exact snapshot name from vm_list_snapshots output. | |
| remove_children | No | False (default) = children are kept and consolidated; True = delete the entire snapshot subtree below this one as well. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, but the description adds crucial behavioral detail: confirm=False only previews, confirm=True actually deletes, the VM state is not changed, deletion works while powered on, and refusal cases are enumerated. It also warns that consolidation is slow and that wait=False returns a task id immediately. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every paragraph earns its place: safe preview behavior, error conditions, performance warning, and return value mapping. It is front-loaded with the critical [WRITE] marker and the 'do not confirm without showing preview' directive, and the structure makes the information scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, yet the description compensates by stating the action values (preview, snapshot_delete_started, snapshot_deleted), the blast_radius contents, and how to poll the result. It also covers failure modesamentations and performance expectations. An agent has everything needed to call this safely and correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description clearly adds value beyond the schema: it explains the consequence of confirm=False (preview with blast_radius), the meaning of wait=False (async task id), and the effect of remove_children on the snapshot tree. This elevates the parameter understanding beyond raw field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Permanently delete a named snapshot, consolidating its delta disk into the parent.' It also distinguishes itself from vm_revert_snapshot, making its purpose and scope immediately clear. An agent can confidently separate this from sibling tools like vm_create_snapshot and vm_list_snapshots.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use and when-not-to-use guidance: run vm_list_snapshots first for exact names, never set confirm=True before the user has seen the preview, and prefer vm_task_status polling over blocking. It even names the sibling that behaves differently (vm_revert_snapshot). This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_guest_downloadA
[WRITE] Download a file from a VM and write it to a local path.
Reads from the guest, writes the local filesystem — the write is why this is not a read tool. Returns a status string. Requires VMware Tools running in the guest OS. Use vm_guest_upload for the reverse direction; to capture command output use vm_guest_exec_output instead — it redirects and downloads for you.
Refuses a destination that already exists unless overwrite=True, and never writes through a symlink or over a directory. Pick a path that does not exist yet rather than passing overwrite=True by default.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| vm_name | Yes | Target VM name. | |
| password | No | Guest OS password. | |
| username | Yes | Guest OS account to run as. Required — there is no default, so a call can never act as root without choosing root. | |
| overwrite | No | True replaces an existing file at local_path (default False). | |
| guest_path | Yes | File path inside the guest to download. | |
| local_path | Yes | Local destination path, including the file name. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond annotations: it discloses the local-filesystem write side effect, the refusal of existing destinations unless overwrite=True, refusal to write through symlinks or over directories, and the fact that it returns a status string. This matches the annotations (readOnlyHint=false) and adds meaningful behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a one-line purpose and a [WRITE] marker, then groups alternatives, prerequisites, and safety semantics. Every sentence earns its place; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers operation, direction, prerequisites, alternatives, return type, and important safety constraints. Combined with the schema and annotations, an agent has everything needed to invoke the tool correctly. An output schema exists, so detailed return-value documentation is not required here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description adds valuable semantics for local_path and overwrite by explaining refusal behavior, symlink/directory handling, and recommending a non-existent destination rather than overwrite=True by default. This exceeds the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Download a file from a VM and write it to a local path.' It also explicitly contrasts itself with vm_guest_upload and vm_guest_exec_output, so an agent can distinguish this tool from nearby siblings without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use context by naming vm_guest_upload as the reverse direction and vm_guest_exec_output for capturing command output. It also states a prerequisite ('Requires VMware Tools running in the guest OS') and advises against defaulting to overwrite=True, giving actionable selection and invocation guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_guest_execADestructive
[WRITE] Execute a command inside a VM via VMware Tools.
Requires VMware Tools running in the guest OS. Returns exit_code, stdout, stderr, timed_out, and blast_radius.
Without confirm=True this only previews: blast_radius names the VM (name, instance UUID), the guest account, the exact command, arguments and working directory, VMware Tools status and any blockers; nothing runs. Show it to the user. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet. Refused outright: VM not powered on, VMware Tools not running, an identity or status that cannot be read, or a name that matches more than one VM.
Note: the Guest Ops API does not capture stdout/stderr directly, so use this only for fire-and-forget commands — prefer vm_guest_exec_output whenever you need the output.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| command | Yes | Full path to program (e.g. "/bin/bash", "C:\Windows\System32\cmd.exe"). | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| vm_name | Yes | Target VM name. | |
| password | No | Guest OS password. | |
| username | Yes | Guest OS account to run as. Required — there is no default, so a call can never act as root without choosing root. | |
| arguments | No | Command arguments (e.g. "-c 'whoami'"). | |
| working_directory | No | Working directory inside guest (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by explaining the confirmation-gated preview behavior, required VMware Tools, what the blast_radius contains, refusal conditions, and the limitation that the Guest Ops API does not capture stdout/stderr directly. Since annotations only mark readOnlyHint=false and destructiveHint=true, the description carries the behavioral burden and does so thoroughly. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every section earns its place: purpose, prerequisite, return fields, confirmation policy, refusals, and the alternative sibling. The [WRITE] tag and first line front-load the core purpose. It is slightly dense but well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the returned fields (exit_code, stdout, stderr, timed_out, blast_radius), explaining behavior, listing refusal conditions, and naming the preferred sibling for output capture. The 8 parameters are all documented in the schema, so no critical gap remains for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds context around the confirm parameter and blast_radius output, but most parameter-level meaning already exists in the schema. It does not materially deepen per-parameter semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Execute a command inside a VM via VMware Tools.' It also distinguishes itself from the sibling vm_guest_exec_output by explicitly noting this tool is for fire-and-forget commands. An agent can clearly understand what this tool does and how it differs from nearby alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'use this only for fire-and-forget commands — prefer vm_guest_exec_output whenever you need the output.' It also provides clear operational rules around confirm=True, saying not to set it on the agent's own initiative and to show the preview to the user. Refusal conditions are enumerated, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_guest_exec_outputADestructive
[WRITE] Execute a shell command inside a VM and capture stdout + stderr.
Automatically detects guest OS (Linux/Windows) and selects the correct shell. Output is captured by redirecting to a temp file, downloading it, then cleaning up — no manual redirection needed. Prefer this over vm_guest_exec whenever you need the output. Requires VMware Tools running and a writable temp directory in the guest.
Returns exit_code, stdout, stderr, timed_out, os_family, and blast_radius.
Without confirm=True this only previews: blast_radius names the VM (name, instance UUID), the guest account, the exact command and the shell it runs through, VMware Tools status and any blockers; nothing runs. Show it to the user. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet. Refused outright: VM not powered on, VMware Tools not running, an unreadable identity or status, or a name that matches more than one VM.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| command | Yes | Shell command (e.g. "df -h", "ls /etc", "ipconfig"). | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| timeout | No | Max wait seconds (default 300). | |
| vm_name | Yes | Target VM name. | |
| password | No | Guest OS password. | |
| username | Yes | Guest OS account to run as. Required — there is no default, so a call can never act as root without choosing root. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations: it explains the [WRITE] nature, the temp-file output capture and cleanup mechanism, the non-destructive preview when confirm=False, what blast_radius contains, and exactly when the tool refuses to run. Annotations already mark destructiveHint=true, and the description enriches this with concrete safety behavior without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although lengthy, the description is dense and well-structured: it leads with the core action and result, then mechanism, then selection guidance, then preview/refusal behavior. Every sentence adds operational or safety information that matters for invoking the tool correctly. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description compensates by naming the return fields (exit_code, stdout, stderr, timed_out, os_family, blast_radius). It also covers prerequisites, preview contents, refusal conditions, and the approved workflow of showing the preview to the user. This is complete for a complex, potentially destructive tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers all 7 parameters (100% coverage), so the baseline is 3. The description adds critical semantics beyond the schema, especially the confirm parameter: 'Without confirm=True this only previews... nothing runs' and the instruction not to set confirm=True on the agent's own initiative. This materially affects correct usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Execute a shell command inside a VM and capture stdout + stderr.' It clearly distinguishes itself from the sibling tool vm_guest_exec by emphasizing output capture and by explicitly naming vm_guest_exec as the alternative. This removes ambiguity at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit selection guidance: 'Prefer this over vm_guest_exec whenever you need the output.' It also states prerequisites (VMware Tools running, writable temp directory) and refusal conditions (VM not powered on, Tools not running, ambiguous name). This is strong, actionable usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_guest_provisionADestructive
[WRITE] Provision a VM by running an ordered sequence of guest operations.
Prefer this over repeated vm_guest_exec / vm_guest_upload calls when the steps form one provisioning run. Steps stop on the first failure, so a partial run leaves the guest half-configured. Requires VMware Tools running in the guest.
Without confirm=True this only previews: blast_radius names the VM (name, instance UUID), the guest account and every step it would run, in order, with counts per type, local file sizes and any blockers; nothing runs. Show it to the user. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet. Refused outright: VM not powered on, VMware Tools not running, an empty step list, a step with an unknown type or a missing key, an upload whose local file is missing or unreadable, a service step on a Windows guest, an unreadable identity or status, or a name that matches more than one VM.
Step types:
exec: {"type": "exec", "command": "apt-get install -y nginx"}
upload: {"type": "upload", "local_path": "...", "guest_path": "..."}
service: {"type": "service", "name": "nginx", "action": "start"}
Returns: dict with success, completed_steps, total_steps, results, error, and blast_radius.
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | Ordered list of step dicts (see Step types). | |
| target | No | Optional vCenter/ESXi target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| timeout | No | Per-step timeout in seconds (default 300). | |
| vm_name | Yes | Target VM name. | |
| password | Yes | Guest OS password. | |
| username | Yes | Guest OS username. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a destructive, non-read-only, non-idempotent operation. The description adds substantial behavioral context: steps stop on first failure leaving the guest half-configured, confirm=False only previews while nothing runs, and the full refusal list is disclosed. This goes well beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though long, the description is tightly organized and every section earns its place: overview, usage guidance, preview semantics, refusal conditions, step-type examples, and return keys. The most important safety information is front-loaded, and the formatting makes it scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no output schema, the description is remarkably complete. It explains when to use it, prerequisites, failure behavior, preview mode, all refusal conditions, step schema, and return value keys. An agent has enough information to invoke it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameter names and defaults are already documented. The description adds critical extra meaning: concrete step type schemas with JSON examples, the semantics of confirm (preview vs apply), and the ordered-list behavior for steps. This meaningfully supplements the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Provision a VM by running an ordered sequence of guest operations.' It also distinguishes itself from sibling tools by explicitly preferring this over repeated vm_guest_exec / vm_guest_upload calls when steps form one provisioning run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance ('Prefer this over repeated vm_guest_exec / vm_guest_upload calls when the steps form one provisioning run'), names the alternatives, and states a key prerequisite (VMware Tools running in the guest). It also explains when not to apply directly via confirm=True, since the user hasn't seen the preview yet.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_guest_uploadADestructive
[WRITE] Upload a file from local machine to a VM via VMware Tools.
Returns a dict with action, message and blast_radius (a status string before the confirmation gate). Requires VMware Tools running in the guest OS. An existing file at guest_path is replaced. Use vm_guest_download for the reverse direction, and vm_guest_provision instead when uploads and commands belong to one ordered provisioning run.
Without confirm=True this only previews: blast_radius names the VM (name, instance UUID), the guest account, the local path and size, the guest path, VMware Tools status and any blockers; nothing is transferred. Show it to the user. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet. Refused outright: VM not powered on, VMware Tools not running, a local file that is missing or unreadable, an unreadable identity or status, or a name that matches more than one VM.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| vm_name | Yes | Target VM name. | |
| password | No | Guest OS password. | |
| username | Yes | Guest OS account to run as. Required — there is no default, so a call can never act as root without choosing root. | |
| guest_path | Yes | Destination path inside the guest. | |
| local_path | Yes | Local file path to upload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, and the description adds substantial behavioral context: the confirmation gate behavior, that an existing file at guest_path is replaced, that without confirm=True nothing is transferred, and the exact refusal conditions. It also explicitly warns the agent not to set confirm=True on its own because the user hasn't seen the preview yet. This is rich, honest behavioral disclosure beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized: it front-loads the core action, then covers the confirmation gate, prerequisites, and refusal conditions. Every sentence earns its place. It is longer than the typical description, but the complexity of the confirmation-gate behavior justifies the length. Slightly verbose in the refusal list, but still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, confirmation-gated write tool with no output schema, the description covers everything an agent needs: what it does, prerequisites, what the preview returns, when to confirm, when to refuse, and how it differs from siblings. The refusal conditions are enumerated. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters. The description adds meaning beyond the schema by explaining the confirm parameter's role in the preview/apply flow, and by clarifying the username requirement ('there is no default, so a call can never act as root without choosing root'). It doesn't add per-parameter detail for every field, but the schema already covers those, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource: 'Upload a file from local machine to a VM via VMware Tools.' It also distinguishes itself from siblings by naming vm_guest_download (reverse direction) and vm_guest_provision (ordered provisioning runs). This is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool vs alternatives: use vm_guest_download for the reverse direction, and vm_guest_provision when uploads and commands belong to one ordered provisioning run. It also gives clear prerequisites (VMware Tools running, VM powered on) and refusal conditions. This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_investigation_bundleARead-onlyIdempotent
[READ] "What is happening around this VM?" — one correlated drill-down.
Correlates everything around one VM so you don't stitch it together yourself: the VM's state, its host, cluster context, backing datastores, snapshots, triggered alarms, live performance, and a merged event timeline across VM, host, cluster and datastores (newest first). Batched, cheap even on large fleets. Delegates to the vmware-monitor library (read-only). Explain the result in operational language; do not dump it raw.
Use this AFTER cluster_health_summary points at a problem VM. Point-in-time.
| Name | Required | Description | Default |
|---|---|---|---|
| hours | No | Event-timeline look-back window in hours (default 24). | |
| target | No | Optional vCenter/ESXi target name from config (default if omitted). | |
| vm_name | Yes | Exact VM name. Unknown names return a teaching error (list VMs first). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description reinforces this by stating 'Delegates to the vmware-monitor library (read-only)' and adds performance context with 'Batched, cheap even on large fleets.' It also instructs the agent to 'Explain the result in operational language; do not dump it raw,' which is a behavioral guideline beyond annotations. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded: a one-line purpose, a bullet-like list of contents, then batching/read-only info, and finally usage guidance. Every sentence contributes value without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex correlation tool, the description covers the included data types, performance characteristics, read-only nature, and output handling (explain operationally). It also gives usage context and a point-in-time note. With no output schema, the description adequately conveys what to expect and how to present results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions for all three parameters (vm_name, hours, target). The tool description does not add parameter-specific details, but it indirectly informs usage via the 'AFTER cluster_health_summary' guidance. Since the schema already documents parameters fully, a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear one-liner: 'What is happening around this VM?' and then enumerates the exact data correlated: VM state, host, cluster context, datastores, snapshots, alarms, performance, and an event timeline. This precise resource and scope distinguish it from sibling bundles for hosts, datastores, and cluster_health_summary, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to use: 'Use this AFTER cluster_health_summary points at a problem VM.' This names the trigger and the preceding tool, and 'Point-in-time' clarifies it is a snapshot. While it doesn't list when not to use, the condition is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vmk_pingARead-onlyIdempotent
[READ] DF-bit-capable ping sourced from a host vmk - MTU path validation.
Runs esxcli network diag ping on the ESXi host through the vSphere API
(no host SSH). df=True sets Don't-Fragment so an oversized packet FAILS
instead of fragmenting - that failure is the diagnostic:
df=True size=1572 proves a >=1600 MTU path (overlay/TEP floor)
df=True size=8972 proves full jumbo (9000 minus 28 bytes overhead) A too-big result reports 'Message too long' in the fault field rather than erroring - read success + fault together.
Returns: Dict with request, success, and summary (transmitted/received/loss/ rtt) or fault (the esxcli failure text). Errors return "error" + hint.
| Name | Required | Description | Default |
|---|---|---|---|
| df | No | True sets the Don't-Fragment bit (the MTU probe mode). | |
| size | No | ICMP payload bytes (default 56). Path proves size+28 MTU. | |
| count | No | Packets to send (default 3, max 60). | |
| target | No | vCenter target name from config.yaml; omit to use the default target. | |
| dest_ip | Yes | IPv4 address to ping. | |
| netstack | No | Optional netstack instance (e.g. "vxlan" for real TEP vmks). | |
| host_name | Yes | ESXi host to source the ping from. | |
| source_vmk | Yes | VMkernel device to source from (e.g. "vmk2"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and non-destructive behavior, so the bar is lower, but the description adds substantial behavioral detail: DF behavior, the 'Message too long' fault semantics, the need to read success and fault together, and the exact return shape. It even explains error handling ('error' + hint). This goes well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficiently organized: a one-line purpose header, a mechanism note, bulleted diagnostic semantics, and a compact return summary. Every sentence adds operational value, with no filler or tautology. The format is front-loaded with the tool's identity and safety posture.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 8 parameters and no output schema, the description covers return structure, fault behavior, error handling, and the diagnostic meaning of key parameter combinations. The schema handles the remaining parameter definitions. An agent has enough context to invoke the tool correctly and interpret results without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description adds meaningful interpretive value beyond the schema. It explains that size proves size+28 MTU, that df=True makes oversized packets fail instead of fragmenting, and gives concrete examples for source_vmk and netstack. The description does not redundantly restate every schema property, but it enriches the most diagnostic-critical ones.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: a DF-bit-capable ping sourced from a host vmk for MTU path validation. It clearly identifies the underlying mechanism (esxcli network diag ping through the vSphere API) and distinguishes itself from sibling tools like list_host_vmks. The READ prefix and the diagnostic framing make the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear use context: MTU path validation with specific probe sizes (1572 for >=1600 MTU, 8972 for jumbo) and explains the failure-mode interpretation. It explicitly contrasts with no-host-SSH execution, which helps an agent choose this over an SSH-based approach. It does not name an alternative ping tool, but none exists among siblings, so the contextual guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_list_plansARead-onlyIdempotent
[READ] List all pending/failed plans.
Use this first to find a plan_id for vm_apply_plan or vm_rollback_plan. Returns the list envelope: 'items' holds plan summaries (plan_id, created_at, status, steps count, VMs affected), and 'returned'/'total'/ 'truncated' state whether the listing is complete. Every plan file is read, so truncated is always false. Listing never deletes: stale plans (>24h) are swept by vm_create_plan, not by this tool.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool readOnly, idempotent, and non-destructive, and the description reinforces these with concrete behavioral guarantees: 'Listing never deletes' and 'Every plan file is read, so truncated is always false.' It also describes the return envelope structure, which is valuable since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then adds usage guidance and the return envelope in a compact, well-ordered manner. Every sentence contributes meaningful information without repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description fully covers invocation intent, return format, listing-completeness semantics, and lifecycle ownership of stale plans. An agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the input schema is already fully self-explanatory, so the baseline of 4 applies. The description wisely focuses on output and behavior instead of inventing parameter details that don't exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with '[READ] List all pending/failed plans', naming a specific verb and resource with clear scope. It also distinguishes itself from related plan tools by mentioning apply/rollback and by explicitly stating that stale-plan sweeping belongs to vm_create_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs the agent to use this tool first to obtain a plan_id for vm_apply_plan or vm_rollback_plan. It also clarifies what this tool is not for by stating that stale plans are swept by vm_create_plan, not by this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_list_snapshotsARead-onlyIdempotent
[READ] List the full snapshot tree of a VM, including nested child snapshots.
Read-only, no side effects. Call this before vm_revert_snapshot, vm_delete_snapshot, or deploy_linked_clone to get exact snapshot names. 'items' is empty when the VM has no snapshots.
Returns: The list envelope. 'items' is one dict per snapshot: name, description, created, state (power state at snapshot time), level (0 = root). The whole tree is walked, so 'total' is the real count and 'truncated' is always false.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config.yaml; omit to use the default target. | |
| vm_name | Yes | Exact VM name as shown in vCenter inventory. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint false. The description adds useful behavioral details: the whole tree is walked, 'total' is the real count, 'truncated' is always false, and 'items' is empty when no snapshots exist. This goes beyond the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the operation and key usage, then gives only necessary behavioral and response details. Every sentence adds value and the returns section is clearly separated. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully explains return values: one dict per snapshot, fields included, nesting behavior, and the empty case. It also explains when to use the tool and what it guarantees about pagination. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents vm_name and target. The description does not add parameter-level details beyond what the schema provides, but no compensation is needed at full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List the full snapshot tree of a VM, including nested child snapshots.' This clearly distinguishes it from snapshot creation, deletion, and revert tools by emphasizing the read-only listing of the entire tree.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call this before vm_revert_snapshot, vm_delete_snapshot, or deploy_linked_clone to get exact snapshot names.' This gives direct, actionable guidance on when this tool is the right choice compared to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_list_ttlARead-onlyIdempotent
[READ] List all VMs with TTLs registered, including expiry time and status.
Use this first to find the exact vm_name for vm_cancel_ttl. Returns the list envelope: 'items' holds TTL entries with remaining_minutes and expired flag, and 'returned'/'total'/'truncated' state whether the listing is complete. The whole TTL store is read, so truncated is always false.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, and the description adds meaningful behavior beyond that: it reads the entire TTL store, describes the list envelope fields, and explicitly explains that truncated is always false. This gives the agent a precise model of what the call will return and how the data behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose first, then the key usage hint, then return format. Every sentence adds useful information and there is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only list tool, this is complete. It names the output envelope, the important fields (remaining_minutes, expired flag, returned/total/truncated), and the invariant that truncated is always false. The lack of an output schema is compensated by the description's return-value detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema already documents this fully, so the description does not need to add parameter guidance. The description instead intelligently emphasizes output semantics, which is the relevant semantic burden for a no-argument list tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'List all VMs with TTLs registered, including expiry time and status.' It clearly distinguishes itself from related TTL manipulation tools by stating it is the read-oriented listing operation and 'Use this first to find the exact vm_name for vm_cancel_ttl.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit use case: use this tool first to discover the exact vm_name needed by vm_cancel_ttl. It provides clear context for when to invoke it, though it does not explicitly state when not to use it or name broader alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_migrateA
[WRITE] Migrate (vMotion) a VM to another host, optionally with storage vMotion.
Without confirm=True this only previews: it returns blast_radius (VM and instance UUID, source host and datastores, target host with its connection and maintenance state, target datastore, and live vMotion vs cold migration) and moves nothing. Show it to the user and get their decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Refused: a target host that is not found, not connected, in maintenance mode or outside a cluster; a target datastore that is not found; a target host that does not mount the VM's datastores when to_datastore is omitted (vCenter rejects cross-host vMotion without shared storage — pass to_datastore); and anything above that cannot be read. The VM's current host with no to_datastore returns action "noop". Run cluster_info first for host names.
Returns: Dict with action (preview, noop, migrated), blast_radius, and result.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| to_host | Yes | Target ESXi host name. | |
| vm_name | Yes | VM to migrate. | |
| to_datastore | No | Target datastore (required for cross-storage hosts). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral detail beyond the annotations: without confirm it changes nothing, confirm=True applies the migration, and the agent must not set confirm on its own before the user has seen the preview. It also discloses refusal conditions and the noop case, which is exactly the kind of context an agent needs for safe invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action, then organized into preview behavior, refusal conditions, and return value. It is somewhat long, but each section provides actionable detail. The refusal list is dense but not redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the return dictionary keys and describing the contents of blast_radius. It also covers edge cases like noop, cross-host shared storage requirements, and preconditions such as running cluster_info. This is sufficient for an agent to invoke the tool safely and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All five parameters already have schema descriptions, so the baseline is 3. The description adds meaningful semantics beyond the schema by tying confirm to the preview/apply lifecycle and explaining that to_datastore is required when vMotion would cross hosts without shared storage. This helps the agent choose parameter values correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Migrate (vMotion) a VM to another host, optionally with storage vMotion.' This distinguishes it from sibling tools like vm_clone or vm_reconfigure. It also clarifies the preview/noop/migrated action states, so the tool's core behavior is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear operational context: preview with confirm=False, apply with confirm=True, and the instruction to show the preview to the user before deciding. It also says to run cluster_info first for host names. It does not explicitly compare against alternatives such as vm_clone, but the migration-specific language makes the intended use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_power_offADestructive
[WRITE] Power off a VM — graceful guest shutdown by default, hard power-off with force=True.
Without confirm=True this only previews: it returns blast_radius (VM name and instance UUID, host, power state, VMware Tools status, and whether this is a guest shutdown or a hard power-off) and changes nothing. Show it to the user and get their decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Graceful mode calls VMware Tools guest shutdown and waits up to 120s; if it does not finish, action is "still_running". Refused: a graceful shutdown when Tools is not running or the VM is suspended (preview force=True instead), and a VM whose identity or power state cannot be read. An already-off VM returns action "noop". Use vm_power_on to start a VM; vm_delete requires it off first.
Returns: Dict with action (preview, noop, powered_off, still_running), blast_radius, and the executor's message under result.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | False (default) = graceful guest shutdown via VMware Tools; True = immediate hard power-off (risks guest filesystem damage). | |
| target | No | vCenter/ESXi target from config.yaml; omit for the default target. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| vm_name | Yes | Exact VM name as shown in vCenter inventory (case-sensitive). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate destructiveHint=true, but the description adds substantial behavioral context: preview mode changes nothing, graceful shutdown waits up to 120s and may return 'still_running', an already-off VM returns 'noop', and certain states are refused. This goes well beyond the annotation fields and tells the agent exactly what the tool will and will not do.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place: mode explanation, confirmation contract, refusal conditions, sibling routing, and return values. It is front-loaded with the core action and uses short paragraphs to keep the information scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description enumerates possible action values (preview, noop, powered_off, still_running) and describes blast_radius contents. It covers safety, sequencing, failure modes, and related tools, leaving no critical gap for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the interaction between force, confirm, and execution behavior, and by clarifying that force=True should be previewed when graceful shutdown is refused. This is more than the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Power off a VM', then immediately distinguishes graceful vs hard power-off. It also names sibling tools (vm_power_on, vm_delete) to clarify boundaries, so an agent can tell it apart from related operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use and when-not-to-use guidance: preview with confirm=False, do not set confirm=True on the user's earlier request, use vm_power_on to start a VM, and vm_delete requires the VM off first. It also lists refusal cases, such as graceful shutdown when VMware Tools is not running or the VM is suspended.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_power_onAIdempotent
[WRITE] Power on a virtual machine.
Returns a status string; an already-on VM is a no-op. Reverse with vm_power_off. Call this first when a VM is off: guest tools such as vm_guest_exec only work once VMware Tools has finished booting.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. Uses default if omitted. | |
| vm_name | Yes | Exact name of the virtual machine. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses that an already-on VM is a no-op, that it returns a status string, and that its effects are relevant to guest tool readiness. This adds meaningful behavioral context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences carry high-signal information: the operation, the no-op behavior, and the practical usage context. There is no filler or repetition, and the key verb-action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only two parameters, full schema coverage, and strong annotation hints, this description fully covers what an agent needs: what it does, when to call it, how it behaves on already-on VMs, and its relationship to other VM operations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, documenting vm_name as the exact name and target as an optional config target with default behavior. The description adds no new parameter-specific details, so the schema-based baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: "Power on a virtual machine." It also clearly distinguishes this from its reverse operation by naming vm_power_off, so an agent can identify the tool's role among many VM-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: "Call this first when a VM is off." It also references the relevant alternative (vm_power_off) and provides a concrete usage context by noting that guest tools depend on VMware Tools finishing boot.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_reconfigureA
[WRITE] Change a VM's vCPU count and/or memory.
Pass only the fields you want to change; omitted fields are left untouched. Hot-add of CPU/memory requires it to be enabled on the VM and a running guest; otherwise power the VM off first (vm_power_off).
Returns: Status string describing the applied change, or a VM-not-found error.
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | New vCPU count; omit to leave unchanged. | |
| target | No | vCenter/ESXi target name from config.yaml; omit to use the default target. | |
| vm_name | Yes | Exact name of the VM to reconfigure. | |
| memory_mb | No | New memory in MB; omit to leave unchanged. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses important behavior: partial-update semantics (omitted fields left untouched), hot-add prerequisites, and the need to power off first if hot-add is unavailable. It also states the return value and error case. These add real context beyond readOnlyHint=false and destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action, followed by usage caveats and return behavior. The blank lines and '[WRITE]' tag add minor noise, but every substantive sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter reconfiguration tool, the description covers the core operation, partial-update semantics, hot-add prerequisites, power-off fallback, and return/error behavior. Given the presence of an output schema and straightforward parameters, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter already documents whether it may be omitted. The description reinforces the partial-update behavior but does not add new semantic detail beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation and resource: 'Change a VM's vCPU count and/or memory.' This clearly distinguishes the tool from the many sibling VM operations such as power control, snapshots, cloning, and deletion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: pass only fields to change, omitted fields are untouched, and hot-add requires both a VM-level enablement and a running guest, otherwise power off first with vm_power_off. It does not explicitly name alternatives or when-not-to-use conditions, but the prerequisite guidance is strong and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_revert_snapshotADestructive
[WRITE] Revert a VM to a named snapshot (loses changes since snapshot).
Without confirm=True this only previews: it returns blast_radius (VM and instance UUID, the snapshot's name, id and creation time, total snapshot count, current power state and the power state after the revert) and changes nothing. Show it to the user and get their decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Irreversible — everything written since the snapshot is lost. Refused: a snapshot name that is not found, a name that matches more than one snapshot on the VM (vSphere allows duplicates; rename one first), and a VM whose snapshot tree, identity or power state cannot be read. Run vm_list_snapshots first for exact names. To reclaim space without changing state use vm_delete_snapshot.
Returns: Dict with action (preview, reverted), blast_radius, and result.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| vm_name | Yes | VM to revert. | |
| snapshot_name | Yes | Snapshot to revert to. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive and not read-only, and the description complements this by explaining the preview workflow, irreversible data loss, duplicate-snapshot handling, and the exact refusal cases. It also discloses what the return dictionary contains, which is especially valuable given there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries operational weight: purpose, preview behavior, confirmation safeguard, irreversibility, refusal conditions, prerequisite command, sibling alternative, and return format. It is front-loaded with the most important destructive warning and structured so the agent can quickly extract the key rules.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive VM operation with no output schema, the description is complete: it explains when to use it, what the preview returns, what the final return looks like, what inputs are required, what can go wrong, and what to do first. Nothing essential for correct invocation is left to guesswork.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: confirm=false previews versus true applies, snapshot_name must be exact and unique, and snapshot_name is validated against list_snapshots first. This extra context makes the parameter semantics genuinely more actionable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Revert a VM to a named snapshot (loses changes since snapshot).' It clearly distinguishes itself from sibling tools by noting that vm_delete_snapshot is for reclaiming space without changing state, so an agent can tell them apart immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: run vm_list_snapshots first for exact names, do not set confirm=True before the user has seen the preview, and use vm_delete_snapshot instead for space reclamation without state change. It also lists refusal conditions, leaving no ambiguity about how to safely invoke the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_rollback_planADestructive
[WRITE] Rollback executed steps of a failed plan in reverse order.
Without confirm=True this only previews: it returns blast_radius listing the rollback steps that would run (in order, with the VMs a rollback deletes) and those skipped as irreversible — and runs nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet.
Only call this after vm_apply_plan returns status='failed'; check vm_list_plans first for the plan_id. Irreversible steps (delete_vm, revert_snapshot, etc.) are skipped with a warning. Each destructive rollback step (power off, delete snapshot, cluster delete, host removal) is measured as its own tool measures it, in the preview and again just before it runs; a refused check stops the rollback there and the plan stays 'failed'. Refused on a target other than the one the plan was created against.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | The vCenter/ESXi target the plan was created against. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| plan_id | Yes | The plan ID of the failed plan. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint=true, but the description adds preview-only mode without confirm, irreversible steps skipped with warning, per-step destructive measurement with refusal stopping the rollback, and the plan staying 'failed' on refusal. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded, but the final sentence is a fragment and the safety warnings make it longer than strictly minimal. Still, each sentence earns its place for a destructive rollback tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description explains what the preview returns (blast_radius with ordered steps and skipped irreversible ones), the failure/refusal behavior, and the preconditions. An agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all three params, and the description adds critical semantics: confirm=false previews vs true applies, target must be the plan's original target, and plan_id identifies the failed plan. The warning about not setting confirm on one's own initiative is extra value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Rollback executed steps of a failed plan in reverse order.' This clearly distinguishes it from vm_apply_plan and vm_list_plans and scopes its role in the plan lifecycle.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly conditions use on vm_apply_plan returning status='failed', directs the agent to check vm_list_plans for plan_id, and instructs not to auto-set confirm=True before the user sees the preview. This is actionable when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_set_ttlADestructive
[WRITE] Set a Time-To-Live (TTL) for a VM. The daemon auto-deletes it when expired.
Without confirm=True this only previews: it returns blast_radius — the VM that will be deleted (identity, host, disks, total size, snapshot count), when (expires_at), and any TTL it replaces — and schedules nothing. Show that to the user and get their explicit decision. Do not set confirm=True on your own because the user asked earlier: they have not seen the preview yet. Refused when the VM's identity or disks cannot be read.
Use this for short-lived lab VMs so cleanup is not forgotten; cancel with
vm_cancel_ttl and review pending expiries with vm_list_ttl. The scheduler
daemon must be running (vmware-knight daemon start) or nothing is ever
deleted; it powers a running VM off first. TTLs persist in
~/.vmware-knight/ttl.json. Returns a dict (action, blast_radius).
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Optional vCenter/ESXi target name from config. | |
| confirm | No | False (default) returns the blast radius and changes nothing. True applies it. | |
| minutes | Yes | Minutes until deletion (minimum 1). | |
| vm_name | Yes | Name of the VM to auto-delete. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as destructive and non-read-only, but the description adds important non-obvious behaviors: confirm=False only previews and schedules nothing, confirm=True applies the TTL, the daemon must be running or nothing is deleted, it powers a running VM off first, and it refuses when VM identity or disks cannot be read. It also discloses the return dict and the persistence path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, preview behavior, safety instruction, intended use case, sibling routing, daemon prerequisite, side effects, persistence, and return type. The critical safety caveats are front-loaded, and the write nature is marked with [WRITE].
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly explains the return value as a dict with action and blast_radius, and enumerates the blast_radius contents. It covers confirmation semantics, failure modes, prerequisites, persistence, and side effects, leaving no major information gap for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds substantial meaning beyond the schema for the confirm parameter, explaining exactly what false versus true does and what the returned blast_radius contains. It also clarifies that only confirm=True has destructive side effects. Other parameters like target and minutes are not elaborated, but the schema already documents them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Set a Time-To-Live (TTL) for a VM. The daemon auto-deletes it when expired.' It clearly distinguishes itself from the sibling tools vm_cancel_ttl and vm_list_ttl by describing the scheduling behavior and naming the cancellation/list alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool ('Use this for short-lived lab VMs so cleanup is not forgotten'), names the alternatives for cancellation and review, and provides a firm usage rule: do not set confirm=True on your own because the user has not seen the preview. It also warns about the required daemon, making the conditions of safe use clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vm_task_statusARead-onlyIdempotent
[READ] Poll a long-running vSphere task by its id (from an async vm_delete_snapshot).
Use after vm_delete_snapshot returns a task id, instead of re-running the delete. Returns state (queued/running/success/error/gone), progress percent, and the entity name. 'gone' means vCenter already garbage-collected a completed task — re-list the resource to confirm the final state. A failed task reports its fault under 'task_error'; a top-level 'error' key would mean this poll itself failed.
Returns: Dict with task_id, state, progress_pct, operation, entity, and task_error/note when relevant.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | vCenter/ESXi target name from config.yaml; omit to use the default target. | |
| task_id | Yes | The task id string returned by an async write operation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, idempotentHint=true, and destructiveHint=false, which the description does not contradict. The description adds value by explaining the 'gone' state semantics (vCenter garbage collection) and how to resolve it, which is beyond the raw annotations. It also clarifies the distinction between a task failure ('task_error') and a poll failure ('error'), adding useful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear intro sentence, usage notes, and a return format list. It is moderately concise, but the multiple sentences on 'gone' and the explicit return list add some length. Still, every sentence serves a purpose, so it is strong, though slightly verbose for a simple poll.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simple nature (2 parameters, no schema), the description is complete: it explains when to use, what states may be returned, the meaning of 'gone', and the error semantics. There is no missing information an agent would need to call it correctly, and the output schema is absent but the return fields are listed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (target and task_id) are fully documented in the schema. The description adds minimal extra meaning beyond re-emphasizing the source of task_id (from an async write) and the meaning of 'gone', but it does not provide new syntax or format details for either parameter. Thus, a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with '[READ] Poll a long-running vSphere task by its id', which clearly states the verb (poll), the resource (task), and the method (by id). It further distinguishes itself by specifying it is specifically for tasks from 'vm_delete_snapshot', and the next paragraph clearly marks it as an alternative to re-running the delete. This clearly separates it from siblings that perform the operation itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs when to use: 'Use after vm_delete_snapshot returns a task id, instead of re-running the delete.' It also provides guidance on how to interpret the 'gone' state, advising to 're-list the resource to confirm the final state.' This is a clear definition of its role relative to the delete operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
60 tool updates
v1.12.10- First observed
acknowledge_vcenter_alarm - First observed
add_host_vmk - First observed
attach_iso_to_vm - First observed
batch_clone_vms - First observed
batch_deploy_from_spec - First observed
batch_linked_clone_vms - First observed
browse_datastore - First observed
cluster_add_host - First observed
cluster_configure - First observed
cluster_create - First observed
cluster_delete - First observed
cluster_health_summary - First observed
cluster_info - First observed
cluster_remove_host - First observed
convert_vm_to_template - First observed
create_drs_rule - First observed
create_dvs_portgroup - First observed
cross_vcenter_attention - First observed
datastore_investigation_bundle - First observed
delete_drs_rule - First observed
deploy_linked_clone - First observed
deploy_vm_from_ova - First observed
deploy_vm_from_template - First observed
host_investigation_bundle - First observed
list_drs_rules - First observed
list_dvs_portgroups - First observed
list_host_vmks - First observed
list_vcenter_alarms - First observed
remove_host_vmk - First observed
reset_vcenter_alarm - First observed
scan_datastore_images - First observed
set_drs_rule_enabled - First observed
set_vmk_service - First observed
vm_apply_plan - First observed
vm_cancel_ttl - First observed
vm_clean_slate - First observed
vm_clone - First observed
vm_create - First observed
vm_create_plan - First observed
vm_create_snapshot - First observed
vm_delete - First observed
vm_delete_snapshot - First observed
vm_guest_download - First observed
vm_guest_exec - First observed
vm_guest_exec_output - First observed
vm_guest_provision - First observed
vm_guest_upload - First observed
vm_investigation_bundle - First observed
vm_list_plans - First observed
vm_list_snapshots - First observed
vm_list_ttl - First observed
vm_migrate - First observed
vm_power_off - First observed
vm_power_on - First observed
vm_reconfigure - First observed
vm_revert_snapshot - First observed
vm_rollback_plan - First observed
vm_set_ttl - First observed
vm_task_status - First observed
vmk_ping
TDQS
Scored across 60 tools
Each tool has a clearly distinct purpose, with descriptions that explicitly differentiate similar operations (e.g., deploy_vm_from_ova vs. deploy_vm_from_template vs. vm_clone). Even closely related tools like vm_guest_exec and vm_guest_exec_output are clearly separated by their output-capturing behavior.
All tool names follow a consistent verb_noun snake_case pattern (e.g., list_vcenter_alarms, create_drs_rule, vm_power_on, vm_guest_upload). The naming is uniform and predictable across all 60 tools.
60 tools is high, but the server covers an extensive domain: VM lifecycle, snapshots, cloning, guest operations, cluster management, DRS rules, networking, datastore browsing, monitoring, TTL, and a plan system. Each tool has a specific role, so the count is defensible, though it may feel heavy for simple use cases.
The tool set provides comprehensive coverage of vSphere operations, including CRUD for VMs, snapshots, cloning, guest execution, cluster and DRS management, network portgroups and VMKs, datastore browsing, and monitoring bundles. Minor gaps exist (e.g., no explicit VM rename or vCenter host addition), but they do not block core workflows.
Maintenance
Related MCP Connectors
Provides capabilities that let LLM agents perform a range of infrastructure management tasks.
Deploy, monitor, and manage your OpenClaw AI assistants via natural language.
An AI concierge that turns static forms into adaptive AI conversations. From any MCP client.
Conversational AI Coaching from calls; permissioned Team Dynamics reports in a limited U.S. pilot.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables AI assistants to manage and monitor VergeOS virtualization platforms through natural language, including VM operations, network management, tenant administration, and cluster monitoring.21MIT
- AlicenseAqualityAmaintenanceAI-powered VMware vCenter/ESXi monitoring and operations. 20 MCP tools for inventory queries, health monitoring, VM lifecycle management, fast provisioning (Linked Clone, OVA, template deploy), snapshot management, and datastore browsing. Supports vSphere 6.5–8.0. Works with local models via Ollama/LM Studio.44623 PyPI74MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to interact with Veeam Backup & Replication infrastructure through natural language, allowing monitoring, management, and troubleshooting of backup jobs, sessions, and restore points.MIT
- AlicenseNot gradedqualityCmaintenanceEnables natural language interaction with VMware SDDC Manager and vCenter APIs through MCP tools, allowing users to query workload domains, VMs, clusters, and more.MIT