Skip to main content
Glama

esxi_execute_plan

Destructive

Execute a sequence of ESXi administrative steps with per-step preview, handling asynchronous tasks and optional reverse-order rollback on failure.

Instructions

后台顺序执行明确步骤;每步先预览,异步任务等到成功再下一步。可指定compensate,失败按逆序执行;不承诺原子回滚、不自动重试。

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
stepsYes
dry_runNo
request_idNo
timeout_per_stepNo
rollback_on_failureNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.3.0

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite annotations already marking destructiveHint=true and idempotentHint=false, the description adds substantial behavioral context: background sequential execution, per-step preview, async task waiting, optional compensate in reverse order, and explicit disclaimers of atomic rollback and automatic retry. These details materially inform how the tool behaves and what guarantees it does not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence with semicolons, front-loading the primary action and then layering constraints. Every clause adds useful information (preview, async waiting, compensate, rollback limits, no retry) without redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex (multi-step plan execution, destructive, no atomic rollback) and has an output schema, so return values need not be explained. The description covers behavioral guarantees and limitations well, but given 0% schema coverage for five parameters, it should explain the steps format and key controls like dry_run or rollback_on_failure. That gap makes it only partially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there are 5 parameters. The description mentions 'preview' (possibly related to dry_run) and 'compensate' (possibly related to rollback_on_failure) but does not name or explain any parameter. It leaves steps structure, request_id, timeout_per_step, and rollback_on_failure entirely undocumented, which is inadequate for a tool with zero schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the core action: executing explicit steps sequentially in the background. It specifies key mechanics (preview first, wait for async success before next step, optional compensate in reverse order), which distinguishes it from generic executors. However, it does not explicitly name or contrast with sibling tools like esxi_stage_finalize or esxi_api_invoke, so it misses full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for multi-step plans with preview and compensation, but never states when to choose this tool over alternatives such as esxi_api_invoke or esxi_stage_finalize. There are no explicit when-to-use or when-not-to-use conditions, leaving the agent to infer suitability from behavioral details.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.