ESXi MCP
Provides tools for managing a standalone VMware ESXi host, including VM operations, host resource controllers, guest operations, OVF import/export workflows, background file transfers, durable operation tracking, and optional ESXCLI/SSH administration.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ESXi MCPlist all VMs on the host and show their power state"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ESXi MCP
Version 0.3.0 exposes 63 stdio tools for a single standalone ESXi host using the official MCP Python SDK and pyVmomi. It covers VM operations, 14 host-resource controllers, guest operations, complete OVF workflows, background file transfers, durable operation tracking and an optional ESXCLI/shell backend.
中文 · MIT license · Capabilities · Execution guide · Test evidence
Install and configure
Python 3.10+ is required. This project is not published to PyPI:
python -m venv .venv
# Linux/macOS
.venv/bin/python -m pip install -e '.[dev]'
# Windows PowerShell
.venv/Scripts/python.exe -m pip install -e '.[dev]'Install .[admin] or .[dev,admin] to use the optional ESXi SSH backend. requirements-lock.txt records the earlier Python 3.12 Linux core acceptance environment, excluding the optional SSH dependency; use pyproject resolution for other platforms and extras.
Copy credentials.example.json and config.example.json to a private directory outside the repository. Defaults are ~/.config/esxi-mcp/credentials.json and a sibling config.json. Restrict directory/file permissions to 0700/0600 on Unix and use Windows ACLs. TLS verification is enabled and writes are disabled by default. Set ESXI_CA_FILE for a private CA.
Setting | Meaning |
| Private credentials and deployment settings |
| Credential overrides; avoid publishing them |
| Explicit ordinary write override; default false |
| Additional API/upload/OVF/SSH write opt-in; default false |
|
|
| Exact comma-separated management VM IDs |
| Deployment opt-in; a call must also explicitly allow protected impact |
| Deployment opt-in; a call must also acknowledge disruption |
| Optional ESXi SSH backend and private credentials |
| Private SQLite journal and staged transfer files; default beside config |
| Process-local task/lease retention upper bound, default 86400 seconds |
| TLS validation override or private CA path |
| API HTTP timeout (30s default), private audit location |
Deployment JSON also supports allowed_operations, denied_operations, protected_datastores and protected_networks. Discover actual management VM IDs first; re-check them after re-registration.
Related MCP server: VMWare MCP
Connect an MCP client
Use the installed virtual environment's Python executable:
{
"mcpServers": {
"esxi": {
"command": "/absolute/path/esxi-mcp/.venv/bin/python",
"args": ["-m", "esxi_mcp.server"],
"env": {
"ESXI_CONFIG_FILE": "/private/path/credentials.json",
"ESXI_DEPLOY_CONFIG_FILE": "/private/path/config.json"
}
}
}
}On Windows use the .venv/Scripts/python.exe path. Clients may use their own configuration format. Reload the client after changing configuration. esxi-mcp is an alternative console entrypoint. run-stdio.sh is a portable Unix launcher; remote stdio may use a dedicated SSH key restricted to that launcher. The transport key and optional ESXi administration key are separate credentials.
Control surfaces
Tools | Coverage |
17 common VM/host tools | Inventory, health, power, CPU/memory/allocations, creation, rename, snapshots, disk expansion, deletion, tasks |
inventory / managers / api_schema / api_get / api_invoke | Typed access to declared, server-supported vSphere methods and properties |
14 | network, datastore, storage, service, firewall, account, permissions, license, time, pci, certificate, patch, host, options |
resource_schema / resource_get / dependencies / capabilities | Named action discovery, exact parameters, current state, conservative impact analysis |
operation_status / operation_history | Durable results and reconnectable task tracking |
vm_hardware / vm_register | Explicit VirtualDeviceSpec changes and existing VMX registration |
guest_process / guest_process_status / guest_files | Authenticated Tools-backed guest process/file workflows |
datastore_transfer / api_transfer | Legacy bounded inline datastore/guest/NFC transfer |
transfer_start / status / read_chunk / cancel / resume | Bounded-memory background transfer with progress, SHA256 and explicit resume |
stage_upload / write_chunk / finalize | Private 1MiB chunk staging, identical-chunk replay and final checksum |
ovf_export / ovf_import / execute_plan | Descriptor/disks/manifest/lease workflows and explicit compensated plans |
esxcli_query / esxcli / shell | Optional fixed query catalog and explicit administrator commands |
Execute
Preview by default; use exact VM ID/name, host address or SDK type/ID as the tool requires. Ordinary mutations require writes; advanced operations additionally require advanced writes. An administrator profile enables host-level administrative operations. Shared changes report dependencies: acknowledge every affected VM. Unknown impact requires explicit administrator acknowledgement. Management-disrupting or protected-impact operations require both deployment and call opt-ins. Default protection remains enforced.
Use esxi_get_task / esxi_operation_status to verify completion. Supported request_id parameters prevent duplicate submissions and return existing results/handles; uncertain requests require inspection. The server never implicitly powers off workloads or deletes snapshots. Guest shutdown does not fall back to power-off. Disk growth does not resize guest filesystems. Plans compensate only explicitly specified completed steps, in reverse order; compensation is not an atomic transaction and is skipped when the failing step has an uncertain result. See resources.
For optional SSH administration, install the extra, configure private ssh.json using ssh.example.json, independently verify a host-key fingerprint (or configure a trusted known_hosts file), and enable the SSH backend. SSH writes are treated as affecting the whole host. Shell payloads and output are omitted from durable logs, retaining argument digests.
Transfer and OVF client
scripts/file_client.py accepts a private stdio transport JSON containing command, args and optional env. Upload/import/export commands preview until --execute is supplied. It streams 1MiB chunks and checks SHA256:
python scripts/file_client.py --transport transport.local.json upload --file installer.iso --datastore DATASTORE --path images/installer.iso --execute
python scripts/file_client.py --transport transport.local.json download --datastore DATASTORE --path images/installer.iso --output copied.iso
python scripts/file_client.py --transport transport.local.json export --vm-id VM_ID --expected-name VM_NAME --output-dir bundle --execute
python scripts/file_client.py --transport transport.local.json import --ovf bundle/vm.ovf --name NEW_VM --datastore DATASTORE --network-map network-map.json --expected-host esxi.example.com --executeOVF export requires an already powered-off VM. Import accepts an unpacked OVF bundle and explicit network/file mappings, not a raw OVA-as-VMDK upload. fetch --job-id HANDLE --output FILE can resume a local partial download of a completed handle. Remote download resume uses a matching ETag/Last-Modified and exact HTTP Range; uploads restart and require explicit overwrite when needed. Transfer bounds are explicit (up to 64TiB), with private disk space and a 64MiB reserve required; this maximum has not been tested at scale. Staged files remain private until an operator removes them; no automatic retention service is provided.
Test and publish
python -m pytest -q
python scripts/verify_vmomi_contract.py
python -m buildVersion 0.3.0 is deployed and passed 121 automatic tests on each of Windows and Linux/Python 3.12, nine offline SDK write checks and four ephemeral loopback SSH integration checks. On an ESXi reporting API 8.0.2.0, it passed 15 disposable-VM checks, eight full-workflow checks (including dedicated network writes, request replay, compensation, 12MiB file transfer and a complete OVF export/import roundtrip), and three real ESXCLI queries. Test resources were removed; the 33 existing VMs and original network configuration matched the run baseline. These checks are not exhaustive compatibility or destructive-operation coverage. See test evidence. GitHub Actions passed all four Ubuntu/Windows and Python 3.10/3.12 combinations (121 tests, SDK contracts, SSH loopback checks and package build in each). See the verified run.
Publish the clean source bundle only. Do not commit private configurations, state databases, staged files, signed URLs, logs or raw acceptance evidence. See SECURITY and CONTRIBUTING.
Coverage follows actual ESXi APIs/commands, licenses, privileges, hardware and state. A schema is not proof of server support. This server does not provide physical BMC/BIOS access, power-on after a physical shutdown or multi-host vCenter orchestration. Guest operations require an OS, running VMware Tools and guest credentials. Background OVF/lease work is process-local and cannot automatically survive an MCP process restart. Only stdio is exposed; use trusted clients. Sources: official MCP SDK v1, pyVmomi.
Available Tools
63 toolsesxi_account_manageBDestructive
专用 account 管理:create, update, remove, password_change。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the agent already knows this is a mutating, non-idempotent, destructive operation. The description adds some behavioral context: default preview via dry_run, confirmation via expected_host, write switches, dependency confirmation, and request_id for deduplication. It does not clarify what gets destroyed (e.g., accounts) or permission requirements, but it adds meaningful procedural detail beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the core purpose and actions. However, the second sentence is a dense string of instructions with limited punctuation, making it somewhat hard to parse. It avoids unnecessary filler but could be structured more clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, multi-action tool with 9 parameters (mostly undocumented), the description provides some critical procedural context (preview-first, expected_host, request_id). Still, it omits details about the 'arguments' object, the exact effects of the impact flags, and what each action requires. An output schema exists, so return values are covered, but the input side remains underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should explain parameters. It mentions action, dry_run, expected_host, write switches (likely allow_disruption/allow_unknown_impact/allow_protected_impact), dependency confirmation (acknowledged_vm_ids), and request_id, but it does so in a compressed, non-exhaustive way and does not define the 'arguments' object or the exact meaning of each switch. With 9 parameters and no schema help, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (ESXi account) and the four CRUD-style actions (create, update, remove, password_change), which is clear enough for an agent to know what the tool does. It does not explicitly name sibling tools or differentiate from other '*_manage' tools, but the resource is specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies a workflow: first use resource_schema/get and a default preview, then perform the actual operation with proper confirmations. However, it doesn't explicitly say when to use this tool versus other account management or permission tools, and there are no exclusions or clear alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_api_getCRead-onlyIdempotent
读取发现引用的已声明属性;可查询硬件、PCI、VM 设备、网络、存储、服务、账户信息。秘密字段脱敏。
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| target | Yes | ||
| max_items | No | ||
| properties | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds one useful behavioral trait: secret fields are redacted. It does not mention auth requirements, rate limits, or other operational constraints, so it is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short clauses, front-loading the core action and then the redaction behavior. Every part earns its place, with no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex tool with four parameters, a nested required target object, and an output schema. The description covers the general purpose and redaction, but omits any explanation of how to use target, properties, depth, or max_items, and offers no guidance against alternatives, leaving significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for four parameters, including a required nested target object and a required properties array. The description provides no meaning for target, properties, depth, or max_items, so it fails to compensate for the schema's lack of documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: reads declared properties referenced by discovery, and lists queryable domains (hardware, PCI, VM devices, network, storage, services, accounts). It is clear what the tool does, but it does not distinguish itself from siblings like esxi_api_invoke, esxi_api_schema, or esxi_resource_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use, when-not-to-use, or alternative guidance. It only states what can be queried, leaving the agent to infer when this generic getter is preferable to esxi_api_invoke or resource-specific tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_api_invokeADestructive
高级类型化调用:覆盖服务器支持的 vSphere API 管理方法(网络/存储/硬件/服务/账户/guest/主机)。
先 managers/inventory → schema/get → dry_run。真实调用需要写开关、 ESXI_ENABLE_ADVANCED_WRITES=true、expected_target='准确_type:准确_moId'。 非 VM 对象的调用需 allow_disruption=true,可能影响主机/共享资源/管理链路。 直接 VM 引用仍受 protected_vm_ids 限制;共享资源修改须由调用者审查其依赖。 不执行任意 Python 或 shell。异步返回 task_id 后继续查询状态。
| Name | Required | Description | Default |
|---|---|---|---|
| method | Yes | ||
| target | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_target | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and readOnlyHint=false, and the description adds substantial context beyond them: real writes require an env-var write switch and an exact type:moId expected_target, non-VM objects may impact host/shared resources/management link, protected_vm_ids still restricts direct VM refs, no arbitrary Python/shell is executed, and async calls return a task_id. This is a rich disclosure of side effects and safeguards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by workflow and then a bullet-style set of constraints and limits. Dense and information-bearing with little filler, though the run-on constraints could be more scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, 10-parameter, nested-object mutation tool with no schema descriptions, the description covers the critical safety gates and the invocation workflow. It omits meaning for several parameters, but the presence of an output schema means return values need not be explained, keeping it nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 10 parameters, so the description must carry the burden; it clarifies dry_run sequencing, expected_target's 'exact_type:exact_moId' format, and allow_disruption's role. It leaves method, target, arguments, request_id, acknowledged_vm_ids, allow_unknown_impact, and allow_protected_impact essentially unexplained, so compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource: advanced typed invocation of server-supported vSphere API management methods, and enumerates the domains covered (network/storage/hardware/service/account/guest/host). It clearly reads as the generic escape hatch versus the typed siblings (esxi_network_manage, esxi_storage_manage, etc.), though it never explicitly names an alternative to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit ordered workflow (managers/inventory -> schema/get -> dry_run) and the preconditions for a real call (write switch, ESXI_ENABLE_ADVANCED_WRITES=true, expected_target format, allow_disruption for non-VM objects). It stops short of stating when NOT to use it versus the typed per-domain tools, so it is strong but not fully exclusive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_api_schemaBRead-onlyIdempotent
查看官方 SDK 类型、属性、方法、参数及返回类型。SDK 包含的方法仍可能不被当前 ESXi 支持。
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| method | No | ||
| type_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish this as a safe, idempotent, non-destructive read in a closed world, so the safety profile is covered. The description earns credit by adding a non-obvious behavioral caveat: the SDK schema is authoritative for signatures but does not guarantee the running ESXi host supports a given method — a real gotcha an agent would otherwise miss.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the capability statement front-loaded and the caveat trailing. No filler, though the caveat could be tied to an action rather than stated abstractly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover safety. However, with zero schema coverage the unexplained depth parameter and the lack of routing versus esxi_api_get / esxi_resource_schema leave gaps for a three-tool schema-inspection family.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It loosely maps to type_name and method ('types... methods, parameters') but says nothing about the 'depth' parameter, whose recursion depth default of 2 is the least self-explanatory field. Undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (official SDK schema) and enumerates exactly what it exposes: types, properties, methods, parameters and return types. This distinguishes it from siblings like esxi_api_invoke or esxi_resource_schema by implication, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — an agent is expected to consult this before invoking an SDK call — and the caveat that 'SDK methods may still not be supported by the current ESXi' warns about reliability rather than telling when to prefer this over esxi_api_get or esxi_resource_schema. No explicit when/when-not or alternative routing is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_api_transferADestructive
处理 GuestFileManager 或 HttpNfcLease 返回的 HTTPS URL,完成 guest 文件/OVF-NFC 的 HTTP 阶段。
仅连接当前 ESXi 的 guestFile/NFC 路径。上传默认预览,执行需普通和高级写开关。 lease 传发现的 HttpNfcLease 引用以在下载源文件/上传时刷新进度,最后用 API 调用 Complete/Abort。 OVF 流式 VMDK 通常用 POST 和 application/x-vnd.vmware-streamVmdk;guest 文件用 PUT。 该工具不是一键 OS 安装器;guest 需要 Tools 和认证,OVA 需先取得描述符和独立磁盘数据。
| Name | Required | Description | Default |
|---|---|---|---|
| lease | No | ||
| dry_run | No | ||
| max_bytes | No | ||
| operation | Yes | ||
| source_url | No | ||
| data_base64 | No | ||
| http_method | No | PUT | |
| content_type | No | application/octet-stream | |
| transfer_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, non-idempotent, readOnly=false, closed-world. The description adds real behavior beyond that: the dry-run/preview default, the requirement for normal and advanced write switches to actually execute, the lease-completion/abort lifecycle, and authentication prerequisites. This is meaningful context on top of the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five front-loaded lines that each carry distinct information (mechanism, scoping constraint, write-switch behavior, lease lifecycle, HTTP method mapping) with a closing limits statement. It is appropriately sized for a 9-parameter, destructive transfer tool, with only minor jargon density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, return values need no explanation, and the description covers the transfer flow, HTTP verb selection, lease refresh/complete/abort, write-switch gating, and prerequisites. The remaining gap is that a 9-parameter tool with 0% schema coverage does not describe the operation types or source/data parameters, but overall it is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 9 parameters, so the description must carry the burden. It partially does: it explains lease ('传发现的 HttpNfcLease 引用'), http_method/content_type (OVF streaming uses POST with application/x-vnd.vmware-streamVmdk; guest files use PUT), and touches dry_run (preview default). But operation, source_url, data_base64, and max_bytes receive no explanation, leaving several parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and mechanism: handling HTTPS URLs returned by GuestFileManager/HttpNfcLease to complete the HTTP phase for guest-file and OVF-NFC transfers. It also clarifies what it is not ('not a one-click OS installer'), which sharpens scope. It does not, however, name a sibling tool it supersedes (e.g. esxi_datastore_transfer or esxi_transfer_start), so differentiation is only partly achieved.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives usable context: only connect to the current ESXi guestFile/NFC paths, uploads are preview-by-default while execution needs write switches, and it states prerequisites (guest needs Tools/auth, OVA needs descriptor and disk data first). However, it never explicitly says when to choose this over sibling transfer tools like esxi_datastore_transfer, esxi_transfer_start, or esxi_ovf_import, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_capabilitiesCRead-onlyIdempotent
实际主机版本、许可证、可发现管理器、工具操作目录及部署开关;不把 SDK schema 当作实测支持。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so the safety profile is covered. The clause '不把 SDK schema 当作实测支持' adds a genuine behavioral caveat: results reflect the live host, not the SDK schema. That is useful but thin, and return shape is left to the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single semicolon-joined run-on whose head is a catalog of things rather than an action. The one behavioral caveat is buried after the list instead of front-loaded, so the reader must parse the whole string to learn what the tool actually does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need no prose. But a capability-discovery tool in a 50-sibling namespace must at minimum state that it returns live host ground truth and when to prefer it over schema tools; this description leaves callers unable to route themselves correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. There is nothing for the description to compensate for, and it correctly does not invent parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a run-on list of nouns (actual host version, license, discoverable managers, tool operation catalog, deployment switches) rather than a verb+resource statement. It never says what the tool does — read, return, or enumerate — though the name and readOnly annotation imply a capability query. It also fails to distinguish itself from neighbors like esxi_api_schema or esxi_resource_schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative named. With 50+ siblings including esxi_api_schema and esxi_resource_schema, this is the single most important missing piece and it is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_certificate_manageADestructive
专用 certificate 管理:csr_create, csr_by_dn, install, ca_replace, notify_services。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and non-idempotent, but the description adds real operational context the annotations cannot: the preview-first default, the need for an exact expected_host, explicit disruption/dependency acknowledgements, and request_id-based duplicate-execution protection. The request_id dedupe does not contradict idempotentHint=false (which refers to natural idempotency without a request_id), so no contradiction is flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense clauses, purpose front-loaded before the workflow and safety requirements. Nothing is padding, though the compressed enumeration of actions and preconditions makes it slightly harder to parse than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a destructive, multi-action tool with zero schema description coverage, the description covers the safety preconditions and workflow but says nothing about what distinguishes the five actions or what a failed install/ca_replace implies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must carry the load. It names expected_host, request_id, the default preview (dry_run) and gestures at the write switches and dependency acknowledgements, but the required 'action' parameter has no enum and the free-form 'arguments' object is entirely unexplained, leaving meaningful gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (certificate management) and enumerates the concrete operations it performs (csr_create, csr_by_dn, install, ca_replace, notify_services), which maps directly to the required 'action' parameter. This clearly separates it from the other esxi_*_manage siblings, though the enumeration is a bare list rather than an explanation of each action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a workflow hint — run resource_schema/get first and rely on the default preview (dry_run=true) before executing — and notes that real execution requires expected_host, write switches and dependency confirmation. However it never states when NOT to use this tool or which sibling to prefer for adjacent certificate/host tasks, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_configure_vmBDestructive
修改 CPU/内存(核数、MHz limit/reservation、MB limit/reservation)。未传字段不变。
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | No | ||
| vm_id | Yes | ||
| dry_run | No | ||
| memory_mb | No | ||
| cpu_mhz_limit | No | ||
| expected_name | Yes | ||
| memory_mb_limit | No | ||
| cpu_mhz_reservation | No | ||
| memory_mb_reservation | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds genuinely useful behavioral context beyond that — '未传字段不变' clarifies partial-update semantics — but omits the critical dry_run-defaults-to-true preview behavior and any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and scope, ending with the partial-update rule. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter destructive mutation tool with 0% schema coverage, the description is too thin: it never addresses dry_run (which silently defaults to a preview), expected_name concurrency semantics, or permission requirements. The existing output schema spares it from explaining return values, but the input-side gaps remain significant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It maps to roughly 6 of 9 fields (cpu, cpu_mhz_limit/reservation, memory_mb, memory_mb_limit/reservation) but never explains the critical dry_run flag (default true) or the required vm_id / expected_name parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (修改/modify) and resource (VM CPU/内存) and enumerates the affected fields (核数、MHz limit/reservation、MB limit/reservation). It is distinguishable from siblings like esxi_rename_vm and esxi_power_vm, though it does not explicitly differentiate itself from esxi_vm_hardware.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives, no prerequisites, and no mention that dry_run defaults to true (so the default call is a no-op preview). The only usage-relevant statement is the partial-update rule '未传字段不变'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_create_snapshotDDestructive
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| vm_id | Yes | ||
| memory | No | ||
| dry_run | No | ||
| quiesce | No | ||
| description | No | ||
| expected_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_create_vmADestructive
创建 VM(明确名称/CPU/内存/磁盘/datastore/network/guest_id,使用正确 resourcePool)。
新建 VM 的 name 就是目标本身,无需额外 expected_name。datastore 和 network 必须 从查询结果中明确指定。创建空 VM,不安装操作系统;不包含 clone/migrate。
| Name | Required | Description | Default |
|---|---|---|---|
| cpu | Yes | ||
| name | Yes | ||
| disk_gb | Yes | ||
| dry_run | No | ||
| network | No | ||
| guest_id | No | otherGuest | |
| datastore | No | ||
| memory_mb | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true / readOnlyHint=false, so the safety profile is covered. The description adds real context beyond that: the VM is created empty with no OS installed, and clone/migrate are out of scope. It omits the notable dry_run=true default, which matters for a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and its required inputs, then constraints and exclusions. No filler, though the parenthetical parameter list is dense rather than elegant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations carry the safety profile. The description covers scope, exclusions, and parameter sourcing, leaving only the dry_run default unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry semantics, and it names 7 of the 8 parameters plus the datastore/network sourcing constraint. However it omits units (memory_mb, disk_gb), the guest_id default, and the dry_run flag, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (创建 VM / create VM) and enumerates the required inputs, then explicitly scopes out clone/migrate. An agent can distinguish this from siblings like esxi_rename_vm, esxi_configure_vm, or esxi_ovf_import without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear preconditions: datastore and network must be explicitly supplied from query results (implying the list_datastores/list_networks siblings), and notes the name is the target itself so no expected_name is needed. It does not explicitly name which sibling to call for the lookup, so it falls just short of the top band.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_datastore_manageADestructive
专用 datastore 管理:nfs_create, vmfs_create, vmfs_expand, vmfs_extend, remove, local_create, vvol_create, swap_update。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds real value beyond that by explaining the safety machinery: default dry-run preview, precise expected_host targeting, explicit write switches, dependency acknowledgment, and request_id-based duplicate protection. These are exactly the mutation-safety details annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded: the action roster comes first, then the execution protocol. No filler sentences. The density is a slight trade-off against readability, since the action names are cryptic without the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive multi-action management tool with an output schema already present, the description covers the key operational requirements (preview-first, host precision, dependency confirmation, idempotency). It stops short of defining each action's semantics, but nothing critical to safe invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description carries the explainer burden. It groups and hints at the roles of expected_host, dry_run, request_id, and the write/acknowledgment switches, but never defines the individual parameters or the allow_disruption / allow_unknown_impact / allow_protected_impact distinctions, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (datastore) and enumerates the concrete actions it manages (nfs_create, vmfs_create, vmfs_expand, vmfs_extend, remove, local_create, vvol_create, swap_update), so an agent can distinguish it from the read-only esxi_list_datastores. It does not explicitly differentiate itself from siblings like esxi_datastore_transfer or esxi_storage_manage, which keeps it below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete workflow guidance: run resource_schema/get first, rely on the default preview (dry_run), and complete expected_host / write switches / dependency confirmation before actual execution. It lacks an explicit when-not-to-use statement or direct routing to sibling tools, but the sequencing guidance is clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_datastore_transferADestructive
Datastore 文件 stat/download/upload。可上传 ISO/OVF/VMDK 等文件(不等于部署完成)。
upload 用 base64 或 HTTPS source_url;默认只预览,真实写要求普通和高级写开关。 默认不覆盖,保护管理 VM 目录。URL 上传可流式接收大文件;max_bytes 最大 8GiB。 download 的 base64 响应最多 8MiB。服务端不读取任意本地文件。
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| dry_run | No | ||
| datastore | Yes | ||
| max_bytes | No | ||
| operation | Yes | ||
| overwrite | No | ||
| source_url | No | ||
| data_base64 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavioral context beyond annotations: dry_run defaults to preview-only, real writes require both a normal and advanced write switch, default no-overwrite protects admin VM directories, URL upload streams large files, download base64 capped at 8MiB, server never reads arbitrary local files. Destructive hint is corroborated by the overwrite caveat. Slightly less explicit on exact size limits for uploads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Terse and information-dense, mixing Chinese and English. Front-loads the operations, then behavioral caveats. Some phrasing is compressed to the point of ambiguity (e.g., '普通和高级写开关'), which slightly hurts readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, 0% schema coverage, annotations present, and an output schema, the description supplies the key usage and safety context an agent needs. It lacks explicit routing guidance against the many sibling transfer/stage tools, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden. It explains data_base64 vs source_url, the 8GiB max_bytes for URL streaming, the 8MiB download cap, dry_run preview semantics, and overwrite behavior. It doesn't cover 'path' or 'datastore' formatting, but compensates well for the documented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb+resource: stat/download/upload of datastore files, and clarifies that upload does not equal deployment completion. It distinguishes itself from siblings like esxi_stage_upload and esxi_transfer_start, though the exact boundary with esxi_transfer_start/esxi_stage_upload isn't spelled out.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives context on dry-run default preview, write switches, and no-overwrite default, which implies usage. But with siblings like esxi_transfer_start, esxi_stage_upload, esxi_transfer_read_chunk, it doesn't explicitly say when to choose this tool over those chunked-transfer alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_delete_snapshotDDestructive
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| snapshot_id | Yes | ||
| expected_name | Yes | ||
| remove_children | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_delete_vmBDestructive
删除 VM(必须已关机;dry-run 明示删除磁盘)。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| expected_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is covered. The description adds real value beyond that: the powered-off precondition and the fact that dry-run surfaces disk deletion, which is exactly the kind of side-effect disclosure an agent needs for a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the precondition front-loaded and no filler. Appropriately sized, though the extreme brevity contributes to the parameter gaps noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. But for a destructive, 0%-coverage 3-parameter tool with two required params, the description omits the meaning of expected_name and gives no when-to-use routing, leaving clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It only illuminates dry_run's behavior ('dry-run 明示删除磁盘'); the required expected_name safety-confirmation parameter and vm_id are left entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (delete VM), which is unambiguous and clearly distinct from the create/power/configure siblings. It does not explicitly name a sibling alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one hard prerequisite ('必须已关机' — must be powered off), which is genuine usage guidance. However it never says when to prefer this over related operations such as powering off first, or what to use instead for non-destructive teardown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_dependenciesBRead-onlyIdempotent
VM→网络/存储/文件依赖及管理 VMkernel 连接,供共享资源变更前审查。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the full safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false, openWorldHint=false), so the description only needs to add scope. It does add useful scope (three dependency domains), but the '管理 VMkernel 连接' wording implies mutation and is never reconciled with the read-only annotations, which introduces doubt rather than clarity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the resource (VM dependency domains) before the purpose clause, with no filler. The only inefficiency is the ambiguous VMkernel clause, which spends words without adding usable meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters, rich annotations, and a published output schema, the description need not explain return values, so the baseline burden is low. It still leaves the read-vs-manage ambiguity unresolved, which is the one thing an agent would most want clarified before calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. No parameter-level semantics are required or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a recognizable resource (VM dependencies across network/storage/file) and a use case (review before shared-resource changes), so an agent can broadly tell what it returns. However there is no clear action verb, and the trailing clause about '管理 VMkernel 连接' (managing VMkernel connections) blurs whether this reports dependencies or performs connection administration, leaving the purpose partially ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '供共享资源变更前审查' gives a genuine when-to-use condition: invoke this before modifying shared resources. There is no when-not guidance and no reference to alternatives such as esxi_host_summary, esxi_list_networks, or esxi_list_datastores, so the routing logic is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_esxcliBDestructive
ESXCLI完整argv入口,用于SDK之外的主机配置/补丁/驱动/存储操作。SSH默认关闭,执行需管理员、独立host-key校验与全部影响确认。
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | ||
| timeout | No | ||
| arguments | Yes | ||
| expected_host | No | ||
| allow_disruption | No | ||
| max_output_bytes | No | ||
| acknowledged_vm_ids | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and not idempotent. The description adds valuable context: SSH is disabled by default, execution requires administrator privileges, independent host-key verification, and full impact confirmation. This goes beyond the annotations and clarifies the operational and safety constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense sentence that front-loads the core purpose and then lists key operational constraints. It is appropriately concise for the amount of context provided, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's high risk (destructive, open-ended argv) and 8 parameters with zero schema coverage, the description rightly emphasizes safety prerequisites. However, it omits any explanation of parameter effects (e.g., what dry_run or allow_disruption do) and does not guide the agent on selecting this tool over the numerous specialized siblings, leaving significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all 8 parameters, yet the description provides no information about what any parameter means, including the required arguments array, dry_run, timeout, expected_host, allow_disruption, max_output_bytes, acknowledged_vm_ids, or allow_protected_impact. The description fails to compensate for the complete lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies this as a complete argv entry point for ESXCLI on the host, covering configuration/patch/driver/storage operations outside the SDK. It is a clear purpose but does not explicitly distinguish it from the many sibling management tools or the esxi_esxcli_query tool, leaving ambiguity about when to use this versus those.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage for host operations beyond the SDK, and notes SSH is disabled and admin access is required, which hints at prerequisites. However, it does not state when to use this generic entry point versus the many specific sibling tools like esxi_patch_manage or esxi_storage_manage, nor does it mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_esxcli_queryCRead-onlyIdempotent
可选SSH后端的固定只读ESXCLI查询。查询白名单由服务器判断,不接受调用者自报read_only。
| Name | Required | Description | Default |
|---|---|---|---|
| timeout | No | ||
| arguments | Yes | ||
| max_output_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful context beyond that: the SSH backend is optional, the query whitelist is enforced server-side, and caller-supplied read_only claims are ignored. That whitelist enforcement detail is exactly the kind of behavioral trait annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core constraint (read-only, server-judged whitelist) front-loaded and no filler. It is efficient, though the second sentence about read_only is arguably confusing given no such parameter exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but with 0% schema coverage the three undocumented parameters leave a substantial gap, and the description offers no differentiation from the sibling esxi_esxcli tool. For a query tool whose 'arguments' array is the primary interface, the definition is under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, but it says nothing about the three parameters (timeout, arguments, max_output_bytes). Worse, it discusses read_only, which is not even a parameter of this tool, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it is a fixed read-only ESXCLI query with an optional SSH backend, which is a recognizable verb+resource. However, it does not differentiate this from the sibling esxi_esxcli tool, and it never explains what the query targets or what 'arguments' represent, so the agent cannot tell the two apart from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this versus esxi_esxcli, esxi_shell, or the various esxi_*_manage tools. The only conditional information is the server-side whitelist, which is a safety constraint rather than usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_execute_planADestructive
后台顺序执行明确步骤;每步先预览,异步任务等到成功再下一步。可指定compensate,失败按逆序执行;不承诺原子回滚、不自动重试。
| Name | Required | Description | Default |
|---|---|---|---|
| steps | Yes | ||
| dry_run | No | ||
| request_id | No | ||
| timeout_per_step | No | ||
| rollback_on_failure | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite annotations already marking destructiveHint=true and idempotentHint=false, the description adds substantial behavioral context: background sequential execution, per-step preview, async task waiting, optional compensate in reverse order, and explicit disclaimers of atomic rollback and automatic retry. These details materially inform how the tool behaves and what guarantees it does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with semicolons, front-loading the primary action and then layering constraints. Every clause adds useful information (preview, async waiting, compensate, rollback limits, no retry) without redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex (multi-step plan execution, destructive, no atomic rollback) and has an output schema, so return values need not be explained. The description covers behavioral guarantees and limitations well, but given 0% schema coverage for five parameters, it should explain the steps format and key controls like dry_run or rollback_on_failure. That gap makes it only partially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 5 parameters. The description mentions 'preview' (possibly related to dry_run) and 'compensate' (possibly related to rollback_on_failure) but does not name or explain any parameter. It leaves steps structure, request_id, timeout_per_step, and rollback_on_failure entirely undocumented, which is inadequate for a tool with zero schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the core action: executing explicit steps sequentially in the background. It specifies key mechanics (preview first, wait for async success before next step, optional compensate in reverse order), which distinguishes it from generic executors. However, it does not explicitly name or contrast with sibling tools like esxi_stage_finalize or esxi_api_invoke, so it misses full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for multi-step plans with preview and compensation, but never states when to choose this tool over alternatives such as esxi_api_invoke or esxi_stage_finalize. There are no explicit when-to-use or when-not-to-use conditions, leaving the agent to infer suitability from behavioral details.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_expand_diskCDestructive
仅允许增长,由 device_key 精确定位;有快照时拒绝。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| device_key | Yes | ||
| new_size_gb | Yes | ||
| expected_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=false. The description usefully adds that expansion is growth-only and requires no existing snapshots, which the annotations do not convey. It omits other meaningful behavior, notably that dry_run defaults to true and what a rejected call returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single terse clause-pair with no filler, and the constraints are front-loaded. It is efficient, though the extreme brevity borders on under-specification rather than tight concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, but for a 5-parameter destructive mutation with 0% schema coverage the description is far too thin: no dry_run semantics, no explanation of expected_name or vm_id, and no failure/permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning, but it only touches device_key (precision targeting) and implicitly new_size_gb (growth only). Key parameters vm_id, expected_name and especially dry_run (default true) are left completely unexplained, leaving most of the 5-param surface undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is entirely constraint-based ("growth only", "rejected with snapshots") and never states the core action of expanding a virtual disk; the agent must infer the verb+resource from the tool name alone. It distinguishes the constraint from siblings like esxi_vm_hardware, but the purpose itself is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"有快照时拒绝" conveys a when-not-to-use condition (no snapshots present), which is genuine guidance. However, it names no alternative (e.g., delete snapshot first) and gives no positive context for when to choose this over esxi_vm_hardware or esxi_configure_vm.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_firewall_manageBDestructive
专用 firewall 管理:enable, disable, rules_update, default_policy。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so safety profile is carried by structured data. The description adds real context beyond that: preview defaults, host-pinning via expected_host, write-toggle dependencies, and idempotency via request_id. It still does not explain what the destructive default policy change affects or the semantics of allow_disruption/allow_protected_impact gating, so it is good-but-incomplete rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler and the action list is front-loaded, but the prose packs multiple distinct concepts (preconditions, host constraint, toggles, dedup) into a single run-on clause, which hurts scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return-value explanation is not required, and annotations cover the safety profile. Given 9 params at 0% schema coverage and a destructive mutation, the description is adequate but leaves too many parameters and blast-radius details undocumented for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 9 parameters, so the description must compensate — it names expected_host, write toggles, and request_id but leaves action, dry_run, arguments, and the three allow_* impact flags unexplained. It provides partial compensation, warranting the baseline-ish 3, but falls well short of the full burden for a 9-param destructive tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (firewall management) and enumerates the operations (enable, disable, rules_update, default_policy), which distinguishes it from sibling *_manage tools like esxi_network_manage. However, it never contrasts itself against siblings like esxi_host_manage or esxi_options_manage, and the enumerated actions appear only in prose, not in any schema enum.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions a precondition workflow ('先 resource_schema/get 和默认预览') and repeat-execution protection via request_id, which implies context. But there is no explicit when-not-to-use guidance and no named alternative sibling; usage is implied rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_get_taskBRead-onlyIdempotent
查询异步任务状态:queued/running/success/error。
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description usefully enumerates the possible task states (queued/running/success/error), adding lifecycle context, but says nothing about polling cadence or whether terminal states are sticky.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the verb front-loaded and the state list appended where it is useful. Nothing is wasted, though the extreme terseness leaves little room for the routing context an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover safety. However, for a tool embedded among 60+ siblings, the description omits how task_id is obtained and how this differs from the other *_status tools, leaving the definition minimally viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description says nothing about task_id — notably omitting that it is the identifier returned by other operations. With a single undocumented parameter carrying the entire input contract, the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (查询) and resource (异步任务状态), so the agent knows this polls an async task. It does not distinguish itself from near-identical siblings such as esxi_operation_status, esxi_transfer_status, or esxi_guest_process_status, which also report status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus the many other status/polling tools in the sibling set, nor when in an operation lifecycle it should be used. The agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_get_vmCRead-onlyIdempotent
按 vm_id(MoRef)查询单台 VM 详情。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile fully. The description adds only that lookup is by MoRef, which is modest context but does not raise the bar beyond what annotations provide. No mention of error behavior for invalid or missing vm_id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the identifier and the resource. Nothing is wasted, though there is no elaboration because there is essentially nothing else, and the language mismatch with the surrounding English tool names is a minor structural inconsistency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not explain return values. However, for a tool whose only parameter is undocumented in the schema and whose description provides just the word MoRef, the definition leaves the agent without guidance on obtaining a valid vm_id or on sibling selection. It is thinner than adequate for even a simple lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%: the single required parameter vm_id has no description in the schema. The description partially compensates by clarifying that vm_id is a MoRef, but it offers no format example, no note on where the MoRef comes from (e.g., esxi_list_vms), and no indication of failure when it is omitted or invalid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource (query a single VM's details) and names the key parameter (vm_id/MoRef). However, it is in Chinese while all sibling tool names and the schema are in English, and it makes no effort to distinguish itself from esxi_list_vms or esxi_vm_hardware, which occupy adjacent retrieval territory. The core purpose is understandable but not sharply differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. With siblings like esxi_list_vms and esxi_vm_hardware, an agent has no statement that this tool is for fetching one VM by MoRef rather than enumerating all VMs. The behavior is inferable from the name but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_guest_filesBDestructive
Guest 文件list/mkdir/delete/move/attributes/upload_url/download_url。HTTP传输用后台transfer_start(kind=guest),大文件不经LLM传整个base64。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| action | Yes | ||
| dry_run | No | ||
| arguments | Yes | ||
| request_id | No | ||
| expected_name | Yes | ||
| authentication | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false, so the safety profile is known. The description reinforces destructiveness by listing delete and move, and adds a useful warning about not sending large base64 through the LLM. However, it does not explain what gets destroyed, authentication requirements, or rate limits, so it adds only moderate context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loading the action list and then the transfer guidance. It is highly concise with no wasted words. The mixed Chinese/English phrasing is slightly cryptic but still compact and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters (5 required), nested objects, and 0% schema description coverage, the description is far too sparse. It does not explain required parameters like authentication, expected_name, or the structure of arguments, leaving the agent without enough information to call the tool correctly. The output schema exists, so return values need not be described, but input documentation is severely lacking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all 7 parameters, so the description must compensate. It only implicitly hints at the action parameter's possible values (list, mkdir, delete, move, attributes, upload_url, download_url). It says nothing about vm_id, expected_name, authentication, arguments, dry_run, or request_id, leaving most parameters completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists specific guest file operations (list/mkdir/delete/move/attributes/upload_url/download_url), making the purpose clear as a multi-action guest file management tool. It distinguishes itself from transfer-related siblings by directing HTTP transfers to transfer_start(kind=guest). However, it does not explicitly contrast with other guest tools like esxi_guest_process.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides one clear usage rule: for HTTP transfers use transfer_start(kind=guest) instead of passing large base64 through the LLM. This is a valuable alternative path. But it lacks broader when-to-use guidance, such as when to use this tool versus esxi_datastore_transfer or esxi_guest_process, and gives no prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_guest_processADestructive
Guest 进程start/terminate:显式OS内路径和认证;start返回PID,再process_status查看exitCode。不会自动用ESXi shell替代guest。
| Name | Required | Description | Default |
|---|---|---|---|
| pid | No | ||
| vm_id | Yes | ||
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| environment | No | ||
| program_path | No | ||
| expected_name | Yes | ||
| authentication | Yes | ||
| working_directory | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (destructiveHint=true, idempotentHint=false), and the description adds non-obvious behavior beyond them: explicit in-OS program path and authentication are required, start returns a PID, and there is no automatic ESXi-shell substitution. The destructive nature of terminate is not elaborated, but the added context is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact, front-loaded two-clause statement with essentially no wasted words, leading with the start/terminate action and following with the exitCode hand-off and shell caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 11 params, a nested authentication object, and no schema descriptions, the description is too thin. An output schema exists so return values need not be explained, but the input contract (dry_run defaulting to true, action semantics, path vs arguments) is left undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, so the description must compensate and largely does not. It only alludes conceptually to program path and authentication; fields like dry_run, action, arguments, environment, working_directory and expected_name are left entirely to the schema (which has no descriptions of its own).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (start/terminate a guest process) and names the sibling (process_status) it hands off to for exitCode, so an agent can distinguish it from esxi_guest_process_status and esxi_shell. Clear overall, though terse and partly in a mixed language register.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable context: start then use process_status to read exitCode, and clarifies it will NOT silently fall back to the ESXi shell. This tells the agent when and how to use it versus the shell-based sibling, though no explicit when-not-exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_guest_process_statusBRead-onlyIdempotent
查询guest进程状态、PID、启动/结束时间与exitCode;只读,可检查受保护VM。
| Name | Required | Description | Default |
|---|---|---|---|
| pids | No | ||
| vm_id | Yes | ||
| expected_name | Yes | ||
| authentication | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint and openWorldHint, so the safety profile is covered. The description's '只读' merely restates the annotation, but '可检查受保护VM' adds genuine context about operating against protected VMs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence listing the resource and returned fields, then a short capability clause. No filler, though the readOnly restatement is slightly redundant with the annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed, and the description covers the operation's scope. However, with a nested authentication object and three required params at 0% schema coverage, the definition is only minimally complete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and none of the four parameters (pids, vm_id, expected_name, authentication) are explained in the schema. The description's mention of PID loosely maps to the pids parameter, but the required vm_id/expected_name/authentication params receive no clarification, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (查询guest进程状态) and enumerates the returned fields (PID, start/end time, exitCode), so the operation is identifiable. It does not explicitly distinguish itself from the sibling esxi_guest_process, leaving the status-vs-launch split to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no named alternative among the many siblings. The clause 可检查受保护VM hints at a scenario but is a capability note rather than routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_healthARead-onlyIdempotent
ESXi 主机整体健康状态(连接态/维护模式/overallStatus/uptime/quickStats 即时负载)。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds context by listing the specific health dimensions returned, which goes slightly beyond annotations. However, with an output schema present, return details are redundant, and no additional behavioral context (e.g., freshness, scope) is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, compact sentence that front-loads the resource and lists the covered fields without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters, a present output schema, and rich annotations, the description is sufficiently complete for an agent to call it. The only minor gap is lack of explicit routing versus siblings like esxi_host_summary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. No parameter semantics are needed, and the description correctly omits them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific purpose: retrieving ESXi host overall health status, enumerating the exact fields it covers (connection state, maintenance mode, overallStatus, uptime, quickStats). This distinguishes it from siblings like esxi_host_summary or esxi_inventory, though the differentiation is not explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as esxi_host_summary or esxi_inventory. The description only states what it returns, not when an agent should choose it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_host_manageADestructive
专用 host 管理:maintenance_enter, maintenance_exit, reboot, shutdown, lockdown_enter, lockdown_exit, disconnect, reconnect。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructive=true, readOnly=false, non-idempotent, so the safety profile is known; the description adds real value by explaining the preview-first protocol, the need for an exact expected_host, dependency confirmation, and the request_id dedup guard against repeated execution. It does not spell out per-action consequences (e.g. what reboot does to running VMs), which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action list is front-loaded and the workflow is compressed into two dense clauses with no filler. It is tight, though the run-on second sentence packs several distinct requirements (resource_schema/get, preview, expected_host, switches, dependencies, request_id) without separation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter destructive host operation with an output schema present (so return values need no explanation), the description covers actions, prerequisite discovery, preview semantics, identity pinning and deduplication. Remaining gaps are the per-parameter meaning of the allow_* flags and acknowledged_vm_ids, which the zero-coverage schema leaves undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate: it references expected_host, request_id, dry_run preview, and the 'write switches' (allow_disruption / allow_unknown_impact / allow_protected_impact) plus dependency confirmation. However it never distinguishes the three allow_* switches or acknowledged_vm_ids, and the arguments/action payload structure is left unclear, so the gap is only partially filled.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names the resource (ESXi host) and enumerates the exact eight operations it performs (maintenance enter/exit, reboot, shutdown, lockdown enter/exit, disconnect, reconnect), which clearly distinguishes it from the sibling *_manage tools for network, datastore, storage, service, firewall, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It prescribes a workflow: call resource_schema/get first, rely on the default preview (dry_run defaults true), then execute with precise expected_host, write switches and dependency confirmation. That is strong when-to-use guidance, though it never states when NOT to use it (e.g. use esxi_power_vm for VMs) or names exclusions explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_host_summaryBRead-onlyIdempotent
主机容量摘要。quickstats_* 为物理已用量;VM 配置内存总和另列,勿混淆。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so safety is covered. The description adds useful interpretive context that quickstats_* values are physical usage while VM configured memory is reported separately, which is behavioral nuance not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core purpose leads. The second sentence leans toward explaining return semantics, which partly duplicates the output schema, so it is not maximally economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return field documentation is not required in the description. However, for a summary tool surrounded by many overlapping host/inventory siblings, the absence of any routing guidance leaves a genuine gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is nothing for the description to disambiguate. The baseline for a zero-parameter tool applies, and no schema documentation debt exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and scope: a host capacity summary. That is a clear verb-less but concrete noun phrase distinct from esxi_health or esxi_inventory. It does not, however, explicitly contrast itself with those overlapping siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no named alternative, even though siblings such as esxi_health, esxi_inventory, and esxi_host_manage could plausibly be confused with a host summary. The caution about mixing metrics is interpretive, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_inventoryCRead-onlyIdempotent
列出 SDK 管理实体的准确类型/ID/名称:VM、HostSystem、Datastore、Network、Folder、ResourcePool。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| resource_type | No | vim.ManagedEntity |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered by structured data. The description adds only the set of entity types returned; it says nothing about result size, default behavior of the limit parameter, or whether the inventory is snapshot/current-state based, so it provides modest extra context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the action and the payload entity types with no filler. It is well sized for a simple listing tool, though it spends its words on enumerating types rather than on the ambiguities that matter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but with two undocumented parameters and roughly 65 sibling tools the description leaves the key questions unanswered: how resource_type selects entities and how this differs from the specialized list tools. It is not complete enough for an agent to invoke confidently without inspecting the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, so the description must carry the burden. It loosely hints at the domain of resource_type by naming entity kinds, but never explains that resource_type takes SDK type strings (default 'vim.ManagedEntity') or what limit=200 governs. The compensation is partial at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (列出) and resource (SDK 管理实体) and enumerates the entity kinds it covers (VM、HostSystem、Datastore、Network、Folder、ResourcePool), so the intent is unambiguous. However, it offers no differentiation from the many overlapping siblings such as esxi_list_vms, esxi_list_datastores, and esxi_list_networks that appear to retrieve much of the same data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this generic inventory tool versus the specialized list tools in the sibling set, nor any prerequisite, exclusion, or scope condition. An agent facing esxi_list_vms/esxi_list_datastores/esxi_host_summary cannot infer from this text which one to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_license_manageADestructive
专用 license 管理:add, remove, update, labels_update。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, but the description adds value beyond them: a default preview mode, precision on expected_host, required write/dependency acknowledgements, and request_id to prevent duplicate execution. It does not explain what exactly gets destroyed, but the safety workflow is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the purpose and the operation list front-loaded, followed by the execution requirements. No filler, though the compressed Chinese phrasing packs several distinct behaviors into one clause without breaking them out.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, and annotations cover the safety profile. However, for a 9-parameter destructive tool with 0% schema coverage, the description only partially documents the parameters, leaving a meaningful gap an agent must guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description carries the full burden. It names expected_host, request_id and hints at dry_run via '默认预览', but leaves action values, arguments, acknowledged_vm_ids, allow_unknown_impact and allow_protected_impact entirely unexplained – half the parameters remain opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (license management) and enumerates the supported operations (add, remove, update, labels_update). '专用 license 管理' clearly separates it from sibling management tools like network_manage or storage_manage, though it doesn't explicitly contrast with any specific sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete workflow: call resource_schema/get first and use the default preview, then execute with a precise expected_host, write switch and dependency confirmation. This tells the agent the conditions for safe invocation, though it stops short of explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_list_datastoresDRead-onlyIdempotent
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_list_networksDRead-onlyIdempotent
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_list_snapshotsBRead-onlyIdempotent
递归列出指定 VM 的快照树,输出 snapshot_id。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds two useful behavioral facts beyond that: traversal is recursive and the output key is snapshot_id. It still says nothing about ordering, depth limits, or how the tree is flattened.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the operation and scope with no filler. It is arguably under-specified rather than overly long, but there is zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the return format need not be re-explained (the mention of snapshot_id is redundant). With annotations covering safety, the remaining gaps are the missing vm_id format and the absent usage guidance, which keep it at minimum-viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single vm_id parameter, so the description must carry the meaning. It only refers to '指定 VM' (the specified VM), which adds nothing over the parameter name itself — no identifier format, source, or example is given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list/递归列出) and resource (快照树/snapshot tree) scoped to a given VM, so the agent knows exactly what it returns. It does not, however, differentiate itself from the snapshot siblings (esxi_create_snapshot, esxi_revert_snapshot, esxi_delete_snapshot) beyond the obvious verb difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to prefer this over the other snapshot tools or what preconditions apply (e.g., VM must exist, snapshot support). Usage is only implied by the verb 'list'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_list_vmsARead-onlyIdempotent
列出所有 VM(可按名称/IP 过滤)。内存字段为配置值,非物理内存占用。
| Name | Required | Description | Default |
|---|---|---|---|
| ip_filter | No | ||
| name_filter | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely non-obvious context not present in structured fields: that memory fields are configured values, not physical usage, which prevents misinterpretation of returned data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the primary action front-loaded and the data-interpretation caveat following. Nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure needn't be explained. For a read-only list tool with annotations carrying the safety profile, the description covers purpose, filtering, and the key data caveat; only explicit sibling routing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the two parameters are self-describing (ip_filter, name_filter) with titles. The description confirms both filters exist and apply to the listing, adding marginal meaning, but gives no format or matching-behavior detail (substring vs exact, wildcards).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (列出) and resource (VM), plus the filtering scope. It is clearly distinguishable from mutation siblings like esxi_power_vm or esxi_delete_vm, though it does not explicitly contrast with read siblings such as esxi_get_vm or esxi_inventory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: list all VMs, optionally narrowed by name/IP. There is no explicit when-to-use-this-vs-alternative guidance (e.g. use esxi_get_vm for a single VM's details), so the agent must infer routing from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_managersBRead-onlyIdempotent
发现服务与 HostConfigManager 的网络、存储、服务、防火墙、账户、硬件等管理器引用。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety and idempotency are covered. The description adds no further behavioral context such as return format, size, or caching behavior. With annotations carrying the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that lists the manager categories. It is efficient and wastes no words, though the list is slightly dense. Overall, it is appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple: no inputs, read-only, idempotent, and has an output schema. The description identifies the managers discovered, which is sufficient given the output schema will document return values. However, it could better differentiate from sibling inventory tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document. Per the rubric, zero parameters yields a baseline of 4. The description does not need to add parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the resource (HostConfigManager manager references) and the verb (discover), but it is framed as a vague discovery operation. It does not clearly distinguish itself from sibling tools like esxi_host_summary or esxi_inventory, which likely also expose host manager information. An agent could confuse this with other discovery-oriented tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as esxi_host_summary, esxi_api_schema, or esxi_inventory. There are no conditions, exclusions, or named alternatives, leaving the agent to infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_network_manageADestructive
专用 network 管理:switch_create, switch_update, switch_remove, portgroup_create, portgroup_update, portgroup_remove, vmkernel_create, vmkernel_update, vmkernel_remove, dns_update, route_update, routes_update, physical_link_update, configuration_update。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, idempotentHint=false, so the agent knows this is a disruptive write operation. The description adds meaningful behavioral context: use resource_schema/get and default preview first, specify expected_host for precision, enable write switches, confirm dependencies, and use request_id to prevent duplicate execution. These are operational behaviors beyond the annotations. However, it doesn't explain what gets destroyed or the consequences of the write switches, and the description is more of a workflow instruction than a disclosure of side effects. A 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph listing many actions and then workflow instructions. It is front-loaded with the action list, which is useful, but the long enumeration is somewhat unwieldy and the workflow sentence is packed. It is reasonably concise but could be better structured (e.g., separated into purpose, prerequisites, and execution notes).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 9 parameters, a nested 'arguments' object, and complex multi-action behavior, the description is only partially complete. It covers the general workflow and some parameters but leaves many details (action validation, argument structure, individual switch semantics) unaddressed. An output schema exists, so return values need not be explained. Missing parameter semantics and edge cases keep this at 3.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, but it only names a few parameters implicitly ('expected_host', '写开关' for allow_disruption/allow_unknown_impact/allow_protected_impact, 'request_id', 'dry_run' via 默认预览). It does not explain the 'action' enum values, the 'arguments' object structure, or the meanings of the allow_* switches. With 9 parameters and zero schema descriptions, this is a significant gap, but the description does mention some key parameters and their workflow role, meriting a 3 rather than lower.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a clear resource ('network 管理') and enumerates the specific operations covered (switch_create, portgroup_create, vmkernel_create, dns_update, route_update, etc.). This is a broad multi-action tool, and the action list gives good purpose clarity. It distinguishes itself from siblings like esxi_list_networks (read) and esxi_firewall_manage by covering the specific network mutation actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit workflow guidance: '先 resource_schema/get 和默认预览' (first use resource_schema/get and default preview), plus '实际执行精确 expected_host、写开关与依赖确认' (for actual execution, specify expected_host, write switches, and dependency confirmation). This tells the agent the prerequisite steps and when to set specific parameters. It lacks explicit when-not-to-use guidance or alternatives, so 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_operation_historyBRead-onlyIdempotent
查询当前主机/账户的持久化操作清单;不会输出密码或 signed URL。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: the persisted operation history will not expose passwords or signed URLs, which is a useful output-security disclosure beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the core purpose and appends the security caveat. No filler, nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return structure need not be explained, and annotations cover the safety profile. However, for a history/listing tool the description omits sibling differentiation and the meaning of its one parameter, leaving gaps an agent would otherwise need to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter (limit) exists with 0% schema description coverage and only a default value of 50. The description says nothing about it, so neither the schema nor the description explains what limit controls (count, pagination size, etc.).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (查询/query) and resource (持久化操作清单/persisted operation list) scoped to the current host/account. It is clear what the tool returns, but it never distinguishes itself from the sibling esxi_operation_status, which appears to occupy overlapping territory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of the closely related esxi_operation_status or esxi_get_task alternatives. The agent must infer from the name alone when this tool is preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_operation_statusBRead-onlyIdempotent
持久化操作记录:submitted 任务可重新连接查询;中断的 executing 操作不自动重复。
| Name | Required | Description | Default |
|---|---|---|---|
| operation_id | Yes | ||
| refresh_task | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the description earns credit for adding real behavioral context: records are persisted, submitted tasks reattach, and interrupted executing work is not retried. This is meaningful beyond the structured hints, though it omits return/pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with two clauses; no filler. It is terse and slightly telegraphic, but front-loads the core notion of persistent records efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. However, with 0% parameter coverage and no sibling disambiguation against esxi_operation_history, the description is only minimally sufficient for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description names neither operation_id nor refresh_task. Only a faint implication of re-querying exists; the meaning of refresh_task (default true) is entirely unexplained, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description establishes that this concerns persistent operation records but never states an explicit verb such as 'query the status of an operation by id'. The reader must infer the purpose from the tool name. It does not distinguish itself from the close sibling esxi_operation_history or esxi_get_task.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It notes that submitted tasks can be re-queried and that interrupted executing operations are not auto-retried, but gives no when-to-use guidance relative to esxi_operation_history or esxi_get_task. No prerequisite or selection criteria are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_options_manageBDestructive
专用 options 管理:update。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, idempotentHint=false, so the safety profile is partly covered; the description adds real context beyond that: dry_run defaults to a preview, real execution demands exact expected_host, explicit write switches and dependency confirmation, and request_id guards against duplicate execution. It stops short of saying what gets destroyed or which host objects are affected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded, with the purpose in the first clause and the workflow second, so nothing is padding. The telegraphic, jargon-dense phrasing ('写开关与依赖确认') makes it harder to parse than its length warrants.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, 9-parameter tool with 0% schema description coverage, the description is too thin: it leaves action/arguments semantics and the meaning of the three allow_* impact flags and acknowledged_vm_ids largely to inference. The output schema exists, so return values need not be covered, but input-side completeness is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate and only partly does: it hints at dry_run-based preview, expected_host, write switches and request_id. The required 'action' and the free-form 'arguments' object – the actual payload of the update – are left completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb (update) and a resource area ('专用 options 管理'), but 'options' is never grounded in ESXi terms (advanced settings / host options), and with ~65 sibling tools the agent cannot tell which namespace this touches. It does gesture at the resource_schema/resource_get workflow but never names a competing tool it should be picked over.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a usable procedural hint: fetch schema/get first, preview by default, then execute with precise expected_host, write switches and dependency confirmation. However, it never states when to choose this over esxi_host_manage, esxi_api_invoke, or esxi_resource_schema for the same options, so selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_ovf_exportADestructive
完整OVF导出:VM须已关机,后台导出磁盘、生成OVF与SHA256 manifest,并Complete/Abort租约。operation_status查询,transfer_read_chunk取回files句柄。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| request_id | No | ||
| expected_name | Yes | ||
| max_disk_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readOnly, destructive, non-idempotent, but the description adds traits annotations cannot express: the operation is asynchronous/background, it produces a SHA256 manifest, and it manages a lease lifecycle with Complete/Abort outcomes. It does not explain what the destructive side effects actually are or the failure/cleanup semantics in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense, front-loaded sentences: scope and precondition first, then the follow-up tool chain. The telegraphic second sentence ('operation_status查询,transfer_read_chunk取回files句柄') is terse but earns its place by naming the exact continuation tools.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the safety profile. However, with five parameters at 0% schema coverage and dry_run defaulting to true, a caller lacks the information needed to invoke this mutation correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across five parameters, so the description carries the full burden, yet it names no parameter. Critically, the dry_run flag defaults to true and max_disk_bytes defaults to 8 GiB, and neither their meaning nor expected_name's role is explained anywhere the agent can see.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('完整OVF导出') and enumerates the concrete work: exporting disks, generating the OVF and SHA256 manifest, and completing/aborting the lease. This is unmistakable next to siblings like esxi_ovf_import or esxi_transfer_start.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the key prerequisite ('VM须已关机') and routes the agent to the correct follow-up tools (operation_status for polling, transfer_read_chunk for retrieving the files handle). It never states when NOT to use it or contrasts it with alternative export/transfer paths, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_ovf_importADestructive
完整OVF导入:验证描述符/网络/数据存储/文件映射,后台NFC上传、完成租约;失败Abort,不删除已有VM。disks按OVF文件path映射source_job_id或HTTPS source_url。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| disks | Yes | ||
| dry_run | No | ||
| power_on | No | ||
| datastore | Yes | ||
| properties | No | ||
| request_id | No | ||
| expected_host | No | ||
| ovf_descriptor | Yes | ||
| accept_warnings | No | ||
| network_mappings | Yes | ||
| deployment_option | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so safety is partly covered. The description adds genuine behavioral context beyond that: it validates mappings before transfer, uploads via background NFC, completes the lease, and on failure issues Abort while leaving existing VMs intact. This rollback/abort behavior is exactly the kind of nuance annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense, front-loaded sentence covering the pipeline and the disks mapping. No filler, though the compressed style borders on telegraphic and could be slightly clearer about the destructive/rollback nuance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers the pipeline plus rollback behavior. However, with 12 params at 0% schema coverage, the semantics of most inputs remain undocumented, leaving a real gap for a complex, non-idempotent, destructive operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters (5 required, nested objects), so the description carries the full burden. It explains only the disks mapping (OVF path -> source_job_id or HTTPS source_url); the remaining eleven parameters, including dry_run, accept_warnings, deployment_option, expected_host, and network_mappings, get no semantic explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('完整OVF导入' – full OVF import) and enumerates the internal stages it performs (descriptor/network/datastore/file-mapping validation, NFC upload, lease completion). It is cleanly distinguishable from the sibling esxi_ovf_export.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (importing an OVF/OVA appliance) and notes failure handling, but never says when to prefer this over siblings like esxi_stage_upload, esxi_transfer_start, or esxi_vm_register, nor states prerequisites such as required permissions or a prior staging step. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_patch_manageADestructive
专用 patch 管理:check, scan, stage, install, uninstall, query。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is partly covered. The description adds meaningful context beyond them: it defaults to preview, demands a precise expected_host, requires write switches and dependency acknowledgement, and supports request_id for dedup — the kind of operational detail annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence, front-loaded with the action list followed by the execution prerequisites, with no filler. It is information-dense but reads as a compressed run-on rather than well-chunked guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the destructive-execution workflow reasonably. For a 9-parameter destructive multi-action tool, though, the unlisted parameters and the lack of per-action distinctions leave gaps an agent would need to fill elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage across 9 parameters, the description must carry the burden and only partly does: it maps to dry_run (preview), expected_host, the write switches, dependency acknowledgement and request_id. The generic 'arguments' object and the individual switch semantics (allow_disruption vs allow_protected_impact, acknowledged_vm_ids) are not explained, leaving several parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (patch) and enumerates the concrete operations (check, scan, stage, install, uninstall, query), so an agent knows the domain precisely. It does not, however, distinguish itself from adjacent workflow tools such as esxi_stage_finalize or esxi_execute_plan, so sibling differentiation is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a real prerequisite flow — go through resource_schema/get and a default preview before executing — which implies when to escalate to real execution. It never names an alternative tool or states when NOT to use this one, so guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_pci_manageADestructive
专用 pci 管理:configuration_update, refresh。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the safety profile is given; the description adds real value by disclosing the required write switches, dependency confirmation, exact expected_host matching, and request_id deduplication for replay safety. This goes meaningfully beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense, front-loaded sentence whose clauses each convey distinct information (purpose, actions, pre-flight, execution requirements). It is efficient but runs together several ideas, making it slightly hard to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description does cover the destructive-execution workflow. However, for a 9-parameter, destructive, non-idempotent tool it leaves the 'arguments' payload unexplained and the dependency-confirmation flow only half-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 9 parameters, so the description carries the full burden. It adds meaning for expected_host, request_id, dry_run ('默认预览') and the disruption switches, but the central 'arguments' object (and acknowledged_vm_ids) is never explained, leaving key parameters undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource ('专用 pci 管理') and enumerates the supported operations (configuration_update, refresh), which lets an agent separate it from the numerous generic *_manage siblings. It stops short of explaining what 'configuration_update' actually mutates, but the verb+resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It prescribes a concrete pre-flight workflow – '先 resource_schema/get 和默认预览' – pointing at sibling tools (esxi_resource_schema/get) and the default dry-run preview before real execution. This is clear context for when/how to use it, though no explicit when-not or alternative tool is named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_permissions_manageBDestructive
专用 permissions 管理:role_create, role_update, role_remove, permission_set, permission_remove。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true, so the mutation risk is declared. The description adds real behavioral context beyond that: a default dry-run preview step, exact expected_host pinning, write-switch and dependency confirmation, and request_id for deduplication. These are meaningful operational traits (idempotency, guardrails) not present in the annotations, though it doesn't spell out what gets destroyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It's a single dense sentence covering actions and workflow, appropriately brief. But the mixed Chinese/English phrasing and run-on structure reduce readability, and the most important safety info (irreversibility) is buried rather than front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, 9-parameter permission-management tool with an output schema present, the description is thin. It never explains the multi-parameter safety flags, the containment relationship between preview and execution, or how request_id interacts with the already-defaulted dry_run. An output schema exists, so return values needn't be explained, but the input contract is largely missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must carry the load. It explains request_id's dedup purpose, expected_host pinning, and the write/dependency confirmation implied by the allow_* flags, but never addresses dry_run, arguments, acknowledged_vm_ids, allow_disruption, allow_unknown_impact, or allow_protected_impact. Most parameters remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (permissions) and lists the sub-actions (role_create, role_update, role_remove, permission_set, permission_remove), which tells an agent what domain this covers. However, it's opaque about what a 'permissions manage' operation actually does at the call level and doesn't differentiate clearly from siblings like esxi_account_manage. The mix of Chinese and English also hurts clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It hints at a workflow ('先 resource_schema/get 和默认预览' – first call resource_schema/get and default preview), which implies when to use this tool versus raw esxi_api_invoke. But it gives no explicit when-not guidance or alternative comparison against the many sibling *_manage tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_power_vmADestructive
电源操作:power_on / shutdown(优雅,需Tools) / power_off / reboot_guest(需Tools) / reset。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| operation | Yes | ||
| expected_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and non-idempotent, so the safety profile is covered. The description adds genuinely new behavioral context beyond that: shutdown is '优雅' (graceful) and both it and reboot_guest depend on VMware Tools, implicitly warning that they can fail or behave differently from the hard power_off/reset paths.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, densely packed sentence that front-loads the resource and then the operation list with inline constraints. Nothing is wasted, though the parenthetical notation is terse enough to be slightly cryptic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, and the graceful/hard distinction is a useful addition for a destructive tool. Still missing: what expected_name guards against (apparently a safety confirmation), what dry_run=true actually does, and which operations are irreversible.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It partially does by supplying the valid values for 'operation' (power_on / shutdown / power_off / reboot_guest / reset), which the schema lacks as an enum, but vm_id, expected_name, and especially dry_run (defaulting to true) are left completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource (power operations on an ESXi VM) and enumerates the five supported operations, which is the key discriminator versus siblings like esxi_configure_vm or esxi_delete_vm. It does not, however, explicitly contrast itself with any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '需Tools' (requires VMware Tools) on shutdown and reboot_guest is a real prerequisite hint that tells the agent those two paths need guest tooling, unlike power_off/reset. But there is no guidance on when to prefer graceful shutdown over power_off, nor any exclusions or alternatives named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_rename_vmDDestructive
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| new_name | Yes | ||
| expected_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_resource_getCRead-onlyIdempotent
按资源域直接查询状态,无需手工拼管理器 ID;只允许已声明属性。
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| domain | Yes | ||
| max_items | No | ||
| properties | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world behavior, so the safety profile is covered. The description adds one genuinely useful rule beyond that: only declared properties are permitted ('只允许已声明属性'). It does not disclose pagination/truncation behavior implied by max_items and depth defaults.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, front-loaded with the core capability (query by domain) before the constraints. No filler, though it is perhaps too terse given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, but with 4 params at 0% schema coverage and no usage guidance, the description does too little for a tool sitting among many similar query/schema siblings. Undefined 'resource domain' and silent depth/max_items leave real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It maps only two of four params loosely ('资源域'→domain, '已声明属性'→properties) and says nothing about depth or max_items, leaving half the interface unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a verb and abstract resource ('按资源域直接查询状态') but '资源域' is undefined and it does not distinguish itself from conceptual siblings like esxi_resource_schema, esxi_api_get, or esxi_inventory. An agent cannot confidently tell what concrete data this returns versus those tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '无需手工拼管理器 ID' hints that this is a convenience alternative to manually assembling IDs (e.g., via esxi_api_get), but there is no explicit when-to-use, when-not, or naming of the alternative. No usage context for the required vs optional params is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_resource_schemaCRead-onlyIdempotent
专用资源操作目录和准确参数:network/datastore/storage/service/firewall/account/permissions/license/time/pci/certificate/patch/host。
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| action | No | ||
| domain | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint false, and openWorldHint false, so the safety profile is clear. The description adds that the tool returns a catalog of resource operations and accurate parameters, reinforcing its read-only discovery role, but it offers no further behavioral context such as caching, permissions, or output format (output schema exists).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with a colon-separated list; it is front-loaded with the tool's role and avoids waste. The fragmented noun-phrase style slightly reduces clarity, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has annotations and an output schema, reducing the description's burden, and the domain list covers the key input. However, it omits usage guidance and leaves the 'action' and 'depth' parameters unexplained, so it is only partially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It lists valid domain values (network, datastore, storage, etc.), which directly maps to the required 'domain' parameter, but it says nothing about the optional 'action' or 'depth' parameters. Partial compensation justifies a middle score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource domains it covers and calls itself a 'resource operation catalog and accurate parameters,' but it lacks a clear verb (e.g., 'retrieve schema') and does not distinguish itself from siblings like esxi_api_schema or esxi_resource_get. An agent can guess it provides schema details for the listed domains, but the purpose is not stated with the specificity of a verb+resource pair.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, prerequisites, or alternatives are given. It does not say when to call this schema tool instead of esxi_api_schema, esxi_resource_get, or the domain-specific manage tools, leaving the agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_revert_snapshotDDestructive
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| snapshot_id | Yes | ||
| expected_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_service_manageADestructive
专用 service 管理:start, stop, restart, startup_policy。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=false, so the danger is covered. The description adds value beyond them: dry_run defaults to preview, real execution requires an exact expected_host plus a write switch and dependency confirmation, and request_id guards against duplicate execution. These are concrete operational traits not present in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the purpose (dedicated service management) and the supported actions before the workflow caveats. It is compact and wastes little, though the compressed phrasing slightly hurts readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and annotations carry the safety profile. However, for a 9-parameter destructive tool, several impact-control parameters are left unexplained and the guidance on dry_run vs live execution is only implicit, leaving gaps an agent must fill by guesswork.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning. It does explain action values, dry_run preview, expected_host exactness, request_id dedup, and alludes to dependency/impact acknowledgement, but it leaves allow_disruption, acknowledged_vm_ids, allow_unknown_impact and allow_protected_impact unaddressed, so roughly half the parameters remain opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (service) and enumerates the concrete verbs: start, stop, restart, startup_policy. This clearly separates it from the many other *_manage siblings (network, datastore, storage, firewall, etc.). It is clear but does not explicitly name a sibling it should not be confused with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It advises a workflow: first call resource_schema/get and rely on the default preview before real execution. That is useful implied sequencing, but there is no explicit when-to-use / when-not-to-use vs alternatives, and no statement of prerequisites beyond a generic 'dependency confirmation'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_shellBDestructive
显式整台ESXi shell管理入口,仅administrator+SSH独立开关+管理链路/保护影响双重允许。任意shell无法做完整依赖推断,不自动重试。
| Name | Required | Description | Default |
|---|---|---|---|
| command | Yes | ||
| dry_run | No | ||
| timeout | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| max_output_bytes | No | ||
| acknowledged_vm_ids | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and readOnlyHint=false, but the description adds meaningful beyond-annotation context: administrator gating, the separate SSH requirement, dual approval via management-link and protection-impact flags, no automatic retry, and that full dependency inference is impossible. These are genuine operational traits. It stops short of what gets destroyed or a recovery path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single dense sentence with the gating constraints front-loaded and no filler. It is compact and information-rich, though the density and mixed metaphors ('management entry', 'link/protection') make it slightly hard to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description covers the critical safety gating for a hazardous execution tool. It is still incomplete: no guidance on the dry_run default or on how to satisfy the dual-approval and acknowledgement parameters an agent must set to proceed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 8 parameters, so the description must carry the full load; it only obliquely maps to allow_disruption/allow_protected_impact via 'dual allowance.' dry_run, timeout, expected_host, max_output_bytes, and acknowledged_vm_ids are never explained, leaving their semantics undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description calls itself the explicit host-wide ESXi shell entry point, which implies arbitrary shell command execution on the host, but the framing ('shell management entry') is abstract rather than a crisp verb+resource. It does not clearly state 'execute a command on the ESXi host' nor contrast itself with close siblings like esxi_esxcli or esxi_guest_process. Purpose is inferable but not sharply defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It lists real prerequisites: administrator-only access, SSH enabled as an independent switch, and dual allowance of management link and protection impact. However it gives no explicit when-to-use/when-not-to-use against alternatives such as esxi_esxcli or esxi_api_invoke, so routing among the many shell/API siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_stage_finalizeCDestructive
核对完整字节数及SHA256,完成暂存后可作为transfer_start的source_job_id。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, but the description does not acknowledge that finalization commits or mutates anything irreversible, nor what happens to the staged data afterwards. Critically, it never mentions that dry_run is a parameter that defaults to true, so an agent may believe the operation has executed when it has not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with no filler, front-loading the verification action and then the downstream consequence. Efficient, though very terse for a destructive, non-idempotent tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a destructive, non-idempotent finalize step the description omits the dry_run default, failure handling on hash mismatch, and any post-finalization side effects. Given 0% schema coverage, the definition is not complete enough to call safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with two parameters, so the description must compensate. It hints at job identity via 'source_job_id' but never explains what job_id refers to (a staging job handle?) or what dry_run does, leaving the default-true semantics completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (verify full byte count and SHA256, then finalize the staging job) and ties the result to a concrete consumer, transfer_start's source_job_id, which distinguishes it from generic staging siblings. It is clear but never contrasts itself with esxi_stage_upload, esxi_stage_write_chunk, or esxi_transfer_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: call this after chunks are staged so the job can be fed to transfer_start. There is no explicit 'when not to use', no prerequisite list (e.g. must all chunks be written first?), and no guidance on what to do if the SHA256 check fails.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_stage_uploadBDestructive
创建私有大文件暂存句柄;客户端随后write_chunk并finalize,不接收任意服务器本地路径。
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | ||
| filename | Yes | ||
| total_bytes | Yes | ||
| expected_sha256 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false. The description adds that it creates a private staging handle, that write_chunk and finalize follow, and that arbitrary server local paths are rejected. It does not explain handle lifetime, cleanup, or auth requirements, but it adds meaningful workflow and safety context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence: purpose first, then follow-up workflow, then an exclusion. It is dense but has no filler and every clause contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return value explanation is not needed. However, with 0% schema description coverage on four parameters and no explicit usage alternatives, the definition is incomplete for correct invocation. The agent lacks enough detail about how to call it properly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for four parameters (filename, total_bytes, dry_run, expected_sha256), so the description carries the full burden. It mentions none of them, leaving the agent unable to infer required formats, dry_run behavior, or hash semantics. This fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('创建私有大文件暂存句柄') and names the follow-up tools write_chunk and finalize, clarifying the workflow. It also constrains the input by saying no arbitrary server local paths are accepted. However, it does not explicitly contrast with sibling tools like esxi_transfer_start.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies this is the first step in a staged upload because it mentions subsequent write_chunk and finalize calls. It does not state when to prefer this over alternatives such as esxi_transfer_start or esxi_datastore_transfer, nor does it list prerequisites. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_stage_write_chunkCDestructive
追加<=1MiB块;相同offset/内容的重试幂等,拒绝空洞或覆盖不同内容。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| offset | Yes | ||
| dry_run | No | ||
| data_base64 | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description claims retries with the same offset and content are idempotent, directly contradicting the idempotentHint=false annotation. Although it adds useful safety constraints (rejects holes and overwriting different content), the explicit contradiction triggers a score of 1 per the rules.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence, front-loaded with the core action and size limit. Each clause adds a distinct constraint with no filler, making it highly economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists so return values need not be described, but for a 4-parameter write tool with 0% schema descriptions, the description omits dry_run (default true, meaning no actual write unless disabled) and job_id context. Workflow positioning relative to stage_upload/finalize is also absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all parameter meaning. It explains offset behavior (no holes, no differing content) and data size (<=1MiB), but omits job_id and dry_run entirely, leaving two of four parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action: '追加<=1MiB块' (append a chunk up to 1MiB). The core purpose is clear, but it does not name sibling tools like esxi_stage_upload or esxi_stage_finalize to distinguish the staging workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful invocation constraints: max 1MiB, idempotent retry on same offset/content, and rejection of holes or overwriting different content. However, it does not say when to use this tool versus stage_upload/finalize or other transfer alternatives, leaving the workflow context implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_storage_manageBDestructive
专用 storage 管理:rescan, rescan_vmfs, vmfs_format, partition_update, vmfs_mount, vmfs_unmount, lun_attach, lun_detach, iscsi_enable, iscsi_discovery, iscsi_authentication, iscsi_send_targets_add, iscsi_send_targets_remove, multipath_policy, nvme_connect, nvme_disconnect。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and non-idempotent. The description adds non-obvious behavior: a preview-first workflow, the need for expected_host, a write switch, dependency confirmation, and request_id for idempotency. It does not explain what each of the 16 actions actually does, which matters given a destructive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is compact but the opening is a 16-item action dump that crowds out the workflow guidance in the second sentence. Content is front-loaded enough to be usable yet the run-on enumeration hurts readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a destructive 9-parameter tool with 16 sub-actions the description should convey per-action impact and which parameters gate destructive operations. It covers the safety workflow but leaves the action-to-argument mapping and the impact flags under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, and the description leaves the action string with no enumeration of valid values even though the tool supports 16 named actions. It mentions expected_host, request_id, and the preview flag but does not explain dry_run semantics, allow_disruption, acknowledged_vm_ids, or the free-form arguments object, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this is dedicated storage management and enumerates the 16 supported actions (rescan, vmfs_format, lun_attach, iscsi_* etc.), which precisely identifies the verb family and resource. It does not explicitly contrast with the many siblings, but the enumerated action list separates it from esxi_datastore_manage and esxi_esxcli.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It advises running resource_schema/get and the default preview first, and requires expected_host, write switch, and dependency confirmation before execution. That gives real when-to-use context but never names when a sibling (e.g. esxi_datastore_manage) should be chosen instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_time_manageBDestructive
专用 time 管理:configuration_update, clock_update。先 resource_schema/get 和默认预览;实际执行精确 expected_host、写开关与依赖确认,支持 request_id 防止重复执行。
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| dry_run | No | ||
| arguments | No | ||
| request_id | No | ||
| expected_host | No | ||
| allow_disruption | No | ||
| acknowledged_vm_ids | No | ||
| allow_unknown_impact | No | ||
| allow_protected_impact | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, non-idempotent, and closed-world, so the safety bar is partly met. The description usefully adds that execution needs an exact expected_host, explicit write switches, dependency acknowledgment, and a request_id for de-duplication, but the four allow_* / acknowledged_vm_ids safety gates are never explained, leaving real destructive-side behavior opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense, front-loaded block that leads with the managed resource and operations before moving to prerequisites and execution requirements. It wastes no sentences, though the compressed phrasing borders on cryptic in places.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the destructive profile. But for a 9-parameter destructive mutation tool with zero schema documentation, the description omits the meaning of most safety switches and the accepted action values' semantics, leaving the agent under-equipped to call it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description carries the full burden. It only illuminates expected_host, request_id, and the presence of 'write switches' and 'dependency confirmation'; dry_run, arguments, and the four impact-permission flags are left entirely unmapped, so an agent cannot know what allow_disruption or allow_protected_impact actually gate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete resource ('time 管理') and enumerates the two operations it exposes (configuration_update, clock_update), so an agent can tell this is host time/NTP configuration rather than the sibling service_manage or host_manage tools. It stops short of explicitly contrasting itself with those neighbors, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real workflow guidance: consult resource_schema/get first, rely on the default preview, and only then execute with expected_host, write switches, and dependency confirmation. However, it never states when NOT to use the tool or how it differs from siblings that also mutate host settings, so usage is only partially specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_transfer_cancelADestructive
请求停止后台传输,保留部分数据;不会自动删除/覆盖远端文件。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as non-read-only and destructive. The description adds meaningful context beyond that: it retains partial data and will not automatically delete or overwrite remote files. It still omits details like whether the job can be resumed or what auth is needed, but the added data-handling behavior is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler. The key behavioral constraint is front-loaded after the primary action, and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with a required job_id, an undocumented dry_run parameter, and no usage guidance, the description is too thin. An output schema exists so return values need not be explained, but the agent still lacks parameter semantics and job-context details needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters, including the required job_id and the dry_run flag. The description never mentions either parameter, so it provides no meaning beyond the bare schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (停止) and resource (后台传输), with added scope about data retention and remote file safety. It clearly distinguishes this from sibling operations like transfer_start, transfer_status, and transfer_resume.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to stop a background transfer) but does not state when to prefer cancel over alternatives such as resume or letting it finish, nor any prerequisites. Implied usage only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_transfer_read_chunkARead-onlyIdempotent
读取已完成的下载/暂存文件;每块最多1MiB,用offset拼接并核对最终SHA256。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| length | No | ||
| offset | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description goes beyond them usefully by disclosing the 1MiB per-chunk ceiling and the intended concatenation-and-SHA256-verification workflow. It still omits what happens on an incomplete job, bad offset, or checksum mismatch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core read action front-loaded, followed by the chunk constraint and the workflow. Dense but every clause carries information; minor crypticness in the offset phrasing keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and annotations cover safety. The definition supplies the chunking contract an agent needs to reconstruct the file correctly; only edge-case/error behavior for incomplete or mismatched chunks is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It partially compensates by explaining offset's role in reassembling chunks and confirming the 1MiB block size (matching the length default), but job_id is never explained and the length/offset defaults are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: reading already-completed download/staged files as chunks. It also pins the chunk size (1MiB), which distinguishes it from the write-side sibling esxi_stage_write_chunk. It stops short of naming those siblings explicitly, so 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Saying the file must be '已完成' (already completed) implies the tool is only for post-download reads and hints at a start-then-read workflow, but no alternative tool is named and no when-not-to-use case (e.g. use esxi_stage_write_chunk for uploads) is given. Usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_transfer_resumeADestructive
显式恢复失败/取消/中断传输。下载验证ETag/Last-Modified与Range;上传重新发送,远端已有文件需明确overwrite。NFC需原租约仍活着。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| dry_run | No | ||
| overwrite | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint=true and idempotentHint=false; the description goes further by explaining the mechanism: downloads validate ETag/Last-Modified with Range, uploads re-send and require explicit overwrite when the remote file exists, and NFC needs the original lease alive. This is meaningful behavioral context beyond the structured hints, though it does not address auth or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, well-ordered sentences with the core purpose front-loaded, followed by download/upload specifics and the NFC constraint. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent resume operation with an output schema handling return values, the description covers the key behavioral distinctions (download vs upload, NFC lease, overwrite). The parameters job_id and dry_run remain unspecified, leaving a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it only clarifies overwrite ('remote existing file requires explicit overwrite'). job_id and dry_run are left unexplained, so compensation is partial at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resume) and resource (transfers) with qualifying states (failed/cancelled/interrupted). The purpose is clearly distinguishable from siblings like esxi_transfer_start or esxi_transfer_cancel, though no sibling is named explicitly to reinforce the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger condition for use (failed/cancelled/interrupted transfers) and differentiates download vs upload resume behavior. It stops short of naming alternatives (e.g., when to use transfer_start instead) or stating exclusions, so it does not reach a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_transfer_startADestructive
后台流式 datastore/guest/NFC 下载或上传(最多64TiB显式上限,受暂存磁盘空间约束)。立即返回 job_id,status 查询,read_chunk 分块取回。
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| path | No | ||
| lease | No | ||
| dry_run | No | ||
| datastore | No | ||
| max_bytes | No | ||
| operation | Yes | ||
| overwrite | No | ||
| request_id | No | ||
| source_url | No | ||
| http_method | No | PUT | |
| content_type | No | application/octet-stream | |
| transfer_url | No | ||
| source_job_id | No | ||
| expected_sha256 | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive=true, readOnly=false, and non-idempotent, so the description is not burdened with safety disclosure. It adds meaningful operational context: the transfer runs in the background, returns a job_id immediately, has a 64TiB upper bound, and is constrained by staging disk space.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded and free of filler. The first sentence states the core operation and constraints; the second states the return value and follow-up actions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations and an output schema cover safety and return structure, and the description adds async job semantics plus size/staging constraints. However, for a 15-parameter mutation tool with 0% schema description coverage, the definition does not provide enough parameter-level detail to guide invocation confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 15 parameters, so the description must compensate. It only hints at kind/operation through 'datastore/guest/NFC download or upload' and mentions a 64TiB limit, but leaves parameters like max_bytes, dry_run, lease, source_url, expected_sha256, and overwrite unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (background streaming download/upload), resource types (datastore/guest/NFC), and the immediate outcome (returns job_id). It distinguishes itself from follow-up siblings by naming status query and read_chunk retrieval, though it does not explicitly say 'start' or contrast with esxi_datastore_transfer / esxi_api_transfer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly implies the workflow: start here, then use status to query and read_chunk to retrieve data. This gives strong post-invocation guidance, but it does not state when to choose this tool over alternatives like esxi_datastore_transfer or esxi_api_transfer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_transfer_statusBRead-onlyIdempotent
传输/暂存进度、字节数、校验和、错误;不输出 signed URL 或服务器任意路径。
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, idempotentHint=true, destructiveHint=false), and the description adds useful behavioral context by specifying what is returned and explicitly excluding signed URLs or arbitrary server paths. This negative disclosure is valuable for an agent deciding whether this tool is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that lists the key returned information and then states a clear exclusion. It is compact and contains no filler, though its terse list format could be slightly more structured for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one required parameter) and has a rich annotation set plus an output schema, so the description need not explain return values. However, it omits any usage context and does not clarify the required job_id, leaving clear gaps for an agent to fill by inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the only parameter, job_id, is not mentioned or explained anywhere in the description. The description therefore does not compensate for the missing schema-level parameter documentation, leaving an agent to infer what job_id refers to and where it comes from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the resource ('transfer/staging') and enumerates what it returns (progress, byte counts, checksums, errors), making the tool's purpose clear. It does not explicitly differentiate itself from sibling status tools such as esxi_operation_status or esxi_guest_process_status, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no positive guidance on when to use this tool versus alternatives like esxi_transfer_start, esxi_transfer_read_chunk, or esxi_operation_status. The only boundary provided is a negative output constraint ('does not output signed URL or arbitrary server paths'), which is not enough to constitute usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_vm_hardwareCDestructive
增删/编辑虚拟磁盘、网卡、光驱、控制器、PCI 设备。device_changes 使用官方 VirtualDeviceSpec,保留其余 VM 配置。
| Name | Required | Description | Default |
|---|---|---|---|
| vm_id | Yes | ||
| dry_run | No | ||
| request_id | No | ||
| expected_name | Yes | ||
| device_changes | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the mutation risk is covered structurally. The description adds genuinely useful behavior ('uses official VirtualDeviceSpec', 'preserves the rest of the VM configuration' – i.e. a partial merge), but omits a critical trait: dry_run defaults to true, so changes are staged rather than applied unless explicitly disabled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the action and affected resources front-loaded, followed by the one behavioral caveat worth stating. No padding, though it could be slightly more explicit about the staging semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but this is a destructive mutation tool with 0% schema coverage and no annotation-explained staging behavior. The default-on dry_run and the expected_name guard are prerequisites an agent must know to call it correctly, and neither is stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, so the description must carry the burden. It only clarifies device_changes (VirtualDeviceSpec format) and says nothing about vm_id, expected_name (evidently a concurrency/identity guard), dry_run, or request_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names specific verbs (add/remove/edit) and specific resources (virtual disks, NICs, CD-ROM, controllers, PCI devices), so the agent knows exactly what the tool manipulates. It does not, however, differentiate itself from overlapping siblings like esxi_configure_vm, esxi_expand_disk, or esxi_pci_manage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. The agent is not told how this tool relates to esxi_configure_vm or esxi_expand_disk when modifying disks, nor that device_changes must be built before calling. Usage is only implied by the resource list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
esxi_vm_registerBDestructive
将已有 VMX 注册到该 ESXi 的 VM 文件夹与资源池;不创建或覆盖 VM 文件。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| dry_run | No | ||
| vmx_path | Yes | ||
| datastore | Yes | ||
| request_id | No | ||
| expected_host | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and destructiveHint=true, so the safety profile is covered. The description usefully scopes the mutation (inventory registration only, no file creation or overwrite), but says nothing about permission requirements, duplicate-name handling, or why the operation is flagged destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a clarifying clause; no padding or repetition. It is terse, arguably to the point of under-specification, but there is no wasted language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but for a destructive-flagged, non-idempotent mutation tool with 6 undocumented parameters, the description leaves too much unspecified — notably the dry_run preview default and path/identifier formats.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 6 parameters and 0% schema description coverage, the description must compensate and largely does not. The critical dry_run parameter (defaulting to true) is never mentioned, nor are vmx_path format, datastore, expected_host, or request_id semantics; only the general notion of a VMX path is implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: registering an existing VMX into the ESXi VM folder and resource pool. The clause '不创建或覆盖 VM 文件' implicitly separates it from esxi_create_vm/esxi_ovf_import, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '将已有 VMX' (an already-existing VMX) implies the precondition and distinguishes it from tools that create VM files, but there is no explicit when-to-use guidance or named alternative such as esxi_create_vm or esxi_ovf_import.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
63 tool updates
v0.3.0- First observed
esxi_account_manage - First observed
esxi_api_get - First observed
esxi_api_invoke - First observed
esxi_api_schema - First observed
esxi_api_transfer - First observed
esxi_capabilities - First observed
esxi_certificate_manage - First observed
esxi_configure_vm - First observed
esxi_create_snapshot - First observed
esxi_create_vm - First observed
esxi_datastore_manage - First observed
esxi_datastore_transfer - First observed
esxi_delete_snapshot - First observed
esxi_delete_vm - First observed
esxi_dependencies - First observed
esxi_esxcli - First observed
esxi_esxcli_query - First observed
esxi_execute_plan - First observed
esxi_expand_disk - First observed
esxi_firewall_manage - First observed
esxi_get_task - First observed
esxi_get_vm - First observed
esxi_guest_files - First observed
esxi_guest_process - First observed
esxi_guest_process_status - First observed
esxi_health - First observed
esxi_host_manage - First observed
esxi_host_summary - First observed
esxi_inventory - First observed
esxi_license_manage - First observed
esxi_list_datastores - First observed
esxi_list_networks - First observed
esxi_list_snapshots - First observed
esxi_list_vms - First observed
esxi_managers - First observed
esxi_network_manage - First observed
esxi_operation_history - First observed
esxi_operation_status - First observed
esxi_options_manage - First observed
esxi_ovf_export - First observed
esxi_ovf_import - First observed
esxi_patch_manage - First observed
esxi_pci_manage - First observed
esxi_permissions_manage - First observed
esxi_power_vm - First observed
esxi_rename_vm - First observed
esxi_resource_get - First observed
esxi_resource_schema - First observed
esxi_revert_snapshot - First observed
esxi_service_manage - First observed
esxi_shell - First observed
esxi_stage_finalize - First observed
esxi_stage_upload - First observed
esxi_stage_write_chunk - First observed
esxi_storage_manage - First observed
esxi_time_manage - First observed
esxi_transfer_cancel - First observed
esxi_transfer_read_chunk - First observed
esxi_transfer_resume - First observed
esxi_transfer_start - First observed
esxi_transfer_status - First observed
esxi_vm_hardware - First observed
esxi_vm_register
TDQS
Scored across 63 tools
Most tools are named for a distinct domain (firewall, license, time, snapshots), but there is real overlap between the generic catch-alls (esxi_api_get, esxi_api_invoke, esxi_resource_get) and the specialized *_manage tools, plus four or more file-transfer paths (datastore_transfer, api_transfer, transfer_start/stage_upload, guest_files upload_url/download_url). Descriptions provide guidance ('managers/inventory → schema/get → dry_run'), which helps, but an agent can still reasonably misselect between the generic and domain-specific layers.
All names share the esxi_ prefix and snake_case, but the verb/noun ordering is inconsistent: verb_noun (esxi_list_vms, esxi_create_snapshot), noun_verb (esxi_transfer_start, esxi_stage_upload), noun-only (esxi_inventory, esxi_capabilities, esxi_managers), and bare verbs (esxi_shell, esxi_esxcli). Readable overall, but the mixed conventions reduce predictability.
63 tools is far above the practical range for reliable tool selection, and several clusters (transfer, API access, plans) could be consolidated. The breadth of ESXi management justifies many tools, but the count still creates a heavy selection surface with substantial redundancy.
The surface covers VM lifecycle (create/configure/power/delete/register/hardware), snapshots, datastores, networks, storage, services, firewall, accounts, permissions, license, time, PCI, certificates, patching, guest processes/files, OVF import/export, and raw esxcli/shell access. Only niche operations like VM clone/migrate are explicitly excluded, so coverage is effectively complete for standalone ESXi management.
Maintenance
Related MCP Connectors
61 text, security, converter, calculator, and PDF tools -- callable via MCP on one host.
Securely control computers you explicitly pair through files, terminals, processes, screenshots, desktop UI/input, clipboard, browser automation, diagnostics, and document tools.
Official Sevalla MCP — full PaaS API access through just 2 tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables comprehensive management of Proxmox Virtual Environment through 35 tools for VMs, containers, storage, backups, snapshots, and cluster operations via the Proxmox VE API.16 npm22MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to manage VMware vSphere infrastructure through 55 typed tools built on the govc CLI. It supports comprehensive operations including VM lifecycle management, snapshot control, datastore navigation, and networking configuration.36 npm3MIT
- AlicenseAqualityDmaintenanceProvides tools to list, create, power on/off, and delete VMs on vCenter or ESXi hosts using natural language.5Apache 2.0
- FlicenseNot gradedqualityDmaintenanceEnables managing VMware Workstation Pro VMs via MCP tools, including power operations, snapshots, guest processes, and network configuration through the vmrest, vmrun, and vmcli interfaces.-