esxi-readonly-mcp
Provides read-only diagnostic tools for VMware ESXi hosts, enabling inspection of host hardware health, CPU/memory performance, datastores, VMs, snapshots, events, overcommit, and datastore files without exposing any write operations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@esxi-readonly-mcpHacé un diagnóstico completo del host y decime cuál es el cuello de botella"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🔍 esxi-readonly-mcp
Diagnóstico completo de VMware ESXi desde tu asistente de IA, sin poder romper nada.
Un servidor MCP que le da a Claude (o a cualquier cliente MCP) visibilidad total sobre tu host ESXi — CPU Ready, latencia de disco, snapshots, espacio real, salud del hardware, eventos — sin una sola herramienta de escritura.
¿Por qué?
Tenés un ESXi standalone (sin vCenter), el ERP anda lento y alguien propone gastar miles de dólares en hardware. Antes de comprar, necesitás saber qué recurso es el cuello de botella de verdad. Eso implica mirar métricas que el Host Client esconde o no guarda (CPU Ready, Co-Stop, latencia por disco virtual), cruzar el espacio provisionado contra el real y revisar snapshots, logs y sensores.
Con este MCP le preguntás a Claude "¿por qué está lenta la VM del ERP?" y él mismo consulta el host, cruza los datos y te responde con números.
Y como vas a darle acceso a un asistente de IA a tu infraestructura de producción, el servidor no expone ninguna operación que modifique algo: no puede apagar, borrar, crear, consolidar ni reconfigurar nada. Combinado con un usuario de ESXi con rol de solo lectura, la garantía es doble.
Related MCP server: MCP SSH SRE
✨ Qué podés preguntarle
"Hacé un diagnóstico completo del host y decime cuál es el cuello de botella."
"¿Qué VM tiene más CPU Ready en la última hora?"
"¿Cuánto ocupa realmente cada VM en cada datastore?"
"¿Hay snapshots de más de 3 días?"
"¿Qué archivos grandes hay en los datastores que no pertenecen a ninguna VM?"
"¿Algún sensor de hardware en amarillo o rojo? ¿Hubo errores de disco esta semana?"
"¿Tengo overcommit de CPU? ¿Cuántas vCPU tengo por core físico?"
"¿Conviene más comprar RAM o cambiar el procesador?"
🧰 Herramientas
Herramienta | Qué devuelve |
| Todo lo de abajo en una sola llamada — el punto de partida ideal |
| Fabricante, modelo, serie, BIOS, CPU (cores / hilos / HT), RAM total y usada, versión y build de ESXi, uptime, política de energía, controladoras de storage |
| Sensores (temperaturas, ventiladores, fuentes, voltajes), memoria, CPU y estado de RAID/discos físicos si el host tiene el proveedor CIM del fabricante. Lista primero lo que no está en verde |
| Licencia (clave enmascarada) y vencimiento, NTP y desfase real del reloj, SSH/Shell, servicios activos, perfil de imagen, syslog, placas de red con velocidad de enlace, vSwitches, portgroups/VLANs, VMkernel |
| Capacidad, usado, libre, % de uso, % provisionado (sobre-asignación thin), disco físico detrás de cada datastore y VMs que lo usan |
| Por cada VM: estado, vCPU, RAM asignada / activa / consumida, ballooning, swap, reservas, límites, shares, VMware Tools, heartbeat, sincronización horaria, NICs (e1000 vs vmxnet3), controladoras, discos (thin/thick, usado real), espacio libre dentro del guest |
| Todos los snapshots con fecha, antigüedad, tamaño real y marca de los que superan N días |
| CPU %, CPU Ready %, Co-Stop %, latencia de CPU, memoria activa, ballooning, swap, latencia por disco virtual y por datastore, IOPS. Devuelve promedio, p95 y máximo con el momento en que ocurrió |
| Las mismas métricas sobre semanas, filtrables por horario laboral (ver Historial) |
| Errores, advertencias, alarmas, tareas y logins fallidos, con la hora corregida por el desfase del reloj de ESXi |
| Los archivos más grandes de cada datastore, a qué VM pertenecen y posibles huérfanos (vmdk sin VM registrada) |
| vCPU encendidas vs hilos físicos, RAM asignada vs física, VM más grande |
| Con qué usuario se conecta y si ese usuario tiene privilegios de escritura en ESXi |
Todas las respuestas vienen con valores de referencia (por ejemplo: CPU Ready > 10 % = contención) para que el asistente pueda interpretarlas sin adivinar.
🔒 Seguridad: cómo se garantiza que es solo lectura
1. En el código. El servidor solo lee propiedades y llama a estos métodos de la API de vSphere, todos de consulta:
Método | Para qué |
| Vista temporal de la propia sesión para listar objetos |
| Métricas de performance |
| Listar archivos de los datastores |
| Leer eventos (el colector es de la sesión y se destruye al terminar) |
| Verificar los permisos del propio usuario |
| Leer opciones y perfil de imagen |
No hay ninguna herramienta MCP para modificar nada, y el servidor le indica al asistente que no existen.
2. En ESXi. Usá un usuario dedicado con un rol de solo lectura (paso 1 de la instalación). Aunque algo
intentara escribir, ESXi lo rechazaría. verificar_permisos te confirma que quedó bien configurado:
{ "usuario": "mcp-readonly", "roles": ["ReadOnly+Browse"],
"privilegios_escritura_detectados": [], "solo_lectura_garantizado_por_ESXi": true }3. Credenciales. La contraseña nunca va en archivos de configuración: se guarda en el almacén de credenciales
del sistema operativo (Administrador de credenciales de Windows, Keychain en macOS, Secret Service en Linux)
mediante keyring. Las claves de licencia se muestran enmascaradas.
🚀 Instalación
Requisitos
uv(se encarga de Python y de las dependencias)Acceso por red al puerto 443 del ESXi
Un cliente MCP: Claude Code, Claude Desktop u otro
1. Crear un usuario de solo lectura en ESXi
En el Host Client (https://<ip-esxi>/ui):
Manage → Security & users → Roles → Add role — nombre
ReadOnly+Browse:✅ System (Anonymous, Read, View)
✅ Datastore → Browse datastore — solo ese; nada de Delete, FileManagement ni AllocateSpace
Manage → Security & users → Users → Add user — por ejemplo
mcp-readonly. Dejá Enable shell access desmarcado.Host → Actions → Permissions → Add user — elegí
mcp-readonly, el rolReadOnly+Browsey marcá Propagate to all children.
Sin Browse datastore el MCP funciona igual, pero
top_archivossolo ve archivos de VMs registradas (no ISOs ni huérfanos) y el espacio "usado" de discos thin es aproximado.
2. Guardar la contraseña en el almacén del sistema
uvx esxi-readonly-mcp --set-passwordTe pide el usuario de ESXi y la contraseña (sin mostrarla) y la guarda en el almacén de credenciales del sistema.
No hace falta clonar nada: uvx descarga el paquete de PyPI y lo ejecuta.
3. Registrar el servidor en tu cliente MCP
Claude Code
claude mcp add esxi-readonly --scope user \
-e ESXI_HOST=192.0.2.10 -e ESXI_USER=mcp-readonly \
-- uvx esxi-readonly-mcpClaude Desktop u otro cliente (en claude_desktop_config.json o equivalente):
{
"mcpServers": {
"esxi-readonly": {
"command": "uvx",
"args": ["esxi-readonly-mcp"],
"env": { "ESXI_HOST": "192.0.2.10", "ESXI_USER": "mcp-readonly" }
}
}
}En Windows usá la ruta completa a
uvx.exesi el cliente no lo encuentra en elPATH(por ejemploC:\\Users\\<usuario>\\.local\\bin\\uvx.exe).
Reiniciá el cliente y pedile: "verificá los permisos del MCP de ESXi".
Desde el código fuente
git clone https://github.com/nicolasboattini/esxi-readonly-mcp.git
cd esxi-readonly-mcp
uv sync
uv run esxi-readonly-mcp --set-passwordY en el cliente MCP usá uv run --directory /ruta/a/esxi-readonly-mcp esxi-readonly-mcp como comando.
Probar sin cliente MCP
ESXI_HOST=192.0.2.10 ESXI_USER=mcp-readonly uv run python -c "from esxi_readonly_mcp import server; import json; print(json.dumps(server._host_info(), indent=2))"(En PowerShell: $env:ESXI_HOST="192.0.2.10"; $env:ESXI_USER="mcp-readonly"; uv run python -c "...")
⚙️ Configuración
Variable | Obligatoria | Default | Descripción |
| ✅ | — | IP o nombre del ESXi (o del vCenter) |
| ✅ | — | Usuario de solo lectura |
| — | Solo si no podés usar | |
|
| Puerto de la API | |
|
|
| |
|
| Ruta de la base SQLite del historial de performance |
📈 Historial de performance
Un ESXi sin vCenter guarda solo la última hora de métricas (muestras cada 20 s). Para ver los picos reales de las últimas semanas en horario laboral, el servidor trae un modo recolector que guarda esa hora en una base SQLite local:
uvx esxi-readonly-mcp --collectProgramalo cada hora y después pedile a Claude, por ejemplo: "mostrame el historial de performance de las
últimas 4 semanas en horario laboral" (historial_perf(dias=28, solo_horario_laboral=True)).
Windows (Programador de tareas)
[Environment]::SetEnvironmentVariable("ESXI_HOST", "192.0.2.10", "User")
[Environment]::SetEnvironmentVariable("ESXI_USER", "mcp-readonly", "User")
$uvx = (Get-Command uvx).Source
$a = New-ScheduledTaskAction -Execute $uvx -Argument 'esxi-readonly-mcp --collect'
$t = New-ScheduledTaskTrigger -Daily -At 7:05am
$t.Repetition = (New-ScheduledTaskTrigger -Once -At 7:05am -RepetitionInterval (New-TimeSpan -Hours 1) -RepetitionDuration (New-TimeSpan -Hours 13)).Repetition
Register-ScheduledTask -TaskName "ESXi perf collect" -Action $a -Trigger $t -User $env:USERNAMELinux / macOS (cron)
5 7-20 * * 1-5 ESXI_HOST=192.0.2.10 ESXI_USER=mcp-readonly $HOME/.local/bin/uvx esxi-readonly-mcp --collectLa base contiene nombres de VMs y métricas. Por defecto vive en tu carpeta de usuario, fuera de cualquier repo.
🧪 Compatibilidad
Estado | |
ESXi 6.7 U3 standalone | ✅ Probado en producción |
ESXi 7.x / 8.x standalone | Debería funcionar (misma API); reportá cualquier problema |
vCenter | Soportado en el código (usa |
pyVmomi 9.x + mcp 1.x | ✅ Probado ( |
⚠️ Limitaciones conocidas
Métricas históricas: sin vCenter,
performancesolo ve ~1 hora. Usá--collect+historial_perf.Eventos: ESXi standalone guarda los últimos ~1000 eventos. Si un sistema de monitoreo abre sesiones constantemente (por ejemplo, Zabbix mal configurado), eso puede cubrir solo unas horas.
RAID y discos físicos:
salud_hardwarelos muestra solo si ESXi tiene el proveedor CIM del fabricante (Dell, HPE, Lenovo). Si no, revisalos en iDRAC / iLO / XCC.Lista de VIBs/drivers: requiere
Host.Config.Image, que también permite modificar la imagen; el MCP no lo pide. Alternativa:esxcli software vib listpor SSH.Dentro del guest: el MCP ve lo que reporta VMware Tools (unidades, espacio libre, IP), pero no archivos ni procesos dentro de la VM.
🛠️ Solución de problemas
Síntoma | Causa probable |
| Falta |
| Usuario o contraseña incorrectos, o el usuario no tiene permiso asignado en el host |
| Al rol le falta Datastore → Browse datastore |
| Normal en ESXi sin vCenter: usá |
| Se instaló |
El cliente no encuentra | Poné la ruta absoluta a |
Horas raras en eventos | El reloj de ESXi está desfasado (sin NTP); |
📁 Estructura
esxi-readonly-mcp/
├── src/esxi_readonly_mcp/
│ └── server.py # servidor MCP + --collect + --set-password
├── pyproject.toml # paquete PyPI (mcp<2, pyvmomi, keyring)
├── server.json # ficha para el registro oficial de MCP
├── uv.lock
└── README.md🤝 Contribuir
Issues y pull requests bienvenidos, sobre todo:
Pruebas en ESXi 7/8 y vCenter
Nuevas métricas o chequeos de solo lectura
La regla del proyecto es una sola: ninguna herramienta puede modificar el entorno. Un PR que agregue operaciones de escritura no se va a aceptar, aunque sea "opcional".
📄 Licencia
Available Tools
13 toolsconfig_hostC
Configuracion del host: licencia (clave enmascarada) y vencimiento, NTP y desfase del reloj, SSH/Shell, servicios activos, perfil de imagen y VIBs/drivers de terceros, syslog, placas de red (velocidad de enlace), vSwitches, portgroups y VMkernel.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses none. It does not state that this is a read-only operation, whether any permissions are required, or whether the listed items are always returned. Listing content categories is not the same as disclosing behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded sentence with the resource named first, which is good structure. However, the long comma-separated enumeration becomes dense to parse and includes low-level detail (link speed, masked key) that adds bulk without direction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does useful work by listing the configuration domains returned, which partially compensates. But it omits the verb, any usage context, and any behavioral or permission details, leaving the definition only minimally viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. The description correctly requires no argument explanation; the metadata is trivially complete on the input side.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates what the tool covers (license, NTP, SSH, services, VIBs, syslog, NICs, vSwitches, portgroups, VMkernel) under the label 'Configuracion del host', so the resource is clear. But it never states a verb or the action ('retrieve/report configuration'), leaving the actual operation inferred. It also does not differentiate itself from close siblings like host_info or salud_hardware.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. Sibling tools such as host_info, salud_hardware, and diagnostico_completo overlap heavily, yet none is referenced or excluded. The agent must guess the retrieval context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
datastoresB
Datastores: capacidad, usado, libre, % de uso, % provisionado (thin), disco fisico donde esta cada uno y VMs.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. Listing read-only metrics (capacity, used, free, %) implicitly signals a non-mutating query, but it never states read-only behavior, permission requirements, or response shape explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, densely packed sentence with the resource named first and the returned metrics listed after. No filler, though the run-on comma list is slightly harder to scan than a structured breakdown.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the field enumeration usefully tells the agent what comes back, which is the right compensation. However, with no annotations either, the description omits any behavioral or scoping context (read-only nature, coverage across hosts/clusters), leaving a gap for a monitoring tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate at the parameter level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (datastores) and enumerates the exact data it surfaces — capacity, used, free, % usage, thin provisioning %, physical disk, and VMs. That field list clearly differentiates it from siblings like vms, snapshots, and host_info, though no explicit verb (list/retrieve) is stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of any alternative or sibling tool. The agent must infer that this is the tool for datastore capacity questions purely from the field list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnostico_completoB
Ejecuta todo junto: permisos, host, salud de hardware, configuracion, datastores, VMs, snapshots, performance realtime, eventos (warning/error 7 dias), top archivos y overcommit.
| Name | Required | Description | Default |
|---|---|---|---|
| dias_alerta_snapshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It usefully discloses scope constants such as the 7-day event window and, via the schema, the snapshot alert window, but it says nothing about the read-only nature, the cost/latency of running everything at once, or the failure behavior of a failed sub-check – which matters greatly for an all-in-one aggregator.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the verb and then the covered modules. Every item earns its place as the tool's scope, though the bare enumeration could have been trimmed in favor of one clarifying clause about intended use.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should carry more, and it partially does by listing coverage. However it leaves the sole parameter unexplained and gives no sense of return structure or runtime cost, which an agent needs before choosing this broad tool over a targeted sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (dias_alerta_snapshot, default 3) with 0% schema description coverage, and the description never mentions or explains it. The 0-parameter baseline of 4 does not apply since a parameter exists and is undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Ejecuta') and enumerates the exact resources covered (permisos, host, hardware, datastores, VMs, snapshots, performance, eventos, top archivos, overcommit), making it unmistakably the aggregate of the sibling tools. It does not explicitly name itself as the 'run-all' alternative to those siblings, but the enumeration does the differentiating work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance: it never states when to prefer this aggregate over the individual siblings (e.g. first-pass triage vs. targeted follow-up), nor any prerequisites or exclusions. The implication that it is a full sweep is present but left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
eventosB
Eventos de ESXi (errores de disco, reinicios de VMs, alarmas, tareas, logins fallidos). nivel: todos | warning (warning y error) | error. Fechas corregidas por el desfase del reloj de ESXi.
| Name | Required | Description | Default |
|---|---|---|---|
| dias | No | ||
| nivel | No | todos | |
| max_eventos | No | ||
| incluir_sesiones | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does add real value by disclosing that dates are corrected for ESXi clock skew and that events span disk errors, reboots, alarms, tasks, and failed logins. It does not state that this is a read-only query, nor anything about return volume or pagination. Useful context, but not enough for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight paragraphs, front-loaded with the resource definition before the filtering detail. No wasted words. The parenthetical examples and the clock-skew note both earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter read query with no annotations and no output schema, the description should explain the time window ('dias'), the result cap ('max_eventos'), and the session-inclusion flag ('incluir_sesiones'). It covers the event taxonomy and clock-skew correction but leaves those three defaults uninterpretable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all four parameters. It partially does: 'nivel' values (todos | warning | error) are explained, including that warning includes errors. But 'dias', 'max_eventos', and 'incluir_sesiones' are never clarified, leaving three of four parameters undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (ESXi events) and enumerates the kinds of events returned (disk errors, VM reboots, alarms, tasks, failed logins), which makes the tool's scope concrete. It does not, however, distinguish itself from monitoring siblings like salud_hardware or diagnostico_completo. Purpose is clear but sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus the many other diagnostic siblings. The only usage-adjacent content is the enumeration of 'nivel' values, which is filtering semantics rather than routing guidance. An agent receives no explicit when-to-use or when-not-to-use context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
historial_perfC
Resumen de las metricas guardadas localmente por 'server.py --collect' (semanas de historia, aunque ESXi solo guarde 1 h). Filtra por horario laboral (lunes a viernes, hora local).
| Name | Required | Description | Default |
|---|---|---|---|
| dias | No | ||
| entidad | No | ||
| hora_fin | No | ||
| hora_inicio | No | ||
| solo_horario_laboral | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load. It adds genuinely useful behavior: metrics are local, retention spans weeks while ESXi keeps only 1 h, and by default results are filtered to weekday working hours in local time. It still omits read-only confirmation, return format, and any auth/execution requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler; the data-source framing comes first and the filtering semantics second. The parenthetical about ESXi's 1 h retention is a slight detour but still earns its place as context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, and five parameters at 0% schema coverage. The description covers the data source and time-filtering concept but leaves parameter meaning, defaults, the 'entidad' selector, and the shape of the returned summary entirely unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implicitly maps to several params (history span -> dias, working-hours filtering -> solo_horario_laboral/hora_inicio/hora_fin) but never explicitly names them, gives no units or defaults, and completely ignores 'entidad'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
'Resumen de las metricas guardadas localmente por server.py --collect' identifies the resource (locally stored performance metrics) and the data source, but 'Resumen' is a weak verb and the definition never names or contrasts with the sibling 'performance' tool, leaving the historical-vs-live distinction to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. The mention that ESXi only retains 1 h versus weeks locally implies this tool is for longer/historical windows, but the agent must infer that rather than being routed to it over 'performance'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
host_infoB
Host: fabricante, modelo, CPU (cores/hilos/HT), RAM total y usada, version y build de ESXi, uptime, politica de energia y controladoras de storage.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It lists return fields but does not explicitly state that this is a read-only operation, whether authentication or special permissions are required, or any side effects or rate limits. Only the content of the response is disclosed, not the operation's safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single colon-led list of fields, with no wasted words. It is front-loaded with the resource name 'Host:' and then enumerates details efficiently, though the list format is dense rather than structured for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description is the main source of truth. It does list the fields returned, which partially compensates for the missing output schema, but it omits any behavioral context such as read-only status, permissions, or output format. For a zero-parameter info tool, this is minimally adequate but leaves clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so parameter semantics are not applicable and the baseline score of 4 applies per the rules. The description correctly adds no parameter information because none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description enumerates exactly what host information is returned (fabricante, modelo, CPU cores/hilos/HT, RAM total/usada, versión/build ESXi, uptime, política de energía, controladoras de storage), which clearly identifies the resource and scope. It does not explicitly differentiate from siblings like salud_hardware or config_host, but the field list distinguishes it as an inventory-style tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as salud_hardware, config_host, or performance. The description only says what data is available, leaving the agent to infer usage without any exclusions or preferred contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
overcommitB
Suma de vCPU vs cores/hilos fisicos y RAM asignada vs fisica (ratios de overcommit).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It implies a read-only aggregation of capacity data, which is intrinsically low risk, but it does not say what is returned, whether ratios are per-host or per-cluster, or how zero-physical-resource cases are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core comparison front-loaded and no filler. It is appropriately sized for a zero-argument aggregation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description should ideally state what the ratios look like or what host scope is analyzed. It conveys the concept adequately but leaves the return shape and scope unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty, so there is nothing for the description to disambiguate. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific computation: summing vCPU against physical cores/threads and assigned RAM against physical RAM to derive overcommit ratios. That is a clear verb+resource. It does not distinguish itself from siblings like host_info or performance, but those siblings cover different domains so ambiguity is low.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus host_info, performance, or diagnostico_completo, and no prerequisites or exclusions are stated. The agent must infer usage purely from the topic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
performanceA
Metricas de host y VMs: CPU uso, CPU Ready %, Co-Stop %, memoria activa/ballooning/swap, latencia de disco (total, por disco virtual y por datastore) e IOPS. Devuelve promedio, p95 y maximo con su momento. intervalo: realtime (20 s, ~1 h en ESXi standalone), dia, semana, mes, anio (estos requieren vCenter). max_muestras: 180 = 1 h en realtime. dias: rango hacia atras (alternativa a max_muestras).
| Name | Required | Description | Default |
|---|---|---|---|
| vm | No | ||
| dias | No | ||
| hora_fin | No | ||
| intervalo | No | realtime | |
| hora_inicio | No | ||
| max_muestras | No | ||
| solo_horario_laboral | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does a decent job: it discloses the return shape (promedio, p95 y maximo con su momento) and an infrastructure prerequisite (intervalos largos requieren vCenter). It still omits the time-zone/units context of hora_inicio/hora_fin and any rate-limit or cost notes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The metric enumeration is front-loaded, and the following lines are dense with parameter-level facts rather than filler. It is slightly long but every clause conveys new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, annotation-free tool with no output schema, the description covers the metric set and three parameters well but leaves the VM selector, hour window, and labor-hour flag unaddressed, and does not clarify the relationship to historial_perf.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all 7 parameters. It explains intervalo (with allowed values), max_muestras, and dias (as an alternative), but vm, hora_inicio, hora_fin, and solo_horario_laboral are never explained beyond the schema's bare titles and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (host and VM metrics) and enumerates the exact metric families returned (CPU uso, CPU Ready, Co-Stop, memoria activa/ballooning/swap, latencia de disco, IOPS), which is far more than a tautology. It does not, however, differentiate itself from the sibling historial_perf, which plausibly covers similar ground.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives real operational guidance at the parameter level: intervalo values, the fact that day/week/month/year require vCenter, that max_muestras=180 equals 1 h in realtime, and that dias is an alternative to max_muestras. But there is no explicit when-to-use-this-vs-alternatives statement, and the overlap with historial_perf is left unresolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
salud_hardwareA
Salud del hardware segun ESXi: estado de la RAID y discos fisicos (PERC), sensores de temperatura, ventiladores, fuentes, memoria y CPU. Lista primero los elementos que no estan en verde.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does usefully disclose the output prioritization behavior ('lists items not in green first'), which is real added context, but it never states that the call is read-only, nor whether it requires vCenter/ESXi credentials or how long a full sensor sweep takes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the scope and followed immediately by the most actionable behavior (non-green items first). No filler or repetition; the embedded line breaks are cosmetic only.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should do more to convey what comes back — e.g. whether each component returns a green/yellow/red verdict, thresholds, or error behavior when ESXi is unreachable. It covers scope and ordering but leaves the response shape largely unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter guidance, and none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (ESXi hardware health) and enumerates the exact subsystems inspected: RAID/PERC disks, temperature sensors, fans, power supplies, memory and CPU. This is clearly distinguishable from siblings like host_info (configuration) and performance (metrics), though no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the resource — an agent can infer this is a read-only hardware diagnostic to run when investigating physical-layer issues. There is no explicit statement of when to use this versus host_info, performance or diagnostico_completo, and no prerequisites or exclusions given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
snapshotsA
Todos los snapshots con fecha, antiguedad, tamano aproximado y marca de los que superan dias_alerta.
| Name | Required | Description | Default |
|---|---|---|---|
| dias_alerta | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the returned content (date, age, approximate size, alert flag), which is genuinely useful, but it never states that this is a read-only listing, whether permissions are required, or how pagination/volume is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the resource and then the returned fields, with zero redundant or filler text. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter listing tool with no output schema and no annotations, the description covers both the returned shape and the parameter's role. It is essentially complete, only missing explicit read-only/permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (the single parameter has only a default of 3 and no description), so the description must compensate. It does: dias_alerta is explained as the threshold above which a snapshot is marked, giving the parameter concrete meaning beyond the bare integer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (snapshots) and enumerates the fields it returns: date, age, approximate size, and an alert mark. That is far more than a tautology, so an agent knows exactly what it gets. It does not need sibling differentiation since no other sibling tool deals with snapshots.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the tool exists to inventory snapshots and flag the ones older than dias_alerta. There is no explicit when-to-use statement, no prerequisite, and no named alternative, but the alert-mark semantics hint at the diagnostic context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
top_archivosB
Archivos mas grandes de cada datastore (o de uno), con la VM a la que pertenecen y marca de posibles huerfanos.
| Name | Required | Description | Default |
|---|---|---|---|
| top | No | ||
| datastore | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose output content (VM ownership, orphan flagging), which is useful behavioral context, but says nothing about read-only status, required permissions, cost, or ordering — reasonable inference is a read tool, but it is never stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the resource and enumerates the appended fields; no filler. Slightly parenthetical-heavy but efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param, no-annotation, no-output-schema tool the description covers what is returned but omits the default behavior (all datastores, top 15), the meaning of 'top', and any read-only confirmation. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. '(o de uno)' explains that 'datastore' optionally narrows to a single datastore, and the 'largest files' framing hints at rank-based ordering, but the 'top' parameter is never explained as a count and the default of 15 is not mentioned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear resource (largest files) with scope (each datastore, or one) and the extra data returned (owning VM, orphan marking). An agent can tell what it produces, though it never distinguishes itself from the sibling 'datastores' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is given, and no alternative is named despite the closely related 'datastores' sibling. The optional datastore filter is implied by '(o de uno)' but not framed as a usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verificar_permisosA
Muestra con que usuario se conecta y si ese usuario tiene privilegios de escritura en ESXi.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the two pieces of information returned (connected user and write-privilege status), which implies a non-mutating check, but it says nothing about required credentials, whether it makes any changes, or how the privilege check is performed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the dual output (user identity and write privileges). Nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey the return content, and it does so for both returned facts. It falls short only on operational context such as side effects or failure behavior, which is a minor gap for a zero-parameter read/check tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing additional the description could clarify about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it shows which user the connection is authenticated as and whether that user has write privileges on ESXi. That is concrete and distinguishable from data-listing siblings such as host_info, datastores or config_host, though it never explicitly names a sibling to differentiate itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this tool versus the many diagnostic siblings (salud_hardware, diagnostico_completo, config_host). The use case of validating access before attempting writes is only weakly implied by the word 'permisos', with no explicit when/when-not statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vmsC
Todas las VMs: estado, vCPU, RAM asignada/activa/consumida, ballooning, swap, reservas/limites, VMware Tools, discos (provisionado vs usado, thin/thick, datastore, controladora) y espacio libre dentro del guest.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it discloses nothing beyond return content: it does not state that it is read-only, whether it requires authentication, whether it queries live or cached data, or its cost/latency. It is a data enumeration, not a behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single sentence with no obvious filler, and the resource scope is front-loaded. But the dense mid-sentence enumeration of a dozen metrics reads as a data dump rather than structured, scannable guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, the description usefully lists the fields the tool returns (state, vCPU, RAM, ballooning, disks, guest free space), which is genuinely valuable for a no-arg read tool. Still, it omits any usage context and the read-only safety profile an agent would want before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to explain and the baseline of 4 applies. The description correctly implies a global, unfiltered scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('Todas las VMs') and enumerates the attributes it reports, so an agent can infer this returns a detailed VM inventory. However, it never states the action (list/get/report) and offers no explicit differentiation from siblings like host_info or performance, leaving the purpose implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no reference to alternatives such as performance or datastores. The agent is given no signal for choosing this tool over its many siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.1.0- First observed
config_host - First observed
datastores - First observed
diagnostico_completo - First observed
eventos - First observed
historial_perf - First observed
host_info - First observed
overcommit - First observed
performance - First observed
salud_hardware - First observed
snapshots - First observed
top_archivos - First observed
verificar_permisos - First observed
vms
TDQS
Scored across 13 tools
Each tool targets a distinct ESXi monitoring area: permissions, host hardware, datastores, VMs, snapshots, performance, historical performance, largest files, overcommit, hardware health, events, host config, and a full diagnostic aggregator. Overlaps are minor and well explained, especially performance vs. historial_perf and the all-in-one diagnostico_completo. An agent can reliably select the right tool for a specific monitoring question.
All names use lowercase snake_case, but they mix English and Spanish terms (e.g., host_info, datastores, eventos, historial_perf) and mix noun phrases with one verb phrase (verificar_permisos). The pattern is readable but not predictable or consistent in language or grammatical style. A more uniform convention would improve clarity.
With 13 tools, the server is well-scoped for read-only ESXi diagnostics. Each tool covers a meaningful monitoring concern, and the aggregate diagnostico_completo tool provides a convenient full-scan option without bloating the set. This is within the ideal 3-15 range for an MCP server.
Coverage is broad: host, hardware health, config, datastores, VMs, snapshots, performance, events, overcommit, top files, and permissions are all represented. Minor gaps exist, such as explicit per-VM network details and local user/role enumeration beyond the current connection. These gaps are workable for most read-only diagnostic workflows.
Maintenance
Related MCP Connectors
Read-only local AI advice, shared reports and website audits. No PC scan or local actions.
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
- HAVNOAuthapp.havnre
Read-only AI access to HAVN properties, leads, tasks, files, media, and analytics.
Assess AI-discovery readiness, plan visibility fixes, and summarize scan evidence. Read-only.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to manage VMware vSphere virtual infrastructure through comprehensive operations including VM power control, snapshot management, resource monitoring, performance analytics, and bulk operations with built-in safety confirmations for destructive actions.-
- AlicenseNot gradedqualityFmaintenanceProvides read-only server monitoring and diagnostic tools for AI assistants to manage Linux and Unraid systems via SSH. It enables natural language interactions for container management, storage health checks, and system log analysis while keeping credentials secure.17ISC
- AlicenseAqualityAmaintenanceRead-only VMware vCenter/ESXi monitoring. 8 MCP tools for VM inventory, host status, datastore capacity, cluster info, alarms, events, and VM details. Code-level enforced safety — no destructive operations exist in the codebase. Supports vSphere 6.5–8.0. Works with local models via Ollama/LM Studio.3212MIT
- AlicenseNot gradedqualityBmaintenanceEnables LLMs and AI agents to perform defensive security posture assessments, privilege escalation surface audits, and post-quantum cryptography readiness checks through read-only diagnostic tools.MIT