An MCP server that enables AI agents to query real-time NVIDIA GPU metrics like utilization, memory, temperature, and power without external monitoring tools.
Enables AI agents to monitor server health and capacity, diagnose outages, inspect Docker and deployment status, and perform safe, bounded recovery actions via MCP without granting unrestricted shell or SSH access.
Enables listing, inspecting, chatting with, and waking deployed Voight Agents from any MCP client, including checking agent state, usage, tasks, and GPU status.
Enables MCP clients to scan local GGUF models, estimate VRAM and suggest GPU offload layers, manage llama-server lifecycle, and proxy OpenAI-format chats with idle auto-unload.
Enables AI assistants to query system health, metrics, service reconnaissance, run multi-node patrols, and send incident alerts across distributed infrastructure through MCP tools.