Enables AI assistants to autonomously control Windows 11 and 10 desktops via sub-10ms screen capture, native UI Automation element inspection, and zero-lag keyboard and mouse input. Combines a visual plane with a semantic plane and stall detection so agents can operate real applications reliably without vision-only guessing.
Enables AI agents to interact with the Windows desktop environment, including browser control, clipboard, file management, GitHub, Roblox Studio, OCR, and more, with a privileged approval system for risky actions.
Enables AI agents to capture multi-monitor screenshots with coordinate grids and perform mouse, keyboard, scrolling, dragging, and window-management actions on Windows via MCP tools or CLI. It provides pixel-accurate desktop automation for computer-use agents.