Enables AI agents to see, locate UI elements, and operate any Windows desktop app through natural language, using accessibility-tree matching with optional vision-model fallback, plus an autonomous visual loop with introspection and meta-learning.
Enables AI agents to operate Windows computers with parallel, human-like input streams, including held keys, concurrent mouse/keyboard actions, UIA/OCR targeting, and a knowledge base for engineering software.
Enables AI clients to see and control a Windows desktop through screenshots, built-in OCR, mouse and keyboard input, and accessibility trees read via UI Automation, MSAA and Chrome DevTools Protocol. It also manages windows, virtual desktops and element-level actions such as invoking controls or setting values, running on Node.js with the PowerShell that ships with Windows.
Enables AI agents to control Windows GUI applications like a human using screen capture, OCR, mouse and keyboard input, and window management, with safety levels and memory.