Gives AI agents a compact, semantic interface to the browser, returning structured page snapshots with stable element IDs instead of raw DOM. Enables agents to navigate, interact, and extract information from web pages efficiently.
Enables AI agents to navigate and interact with web pages through compressed semantic graphs, using element IDs and coordinates for clicking, typing, selecting, scrolling, and screenshots without raw HTML or CSS selectors.
Enables AI agents to navigate the web visually using screenshot-based interaction and Set-of-Mark labeling for interactive elements. It supports humanized browsing behaviors, anti-detection measures, and complex tasks like multi-click CAPTCHA solving.