observe_page
Analyze any web page in one request to return interactive elements, page type, and suggested actions for AI agents, with optional content, ARIA tree, or screenshot.
Instructions
Get a compact, token-budgeted "observation" of any web page, purpose-built for AI agents. In ONE request it returns: id-indexed interactive elements (role, name, CSS selector, state), a heuristic page-type classification (login, signup, search, article, form, generic), and grouped "suggested actions" (login flow, search, primary buttons, navigation). Optionally include readable content (Markdown), the ARIA tree, and a screenshot. This is the fastest way for an agent to understand and act on an un-instrumented page — far more token-efficient than a raw screenshot or full DOM. Use the returned selectors with run_sequence to act. Costs 1 API request.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to observe (required if no html) | |
| html | No | Raw HTML to observe (required if no url) | |
| width | No | Viewport width in pixels (default: 1280) | |
| format | No | Observation representation. "json" (default) returns the id-indexed "elements" array. "flatdomtree" returns "dom_text" — the indexed plain-text DOM used by browser-use / Alibaba page-agent (e.g. `[1]<button>Sign in</button>`) — plus a "selectors" map ({"1":"#signin"}) INSTEAD of the elements array. Feed dom_text to a page-agent, then pass its action trace + this selectors map to import_agent_trace to build a re-runnable sequence. | |
| height | No | Viewport height in pixels (default: 720) | |
| cookies | No | Cookies to set — array of "name=value" strings or { name, value, domain? } objects | |
| headers | No | Extra HTTP headers to send with the request | |
| blockAds | No | Block advertisements on the page | |
| darkMode | No | Emulate dark color scheme (default: false) | |
| timeZone | No | Override browser timezone | |
| bypassCSP | No | Bypass Content-Security-Policy on the page | |
| userAgent | No | Override the browser User-Agent string | |
| waitUntil | No | When to consider navigation finished (default: networkidle2) | |
| blockChats | No | Block live chat widgets | |
| session_id | No | Observe the LIVE state of a persistent session (Starter+; create with create_session) instead of a fresh page load. Omit url to observe the page exactly as the last run_sequence/take_screenshot left it; pass url to navigate within the session first. This is the recommended way to re-perceive between agent actions and recover from popovers/redirects. | |
| maxElements | No | Cap on interactive elements returned (default 40, max 150). Lower = fewer tokens. | |
| blockBanners | No | Hide cookie consent banners (default: false) | |
| includeRects | No | Include bounding boxes {x,y,w,h} per element (default false — omit to save tokens) | |
| authorization | No | Authorization header value (e.g. "Bearer <token>") | |
| blockTrackers | No | Block tracking scripts | |
| includeConsole | No | Also capture browser console output (console.log/info/warn/error/debug) and uncaught page errors emitted during load (default false). Adds a "Console" section — useful for debugging the page's runtime behavior alongside its structure. | |
| includeContent | No | Also extract the main readable content as Markdown (default false) | |
| viewportDevice | No | Device preset for viewport emulation (e.g. "iphone_14_pro"). Use list_devices to see all presets. | |
| includeAriaTree | No | Also include the interesting-only ARIA accessibility tree (default false) | |
| waitForSelector | No | Wait for this CSS selector to appear before observing | |
| screenshotFormat | No | Screenshot format when includeScreenshot is true (default jpeg) | |
| deviceScaleFactor | No | Device pixel ratio (default: 1) | |
| includeScreenshot | No | Also capture a screenshot in the same page load (default false) | |
| navigationTimeout | No | Navigation timeout in ms (default: 25000) | |
| screenshotFullPage | No | Capture the full scrollable page for the screenshot (default false) |