browser.observe
Capture the current browser state—interactables, tabs, console, and page content—to provide context for the next automation step. Choose presets for text-only, screenshot-only, or combined detail.
Instructions
Capture the current browser observation: interactables, tabs, console, and a perception summary. Presets: 'text' — no screenshot or OCR: interactables, accessibility tree and the first 2,000 characters of page text; for text-only models. 'fast' — screenshot only, returned as an image, no text/accessibility extraction; for vision models. 'normal' (default) — text plus a screenshot URL and OCR. 'rich' — normal with twice the interactables and 4,000 characters of text. To read a whole page, use browser.get_html with text_only=true.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Most interactable elements to return. | |
| detail | No | 'compact' (default) leaves out what choosing the next step does not need: the full session record (see browser.get_session), remote-access diagnostics and, for actions, the pre-action snapshot. 'full' returns the complete payload. | compact |
| preset | No | text: page text and interactables, no screenshot. fast: screenshot and title only. normal: both. rich: more text and twice the interactables. Omitted: the deployment default. | |
| session_id | No | Target session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is. |