Skip to main content
Glama

observe_page

Analyze any web page in one request to return interactive elements, page type, and suggested actions for AI agents, with optional content, ARIA tree, or screenshot.

Instructions

Get a compact, token-budgeted "observation" of any web page, purpose-built for AI agents. In ONE request it returns: id-indexed interactive elements (role, name, CSS selector, state), a heuristic page-type classification (login, signup, search, article, form, generic), and grouped "suggested actions" (login flow, search, primary buttons, navigation). Optionally include readable content (Markdown), the ARIA tree, and a screenshot. This is the fastest way for an agent to understand and act on an un-instrumented page — far more token-efficient than a raw screenshot or full DOM. Use the returned selectors with run_sequence to act. Costs 1 API request.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoURL to observe (required if no html)
htmlNoRaw HTML to observe (required if no url)
widthNoViewport width in pixels (default: 1280)
formatNoObservation representation. "json" (default) returns the id-indexed "elements" array. "flatdomtree" returns "dom_text" — the indexed plain-text DOM used by browser-use / Alibaba page-agent (e.g. `[1]<button>Sign in</button>`) — plus a "selectors" map ({"1":"#signin"}) INSTEAD of the elements array. Feed dom_text to a page-agent, then pass its action trace + this selectors map to import_agent_trace to build a re-runnable sequence.
heightNoViewport height in pixels (default: 720)
cookiesNoCookies to set — array of "name=value" strings or { name, value, domain? } objects
headersNoExtra HTTP headers to send with the request
blockAdsNoBlock advertisements on the page
darkModeNoEmulate dark color scheme (default: false)
timeZoneNoOverride browser timezone
bypassCSPNoBypass Content-Security-Policy on the page
userAgentNoOverride the browser User-Agent string
waitUntilNoWhen to consider navigation finished (default: networkidle2)
blockChatsNoBlock live chat widgets
session_idNoObserve the LIVE state of a persistent session (Starter+; create with create_session) instead of a fresh page load. Omit url to observe the page exactly as the last run_sequence/take_screenshot left it; pass url to navigate within the session first. This is the recommended way to re-perceive between agent actions and recover from popovers/redirects.
maxElementsNoCap on interactive elements returned (default 40, max 150). Lower = fewer tokens.
blockBannersNoHide cookie consent banners (default: false)
includeRectsNoInclude bounding boxes {x,y,w,h} per element (default false — omit to save tokens)
authorizationNoAuthorization header value (e.g. "Bearer <token>")
blockTrackersNoBlock tracking scripts
includeConsoleNoAlso capture browser console output (console.log/info/warn/error/debug) and uncaught page errors emitted during load (default false). Adds a "Console" section — useful for debugging the page's runtime behavior alongside its structure.
includeContentNoAlso extract the main readable content as Markdown (default false)
viewportDeviceNoDevice preset for viewport emulation (e.g. "iphone_14_pro"). Use list_devices to see all presets.
includeAriaTreeNoAlso include the interesting-only ARIA accessibility tree (default false)
waitForSelectorNoWait for this CSS selector to appear before observing
screenshotFormatNoScreenshot format when includeScreenshot is true (default jpeg)
deviceScaleFactorNoDevice pixel ratio (default: 1)
includeScreenshotNoAlso capture a screenshot in the same page load (default false)
navigationTimeoutNoNavigation timeout in ms (default: 25000)
screenshotFullPageNoCapture the full scrollable page for the screenshot (default false)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.17.0

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and does well: it discloses the cost ('1 API request'), the token-budgeting behavior, and that the page is loaded/emulated. It implies a read-only, side-effect-free perception step but does not state auth requirements or rate limits explicitly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and return shape, then routing and cost. Four dense sentences, all earning their place, with only mild redundancy between 'token-budgeted' and 'token-efficient'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing the returned structure and the opt-in extras. It names run_sequence as the follow-up tool, though it could say more about how the classification/suggested-actions groups map to next actions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description groups the optional output toggles ('readable content (Markdown), the ARIA tree, and a screenshot') but adds no syntax or format meaning beyond what the schema already documents for each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (observe/get) and resource (web page) and explicitly enumerates the return payload (id-indexed elements, page-type classification, suggested actions). It also positions itself against the obvious siblings (take_screenshot, inspect_page) by contrasting with 'a raw screenshot or full DOM', so an agent can route correctly without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Says to use the returned selectors with run_sequence to act, and frames itself as the fastest way to understand an un-instrumented page. It implies when it beats a screenshot/DOM but never gives an explicit use-this-not-that rule against inspect_page or take_screenshot.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.