Skip to main content
Glama

browser.observe

Read-only

Capture the current browser state—interactables, tabs, console, and page content—to provide context for the next automation step. Choose presets for text-only, screenshot-only, or combined detail.

Instructions

Capture the current browser observation: interactables, tabs, console, and a perception summary. Presets: 'text' — no screenshot or OCR: interactables, accessibility tree and the first 2,000 characters of page text; for text-only models. 'fast' — screenshot only, returned as an image, no text/accessibility extraction; for vision models. 'normal' (default) — text plus a screenshot URL and OCR. 'rich' — normal with twice the interactables and 4,000 characters of text. To read a whole page, use browser.get_html with text_only=true.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoMost interactable elements to return.
detailNo'compact' (default) leaves out what choosing the next step does not need: the full session record (see browser.get_session), remote-access diagnostics and, for actions, the pre-action snapshot. 'full' returns the complete payload.compact
presetNotext: page text and interactables, no screenshot. fast: screenshot and title only. normal: both. rich: more text and twice the interactables. Omitted: the deployment default.
session_idNoTarget session id, from browser.create_session or browser.list_sessions. Optional when exactly one session is live; observe and act tools create one when none is.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.8.1
  2. Removedv1.4.2
  3. Changed1 schema field changedv1.4.2
    • addedInput schema / properties / session_id / description
      Added value: +"ID of the target browser session, as returned by browser.create_session or listed by browser.list_sessions."
  4. Changed5 schema fields changedv1.0.3
    • addedInput schema / additionalProperties
      Added value: +false
    • changedInput schema / properties / limit / maximum
      Previous value: -100New value: +200
    • addedInput schema / properties / preset
      Added value: +{
      +  "default": "normal",
      +  "enum": [
      +    "fast",
      +    "normal",
      +    "rich"
      +  ],
      +  "title": "Preset",
      +  "type": "string"
      +}
    • addedInput schema / properties / session_id / maxLength
      Added value: +120
    • addedInput schema / properties / session_id / minLength
      Added value: +1
  5. First observedv0.5.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint/openWorldHint annotations by disclosing what each preset returns, including OCR, accessibility tree, screenshot URL vs. image-only output, and character limits. It also explains the relationship between text and rich modes, giving the agent accurate expectations about output content without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description front-loads the core purpose and then uses a dense but scannable preset breakdown; every sentence carries decision-relevant information. The alternative routing is placed at the end without redundant restatement of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description specifies the return composition for each preset, including whether output is an image, a URL, OCR, or text, and it covers the main decisions an agent needs to make. Given the 100% parameter schema coverage, nothing essential to selecting or invoking the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds real value by explaining the semantics of the preset enum and the trade-offs among text/fast/normal/rich, plus the recommendation to use get_html for full-page reading. Limit and detail are already self-documenting in the schema, so no further compensation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource ('Capture the current browser observation') and enumerates its contents (interactables, tabs, console, perception summary), so an agent can tell it apart from sibling read tools. The closing sentence explicitly distinguishes it from browser.get_html for full-page reading.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete selection cues by mapping presets to model types ('text' for text-only models, 'fast' for vision models) and names browser.get_html as the alternative for reading a whole page. It does not explicitly state when observe should be avoided in favor of browser.screenshot or browser.find_elements, so it falls just short of full when/when-not coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.