Skip to main content
Glama

browser_extract

Extract current visible content and structured page state from a tab, then query CSS selectors, read logs, inspect elements, or audit pages.

Instructions

Extract the complete current visible content and structured field/page state from a tab. Pass target_ref to scope a fresh read to an observed content block, control or the ARIA region it owns. This is a fresh high-level read, not a change-only observation. If a large result returns evidence_ref and next_cursor from browser_extract, continue it with this same tool. Do not continue a managed-output reference from browser_open, browser_observe, or browser_act; start a fresh extraction with tab_id and instruction instead. Pass selector for an instant CSS query instead of a full read: it counts and lists matching elements (tag, text, requested attributes, and the ref of any match already in the current observation) without building a snapshot, e.g. every product link's href, how many rows a table has, or each card's price. read=console|network reads the tab's logs (with same-site error bodies); read=inspect with target_ref says why a click is refused or a control will not take a value: what receives the click there, what covers or clips it, disabled state, styles. read=design returns how the page looks in CSS terms: variables, colors by use, type scale, radii, shadows, spacing and layout regions with sizes; target_ref limits it to one component. read=audit audits the page you are on (a deployed or localhost site): accessibility violations (axe-core) with the ref of each element, load timings and weight, title/lang/description/headings/alt gaps, broken same-origin links and console errors, worst first.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
findNoReturn every passage that mentions this, with the words around it, instead of the page's text. Use it when the page is long and the question is narrow (a deadline, a price, a name).
readNo
levelNo
limitNoContinuation only: maximum characters for this slice; defaults to 8000.
typesNoread=network: e.g. ['fetch','xhr'].
checksNoread=audit: which checks to run; omit for all four.
cursorNoWith evidence_ref: character offset from next_cursor. With selector or log reads: entry offset from next_cursor. Defaults to 0.
filterNoread=network: URL regex, e.g. '/api/'.
tab_idYes
from_endNoRead the end of the page rather than the beginning. Footers, totals and closing dates live there, and paging forward charges for everything above them.
selectorNoCSS selector to query instead of reading the page (e.g. 'table tbody tr', 'a.product-link', 'article h2'). Returns total, and for each match up to max_results: tag, text, attrs, children_count, visible, and ref when the element is in the current observation. An invalid selector is an error; no match is an empty result, not an error.
attributesNoselector only: attributes to return per match, e.g. ['href'] or ['src', 'alt']. href/src are complete absolute URLs, or explicitly omitted when longer than 8192 characters. Other attributes are bounded previews. value is the live field value (secrets read as [redacted]).
target_refNoFresh extraction, read=inspect or read=design: exact ref from the current observation (read=design also takes a region id r…). Reads that element, or the uniquely related region named by aria-controls, without executing page JavaScript in the model loop.
failed_onlyNoread=network: failed or 4xx/5xx only.
instructionNoFresh extraction: optional objective used to focus field state and report an explicit no-match result. Continuation: accepted as a harmless repeated hint but archived evidence is returned unchanged.
max_resultsNoselector or log reads: entries per page (default 50). total always counts every match; continue with cursor=next_cursor.
navigationsNo
evidence_refNoContinuation only: opaque evidence_ref returned by an earlier browser_extract result in this session, accompanied by next_cursor.
include_textNoselector only: include each match's text (default true). Set false when only attributes or the count matter.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it distinguishes fresh reads from change-only observations, explains continuation semantics (evidence_ref + next_cursor, cursor offset, limit default 8000), notes that selector queries avoid building a snapshot, states invalid selector is an error while no match is an empty result, and discloses redaction of secret values and bounded attribute previews.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then ordered into scoping, continuation, selector, and mode-specific paragraphs. It is dense and almost every sentence adds operative detail, though the sheer length (especially the read=audit and read=design clauses) makes it slightly sprawling for an agent to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 19-param multi-mode tool with no output schema and no annotations, the description covers the major modes, continuation behavior, and error semantics well. It is not fully complete: the level and navigations parameters remain unexplained, and the default page-read return shape is only partially implied through continuation limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 79%, and the description adds substantial meaning for target_ref, selector, attributes, find, from_end, evidence_ref, cursor, instruction, and each read mode. However, params level and navigations have empty schema descriptions and are never addressed in the description, leaving real gaps for a 19-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource (extract visible content and structured state from a tab) and explicitly differentiates itself from siblings: 'fresh high-level read, not a change-only observation' versus browser_observe, and it warns not to continue managed-output refs from browser_open/browser_observe/browser_act. An agent can distinguish this tool's role without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Extremely explicit routing: it says when to pass target_ref, when to use selector instead of a full read, when to use find, from_end, and each read= mode (console/network/inspect/design/audit). It also gives a clear when-not instruction ('Do not continue a managed-output reference from browser_open, browser_observe, or browser_act; start a fresh extraction').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.