Skip to main content
Glama

extract_page_data

Scrape a URL using ScrapingBee and extract specific data using CSS or XPath selectors.

Scope: one URL per call, with the result returned into the conversation. For many URLs, a whole site, output written to a file or directory, resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead — scrapingbee scrape --input-file urls.txt --output-dir results, scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every. This server has no batch, crawl, file-output or scheduling equivalent.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL of the page to scrape.
waitNoMilliseconds to wait after page load (0–35000).
deviceNodesktop (default) or mobile.desktop
cookiesNoSemicolon-separated cookies to send to the target.
timeoutNoRequest timeout in milliseconds (1000–140000, default 140000).
max_costNoOptional credit ceiling for auto_mode (integer >= 1). The request will not escalate to a configuration costing more than this many credits. 0 means no ceiling. Ignored unless auto_mode is True.
wait_forNoCSS or XPath selector to wait for before returning.
auto_modeNoOn by default — ScrapingBee automatically picks the cheapest configuration that successfully fetches the page, escalating proxy strength only as needed (you are charged only for the winning config). To choose a configuration yourself instead, use one of: auto_mode=False for the classic tier (no proxy, cheapest), premium_proxy=True, or stealth_proxy=True — the two proxy flags override auto_mode on their own, so auto_mode=False is only needed for the classic tier. Setting render_js also switches to manual.
render_jsNoWhether to use headless-browser rendering. Setting this explicitly switches the request to manual configuration (disables auto_mode).
session_idNoInteger 0–10000000 to reuse an IP for up to 5 minutes.
ai_selectorNoCSS selector to focus AI extraction on part of the page.
js_scenarioNoStringified JSON object of browser interaction instructions.
country_codeNoISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy or stealth_proxy is True.
wait_browserNodomcontentloaded (default), load, networkidle0, networkidle2.domcontentloaded
window_widthNoViewport width (default 1920).
custom_googleNoSet to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc)
extract_rulesYesA JSON string defining extraction rules. SIMPLE SYNTAX: Use {"key_name": "css_or_xpath_selector"} format. - CSS selector example: {"title": "h1", "price": "span.price"} - XPath selector example: {"title": "//h1", "items": "//div[@class='item']"} - Extract attribute value of an element using @<attribute> or if you are using extended syntax <element>@<attribute>, for example: {"image": "img@src"}, "link": {"selector": "a","output": "@href"} Note: Selectors starting with "/" are treated as XPath, otherwise CSS. EXTENDED SYNTAX: For more control, use a dict with these options: - "selector": CSS or XPath selector string (required) - "type": "item" (single element, default) or "list" (multiple elements) - "output": It is also possible to add extraction rules inside the output option in order to create powerful extractors. Extended example: { "title" : "h1", "subtitle" : "#subtitle", "articles": { "selector": ".card", "type": "list", "output": { "title": ".post-title", "link": { "selector": ".post-title", "output": "@href" }, "description": ".post-description" } } }
json_responseNoReturn the full JSON envelope instead of the raw body.
premium_proxyNoManual option: use a premium proxy (middle tier of classic/premium/stealth). Setting this disables auto_mode.
stealth_proxyNoManual option: use a stealth proxy (strongest tier of classic/premium/stealth). Setting this disables auto_mode.
window_heightNoViewport height (default 1080).
block_resourcesNoBlock images/CSS to speed up rendering.
ai_extract_rulesNoJSON string of AI extraction rules.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. It adds the one-URL-per-call scope and the fact that results return into the conversation, which are useful. But it doesn't disclose cost implications, rate limits, auth requirements, or the auto_mode escalation behavior for an operation that can incur credits — gaps left mostly to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, then a clearly delimited Scope section with alternatives. Some redundancy in the CLI example list, but every line serves routing or scope; no wasted prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity 23-param tool with an output schema present (so return values needn't be explained) and 100% schema coverage, the description covers the key decision boundary (single vs batch/crawl) well. Minor gap: no note on credit/cost behavior or auth, but coverage is otherwise sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents all 23 parameters (including auto_mode, max_cost, and the proxy flags with their own trade-offs). The description adds no parameter syntax or format detail beyond the structured fields, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scrape) + resource (a URL) and the mechanism (ScrapingBee) plus the differentiated capability (extract specific data via CSS/XPath selectors). This clearly separates it from siblings like get_page_html and get_page_text, which presumably return whole-page content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit scope statement ('one URL per call, result returned into the conversation') and a detailed when-not list: many URLs, whole site, file output, resumable jobs, scheduled re-runs — each routed to a named CLI alternative with concrete command examples. Nothing left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources