get_page_text
Scrape a URL using ScrapingBee and return the page content in Markdown or Text format.
Scope: one URL per call, with the result returned into the conversation.
For many URLs, a whole site, output written to a file or directory,
resumable jobs, or a scheduled re-run, use the ScrapingBee CLI instead —
scrapingbee scrape --input-file urls.txt --output-dir results,
scrapingbee crawl URL --save-pattern ..., scrapingbee schedule --every.
This server has no batch, crawl, file-output or scheduling equivalent.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL of the page to scrape. | |
| wait | No | Milliseconds to wait after page load (0–35000). | |
| device | No | desktop (default) or mobile. | desktop |
| cookies | No | Semicolon-separated cookies to send to the target. | |
| timeout | No | Request timeout in milliseconds (1000–140000, default 140000). | |
| ai_query | No | Natural-language question for AI extraction. | |
| max_cost | No | Optional credit ceiling for auto_mode (integer >= 1). The request will not escalate to a configuration costing more than this many credits. 0 means no ceiling. Ignored unless auto_mode is True. | |
| wait_for | No | CSS or XPath selector to wait for before returning. | |
| auto_mode | No | On by default — ScrapingBee automatically picks the cheapest configuration that successfully fetches the page, escalating proxy strength only as needed (you are charged only for the winning config). To choose a configuration yourself instead, use one of: auto_mode=False for the classic tier (no proxy, cheapest), premium_proxy=True, or stealth_proxy=True — the two proxy flags override auto_mode on their own, so auto_mode=False is only needed for the classic tier. Setting render_js also switches to manual. | |
| render_js | No | Whether to use headless-browser rendering. Omit to use the API default (true). Setting this explicitly switches the request to manual configuration (disables auto_mode). | |
| session_id | No | Integer 0–10000000 to reuse an IP for up to 5 minutes. | |
| ai_selector | No | CSS selector to focus AI extraction on part of the page. | |
| js_scenario | No | Stringified JSON object of browser interaction instructions. | |
| country_code | No | ISO 3166-1 alpha-2 country code (e.g. "fr", "us") to route the request through an IP in that country. Only takes effect when premium_proxy or stealth_proxy is True. | |
| wait_browser | No | domcontentloaded (default), load, networkidle0, networkidle2. | domcontentloaded |
| window_width | No | Viewport width (default 1920). | |
| custom_google | No | Set to True if the url is a google domain in the following format: <subdomain>.google.<top-level-domain> (Example: translate.google.com, mail.google.com, www.google.co.in etc) | |
| premium_proxy | No | Manual option: use a premium proxy (middle tier of classic/premium/stealth). Setting this disables auto_mode. | |
| stealth_proxy | No | Manual option: use a stealth proxy (strongest tier of classic/premium/stealth). Setting this disables auto_mode. | |
| window_height | No | Viewport height (default 1080). | |
| block_resources | No | Block images/CSS to speed up rendering. | |
| return_page_text | No | Whether to return the page content in Text format. | |
| return_page_markdown | No | Whether to return the page content in Markdown format. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |