web_scrape_page
Scrape Page — Extract structured data from a public web page at a user-provided http(s) URL: CSS-selector mode returns text per selector; readability mode returns the main article as clean markdown. [category: web]
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | http(s) URL to scrape | |
| mode | No | Readability gives you the page's main article as clean text. Selectors gives you only the specific bits you name below. | readability |
| selectors | No | One CSS selector per thing you want, named: {"headline": "h1", "price": ".price"}. Needed only when you pick Selectors above. | |
| timeout_ms | No | How long to wait for the page before giving up, in milliseconds (30000 = 30 seconds). | |
| user_agent | No | Advanced: how we introduce ourselves to the site. Left blank we identify as JohnsEssentialsBot. | |
| acknowledge_robots | No | Business plan: scrape the page even when the site's robots.txt asks bots to stay away. On any other plan this switch does nothing and the page is still refused. |