web_fetch
Fetch URL — Fetch the raw (decoded) HTML of a public web page at a user-provided http(s) URL. SSRF-guarded, http/https only, 10 MB body cap, robots-aware. [category: web]
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | http(s) URL to fetch | |
| timeout_ms | No | How long to wait for the page before giving up, in milliseconds (30000 = 30 seconds). | |
| user_agent | No | What to tell the site we are. Leave blank and we identify honestly as JohnsEssentialsBot/1.0. | |
| acknowledge_robots | No | Business plan only: fetch the page even when the site's rules file (robots.txt) asks crawlers to stay away. On other plans this switch has no effect. |