fetch_url
Fetch any URL with automatic JS rendering and bot-protection handling, then return body, headers, and cleaned HTML to reduce LLM token costs.
Instructions
Fetch any URL with automatic JS-rendering and common bot-protection handling — advanced behavioral fingerprinting may still block header retrieval (surfaced via headersAvailable: false). Returns body, headers, cleanStats. Optional cleanHtml strips HTML noise while preserving text content — token-cost win for LLM consumption.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Target URL to fetch. Must use http:// or https:// and resolve to a public host. | |
| cleanHtml | No | When `true` and the response content-type is `text/html`, strip HTML noise (scripts, styles, comments) while preserving text content. Significant token-cost reduction for LLM consumption — per-request reduction reported in `cleanStats`. Requires `bodyNeeded`. | |
| bodyNeeded | No | Include `body` and `contentType` in the response. Defaults to service-controlled value when omitted. | |
| bodyMaxBytes | No | Per-request response body cap in bytes. Accepted range 1024–104857600 (1 KiB – 100 MiB). Oversize responses are rejected pre-buffer. Defaults to service-controlled value when omitted. | |
| maxTimeoutMs | No | Caller-side timeout budget in milliseconds. Accepted range 1000–120000. Defaults to service-controlled value when omitted. | |
| headersNeeded | No | Include `headers` and `headersAvailable` in the response. Defaults to service-controlled value when omitted. When the target site uses advanced behavioral fingerprinting, `headersAvailable` is `false` and `headers` is an empty object — present, not missing. Branch on `headersAvailable`, never on whether `headers` exists. |