web_crawl
Follows a site's links and reads every page. Returns a jobId; poll it with web_crawl_status. Charged up front for the pages it is allowed to read (limit), and the pages it never reads are refunded. Use web_map first if you only need the urls.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Where to start. | |
| limit | No | Pages it may read (default 25, max 200). | |
| maxAge | No | Accept an answer up to this many milliseconds old. A cache hit costs $0.0002 instead of the format price. Leave it out to force a fresh render. | |
| formats | No | Outputs you want in the same response. `controls` is the map of what can be clicked; `elements` needs `selectors`; `json` needs `json.prompt`. | |
| maxDepth | No | How far to follow links (default 2, max 5). | |
| excludePaths | No | ||
| includePaths | No | ||
| includeSubdomains | No |