Scrape a URL asynchronously
scrape_asyncSubmit a scrape job asynchronously, receive a job_id instantly for long-running tasks like JS rendering or full-page screenshots. Poll for status or get a webhook on completion.
Instructions
Submit a scrape job to run ASYNCHRONOUSLY and return a job_id immediately, instead of holding the connection open. Use this for long-running scrapes (heavy JS rendering, full-page screenshots of long pages) — for a quick one-shot fetch prefer the synchronous scrape tool, which blocks and returns the page directly. Poll the job with scrape_status and fetch the page with scrape_result once done. Optionally attach a per-job completion webhook: we deliver a single signed scrape.complete POST to your endpoint when the job finishes (HTTPS required in production), and echoes your opaque webhook.metadata back in the delivery (under data.metadata, alongside data.response_meta.usage).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape. Known tracking parameters are removed before the page is fetched, so they are neither sent to the target site nor part of the cache key. Every other query parameter is kept verbatim and is part of the cache key. | |
| cache | No | Cache settings for this request | |
| proxy | No | Proxy tier to use for fetching | auto |
| cleanup | No | What is removed from the page before any output is built. Applies to markdown, cleaned_html, links and images on every engine, and to the screenshot. Never applies to raw_html, which is always the page before we removed anything. | |
| extract | No | Which content formats to extract. Defaults to metadata + cleaned_html. | |
| webhook | No | Completion webhook delivered when this async job finishes. Async-only: the synchronous `scrape` tool does not accept it. | |
| location | No | Optional locale + country emulation for the scrape | |
| require_js | No | Use a headless browser to render JavaScript before scraping |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job identifier — pass it to the `scrape_status` / `scrape_result` tools |