Get an async scrape job result
scrape_resultFetch the result of a finished async scrape job: markdown, cleaned/raw HTML, links, images, screenshot, metadata, and usage. Check scrape_status first to confirm done and avoid pending errors.
Instructions
Fetch the extracted content of a completed async scrape job (the same result shape as the synchronous scrape tool: markdown, cleaned HTML, raw HTML, links, images, screenshot, page metadata in metadata, and response_meta.usage). Errors if the job is still pending/running — check scrape_status first (status done) before calling this. Screenshot URLs are signed download links.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job identifier returned by the `scrape_async` tool |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL that was actually scraped, after any redirects, in normalized form — the base that links, images, and internal labels are computed against | |
| links | No | Links found on the page | |
| images | No | Inline images found on the page | |
| markdown | No | Page content converted to clean Markdown | |
| metadata | No | Extracted page metadata (title, OG tags, etc.) | |
| raw_html | No | Raw, unprocessed HTML of the page | |
| warnings | No | Non-error notices about the scrape. Truncation codes — `screenshot_truncated` (long page exceeded the scrolling-screenshot height cap), `links_truncated` / `inline_images_truncated` (page had more links/images than the per-page extraction caps), `raw_html_truncated` / `metadata_truncated` (rendered HTML exceeded the per-page size budget) — mean the field is present but capped. Unavailability codes — `links_unavailable` / `inline_images_unavailable` / `metadata_unavailable` — mean that optional field could not be extracted and was omitted (null/empty) while the rest of the scrape succeeded, so an empty field carrying one of these does NOT mean the page had none. Stable string codes — clients can switch on them. Warnings are stored with the result: async result fetches and cache hits carry them too, filtered to the fields the request asked for. | |
| screenshot | No | Screenshot of the page, if requested | |
| cleaned_html | No | Cleaned HTML of the main page content | |
| content_type | No | Content-Type header returned by the server | |
| requested_url | Yes | The URL you requested, echoed verbatim — before any redirects | |
| response_meta | Yes | Request-level metadata. `response_meta.usage` reports credits charged, the billed engine, the resolved proxy tier, and any screenshot-slice add-on. | |
| unsupported_fields | No | Extract fields that were requested but are not supported for this content type |