Scrape a URL
scrapeFetches a single URL and returns page content as markdown, cleaned HTML, raw HTML, links, images, screenshot, or metadata.
Instructions
Fetch a single URL via the crawlbrulee scraping API and return the requested content (markdown, cleaned HTML, raw HTML, links, images, screenshot, page metadata in metadata). Use this for one-shot page extraction. For full-site discovery use the map tool first. Screenshot URLs in the response are signed download links — the agent can fetch them when needed. The response also carries response_meta.usage = { credits, engine, proxy (the resolved tier — never auto), screenshot_slices }. engine is the billed engine (http, browser, screenshot, cache); a cache hit is represented by engine: "cache", so you can see what the request cost. Extraction is capped per page: 30,000 links, 10,000 inline images and 10,000,000 characters of HTML. A page past a cap is truncated rather than refused, and the response warnings array names which one (links_truncated, inline_images_truncated, raw_html_truncated) — so treat that output as incomplete. Use cleanup to control what is removed before any output is built: ads_and_popups (on by default) drops ads, cookie banners and chat widgets, and exclude_selectors removes anything else. It shapes markdown, cleaned_html, links, images and the screenshot, and NEVER raw_html — so request raw_html when you need the page exactly as it arrived. warnings also reports a section whose extraction failed outright (links_unavailable, inline_images_unavailable, metadata_unavailable): that field comes back omitted or empty while the rest of the scrape succeeded, so do NOT conclude the page had no links/images/metadata — re-run the scrape instead. An empty field with no such warning does mean the page had none.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape. Known tracking parameters are removed before the page is fetched, so they are neither sent to the target site nor part of the cache key. Every other query parameter is kept verbatim and is part of the cache key. | |
| cache | No | Cache settings for this request | |
| proxy | No | Proxy tier to use for fetching | auto |
| cleanup | No | What is removed from the page before any output is built. Applies to markdown, cleaned_html, links and images on every engine, and to the screenshot. Never applies to raw_html, which is always the page before we removed anything. | |
| extract | No | Which content formats to extract. Defaults to metadata + cleaned_html. | |
| location | No | Optional locale + country emulation for the scrape | |
| require_js | No | Use a headless browser to render JavaScript before scraping |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL that was actually scraped, after any redirects, in normalized form — the base that links, images, and internal labels are computed against | |
| links | No | Links found on the page | |
| images | No | Inline images found on the page | |
| markdown | No | Page content converted to clean Markdown | |
| metadata | No | Extracted page metadata (title, OG tags, etc.) | |
| raw_html | No | Raw, unprocessed HTML of the page | |
| warnings | No | Non-error notices about the scrape. Truncation codes — `screenshot_truncated` (long page exceeded the scrolling-screenshot height cap), `links_truncated` / `inline_images_truncated` (page had more links/images than the per-page extraction caps), `raw_html_truncated` / `metadata_truncated` (rendered HTML exceeded the per-page size budget) — mean the field is present but capped. Unavailability codes — `links_unavailable` / `inline_images_unavailable` / `metadata_unavailable` — mean that optional field could not be extracted and was omitted (null/empty) while the rest of the scrape succeeded, so an empty field carrying one of these does NOT mean the page had none. Stable string codes — clients can switch on them. Warnings are stored with the result: async result fetches and cache hits carry them too, filtered to the fields the request asked for. | |
| screenshot | No | Screenshot of the page, if requested | |
| cleaned_html | No | Cleaned HTML of the main page content | |
| content_type | No | Content-Type header returned by the server | |
| requested_url | Yes | The URL you requested, echoed verbatim — before any redirects | |
| response_meta | Yes | Request-level metadata. `response_meta.usage` reports credits charged, the billed engine, the resolved proxy tier, and any screenshot-slice add-on. | |
| unsupported_fields | No | Extract fields that were requested but are not supported for this content type |