firecrawl_scrape
Scrape a single URL and optionally extract information as markdown, HTML, screenshots, or structured JSON for reading or summarizing a specific webpage.
Instructions
Scrape a single URL and optionally extract information. Use when the user wants to read or summarize a specific webpage. Supports markdown, HTML, screenshots, and structured JSON extraction.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| proxy | No | Specifies the type of proxy to use. | |
| maxAge | No | Returns a cached version of the page if it is younger than this age in milliseconds. The server applies 172800000 (2 days) when this is absent. | |
| minAge | No | When set, the request only checks the cache and never triggers a fresh scrape. | |
| mobile | No | Emulate scraping from a mobile device. | |
| actions | No | Actions to perform on the page before grabbing the content. | |
| formats | No | Output formats to include in the response. Strings or objects. The server applies markdown when this is absent. | |
| headers | No | Headers to send with the request. | |
| parsers | No | Controls how files are processed during scraping. | |
| profile | No | Persistent browser storage across scrape and interact sessions. | |
| timeout | No | Timeout in milliseconds. The server applies 60000 when this is absent. | |
| waitFor | No | Specify a delay in milliseconds before fetching the content. The server applies 0 when this is absent. | |
| blockAds | No | Enables ad-blocking and cookie popup blocking. | |
| location | No | Location settings for the request. | |
| lockdown | No | Serve from cache only and never make an outbound request. On miss, returns 404 SCRAPE_LOCKDOWN_CACHE_MISS. | |
| redactPII | No | Redact personally identifiable information from returned markdown. Pass true for defaults, or an object to tune it. | |
| excludeTags | No | Tags to exclude from the output. | |
| includeTags | No | Tags to include in the output. | |
| storeInCache | No | If true, the page will be stored in the Firecrawl index and cache. | |
| auditMetadata | No | User attribution included with SIEM logging events when SIEM is enabled. | |
| onlyMainContent | No | Only return the main content of the page excluding headers, navs, footers, etc. The server applies true when this is absent. | |
| onlyCleanContent | No | Beta. LLM pass over markdown to remove residual boilerplate that onlyMainContent can miss. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| zeroDataRetention | No | If true, this will enable zero data retention for this scrape. To enable this feature, please contact help@firecrawl.dev | |
| removeBase64Images | No | Removes all base64 images from the markdown output. | |
| skipTlsVerification | No | Skip TLS certificate verification when making requests. |