firecrawl_crawl
Crawl a website recursively and scrape each discovered page to extract full-site content for SEO audits, content mapping, or migration.
Instructions
Recursively crawl a website and scrape each discovered page.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The base URL to start crawling from | |
| delay | No | Delay in seconds between scrapes. Setting this forces concurrency to 1. | |
| limit | No | Maximum number of pages to crawl. The server applies 10000 when this is absent. | |
| prompt | No | Natural language prompt to generate crawler options from. | |
| sitemap | No | Sitemap mode when crawling. The server applies 'include' when this is absent. | |
| webhook | No | Webhook specification for crawl lifecycle events. | |
| excludePaths | No | URL pathname regex patterns that exclude matching URLs from the crawl. | |
| includePaths | No | URL pathname regex patterns that include matching URLs in the crawl. | |
| scrapeOptions | No | Options applied when scraping each crawled page. | |
| maxConcurrency | No | Maximum number of concurrent scrapes for this crawl. | |
| regexOnFullURL | No | Match includePaths and excludePaths against the full URL instead of just the pathname. | |
| allowSubdomains | No | Allow the crawler to follow links to subdomains of the main domain. | |
| ignoreRobotsTxt | No | Ignore the website's robots.txt rules. Enterprise only. | |
| robotsUserAgent | No | Custom User-Agent string for robots.txt evaluation. Enterprise only. | |
| crawlEntireDomain | No | Allow the crawler to follow internal links to sibling or parent URLs, not just child paths. | |
| maxDiscoveryDepth | No | Maximum depth to crawl based on discovery order. | |
| zeroDataRetention | No | If true, this will enable zero data retention for this crawl. | |
| allowExternalLinks | No | Allow the crawler to follow links to external websites (one hop only). | |
| ignoreQueryParameters | No | Do not re-scrape the same path with different query parameters. |