firecrawl_extract
Extract structured data from one or more URLs with an LLM, then poll status to retrieve results. Use it to turn web pages into schema-defined records for marketing workflows.
Instructions
Extract structured data from one or more URLs using an LLM. Poll results with extract_status.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | The URLs to extract data from. URLs should be in glob format. | |
| prompt | No | Prompt to guide the extraction process. | |
| schema | No | Schema to define the structure of the extracted data. Must conform to JSON Schema. | |
| showSources | No | When true, the sources used to extract the data will be included in the response as `sources`. | |
| ignoreSitemap | No | When true, sitemap.xml files will be ignored during website scanning. | |
| scrapeOptions | No | Options applied when scraping pages for extraction. | |
| enableWebSearch | No | When true, the extraction will use web search to find additional data. | |
| threatProtection | No | Per-request threat protection override. Enterprise feature. | |
| ignoreInvalidURLs | No | If invalid URLs are specified, they are ignored and returned in invalidURLs instead of failing the request. The server applies true when this is absent. | |
| includeSubdomains | No | When true, subdomains of the provided URLs will also be scanned. The server applies true when this is absent. |