firecrawl_scrape
Scrape one URL and return its content: markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema. Use it when the request identifies a page and needs its content or defined fields. Use firecrawl_search when additional web sources are needed; on an authenticated session, firecrawl_map lists a site's URLs and firecrawl_crawl collects a set of pages.
Firecrawl may serve recently indexed content; set maxAge: 0 for a live fetch or a smaller maxAge to bound staleness. A successful response does not by itself confirm the page is still current. Browser actions can change the live page when interactive actions are enabled. Authenticated responses can include a metadata.scrapeId for optional scrape feedback.
On an authenticated session with Alexandria access, firecrawl_search with sources unset and firecrawl_find_tools can discover providers for the same fields across several pages; a matching provider returns typed records in one call. Keyless sessions have no provider matches.
Alexandria mode, on an authenticated session with Alexandria access: alexandria selects catalogued capability execution and is mutually exclusive with url.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| proxy | No | ||
| maxAge | No | ||
| mobile | No | ||
| formats | No | ||
| parsers | No | ||
| profile | No | ||
| timeout | No | Execution timeout in milliseconds. | |
| waitFor | No | ||
| location | No | ||
| lockdown | No | ||
| redactPII | No | ||
| requestId | No | Idempotency key bound to one Alexandria execution payload. Generated when omitted and returned with the result. | |
| alexandria | No | Catalogued Alexandria capability invocation, mutually exclusive with url. One {provider, capability, options} object or an array of 1-10, with contracts available through firecrawl_search or firecrawl_find_tools. Each call may include version to pin a published workflow; omitting it uses latest. Only requestId and timeout are supported alongside alexandria. The selected contract marks required inputs and any requiresOneOf groups (at least one member per group); it may include example.request/example.response and response.key (which may differ from records). Where pagination is declared, its fields govern paging with the same filters; catalogue next is separate from provider pagination. Returns per-capability results in data.alexandria with data, records, or an error with a code; individual capabilities can fail even when the outer response succeeds. Requires an API key on a team with Alexandria enabled. Some providers require accepted terms; blocked requests return the applicable requirements. | |
| pdfOptions | No | ||
| toolDetail | No | URL mode only: domain discovery detail, summary by default; compact returns provider/capability/description, full includes contracts. Ignored with alexandria. | |
| domainTools | No | URL mode only: include domain-matched Alexandria tools for the page in tools on the returned document. Ignored with alexandria. | |
| excludeTags | No | ||
| includeTags | No | ||
| jsonOptions | No | ||
| queryOptions | No | ||
| storeInCache | No | ||
| onlyMainContent | No | ||
| screenshotOptions | No | ||
| zeroDataRetention | No | ||
| removeBase64Images | No | ||
| skipTlsVerification | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | Alexandria mode: per-capability results in `data.alexandria`, each with `data`, `records`, or an `error`. | |
| html | No | Processed HTML of the page. | |
| json | No | Structured data matching the requested JSON schema or prompt. | |
| menu | No | Menu data extracted from the page. | |
| audio | No | Audio extracted from the page. | |
| error | No | Error message or error object when the call did not succeed. | |
| links | No | Links found on the page. | |
| pages | No | Physical PDF pages, when `parsers[].pages` is set. | |
| tools | No | Domain-matched Alexandria tools for the page, when `domainTools` is set. | |
| video | No | Video extracted from the page. | |
| answer | No | Targeted answer to the question that was asked of the page. | |
| blocks | No | Typed PDF layout blocks, when `parsers[].blocks` is set. | |
| images | No | Images found on the page. | |
| actions | No | Results of the browser actions that ran during the scrape. | |
| message | No | Guidance that accompanies the result. | |
| product | No | Product data extracted from the page. | |
| rawHtml | No | Unprocessed HTML of the page. | |
| receipt | No | Billing receipt for the execution. | |
| success | No | Whether the API call succeeded. | |
| summary | No | Summary of the page content. | |
| warning | No | Non-fatal warning about the result. | |
| branding | No | Branding data extracted from the page. | |
| delivery | No | `retained` when the full result stayed server-side instead of being inlined. | |
| markdown | No | Page content as markdown. | |
| metadata | No | Page metadata; authenticated responses can include `metadata.scrapeId` for scrape feedback. | |
| nextTool | No | A follow-up tool call (`{name, arguments}`) that continues or inspects this result. | |
| requestId | No | Identifier of this Alexandria execution. | |
| scrape_id | No | Identifier of the underlying scrape. | |
| attributes | No | Values collected by the requested attribute selectors. | |
| highlights | No | Highlighted passages from the page. | |
| screenshot | No | Screenshot of the page. | |
| agent_hints | No | Optional response guidance from the Firecrawl API. | |
| creditsCost | No | Credits this call consumed. | |
| workspaceId | No | Workspace holding a retained result, for inspection through virtual Bash. | |
| feedbackTool | No | Pointer to the feedback tool for reporting how this result served the task. | |
| responseBytes | No | Size of the full result in bytes. | |
| changeTracking | No | Change-tracking comparison against the previous scrape. | |
| idleTtlSeconds | No | Seconds a retained workspace stays available while idle. | |
| estimatedTokens | No | Estimated token cost of the full result. | |
| inlineTokenBudget | No | Token budget above which a result is retained rather than inlined. | |
| tokenEstimateMethod | No | How `estimatedTokens` was derived. |