crawl
Discover URLs from a starting page and scrape each one, enabling bulk extraction of entire site sections or documentation.
Instructions
Multi-page crawl: discover URLs on root, then scrape each.
Use this when the user wants to crawl an entire site section or docs,
or needs multiple pages scraped in bulk. For single pages use scrape;
for research questions use deep_research.
Args:
root: start URL.
max_pages: hard cap on pages scraped.
css_selector: scope each page's html/markdown to the matched element
(non-llm: lxml re-scope of the fetched HTML; llm: native crawl4ai
css_selector).
prefer: "auto" | "fast" | "stealth" | "llm" (llm = Crawl4AI BFS deep-crawl).
include_paths: regex — keep only URLs matching (matched against full URL).
exclude_paths: regex — drop URLs matching (e.g. /tag/|/page/\d+).
max_depth: 0 = flat harvest from the root page's links (default);
>0 = true BFS up to that link depth, honoring the filters.
Returns: {root, pages: [{url, markdown, title, ...}], count, discovered, elapsed_ms} or {error, root} on failure.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | ||
| prefer | No | auto | |
| max_depth | No | ||
| max_pages | No | ||
| css_selector | No | ||
| exclude_paths | No | ||
| include_paths | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||