Crawl a whole site once
crawl_siteCrawl a site once from a start URL -- following its links and reading its sitemap, up to limit pages -- and return what it found, without setting up a project. Use it to read a site or section whose page URLs you do not know; for known URLs use scrape_urls, and to watch a site over time use create_project (or keep_crawl_as_project on this crawl afterwards). Each page costs the credits of the engine that read it (usually 1 to 4) and refused pages are free. It waits for the crawl, minutes for a large limit; hosted, a crawl still going after the time budget comes back as a job for get_job, and repeating the call returns the same crawl. The answer is an index of the pages read plus excerpts inside 60,000 characters; get_job with url reads one page in full and with cursor the next window. The crawl is kept for a day; pass the result's crawl_id to keep_crawl_as_project to keep it for good.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The absolute http(s) URL to start from, e.g. the home page. | |
| limit | No | The most pages to read, 1 to 5,000 (default 50); the plan's page cap also applies. | |
| config | No | Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults. | |
| max_depth | No | How many links deep to follow from the start URL (0 = that page only; default 3). | |
| exclude_paths | No | Globs over the URL path to skip, e.g. ['/tag/*', '*.pdf']; an exclude wins over an include. | |
| include_paths | No | Globs over the URL path to keep, e.g. ['/blog/*'] for a section (its index included); a bare '/blog/' matches only that one page. Omit for the whole site. |