Graph-based website traversal tool using Tavily Crawl.
post_tavily_crawlWalk a site from a root url and return the content of the pages it finds. Steer it with natural-language instructions plus regex path and domain filters, and bound it with max_depth, max_breadth and limit. Returns base_url and results[] with url and raw_content. Measured at about 4.5 seconds for a 3-page limit; cost and time grow with the bounds you set, so set them. Use it for broad coverage of one site — documentation, a catalogue, a competitor's blog. It answers synchronously, which post_firecrawl_crawl does not: that one runs as a background job and suits crawls too large to wait on. For a handful of known pages post_tavily_extract is far cheaper; to size a site before paying to crawl it, run post_tavily_map first.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The root URL to begin the crawl. | |
| limit | No | Total number of links the crawler will process before stopping. | |
| format | No | Format of the extracted web page content. | markdown |
| timeout | No | Maximum time in seconds to wait for the crawl operation. | |
| max_depth | No | Max depth of the crawl. | |
| max_breadth | No | Max number of links to follow per level of the tree. | |
| instructions | No | Natural language instructions for the crawler. | |
| select_paths | No | Regex patterns to select only URLs with specific path patterns. | |
| exclude_paths | No | Regex patterns to exclude URLs with specific path patterns. | |
| extract_depth | No | Depth of the extraction process. | basic |
| include_usage | No | Include credit usage information in the response. | |
| allow_external | No | Include external domain links in the final results list. | |
| include_images | No | Include images in the crawl results. | |
| select_domains | No | Regex patterns to select crawling to specific domains or subdomains. | |
| exclude_domains | No | Regex patterns to exclude specific domains or subdomains from crawling. | |
| include_favicon | No | Include the favicon URL for each result. | |
| chunks_per_source | No | Maximum number of relevant chunks returned per source. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||