site_crawl
Read-only
Polite same-origin crawl of a public site (robots.txt honored; sitemap.xml first, else shallow link BFS). Returns per-page {url, status, title, text_markdown, content_sha256, fetched_at}, totals, and a crawl-level sha256. SSRF-guarded. $0.002/page (min $0.01); failures not billable.
Input Schema
TableJSON Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Public http(s) start URL. | |
| max_pages | No | Pages to fetch (cap 50). | |
| same_origin | No | Only true is supported. |