Crawl a site into a dataset
writ_crawl_siteCrawl a full site or a section into a queryable, change-tracked page dataset you can query, export, or monitor later.
Instructions
COLLECT A SITE (or a section of it) INTO A DATASET — a distributed Dragnet crawl that discovers pages and stores every one as a queryable, change-tracked row. This is the tool for 'crawl ', 'get every page of the docs', 'all products in this category', 'build a dataset of ', or anything that will be queried, exported, monitored or re-run later. NOT FOR: reading a page or a handful of pages right now — that is writ_scrape (url / urls / top_n answers in one call, no dataset); acting on a page (writ_browser_use).
CHOOSE THE MODE — all three fetch pages the same way; they differ in who READS each page:
CLASSIC (default: extract_mode='markdown', executor='regular') — every page becomes clean markdown, no AI spent, fastest. Right for content, docs, articles, discussions (threads keep [top-level]/[reply · depth N] tags). Pick this unless a rule below applies.
SCHEMA (extract_mode='schema' + extract_schema) — every page holds the SAME structured record (a product, a listing row) and you want rows, not prose. Deterministic CSS extraction, no AI.
AI-ASSISTED (executor='ai' + extract_prompt) — the wanted fields need understanding and vary per page (sentiment, pros/cons, a classification, free-form values with no stable selector) across MANY pages. Each page waits on a model call (~10s) and bills 5x the page rate — never use it for a few pages you could read yourself, and never to 'be sure'.
SCOPE IT — an unscoped crawl of a real site collects hundreds of nav, tag and pagination pages and bills for every one. Match the ask to a shape:
A SECTION ('the docs', 'the pricing and blog pages'): pass
intentin plain language — the server derives include/exclude paths and depth from the site's real URLs — andrelevance_threshold≈0.3 to drop off-goal pages.KNOWN PAGES as a dataset:
seed_urls(no discovery). For an immediate answer use writ_scrape(urls) instead.TOP-N of a listing as a dataset (re-run later, monitored):
rank_cap=N. For an immediate answer use writ_scrape(url, top_n) instead.WHOLE SITE ('every page'): the defaults; set
page_budgetto cap the spend.
DELIVERY: a bounded crawl (rank_cap / seed_urls) waits and returns its pages in data.rows in this call; an open site crawl returns a crawl id to poll with writ_crawl_status — results land as a workflow dataset (writ_workflow_data, writ_search_data, writ_export_data). If the ask mentions comments or discussion, set content_spec {"preset": "full", "include_comments": true} (rank_cap crawls do this already). Behind a login: persona_id — never sign in yourself. save_as ONLY when the user will re-run it; writ_saved_crawls lists those — re-run one only when its scope matches the ask.
YOU OWN THE RESPONSE SHAPE. Every answer is Writ's envelope by default (definition + crawl status + a data table whose rows wrap fields in run bookkeeping, and whose records carry page metadata like content_kind/depth). When you are BUILDING AN API on a crawl — anything a program or the user will consume — set output so the answer is THEIR shape: {shape:'record'} for one entity (a usage meter, a dashboard), {shape:'records'} for a list, fields to pick/rename ('percent_used as pct', dotted paths), exclude to drop, key to wrap. Page metadata is stripped unless include_meta=true. Saved with save_as, it becomes the API's default shape (override per call on writ_run_saved_crawl / writ_saved_crawl_data). Do NOT try to prompt the metadata away in extract_prompt — it is added after the model answers; output is the fix.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Seed URL (required). | |
| name | No | ||
| wait | No | Block until the crawl converges and return the collected pages IN THIS CALL (`data.rows`). Default TRUE for a bounded crawl (rank_cap or seed_urls — a few pages, seconds) and false for an open site crawl (returns a crawl id to poll with writ_crawl_status). Past the 75s ceiling you get a 504 that still carries the crawl id. | |
| limit | No | Rows of collected data to return when wait=true (default 50). | |
| speed | No | Throughput tier: slow | normal (default) | fast — how much of your parallel-agent allowance the crawl uses. slow ≈ ¼ at a discounted page rate, normal ≈ ½ at standard rate, fast = all of it at a premium. | |
| device | No | A linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any). | |
| intent | No | Plain-English goal. The server derives include/exclude paths and a depth from it against a sample of the site's real URLs, and ranks the frontier by relevance — so on an unfamiliar site this beats guessing path regexes yourself. | |
| output | No | RESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape. | |
| max_age | No | Only meaningful with `save_as`: if that saved crawl already completed within this many seconds, return its collected data instead of crawling again. 0 always crawls. | |
| save_as | No | ONLY when the user will want to re-run this crawl later (a recurring pull, an API they asked for): saves these settings as a named, callable crawl. A one-off question does NOT need one — saved crawls are listed to every future session as 'already collected', so a saved one-off misleads the next agent. Reusing the same name updates that saved crawl instead of creating a duplicate. | |
| delay_ms | No | Politeness delay between fetches per host (default 250). | |
| executor | No | regular (default) = deterministic crawl, no AI. ai = a fleet of AI agents reads every page against `extract_prompt` and returns structured records — for data with no clean CSS selector. Bills at 5x the page rate. | |
| ocr_mode | No | auto (default) | off | force | |
| rank_cap | No | TOP-N ASK — set this whenever the user wants the top/first N items from a listing (a front page, search results, a category). The server reads the seed page's link order (which IS the ranking), seeds exactly those N item pages, and pins the crawl to them — one page per agent, in parallel. WITHOUT it the same request becomes a breadth crawl that mostly collects nav and pagination and does not answer the question. Pair with include_paths when you know the item-link shape. | |
| max_depth | No | ||
| seed_urls | No | Exact pages to start from, when you already know them — the crawl collects these instead of discovering its own. Cheapest way to scrape a known set. | |
| persona_id | No | Saved identity to crawl AS (list them with writ_personas) — for pages behind a login. Every shard shares the persona's signed-in session and one sticky exit IP. 2FA is minted server-side. A desktop persona ('device:…') crawls on its own desktop, from that machine. | |
| shard_size | No | URLs fetched per shard batch (default 20). | |
| page_budget | No | ||
| render_mode | No | How each page is FETCHED — independent of `executor`, which decides who READS it. auto (default) = plain HTTP first, warm browser only for JS-challenge or near-empty pages; http = never open a browser (fastest, static HTML); browser = warm-render every page (JS/SPA sites). executor=ai works on either lane. | |
| same_domain | No | ||
| content_spec | No | Which ELEMENTS of each page to keep: {preset: 'full'|'main', include_comments: bool, exclude_selectors: [css], include_selectors: [css], keep: {images: bool}}. 'main' = article body only; 'full' = the whole page INCLUDING comment and discussion threads — use 'full' with include_comments when the ask mentions comments, replies or discussion, or they will be stripped out. | |
| extract_mode | No | markdown (default) | schema (uniform records via extract_schema) | html (each page's RAW HTML, for selectors or embedded JSON) | |
| exclude_paths | No | ||
| include_paths | No | ||
| preview_chars | No | Cut each inline page's text cells to this many characters (default 12000; 0 = full pages). Cut rows list the fields under `_truncated`; fetch a full page with writ_workflow_data(workflow_id=<data_workflow_id>, refs=['<run_id>:<record_index>']). | |
| extract_prompt | No | Required with executor=ai: what each agent should extract from each page, in plain language (e.g. 'the product name, price and SKU'). | |
| extract_schema | No | ||
| respect_robots | No | Honor robots.txt (default true). | |
| timeout_seconds | No | Max seconds to hold when wait=true (≤75). | |
| use_residential | No | Route every shard through the platform residential network (premium). Turn on for sites that block datacenter IPs / show a bot wall — a persona crawl forces it on automatically. Costs residential bandwidth; default off. | |
| allow_subdomains | No | ||
| relevance_threshold | No | 0-1. Score every discovered page against `intent` and SKIP anything below the bar, so a broad crawl collects only what the goal needs (≈0.3 for 'the pricing and docs pages'). Leave unset for a whole-site sweep. | |
| residential_country | No | Two-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — pins the exit pool's geo for every shard. Omit for an automatic exit. Ignored unless the session egresses residential. | |
| max_concurrent_shards | No | Explicit parallel-shard cap; overrides the `speed` allocation. |