Skip to main content
Glama

Crawl a site into a dataset

writ_crawl_site
Destructive

Crawl a full site or a section into a queryable, change-tracked page dataset you can query, export, or monitor later.

Instructions

COLLECT A SITE (or a section of it) INTO A DATASET — a distributed Dragnet crawl that discovers pages and stores every one as a queryable, change-tracked row. This is the tool for 'crawl ', 'get every page of the docs', 'all products in this category', 'build a dataset of ', or anything that will be queried, exported, monitored or re-run later. NOT FOR: reading a page or a handful of pages right now — that is writ_scrape (url / urls / top_n answers in one call, no dataset); acting on a page (writ_browser_use).

CHOOSE THE MODE — all three fetch pages the same way; they differ in who READS each page:

  • CLASSIC (default: extract_mode='markdown', executor='regular') — every page becomes clean markdown, no AI spent, fastest. Right for content, docs, articles, discussions (threads keep [top-level]/[reply · depth N] tags). Pick this unless a rule below applies.

  • SCHEMA (extract_mode='schema' + extract_schema) — every page holds the SAME structured record (a product, a listing row) and you want rows, not prose. Deterministic CSS extraction, no AI.

  • AI-ASSISTED (executor='ai' + extract_prompt) — the wanted fields need understanding and vary per page (sentiment, pros/cons, a classification, free-form values with no stable selector) across MANY pages. Each page waits on a model call (~10s) and bills 5x the page rate — never use it for a few pages you could read yourself, and never to 'be sure'.

SCOPE IT — an unscoped crawl of a real site collects hundreds of nav, tag and pagination pages and bills for every one. Match the ask to a shape:

  • A SECTION ('the docs', 'the pricing and blog pages'): pass intent in plain language — the server derives include/exclude paths and depth from the site's real URLs — and relevance_threshold ≈0.3 to drop off-goal pages.

  • KNOWN PAGES as a dataset: seed_urls (no discovery). For an immediate answer use writ_scrape(urls) instead.

  • TOP-N of a listing as a dataset (re-run later, monitored): rank_cap=N. For an immediate answer use writ_scrape(url, top_n) instead.

  • WHOLE SITE ('every page'): the defaults; set page_budget to cap the spend.

DELIVERY: a bounded crawl (rank_cap / seed_urls) waits and returns its pages in data.rows in this call; an open site crawl returns a crawl id to poll with writ_crawl_status — results land as a workflow dataset (writ_workflow_data, writ_search_data, writ_export_data). If the ask mentions comments or discussion, set content_spec {"preset": "full", "include_comments": true} (rank_cap crawls do this already). Behind a login: persona_id — never sign in yourself. save_as ONLY when the user will re-run it; writ_saved_crawls lists those — re-run one only when its scope matches the ask.

YOU OWN THE RESPONSE SHAPE. Every answer is Writ's envelope by default (definition + crawl status + a data table whose rows wrap fields in run bookkeeping, and whose records carry page metadata like content_kind/depth). When you are BUILDING AN API on a crawl — anything a program or the user will consume — set output so the answer is THEIR shape: {shape:'record'} for one entity (a usage meter, a dashboard), {shape:'records'} for a list, fields to pick/rename ('percent_used as pct', dotted paths), exclude to drop, key to wrap. Page metadata is stripped unless include_meta=true. Saved with save_as, it becomes the API's default shape (override per call on writ_run_saved_crawl / writ_saved_crawl_data). Do NOT try to prompt the metadata away in extract_prompt — it is added after the model answers; output is the fix.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesSeed URL (required).
nameNo
waitNoBlock until the crawl converges and return the collected pages IN THIS CALL (`data.rows`). Default TRUE for a bounded crawl (rank_cap or seed_urls — a few pages, seconds) and false for an open site crawl (returns a crawl id to poll with writ_crawl_status). Past the 75s ceiling you get a 504 that still carries the crawl id.
limitNoRows of collected data to return when wait=true (default 50).
speedNoThroughput tier: slow | normal (default) | fast — how much of your parallel-agent allowance the crawl uses. slow ≈ ¼ at a discounted page rate, normal ≈ ½ at standard rate, fast = all of it at a premium.
deviceNoA linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any).
intentNoPlain-English goal. The server derives include/exclude paths and a depth from it against a sample of the site's real URLs, and ranks the frontier by relevance — so on an unfamiliar site this beats guessing path regexes yourself.
outputNoRESPONSE SHAPE — set this whenever the answer is for a program or an API you are building, not for you to read. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone — one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; missing → null so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails are STRIPPED unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is SAVED as the API's default shape.
max_ageNoOnly meaningful with `save_as`: if that saved crawl already completed within this many seconds, return its collected data instead of crawling again. 0 always crawls.
save_asNoONLY when the user will want to re-run this crawl later (a recurring pull, an API they asked for): saves these settings as a named, callable crawl. A one-off question does NOT need one — saved crawls are listed to every future session as 'already collected', so a saved one-off misleads the next agent. Reusing the same name updates that saved crawl instead of creating a duplicate.
delay_msNoPoliteness delay between fetches per host (default 250).
executorNoregular (default) = deterministic crawl, no AI. ai = a fleet of AI agents reads every page against `extract_prompt` and returns structured records — for data with no clean CSS selector. Bills at 5x the page rate.
ocr_modeNoauto (default) | off | force
rank_capNoTOP-N ASK — set this whenever the user wants the top/first N items from a listing (a front page, search results, a category). The server reads the seed page's link order (which IS the ranking), seeds exactly those N item pages, and pins the crawl to them — one page per agent, in parallel. WITHOUT it the same request becomes a breadth crawl that mostly collects nav and pagination and does not answer the question. Pair with include_paths when you know the item-link shape.
max_depthNo
seed_urlsNoExact pages to start from, when you already know them — the crawl collects these instead of discovering its own. Cheapest way to scrape a known set.
persona_idNoSaved identity to crawl AS (list them with writ_personas) — for pages behind a login. Every shard shares the persona's signed-in session and one sticky exit IP. 2FA is minted server-side. A desktop persona ('device:…') crawls on its own desktop, from that machine.
shard_sizeNoURLs fetched per shard batch (default 20).
page_budgetNo
render_modeNoHow each page is FETCHED — independent of `executor`, which decides who READS it. auto (default) = plain HTTP first, warm browser only for JS-challenge or near-empty pages; http = never open a browser (fastest, static HTML); browser = warm-render every page (JS/SPA sites). executor=ai works on either lane.
same_domainNo
content_specNoWhich ELEMENTS of each page to keep: {preset: 'full'|'main', include_comments: bool, exclude_selectors: [css], include_selectors: [css], keep: {images: bool}}. 'main' = article body only; 'full' = the whole page INCLUDING comment and discussion threads — use 'full' with include_comments when the ask mentions comments, replies or discussion, or they will be stripped out.
extract_modeNomarkdown (default) | schema (uniform records via extract_schema) | html (each page's RAW HTML, for selectors or embedded JSON)
exclude_pathsNo
include_pathsNo
preview_charsNoCut each inline page's text cells to this many characters (default 12000; 0 = full pages). Cut rows list the fields under `_truncated`; fetch a full page with writ_workflow_data(workflow_id=<data_workflow_id>, refs=['<run_id>:<record_index>']).
extract_promptNoRequired with executor=ai: what each agent should extract from each page, in plain language (e.g. 'the product name, price and SKU').
extract_schemaNo
respect_robotsNoHonor robots.txt (default true).
timeout_secondsNoMax seconds to hold when wait=true (≤75).
use_residentialNoRoute every shard through the platform residential network (premium). Turn on for sites that block datacenter IPs / show a bot wall — a persona crawl forces it on automatically. Costs residential bandwidth; default off.
allow_subdomainsNo
relevance_thresholdNo0-1. Score every discovered page against `intent` and SKIP anything below the bar, so a broad crawl collects only what the goal needs (≈0.3 for 'the pricing and docs pages'). Leave unset for a whole-site sweep.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — pins the exit pool's geo for every shard. Omit for an automatic exit. Ignored unless the session egresses residential.
max_concurrent_shardsNoExplicit parallel-shard cap; overrides the `speed` allocation.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a write, non-idempotent, open-world, destructive operation, and the description adds substantial context beyond that: cost multipliers (AI executor 5x page rate), spend-capping via page_budget, wait/504 behavior, dataset persistence, login handling through persona_id, and re-run semantics for saved crawls. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly organized with front-loaded caps headings (NOT FOR, CHOOSE THE MODE, SCOPE IT, DELIVERY) that let an agent navigate quickly. Given the tool's 35-parameter complexity the length is largely justified, though some expanded mode explanations could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, the description explains the return envelope, data.rows behavior for bounded vs. open crawls, polling with writ_crawl_status, and how to reshape output for API consumers. This covers what an agent needs to call the tool correctly and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 35 parameters at ~77% schema coverage, the description adds decision-level meaning for the most consequential knobs: intent, rank_cap, seed_urls, extract_mode/executor interplay, output shape configuration, save_as, and content_spec. This meaning is not derivable from the schema descriptions alone and directly shapes correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('COLLECT A SITE ... INTO A DATASET') with clear scope, and explicitly contrasts itself against writ_scrape and writ_browser_use. An agent can identify this as the crawl-to-dataset tool without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use guidance (crawl <site>, build a dataset), when-not-to-use (one-off page reads, page actions), and names the alternatives (writ_scrape, writ_browser_use). It further routes between three execution modes and several scoping shapes with concrete conditions for each.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.