Skip to main content
Glama

Writ Cloud

Crawl a site into a dataset

writ_crawl_site
Destructive

COLLECT A SITE (or a section of it) into a dataset: a distributed Dragnet crawl that discovers pages and stores each one as a queryable, change-tracked row. For: 'crawl ', 'every page of the docs', 'all products in this category', 'a dataset of ', or anything queried, exported, monitored or re-run later. NOT FOR: reading one or a few pages now, which is writ_scrape (url / urls / top_n answer in one call, no dataset); acting on a page, which is writ_browser_use.

Modes: all three fetch pages the same way and differ in who reads each page.

  • CLASSIC (default: extract_mode='markdown', executor='regular'): every page becomes clean markdown, no AI spent, fastest. Fits content, docs, articles, discussions (threads keep [top-level]/[reply · depth N] tags), and any ask the two modes below do not cover.

  • SCHEMA (extract_mode='schema' + extract_schema): every page holds the same structured record (a product, a listing row), returned as rows, not prose. Deterministic CSS extraction, no AI.

  • AI-ASSISTED (executor='ai' + extract_prompt): fields that need understanding and vary per page (sentiment, pros/cons, a classification, free-form values with no stable selector) across many pages. Each page waits on a model call (~10s) and bills 5x the page rate; on a few pages, or on fields the other modes capture, it adds cost and no accuracy.

Scope: an unscoped crawl of a real site collects hundreds of nav, tag and pagination pages and bills for each. Shapes:

  • A section ('the docs', 'the pricing and blog pages'): intent in plain language (the server derives include/exclude paths and depth from the site's real URLs) plus relevance_threshold ≈0.3, which drops off-goal pages.

  • Known pages as a dataset: seed_urls (no discovery). The immediate answer is writ_scrape(urls).

  • Top N of a listing as a dataset (re-run later, monitored): rank_cap=N. The immediate answer is writ_scrape(url, top_n).

  • Whole site ('every page'): the defaults; page_budget caps the spend.

Delivery: a bounded crawl (rank_cap / seed_urls) waits and returns its pages in data.rows in this call; an open site crawl returns a crawl id for writ_crawl_status, and its results land as a workflow dataset (writ_workflow_data, writ_search_data, writ_export_data). Comment and discussion threads are kept by content_spec {"preset": "full", "include_comments": true} (rank_cap crawls set it already). Behind a login: persona_id, whose saved session the crawl uses. save_as keeps the crawl for re-running; writ_saved_crawls lists those, each answering only the asks its scope covers.

Response shape: by default every answer is Writ's envelope (definition + crawl status + a data table whose rows wrap fields in run bookkeeping, and whose records carry page metadata like content_kind/depth). output gives an API built on a crawl its consumer's shape: {shape:'record'} for one entity (a usage meter, a dashboard), {shape:'records'} for a list, fields to pick/rename ('percent_used as pct', dotted paths), exclude to drop, key to wrap. Page metadata is stripped unless include_meta=true. Saved with save_as, it becomes the API's default shape (overridable per call on writ_run_saved_crawl / writ_saved_crawl_data). The metadata is added after the extract_prompt model answers, so output removes it and the prompt cannot.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesSeed URL (required).
nameNo
waitNoHold until the crawl converges and return the collected pages in this call (`data.rows`). Default true for a bounded crawl (rank_cap or seed_urls: a few pages, seconds) and false for an open site crawl (returns a crawl id for writ_crawl_status). Past the 75s ceiling the answer is a 504 that still carries the crawl id.
limitNoRows of collected data to return when wait=true (default 50).
speedNoThroughput tier: slow | normal (default) | fast: the share of the parallel-agent allowance the crawl uses. slow ≈ ¼ at a discounted page rate, normal ≈ ½ at standard rate, fast = all of it at a premium.
deviceNoA linked Writ desktop's agent_id (writ_devices) to act on. Omitted: the desktop this connection chose with writ_devices action='use', if any.
intentNoPlain-English goal. The server derives include/exclude paths and a depth from it against a sample of the site's real URLs, and ranks the frontier by relevance, so on an unfamiliar site it scopes better than guessed path regexes.
outputNoResponse shape, for an answer a program or an API consumes rather than a reader. {shape: 'envelope' (default: Writ's full answer, projected) | 'table' ({columns, rows, total}) | 'records' (bare list of records) | 'record' (the newest record alone: one entity, a usage meter, a dashboard), fields: ['used', 'percent_used as pct', 'items.0.price as first_price'] (ordered pick, renames, dotted paths; a missing path is null, so keys are stable), exclude: ['depth'], include_meta: false (page metadata content_kind/depth/thumbnails is stripped unless true), key: 'usage' (wrap)}. On writ_crawl_site with save_as it is saved as the API's default shape.
max_ageNoOnly meaningful with `save_as`: if that saved crawl already completed within this many seconds, return its collected data instead of crawling again. 0 always crawls.
save_asNoSaves these settings as a named, callable crawl, for one that is re-run later (a recurring pull, an API the user asked for). Saved crawls are listed to every future session as 'already collected', so a saved one-off reads as reusable data to later sessions. Reusing the same name updates that saved crawl instead of creating a duplicate.
delay_msNoPoliteness delay between fetches per host (default 250).
executorNoregular (default) = deterministic crawl, no AI. ai = a fleet of AI agents reads every page against `extract_prompt` and returns structured records, for data with no clean CSS selector. Bills at 5x the page rate.
ocr_modeNoauto (default) | off | force
rank_capNoTop N items of a listing (a front page, search results, a category). The server reads the seed page's link order (which is the ranking), seeds exactly those N item pages, and pins the crawl to them, one page per agent, in parallel. Without it the same request is a breadth crawl that mostly collects nav and pagination and does not answer a top-N ask. include_paths, when given, is the item-link shape.
max_depthNo
seed_urlsNoExact known pages to start from: the crawl collects these instead of discovering its own. The cheapest way to scrape a known set.
persona_idNoSaved identity to crawl as (writ_personas lists them), for pages behind a login. Every shard shares the persona's signed-in session and one sticky exit IP. 2FA is minted server-side. A desktop persona ('device:…') crawls on its own desktop, from that machine.
shard_sizeNoURLs fetched per shard batch (default 20).
page_budgetNo
render_modeNoHow each page is fetched, independent of `executor`, which decides who reads it. auto (default) = plain HTTP first, warm browser only for JS-challenge or near-empty pages; http = never opens a browser (fastest, static HTML); browser = warm-render every page (JS/SPA sites). executor=ai works on either lane.
same_domainNo
content_specNoWhich elements of each page to keep: {preset: 'full'|'main', include_comments: bool, exclude_selectors: [css], include_selectors: [css], keep: {images: bool}}. 'main' = article body only; 'full' = the whole page including comment and discussion threads. Comments, replies and discussion survive only with 'full' plus include_comments; otherwise they are stripped out.
extract_modeNomarkdown (default) | schema (uniform records via extract_schema) | html (each page's raw HTML, for selectors or embedded JSON)
exclude_pathsNo
include_pathsNo
preview_charsNoCut each inline page's text cells to this many characters (default 12000; 0 = full pages). Cut rows list the fields under `_truncated`; a full page comes from writ_workflow_data(workflow_id=<data_workflow_id>, refs=['<run_id>:<record_index>']).
extract_promptNoRequired with executor=ai: what each agent extracts from each page, in plain language (e.g. 'the product name, price and SKU').
extract_schemaNo
respect_robotsNoHonor robots.txt (default true).
timeout_secondsNoMax seconds to hold when wait=true (≤75).
use_residentialNoRoute every shard through the platform residential network (premium), for sites that block datacenter IPs / show a bot wall; a persona crawl forces it on automatically. Costs residential bandwidth; default off.
allow_subdomainsNo
relevance_thresholdNo0-1. Every discovered page is scored against `intent` and skipped below the bar, so a broad crawl collects only what the goal needs (≈0.3 for 'the pricing and docs pages'). Unset for a whole-site sweep.
residential_countryNoTwo-letter ISO country the residential exit should be in (e.g. 'us', 'fr'): pins the exit pool's geo for every shard. Omitted = an automatic exit. Ignored unless the session egresses residential.
max_concurrent_shardsNoExplicit parallel-shard cap; overrides the `speed` allocation.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

Score is being calculated.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.