Read web pages
writ_scrapeREAD PAGE CONTENT NOW: one page, a list of pages, or the top N items of a listing, returned as clean markdown in this call (2-10s). For: 'what does say', 'summarize ', 'the top N posts/products/results of and what's on each', 'fetch these 3 links', including a page a plain fetch cannot read: blocked or empty (403, bot wall), rendered by JavaScript, or behind the user's sign-in (render_mode, use_residential, persona_id). NOT FOR: collecting a whole site or section into a dataset (writ_crawl_site); clicking, typing, signing in or any action on a page (writ_browser_use).
Three shapes, one call each:
url → that page.
urls=[...] (≤20) → all of them, fetched in parallel,
pagesin the order given.url= + top_n=N (≤20) → the listing (
listing) and the N top-ranked item pages it links to (pages, in rank order); e.g. url='https://news.ycombinator.com/', top_n=3 returns the front page and the 3 top stories' discussion pages. include_paths=['item\?id='] gives the item-link shape; without it the server detects it.
Discussion pages keep their comment threads; each comment is tagged [top-level] or [reply · depth N], so 'the top-level comments' is answerable from the text. Long pages are preview-cut at 12000 chars (_truncated); the hint says how to fetch a full page. The caller reads the markdown; no AI is spent here. Behind a login: persona_id. Bot wall: use_residential=true.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | A page to read or, with `top_n`, the listing page whose top items are read. | |
| urls | No | Known pages to read together (max 20), fetched in parallel in one call. | |
| top_n | No | With `url` = a listing/front/search/category page: also read its N top-ranked item pages (the page's link order is the ranking). One call, parallel. | |
| device | No | A linked Writ desktop's agent_id (writ_devices) to act on. Omitted: the desktop this connection chose with writ_devices action='use', if any. | |
| format | No | markdown (default): the page's clean main content. html: the raw HTML as fetched (same egress/persona/render) — for selectors, embedded JSON, anything the cleaned text drops; clipped to preview_chars (default 40000, 0 = whole page). both: html plus the markdown derived from it, one fetch. Single `url` only. | |
| persona_id | No | Saved identity to read as (writ_personas lists them), for pages behind a login. Forces the identity's own residential exit. A desktop persona ('device:…') reads on its own desktop, from that machine. | |
| render_mode | No | auto (default: plain HTTP, browser only if the page needs JS) | http | browser. | |
| include_paths | No | With top_n: regex(es) the item links match (e.g. 'item\\?id=', '/products/'). Optional: the server detects detail links when omitted. | |
| preview_chars | No | Cut each page's text to this many characters (default 12000; 0 = full pages). | |
| respect_robots | No | Apply robots.txt to the explicitly requested page(s). Default false for scrape; writ_crawl_site defaults true for autonomous discovery. | |
| use_residential | No | Fetch through the platform residential network (premium) for a site that blocks datacenter IPs or shows a bot wall. Default off. | |
| residential_country | No | Two-letter ISO country the residential exit should be in (e.g. 'us', 'fr'), also used by the automatic residential retry on a blocked page. Omitted = an automatic exit. Ignored unless the session egresses residential. |