Read web pages
writ_scrapeRead web pages as clean markdown, including multiple URLs or top N items from a listing, even when blocked, JavaScript-rendered, or behind a login.
Instructions
READ PAGE CONTENT NOW — one page, a list of pages, or the top N items of a listing — returned as clean markdown IN THIS CALL (2-10s). This is the tool for 'what does say', 'summarize ', 'the top N posts/products/results of and what's on each', 'fetch these 3 links' — including a page a plain fetch cannot read: blocked or empty (403, bot wall), rendered by JavaScript, or behind the user's sign-in (render_mode, use_residential, persona_id). NOT FOR: collecting a whole site or section into a dataset (writ_crawl_site); clicking, typing, signing in or any action on a page (writ_browser_use).
THREE SHAPES, ONE CALL EACH:
url → that page.
urls=[...] (≤20) → all of them, fetched in parallel,
pagesin the order given.url= + top_n=N (≤20) → the listing (
listing) AND the N top-ranked item pages it links to (pages, in rank order) — e.g. url='https://news.ycombinator.com/', top_n=3 returns the front page and the 3 top stories' discussion pages. Add include_paths=['item\?id='] when you know the item-link shape; the server otherwise detects it.
Discussion pages keep their comment threads; each comment is tagged [top-level] or [reply · depth N], so 'the top-level comments' is answerable from the text. Long pages are preview-cut at 12000 chars (_truncated); the hint tells you how to fetch a full page. YOU read the markdown — no AI is spent here. Behind a login: persona_id. Bot wall: use_residential=true.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | A page to read — or, with `top_n`, the LISTING page whose top items to read. | |
| urls | No | Known pages to read together (max 20), fetched in parallel in one call. | |
| top_n | No | With `url` = a listing/front/search/category page: also read its N top-ranked item pages (the page's link order IS the ranking). One call, parallel. | |
| device | No | A linked Writ desktop's agent_id (writ_devices): act ON it. Omit to use the desktop this connection chose with writ_devices action='use' (if any). | |
| format | No | markdown (default): the page's clean main content. html: the RAW HTML as fetched (same egress/persona/render) — for selectors, embedded JSON, anything the cleaned text drops; clipped to preview_chars (default 40000, 0 = whole page). both: html plus the markdown derived from it, one fetch. Single `url` only. | |
| persona_id | No | Saved identity to read AS (list them with writ_personas) — for pages behind a login. Forces the identity's own residential exit. A desktop persona ('device:…') reads on its own desktop, from that machine. | |
| render_mode | No | auto (default: plain HTTP, browser only if the page needs JS) | http | browser. | |
| include_paths | No | With top_n: regex(es) the item links match (e.g. 'item\\?id=', '/products/'). Optional — the server detects detail links when omitted. | |
| preview_chars | No | Cut each page's text to this many characters (default 12000; 0 = full pages). | |
| respect_robots | No | Apply robots.txt to the explicitly requested page(s). Default false for scrape; writ_crawl_site defaults true for autonomous discovery. | |
| use_residential | No | Fetch through the platform residential network (premium) for a site that blocks datacenter IPs or shows a bot wall. Default off. | |
| residential_country | No | Two-letter ISO country the residential exit should be in (e.g. 'us', 'fr') — also used by the automatic residential retry on a blocked page. Omit for an automatic exit. Ignored unless the session egresses residential. |