Scrape a URL
spicrawl_scrapeRetrieve a web page and return its content as markdown (default), text, HTML, a JSON envelope, or a printed PDF. Handles the fetch vs. browser decision, charset decoding, and PDF-to-text for you. Optionally extract structured data with CSS selectors (extract) or autoparse (a natural-language ai_extract is coming soon), return discovered links, take a screenshot, or scope the DOM with include_tags/exclude_tags. Each successful call is billed in credits (a failure costs 0). Results are cached by default, and a cache hit is billed like the fetch that stored it (the cache saves time, not credits); set cache=false for time-sensitive pages. With the defaults it only reads the page. A method other than GET, HEAD or OPTIONS, and the actions that click, fill or run scripts on the page, can change state on the target site: set them only when the user asks for that. For markdown, text and HTML it returns { content }. With format: "json", or when extract, ai_extract, autoparse, links, screenshot or network_capture is set, it returns the JSON envelope: content, the site's status, credits, engine, warnings and data for extraction. Screenshots come back as image content blocks; each screenshots entry in the JSON keeps its metadata and says which block holds it. A PDF comes back as a resource content block, with its size and the engine and credits in the JSON.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The absolute URL to retrieve. | |
| mode | No | Routing mode. `auto` lets the platform escalate fetch -> browser as needed, billing only the rung that worked. | |
| wait | No | Milliseconds to wait after load before capturing (render only). | |
| cache | No | Serve a recent cached result if fresh (on by default). A hit is billed at the same price as the fetch that stored it: the cache saves time, not credits. Set false only when you need live data. | |
| links | No | Also return the page's discovered links as an absolute, de-duplicated list. | |
| proxy | No | Your own proxy URL to egress through. Takes precedence over the pool; no surcharge. | |
| engine | No | Pin the execution engine instead of letting the router choose. A pinned engine the deployment does not run is refused (ERR::ENGINE::UNAVAILABLE), never substituted. `camoufox` is coming soon: do not pin it yet. | |
| format | No | Output format. 'markdown' (default) is best for feeding to an LLM; 'html' returns the raw document; 'json' returns a structured envelope; 'pdf' prints the page in a browser and returns the file as a resource content block (needs `render` or `engine: "chromium"`). | markdown |
| method | No | HTTP method sent to the target (default GET; no request body can be sent). Anything but GET, HEAD or OPTIONS can change state on the target site: use it only when the user asks. | |
| render | No | Render with a real browser (executes JavaScript). Use for single-page apps or pages whose content is painted client-side. Costs more than the default fetch tier. | |
| actions | No | Browser steps run in order before capture (render only). Each step is {type, ...fields}. Examples: [{"type":"click","selector":"#load-more"},{"type":"wait_for","selector":".item:nth-child(20)","timeout_ms":5000}] · infinite scroll: [{"type":"scroll","to_bottom":true},{"type":"wait_for","ms":1000}] · login: [{"type":"fill","selector":"#user","value":"me"},{"type":"fill","selector":"#pass","value":"…","secret":true},{"type":"click","selector":"button[type=submit]","wait_for_navigation":true}]. A screenshot step switches the response to the JSON envelope (a FORMAT_COERCED warning says so). | |
| extract | No | Selector-based extraction: a map of field -> CSS/XPath selector, e.g. {"title":"h1","price":".amount"}. Returns structured `data`. Precise and cheap when you know the page structure. | |
| stealth | No | Coming soon: not available yet, do not send. Stealth mode: render in the hardened Camoufox browser. Default false. | |
| headless | No | false runs Chromium on a real display (1920x1080 screen). chromium engine only. | |
| max_cost | No | Refuse (rather than run) a request that would cost more than this many credits. | |
| wait_for | No | CSS selector to wait for before capturing (render only). | |
| autoparse | No | Return the structured data the page publishes about itself (JSON-LD, OpenGraph, Twitter Card, microdata). No selectors; survives redesigns. | |
| cache_ttl | No | Max age in seconds of a cached copy to accept (0 forces a fresh fetch; capped at 48h). Freshness only: a hit is billed like a fetch. | |
| parse_pdf | No | Turn a PDF target into text/markdown (default true). false refuses PDFs. | |
| ai_extract | No | Coming soon: not available yet, do not send. Model-driven extraction: describe what you want and let a model find it, no selectors. Refused with ERR::INTERNAL::UNAVAILABLE until it launches; use `extract` or `autoparse`. | |
| screenshot | No | Capture a screenshot (requires render). | |
| session_id | No | Run inside a session from spicrawl_session_create: same exit IP and browser state (cookies, localStorage) across calls. | |
| sticky_key | No | Coming soon: not available yet, do not send. Reuse the same managed-pool exit for every request carrying this key, without a session. | |
| impersonate | No | Fetch tier only: present a real browser's TLS/JA3 + HTTP/2 fingerprint to clear passive bot gates. On by default; set false to send a plain client handshake. | |
| exclude_tags | No | CSS selectors to remove (e.g. ['nav','.ad','#cookie']). | |
| include_tags | No | CSS selectors to KEEP (everything else is dropped). | |
| proxy_verify | No | Verify `proxy` is reachable before using it. | |
| premium_proxy | No | Coming soon: not available yet, do not send. Use a residential exit from Spicrawl's managed pool instead of datacenter. Until then, pass your own `proxy`. | |
| proxy_country | No | Coming soon: not available yet, do not send. ISO-3166 alpha-2 country for a managed-pool exit, lowercase (e.g. 'us', 'de'). | |
| custom_headers | No | Extra request headers sent to the target, e.g. {"Accept-Language":"de"}. | |
| extract_preset | No | Coming soon: not available yet, do not send. A named extraction preset instead of an inline `extract` map; any value fails today. Put the rules in `extract`. | |
| block_resources | No | Resource types the browser should not load, e.g. ["fonts","media","images"] (render only). Faster and cheaper. | |
| network_capture | No | Record the API responses the PAGE fetched while rendering — often cleaner JSON than the DOM. Render only. | |
| original_status | No | Return the target's own HTTP status instead of 200 for a successful scrape. | |
| wait_for_timeout | No | Max milliseconds to wait for `wait_for`. | |
| main_content_only | No | Strip the page to its main article, dropping nav/footer/aside. Defaults on for markdown. | |
| screenshot_format | No | Screenshot image format. | |
| screenshot_quality | No | Screenshot quality 1-100 (lossy formats). | |
| screenshot_fullpage | No | Capture the whole page, not just the viewport. | |
| screenshot_selector | No | Capture only the element matching this CSS selector. | |
| allowed_status_codes | No | Target statuses to treat as success instead of an error (e.g. [404] to capture a not-found page). |