Skip to main content
Glama

Spicrawl

Scrape a URL

spicrawl_scrape

Retrieve a web page and return its content as markdown (default), text, HTML, a JSON envelope, or a printed PDF. Handles the fetch vs. browser decision, charset decoding, and PDF-to-text for you. Optionally extract structured data with CSS selectors (extract) or autoparse (a natural-language ai_extract is coming soon), return discovered links, take a screenshot, or scope the DOM with include_tags/exclude_tags. Each successful call is billed in credits (a failure costs 0). Results are cached by default, and a cache hit is billed like the fetch that stored it (the cache saves time, not credits); set cache=false for time-sensitive pages. With the defaults it only reads the page. A method other than GET, HEAD or OPTIONS, and the actions that click, fill or run scripts on the page, can change state on the target site: set them only when the user asks for that. For markdown, text and HTML it returns { content }. With format: "json", or when extract, ai_extract, autoparse, links, screenshot or network_capture is set, it returns the JSON envelope: content, the site's status, credits, engine, warnings and data for extraction. Screenshots come back as image content blocks; each screenshots entry in the JSON keeps its metadata and says which block holds it. A PDF comes back as a resource content block, with its size and the engine and credits in the JSON.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe absolute URL to retrieve.
modeNoRouting mode. `auto` lets the platform escalate fetch -> browser as needed, billing only the rung that worked.
waitNoMilliseconds to wait after load before capturing (render only).
cacheNoServe a recent cached result if fresh (on by default). A hit is billed at the same price as the fetch that stored it: the cache saves time, not credits. Set false only when you need live data.
linksNoAlso return the page's discovered links as an absolute, de-duplicated list.
proxyNoYour own proxy URL to egress through. Takes precedence over the pool; no surcharge.
engineNoPin the execution engine instead of letting the router choose. A pinned engine the deployment does not run is refused (ERR::ENGINE::UNAVAILABLE), never substituted. `camoufox` is coming soon: do not pin it yet.
formatNoOutput format. 'markdown' (default) is best for feeding to an LLM; 'html' returns the raw document; 'json' returns a structured envelope; 'pdf' prints the page in a browser and returns the file as a resource content block (needs `render` or `engine: "chromium"`).markdown
methodNoHTTP method sent to the target (default GET; no request body can be sent). Anything but GET, HEAD or OPTIONS can change state on the target site: use it only when the user asks.
renderNoRender with a real browser (executes JavaScript). Use for single-page apps or pages whose content is painted client-side. Costs more than the default fetch tier.
actionsNoBrowser steps run in order before capture (render only). Each step is {type, ...fields}. Examples: [{"type":"click","selector":"#load-more"},{"type":"wait_for","selector":".item:nth-child(20)","timeout_ms":5000}] · infinite scroll: [{"type":"scroll","to_bottom":true},{"type":"wait_for","ms":1000}] · login: [{"type":"fill","selector":"#user","value":"me"},{"type":"fill","selector":"#pass","value":"…","secret":true},{"type":"click","selector":"button[type=submit]","wait_for_navigation":true}]. A screenshot step switches the response to the JSON envelope (a FORMAT_COERCED warning says so).
extractNoSelector-based extraction: a map of field -> CSS/XPath selector, e.g. {"title":"h1","price":".amount"}. Returns structured `data`. Precise and cheap when you know the page structure.
stealthNoComing soon: not available yet, do not send. Stealth mode: render in the hardened Camoufox browser. Default false.
headlessNofalse runs Chromium on a real display (1920x1080 screen). chromium engine only.
max_costNoRefuse (rather than run) a request that would cost more than this many credits.
wait_forNoCSS selector to wait for before capturing (render only).
autoparseNoReturn the structured data the page publishes about itself (JSON-LD, OpenGraph, Twitter Card, microdata). No selectors; survives redesigns.
cache_ttlNoMax age in seconds of a cached copy to accept (0 forces a fresh fetch; capped at 48h). Freshness only: a hit is billed like a fetch.
parse_pdfNoTurn a PDF target into text/markdown (default true). false refuses PDFs.
ai_extractNoComing soon: not available yet, do not send. Model-driven extraction: describe what you want and let a model find it, no selectors. Refused with ERR::INTERNAL::UNAVAILABLE until it launches; use `extract` or `autoparse`.
screenshotNoCapture a screenshot (requires render).
session_idNoRun inside a session from spicrawl_session_create: same exit IP and browser state (cookies, localStorage) across calls.
sticky_keyNoComing soon: not available yet, do not send. Reuse the same managed-pool exit for every request carrying this key, without a session.
impersonateNoFetch tier only: present a real browser's TLS/JA3 + HTTP/2 fingerprint to clear passive bot gates. On by default; set false to send a plain client handshake.
exclude_tagsNoCSS selectors to remove (e.g. ['nav','.ad','#cookie']).
include_tagsNoCSS selectors to KEEP (everything else is dropped).
proxy_verifyNoVerify `proxy` is reachable before using it.
premium_proxyNoComing soon: not available yet, do not send. Use a residential exit from Spicrawl's managed pool instead of datacenter. Until then, pass your own `proxy`.
proxy_countryNoComing soon: not available yet, do not send. ISO-3166 alpha-2 country for a managed-pool exit, lowercase (e.g. 'us', 'de').
custom_headersNoExtra request headers sent to the target, e.g. {"Accept-Language":"de"}.
extract_presetNoComing soon: not available yet, do not send. A named extraction preset instead of an inline `extract` map; any value fails today. Put the rules in `extract`.
block_resourcesNoResource types the browser should not load, e.g. ["fonts","media","images"] (render only). Faster and cheaper.
network_captureNoRecord the API responses the PAGE fetched while rendering — often cleaner JSON than the DOM. Render only.
original_statusNoReturn the target's own HTTP status instead of 200 for a successful scrape.
wait_for_timeoutNoMax milliseconds to wait for `wait_for`.
main_content_onlyNoStrip the page to its main article, dropping nav/footer/aside. Defaults on for markdown.
screenshot_formatNoScreenshot image format.
screenshot_qualityNoScreenshot quality 1-100 (lossy formats).
screenshot_fullpageNoCapture the whole page, not just the viewport.
screenshot_selectorNoCapture only the element matching this CSS selector.
allowed_status_codesNoTarget statuses to treat as success instead of an error (e.g. [404] to capture a not-found page).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds the annotations: discloses credit billing (failures cost 0), the counterintuitive cache billing rule (a hit costs the same as the fetch), that pinning an unavailable engine is refused rather than substituted, and that non-GET methods or click/fill/script actions can mutate the target site. It also describes output shapes (image blocks for screenshots, resource block for PDF) that annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded: format/behavior first, then billing, then mutation warnings, then response shapes. Every sentence carries information, and the length is justified by 41 parameters and no output schema, though a few clauses (e.g. the screenshots metadata sentence) are harder to parse than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and 41 parameters, the description carries the return-value burden and does so completely: it defines the { content } shape for text formats, the full JSON envelope fields (content, status, credits, engine, warnings, data), and how screenshots and PDFs are returned. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond the schema: cache billing semantics, that screenshots switch the response to the JSON envelope, and that unset defaults keep the call read-only. It does not re-explain the many nested action fields, but those are well covered in the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (retrieve/scrape) and resource (web page) with the default output (markdown) and the full range of formats and extraction modes. It is unmistakably distinct from sibling tools like spicrawl_batch_submit or spicrawl_session_create, which manage sessions/batches rather than fetching a URL.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use context (render for SPAs, cache=false for time-sensitive pages, method/actions only when the user asks) and clear when-not warnings about state-changing operations. It does not, however, route the agent to alternatives such as the batch tools for multi-URL work, so it falls short of full 5-level guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources