scrape
Convert any URL into clean, LLM-ready markdown. Bypasses anti-bot blocks with automatic fallback methods to extract readable content for analysis, summarization, or extraction.
Instructions
Scrape a single URL → LLM-ready markdown.
Use this when the user shares a URL and wants its content (read, analyze, summarize, extract). Auto-escalates through fast→stealth→llm when blocked.
Args:
url: Target URL (http/https).
prefer: "auto" | "fast" | "stealth" | "llm".
auto = fast first, escalate to stealth on block/short page.
fast = cheap HTTP only (no JS).
stealth = real Chromium + Cloudflare solver.
llm = full Crawl4AI browser + BM25 fit-markdown.
timeout: per-attempt timeout in seconds.
include_html: include raw HTML in the response (large; off by default).
js: (stealth only) JS expression evaluated against the live page
after it settles. The value comes back in meta.js_result.
Use for data that lives in DOM properties (e.g. an input's
.value) rather than in serialized HTML.
wait_for: (stealth only) JS predicate expression polled until truthy
(bounded by timeout). Use to wait for content that arrives
asynchronously after network_idle.
Returns: {url, final_url, status, markdown, title, method, elapsed_ms, meta} or {error, url, method} on failure.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| js | No | ||
| url | Yes | ||
| prefer | No | auto | |
| timeout | No | ||
| wait_for | No | ||
| include_html | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||