Scrape a list of URLs
scrape_urlsFetch known URLs (1 to 500) once and return each page's content, markdown by default; no project is created. For one page with every format or browser steps use extract_url; to find pages by following links use crawl_site; to list a site's URLs without fetching them use map_site. Each page costs the credits of the engine that read it (1 plain fetch, 4 browser render), and a page the site refuses is free. A single URL answers in the same request; a list waits for the batch, and a large one comes back as an index with excerpts inside 60,000 characters. Hosted, a batch still going after the time budget comes back as a job for get_job, and repeating the same call returns that job instead of starting another.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | The pages to fetch: 1 to 500 absolute http(s) URLs. | |
| config | No | Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults. | |
| formats | No | Comma-separated bodies to return per page, from markdown (default), text, cleanHtml and rawHtml, e.g. 'markdown,text'. | markdown |
| parse_documents | No | Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins. |