Submit a batch scrape job
spicrawl_batch_submitSubmit many URLs to be scraped as one asynchronous job, with shared settings applied to every URL. Give EITHER urls or items, not both (at least one URL, unless open is true). Returns the job object (id, status, estimated_credits, progress). Poll with spicrawl_batch_status, read with spicrawl_batch_results or spicrawl_batch_task_content. Use this instead of many spicrawl_scrape calls when you have tens or thousands of URLs. Set open: true to keep the job accepting more URLs via spicrawl_batch_add_items; an open job never finishes until you call spicrawl_batch_close.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | A human label for the job. | |
| open | No | Keep the job accepting items after submission (append with spicrawl_batch_add_items, finish with spicrawl_batch_close). | |
| urls | No | Shorthand: URLs scraped with the shared settings and no per-item overrides. Mutually exclusive with `items`. | |
| wait | No | Milliseconds to wait after load (rendered scrapes). | |
| items | No | Long form: one /v1/scrape request object per item. Besides `url` and `external_id`, any /v1/scrape field (e.g. `js_render`, `wait_for`, `response_format`) overrides the shared setting for that item only. Mutually exclusive with `urls`. | |
| engine | No | Pin the scrape engine for every item: `fetch` or `obscura` (must be allowed by the plan and served by the deployment); `camoufox` is coming soon, do not pin it yet. `chromium` cannot be pinned in a batch — use `render: true`, or spicrawl_scrape. | |
| format | No | Output format for every item (`response_format`). spicrawl_batch_submit defaults to 'markdown', the same as spicrawl_scrape; an item's own `response_format` overrides it. | |
| render | No | Render every URL with a browser (`js_render`). | |
| max_cost | No | Ceiling in credits for ONE item (not the whole job — see credit_budget). | |
| priority | No | Scheduling priority of the job relative to your other jobs. | |
| wait_for | No | CSS selector to wait for before capturing (rendered scrapes). | |
| concurrency | No | Max items in flight at once for this job (may be capped by the server; see warnings). | |
| max_attempts | No | Per-item retry attempts on failure. | |
| credit_budget | No | Ceiling on the WHOLE job, in credits; the run aborts when it would exceed this. | |
| premium_proxy | No | Coming soon: not available yet, do not send. Use residential managed-pool exits for every URL. | |
| proxy_country | No | Coming soon: not available yet, do not send. ISO-3166 alpha-2 exit country; requires premium_proxy. | |
| custom_headers | No | Extra request headers sent with every URL. | |
| block_resources | No | Resource types to block while rendering, e.g. ["image","font","media"]. | |
| failure_threshold | No | Abort the job once this many items have failed. | |
| main_content_only | No | Strip each page to its main article. | |
| webhook_endpoint_id | No | Id of a registered webhook endpoint notified (with the job object) when the job completes. |