scrapy_cloud_get_job_requests
The HTTP requests a Scrapy Cloud job made, in order: url, status, method, response size (rs), duration (ms), time and the parent request index. Pass min_status=400 to see only failed requests, or 'url_contains' to find requests to one path or host. Also returns the job's total request count. Paged with 'count' and 'offset'. Output passes through the conversation, so beyond a small sample use the shub CLI (shub requests) instead.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. | |
| count | No | Maximum number of requests to return. | |
| offset | No | Number of requests to skip, for paging. | |
| min_status | No | Only requests whose response status is this or higher; 400 gives every failed request. | |
| url_contains | No | Only requests whose URL contains this text. |