Skip to main content
Glama
crawlbrulee

@crawlbrulee/mcp

Official
by crawlbrulee

Check an async scrape job status

scrape_status

Retrieve the lifecycle status of an async scrape job by job ID: pending, running, done, or failed. Poll until done to fetch results, or use a completion webhook to skip polling.

Instructions

Look up the current lifecycle status of an async scrape job submitted via scrape_async. Returns the job state (pending, running, done, failed), with an error message when it failed and a response_meta.usage block (credits, billed engine, resolved proxy tier, screenshot_slices) once it is done. Poll this until the status is done, then call scrape_result to fetch the page. If you registered a completion webhook on submit you can skip polling and react to the delivery instead.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
job_idYesJob identifier returned by the `scrape_async` tool

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoError message if the job ended in `failed`
job_idYesThe job identifier
statusYesCurrent state of the job (pending, running, done, failed)
created_atYesISO-8601 UTC timestamp when the job was created
response_metaNoUsage accounting for the finished job. Present only once the job is `done`; `response_meta.usage` reports credits charged, the billed engine, the resolved proxy tier, and any screenshot-slice add-on.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.3

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It richly covers the job states, failure error message, the usage block after completion, and the polling/webhook behavior. It doesn't mention auth, rate limits, or invalid job_id behavior, but for a simple status-read tool the disclosed behavior is quite strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no waste: it states the action, lists the valuable return data, gives the polling workflow, and offers the webhook shortcut. The core purpose is front-loaded and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, read-only status tool with an output schema, the description covers the complete workflow: what to poll for, how to know when done, what to do next, and how to avoid polling entirely. Nothing an agent needs to correctly invoke this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and job_id is already described as the identifier returned by scrape_async in the schema. The description adds provenance ('returned by scrape_async'), reinforcing where the agent obtains the value, which is useful beyond the schema's basic type and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Look up'), a distinct resource ('async scrape job'), and the lifecycle scope ('status'), immediately differentiating it from scrape_result and scrape_async. The state list (pending/running/done/failed) and the note to call scrape_result after done make the tool's specific role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs to poll until done and then call scrape_result, and names the webhook alternative that removes the need for polling. This gives the agent a clear when-to-use / when-not-to-use decision relative to its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.