Skip to main content
Glama

Zyte MCP Server

scrapy_cloud_run_spider

Start a new Scrapy Cloud job for a spider now. Returns the job key and dashboard URL; the job starts in state 'pending' and is picked up by the queue. Fails with 'already scheduled' if an identical job is pending or running. Check progress with scrapy_cloud_get_job. For a recurring schedule use scrapy_cloud_create_periodic_job instead; for a standalone script use scrapy_cloud_run_script.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
tagsNoTags to add to the job.
unitsNoScrapy Cloud units for the job. Default is the project's setting.
spiderYesSpider name, as listed by scrapy_cloud_list_spiders.
job_argsNoSpider arguments, passed as -a name=value; numbers and booleans are sent as strings ('100', 'true').
priorityNoQueue priority, 0 (lowest) to 4 (highest). Default 2.
project_idYesScrapy Cloud project id (the numeric id in the dashboard URL).
job_settingsNoScrapy settings overriding the project's, for example {"CLOSESPIDER_ITEMCOUNT": 100}.

Schema Changelog

Changes observed during successful MCP inspections.

No schema history has been recorded yet.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses the async lifecycle (starts 'pending', picked up by queue), the return payload (job key and dashboard URL), and an important dedup failure mode ('already scheduled'). It omits auth/permission requirements and any expectation about how long a job waits in queue, which is the only meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the action and return value, then failure behavior, then sibling routing. Every sentence carries distinct, actionable information with no repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no output schema and no annotations, the description fills the critical gaps: it says what is returned (job key, dashboard URL), that execution is asynchronous, that duplicate submissions are rejected, and which tool to use instead for adjacent use cases. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (spider, project_id, job_args, units, priority, tags, job_settings) is already documented in the schema, including the job_args stringification rule. The description adds no parameter-level detail beyond 'for a spider', so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a new Scrapy Cloud job for a spider now') and immediately distinguishes itself from the two nearest siblings, scrapy_cloud_create_periodic_job and scrapy_cloud_run_script, by name. An agent can route to it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing rules for all three cases: recurring schedules go to create_periodic_job, standalone scripts go to run_script, and progress checks go to get_job. It also states the condition under which this tool itself fails ('already scheduled' when an identical job is pending or running).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources