Zyte MCP Server
Server Details
Zyte MCP is a remote MCP server that lets an agent use Zyte on your behalf. It can download pages with a plain HTTP request or in a real browser, click and type and take screenshots, and pull structured data out of a page with AI, either as a standard type like product, article or job posting or as fields you describe in plain words. It can also search the web from a chosen country, check whether Zyte API supports a site and what a request costs, and report your usage by domain.
If you have Scrapy Cloud projects, the same server can list spiders, start and cancel jobs, read logs and items, manage schedules and settings, and read and write collections. Those tools can change and delete things, so review what your agent asks to do.
Connect to https://mcp.zyte.com/v1/mcp and sign in to your Zyte account with OAuth. Using the server is free. What your agent does through it is billed at Zyte API prices, on a dedicated mcp_access_key you can track and cap. New accounts get $5 of trial credit for a month.
Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.
If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.
- Status
- Unhealthy
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 34 tools
The four fetch/extract tools are clearly differentiated by browser vs HTTP and raw vs structured, with explicit cross-references in their descriptions. Each Scrapy Cloud tool targets a distinct resource and operation, so no two tools appear to do the same thing.
All tools use snake_case, and the scrapy_cloud_* prefix is applied consistently across 26 tools. The remaining tools mix patterns like extract_from_X, fetch_X, and bare names such as search and user_info, which slightly reduces predictability.
34 tools is heavy, but the server spans two distinct product surfaces (Zyte API and Scrapy Cloud) with many resources requiring full CRUD. The count is large yet mostly warranted by the breadth of the domain.
Scrapy Cloud coverage is thorough across jobs, projects, spiders, collections, periodic jobs, and settings, and Zyte API covers extraction, fetching, search, pricing, and usage. Minor gaps include no spider deployment or project creation, but the surface is otherwise complete.
Available Tools
34 toolsextract_from_browserAInspect
Extract structured data from a web page using Zyte API automatic (AI) extraction, fetching the page with a browser: give it a URL and an extraction type, get typed JSON back (product, article, job posting, ...). Highest quality and the default choice; the only extract tool with browser actions and viewport. Omit extractFrom to let Zyte API choose the browser source (currently 'browserHtml'); 'browserHtmlOnly' skips screenshot signals. For server-rendered pages extract_from_http is faster and cheaper. Choose the type by page kind — detail types for a single entity's page, list types for per-item summaries from one page, navigation types when crawling, pageContent as the generic fallback, webPageInfo for language only. One type per call. product, article, jobPosting and pageContent results carry metadata.probability — the confidence the page matches the type (below ~0.5 means it might not; a signal, not an error). If unsure between two of those types, make one call per type and keep the higher-probability result; list/navigation/forumThread results have no page-level probability. customAttributes adds caller-defined LLM-extracted fields on top of the requested type; for custom attributes alone use type "pageContent". Supports sessions, cookies, geolocation, and IP type. For raw pages (HTML, screenshots, files) use fetch_page or fetch_http instead. Returns a JSON metadata block (final URL, target HTTP status, optional action results/session), then one JSON block with the extracted data keyed by type.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute http(s) URL of the page to extract from (max 8192 chars). The host must be a domain name, not an IP address. | |
| type | Yes | The extraction type to run — one per call. Detail types (product, article, jobPosting, forumThread) suit a page about a single entity; list types (productList, articleList) return per-item summaries from a listing page; navigation types (productNavigation, articleNavigation, jobPostingNavigation) return item/next-page links for crawling; pageContent separates main content from site boilerplate on any page; webPageInfo returns the page language. | |
| ipType | No | Type of IP address to send the request from. Default: Zyte API picks the type that avoids bans for the target site. | |
| actions | No | Browser actions executed after the page loads and before extraction: click, type, scrollBottom, waitForSelector, evaluate, etc. Per-action results are returned in the metadata block. See the Zyte API actions reference (https://docs.zyte.com/zyte-api/usage/reference.html) for behavior details. | |
| viewport | No | Browser viewport size; changes what the extractor sees. | |
| sessionId | No | Client-managed session ID — a version 4 UUID you generate. Requests with the same ID reuse the same session (IP, cookies). Sessions expire 15 minutes after creation, after 2 idle minutes, or after 3 bans. | |
| extractFrom | No | Which browser source the extractor reads. 'browserHtml': rendered HTML plus screenshot signals, highest quality but less robust to rendering issues. 'browserHtmlOnly': rendered HTML without screenshot signals. Omit the field entirely to let Zyte API choose (currently 'browserHtml'). For a raw HTTP fetch use extract_from_http. | |
| geolocation | No | ISO 3166-1 alpha-2 country code to route the request from (e.g. 'US', 'DE'); must be a real country code. Default: Zyte API picks a geolocation that avoids bans and locale surprises for the target site. | |
| organizationId | Yes | Required. The Zyte organization to attribute this call to (max 100 characters, printable ASCII without spaces). If you do not already have an id, call the user_info tool: it lists the organizations your credential belongs to. IMPORTANT: if it lists more than one, ask the user which to use and wait for their answer — this call is billed to whichever organization you name here, so it is the user's choice to make, not yours. Never guess an id, and never fall back to a default. Once the user has chosen, reuse that id across the session unless they ask for a different organization. | |
| requestCookies | No | Cookies to send with the request (max 100). The responseCookies output of a previous fetch_page/fetch_http call can be passed here verbatim. | |
| sessionContext | No | Server-managed session context: up to 10 name/value pairs. Zyte API reuses or creates a session per distinct context. Sessions expire after 4 hours or 3 bans. | |
| enableZeroTrace | No | Keep the URL and other potentially sensitive request data out of Zyte API's request logs, metrics and stats records for this request. Default false. Use it for sensitive targets; it also leaves Zyte support less to go on when investigating the request. | |
| cookieManagement | No | How cookies are handled: 'auto' (default) uses requestCookies if given, otherwise Zyte API's automatic cookies; 'discard' uses requestCookies if given, otherwise no cookies. | |
| customAttributes | No | Ad-hoc fields extracted by a Zyte-operated LLM on top of the requested type (max 20 attributes): attribute name to attribute schema. The requested type scopes the page region fed to the LLM. If you only want custom attributes, use type "pageContent". | |
| sessionContextActions | No | Browser actions run once to initialize a server-managed session for the given sessionContext (e.g. login steps). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the extractFrom default behavior ('browserHtml' when omitted), that probability is a signal not an error, session/cookie/geolocation/IP support, and the two-block return shape. It stops short of billing/permission behavior (organizationId billing is documented in the schema, not here) and any rate-limit or failure-mode context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and mechanism in the first clause, and most sentences carry distinct, non-redundant information. It is, however, a dense block of text where the type-selection, probability, and alternative-routing guidance are all packed together, making it heavier reading than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter tool with nested action definitions and no output schema, the description covers the return structure (metadata block with final URL/status/action results, then a data block keyed by type), the extraction-type taxonomy, and the primary capability differentiators. An agent has enough to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description still adds meaning: it explains what omitting extractFrom does, what customAttributes layers on, and how metadata.probability should be interpreted per type. This interpretive guidance goes beyond the enum/property descriptions in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opening sentence gives a specific verb (extract), resource (structured data from a web page), mechanism (Zyte API automatic extraction via browser), and the input/output contract (URL + extraction type in, typed JSON out). It explicitly distinguishes itself from siblings: extract_from_http for server-rendered pages and fetch_page/fetch_http for raw pages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit selection criteria are given: use extract_from_http for server-rendered pages (faster/cheaper), fetch_page/fetch_http for raw pages, and this tool is 'the default choice' and the only one with browser actions/viewport. It also routes among extraction types (detail vs list vs navigation vs pageContent vs webPageInfo) and gives a tie-break procedure (one call per type, keep higher probability).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extract_from_httpAInspect
Extract structured data from a web page using Zyte API automatic (AI) extraction, downloading the raw HTTP response body — no browser. Usually faster and cheaper; fits server-rendered pages; supports device emulation (desktop/mobile). If the page needs JavaScript rendering, or you need browser actions or viewport, use extract_from_browser. Choose the type by page kind — detail types for a single entity's page, list types for per-item summaries from one page, navigation types when crawling, pageContent as the generic fallback, webPageInfo for language only. One type per call. product, article, jobPosting and pageContent results carry metadata.probability — the confidence the page matches the type (below ~0.5 means it might not; a signal, not an error). If unsure between two of those types, make one call per type and keep the higher-probability result; list/navigation/forumThread results have no page-level probability. customAttributes adds caller-defined LLM-extracted fields on top of the requested type; for custom attributes alone use type "pageContent". Supports sessions, cookies, geolocation, and IP type. For raw pages (HTML, screenshots, files) use fetch_page or fetch_http instead. Returns a JSON metadata block (final URL, target HTTP status, optional action results/session), then one JSON block with the extracted data keyed by type.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute http(s) URL of the page to extract from (max 8192 chars). The host must be a domain name, not an IP address. | |
| type | Yes | The extraction type to run — one per call. Detail types (product, article, jobPosting, forumThread) suit a page about a single entity; list types (productList, articleList) return per-item summaries from a listing page; navigation types (productNavigation, articleNavigation, jobPostingNavigation) return item/next-page links for crawling; pageContent separates main content from site boilerplate on any page; webPageInfo returns the page language. | |
| device | No | Device type to emulate during the HTTP request (default desktop). Browser fetches shape rendering with 'viewport' instead — see extract_from_browser. | |
| ipType | No | Type of IP address to send the request from. Default: Zyte API picks the type that avoids bans for the target site. | |
| sessionId | No | Client-managed session ID — a version 4 UUID you generate. Requests with the same ID reuse the same session (IP, cookies). Sessions expire 15 minutes after creation, after 2 idle minutes, or after 3 bans. | |
| geolocation | No | ISO 3166-1 alpha-2 country code to route the request from (e.g. 'US', 'DE'); must be a real country code. Default: Zyte API picks a geolocation that avoids bans and locale surprises for the target site. | |
| organizationId | Yes | Required. The Zyte organization to attribute this call to (max 100 characters, printable ASCII without spaces). If you do not already have an id, call the user_info tool: it lists the organizations your credential belongs to. IMPORTANT: if it lists more than one, ask the user which to use and wait for their answer — this call is billed to whichever organization you name here, so it is the user's choice to make, not yours. Never guess an id, and never fall back to a default. Once the user has chosen, reuse that id across the session unless they ask for a different organization. | |
| requestCookies | No | Cookies to send with the request (max 100). The responseCookies output of a previous fetch_page/fetch_http call can be passed here verbatim. | |
| sessionContext | No | Server-managed session context: up to 10 name/value pairs. Zyte API reuses or creates a session per distinct context. Sessions expire after 4 hours or 3 bans. | |
| enableZeroTrace | No | Keep the URL and other potentially sensitive request data out of Zyte API's request logs, metrics and stats records for this request. Default false. Use it for sensitive targets; it also leaves Zyte support less to go on when investigating the request. | |
| cookieManagement | No | How cookies are handled: 'auto' (default) uses requestCookies if given, otherwise Zyte API's automatic cookies; 'discard' uses requestCookies if given, otherwise no cookies. | |
| customAttributes | No | Ad-hoc fields extracted by a Zyte-operated LLM on top of the requested type (max 20 attributes): attribute name to attribute schema. The requested type scopes the page region fed to the LLM. If you only want custom attributes, use type "pageContent". | |
| verifyCertificate | No | Validate the target site's TLS certificate and fail with an error instead of extracting from an unverified page. Default false — certificates are not validated. Browser fetches always validate them, so extract_from_browser has no such argument. | |
| sessionContextActions | No | Browser actions run once to initialize a server-managed session for the given sessionContext (e.g. login steps). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full duty and does well: it discloses the no-browser HTTP mechanism, performance/cost profile, device emulation support, session/cookie/geolocation/IP support, and a precise output contract (metadata block then a type-keyed JSON block). It even explains probability semantics and that sub-0.5 is a signal not an error. It omits auth/billing and rate-limit behavior, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, and the dense decision guidance that follows is warranted for a 14-parameter tool with a complex enum. It is long, but nearly every clause provides routing or output semantics rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, high-stakes extraction tool with no annotations, it covers purpose, alternatives, type selection, tie-breaking, custom attributes, transport options, and return format. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents parameters. The description nonetheless adds real value beyond the schema, especially for the type enum (detail vs list vs navigation vs pageContent vs webPageInfo), single-type-per-call constraint, and the customAttributes/pageContent interaction.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Extract structured data from a web page using Zyte API automatic (AI) extraction, downloading the raw HTTP response body — no browser.' It immediately distinguishes itself from extract_from_browser and later from fetch_page/fetch_http.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('server-rendered pages', 'usually faster and cheaper'), when-not ('if the page needs JavaScript rendering, or you need browser actions or viewport, use extract_from_browser'), and names alternative tools for raw pages (fetch_page/fetch_http). It also routes type selection by page kind with a fallback and a tie-breaking strategy for ambiguous types.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_httpAInspect
Fetch a URL with a plain HTTP request through Zyte API: any HTTP method, optional request body, custom headers. Returns the page as Markdown by default; see 'format'. Fast and cheap; never executes JavaScript or renders the page. Use for static pages, form posts, and — with 'format': 'raw' — JSON/GraphQL/XHR API endpoints. For JavaScript-rendered pages, browser actions, or screenshots use fetch_page; for structured data (products, articles, jobs) use extract_from_http, or extract_from_browser when the page needs rendering. Returns a JSON metadata block (final URL, target HTTP status, optional headers/cookies) followed by the content in the chosen format.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute http(s) URL to fetch (max 8192 chars). The host must be a domain name, not an IP address. | |
| tags | No | Arbitrary key/value pairs (string values or null) attached to the request for filtering in the Zyte Stats API. | |
| device | No | Device type to emulate. Default desktop. | |
| format | No | How to return the response: 'markdown' (default; the main content as Markdown, without navigation, headers, footers and sidebars), 'fullPageMarkdown' (the whole page as Markdown) or 'raw' (the response body as received — text, or an image block for images). In the Markdown formats, JSON and other non-HTML text come back in a code block, and a PDF, image or other binary response as a short note; use 'raw' for API endpoints, images, or when the exact bytes matter. HEAD requests always use 'raw'. | markdown |
| ipType | No | Type of IP address to send the request from. Default: Zyte API picks the type that avoids bans for the target site. | |
| method | No | HTTP method. Default GET. | |
| headers | No | Custom HTTP request headers (max 200). The Cookie header is not allowed — use requestCookies. Zyte API sends some headers automatically for ban avoidance and may override or drop custom ones. | |
| bodyText | No | UTF-8 text to send as the request body (max 400000 chars), e.g. a JSON or GraphQL payload. Mutually exclusive with bodyBase64. | |
| sessionId | No | Client-managed session ID — a version 4 UUID you generate. Requests with the same ID reuse the same session (IP, cookies). Sessions expire 15 minutes after creation, after 2 idle minutes, or after 3 bans. | |
| bodyBase64 | No | Base64-encoded bytes to send as the request body (max 400000 chars encoded), for binary or non-UTF-8 content. Mutually exclusive with bodyText. | |
| geolocation | No | ISO 3166-1 alpha-2 country code to route the request from (e.g. 'US', 'DE'). Default: Zyte API picks a geolocation that avoids bans and locale surprises for the target site. | |
| followRedirect | No | Whether to follow HTTP redirects. Default true. | |
| organizationId | Yes | Required. The Zyte organization to attribute this call to (max 100 characters, printable ASCII without spaces). If you do not already have an id, call the user_info tool: it lists the organizations your credential belongs to. IMPORTANT: if it lists more than one, ask the user which to use and wait for their answer — this call is billed to whichever organization you name here, so it is the user's choice to make, not yours. Never guess an id, and never fall back to a default. Once the user has chosen, reuse that id across the session unless they ask for a different organization. | |
| requestCookies | No | Cookies to send with the request (max 100). The responseCookies output of a previous call can be passed here verbatim. | |
| sessionContext | No | Server-managed session context: up to 10 name/value pairs. Zyte API reuses or creates a session per distinct context. Sessions expire after 4 hours or 3 bans. | |
| enableZeroTrace | No | Keep the URL and other potentially sensitive request data out of Zyte API's request logs, metrics and stats records for this request. Default false. Use it for sensitive targets; it also leaves Zyte support less to go on when investigating the request. | |
| responseCookies | No | Return the cookies set during the request in the metadata block. Prefer this over Set-Cookie response headers — it reflects the final cookie state. | |
| cookieManagement | No | How cookies are handled: 'auto' (default) uses requestCookies if given, otherwise Zyte API's automatic cookies; 'discard' uses requestCookies if given, otherwise no cookies. | |
| verifyCertificate | No | Validate the target site's TLS certificate and return an error instead of the page when validation fails. Default false — certificates are not validated. Browser fetches always validate them, so fetch_page has no such argument. | |
| sessionContextActions | No | Browser actions run once to initialize a server-managed session for the given sessionContext (e.g. login steps). Each action takes the parameters documented in the Zyte API actions reference (https://docs.zyte.com/zyte-api/usage/reference.html), e.g. selector {type: css|xpath, value}, timeout, source, url, keyword. Parameters are validated against the Zyte API contract. | |
| includeResponseHeaders | No | Include the target response's HTTP headers in the metadata block. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose meaningful behavior: it is fast/cheap, never executes JavaScript or renders, and returns a JSON metadata block plus content in the chosen format. It omits some operational traits (billing/rate-limit implications of organizationId are only in the schema), but the core behavioral profile is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then the default return, then routing guidance, then return detail. Each sentence is doing distinct work — no filler or repetition despite the dense content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by describing the return shape (metadata block with final URL, status, optional headers/cookies, followed by content in the chosen format). For a 21-parameter tool it relies on the 100%-covered schema for the remaining parameters, which is reasonable, though it says nothing about the billing implications that make organizationId significant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema. The description reiterates a few of them (method, body, headers, format) but adds no syntax or format detail beyond what the schema provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fetch a URL with a plain HTTP request through Zyte API') and scopes it precisely (any method, optional body, custom headers). It explicitly distinguishes itself from the three most relevant siblings (fetch_page for JS rendering, extract_from_http and extract_from_browser for structured data).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use cases (static pages, form posts, JSON/GraphQL/XHR endpoints with format:'raw') and explicit when-not with named alternatives (JS-rendered pages/screenshots -> fetch_page, structured data -> extract_from_http/extract_from_browser). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_pageAInspect
Open a URL in a remote browser using Zyte API and return the rendered page. GET navigation only. Returns Markdown by default; 'output' selects Markdown, the rendered HTML and/or a screenshot. Supports browser actions (click, type, scroll, wait, evaluate, ...), a JavaScript on/off toggle, viewport, referer, sessions, cookies, and geolocation. Slower and costlier than fetch_http — prefer fetch_http for static pages and API endpoints, and extract_from_browser for structured data (products, articles, jobs). Returns a JSON metadata block (final URL, target HTTP status, optional headers/cookies/action results) followed by the requested outputs in this order: Markdown, HTML, screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute http(s) URL to open (max 8192 chars). The host must be a domain name, not an IP address. | |
| tags | No | Arbitrary key/value pairs (string values or null) attached to the request for filtering in the Zyte Stats API. | |
| ipType | No | Type of IP address to send the request from. Default: Zyte API picks the type that avoids bans for the target site. | |
| output | No | What to return: 'markdown' (default; the main content as Markdown, without navigation, headers, footers and sidebars), 'fullPageMarkdown' (the whole page as Markdown), 'html' (the DOM after rendering and actions) and/or 'screenshot'. 'markdown' and 'fullPageMarkdown' can't be combined. | |
| actions | No | Browser actions executed in order after the page loads and before the output is captured: click, type, scrollBottom, waitForSelector, evaluate, etc. Per-action results are returned in the metadata block. Each action takes the parameters documented in the Zyte API actions reference (https://docs.zyte.com/zyte-api/usage/reference.html), e.g. selector {type: css|xpath, value}, timeout, source, url, keyword. Parameters are validated against the Zyte API contract. | |
| referer | No | Referer header for the navigation — the only request header browser requests support. | |
| viewport | No | Browser viewport size; affects rendering and non-fullPage screenshots. Defaults to 1920x1080. | |
| sessionId | No | Client-managed session ID — a version 4 UUID you generate. Requests with the same ID reuse the same session (IP, cookies). Sessions expire 15 minutes after creation, after 2 idle minutes, or after 3 bans. | |
| javascript | No | Force JavaScript execution on (true) or off (false). Default: Zyte API chooses whichever avoids bans for the target site. | |
| geolocation | No | ISO 3166-1 alpha-2 country code to route the request from (e.g. 'US', 'DE'). Default: Zyte API picks a geolocation that avoids bans and locale surprises for the target site. | |
| includeIframes | No | Include iframe content in the returned HTML. Default false. Iframes appear in screenshots regardless. | |
| organizationId | Yes | Required. The Zyte organization to attribute this call to (max 100 characters, printable ASCII without spaces). If you do not already have an id, call the user_info tool: it lists the organizations your credential belongs to. IMPORTANT: if it lists more than one, ask the user which to use and wait for their answer — this call is billed to whichever organization you name here, so it is the user's choice to make, not yours. Never guess an id, and never fall back to a default. Once the user has chosen, reuse that id across the session unless they ask for a different organization. | |
| requestCookies | No | Cookies to send with the request (max 100). The responseCookies output of a previous call can be passed here verbatim. | |
| sessionContext | No | Server-managed session context: up to 10 name/value pairs. Zyte API reuses or creates a session per distinct context. Sessions expire after 4 hours or 3 bans. | |
| enableZeroTrace | No | Keep the URL and other potentially sensitive request data out of Zyte API's request logs, metrics and stats records for this request. Default false. Use it for sensitive targets; it also leaves Zyte support less to go on when investigating the request. | |
| responseCookies | No | Return the cookies set during the request in the metadata block. Prefer this over Set-Cookie response headers — it reflects the final cookie state. | |
| cookieManagement | No | How cookies are handled: 'auto' (default) uses requestCookies if given, otherwise Zyte API's automatic cookies; 'discard' uses requestCookies if given, otherwise no cookies. | |
| screenshotOptions | No | Screenshot tuning; only valid when output includes 'screenshot'. | |
| sessionContextActions | No | Browser actions run once to initialize a server-managed session for the given sessionContext (e.g. login steps). Each action takes the parameters documented in the Zyte API actions reference (https://docs.zyte.com/zyte-api/usage/reference.html), e.g. selector {type: css|xpath, value}, timeout, source, url, keyword. Parameters are validated against the Zyte API contract. | |
| includeResponseHeaders | No | Include the target response's HTTP headers in the metadata block. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose non-obvious traits: relative cost/latency versus fetch_http, the GET-only constraint, and the exact return structure (JSON metadata block followed by outputs in Markdown, HTML, screenshot order). It omits auth/permission requirements and rate limits, but the required-param schema covers billing attribution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences, purpose front-loaded, then constraint, then alternatives, then return shape. No filler; each sentence carries distinct decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 20-parameter tool with nested objects and no output schema, the description supplies the missing return-value contract (metadata block plus ordered Markdown/HTML/screenshot outputs) and the cost-aware routing context, so an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 20 parameters and the baseline is 3. The description enumerates the capability areas (actions, javascript toggle, viewport, referer, sessions, cookies, geolocation) and names the 'output' selector, but adds no syntax or semantics beyond what the schema fields already provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open a URL in a remote browser using Zyte API and return the rendered page') and adds a scope constraint ('GET navigation only'). It explicitly distinguishes itself from fetch_http and extract_from_browser, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and when-not-to-use guidance: 'prefer fetch_http for static pages and API endpoints, and extract_from_browser for structured data (products, articles, jobs)', with the reason (slower and costlier) stated. This is exactly the alternative-routing an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_cancel_jobAInspect
Stop a pending or running Scrapy Cloud job. Cancellation is asynchronous: the job finishes shortly after with close_reason 'cancelled'; poll scrapy_cloud_get_job to confirm. Does nothing for finished or deleted jobs.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that cancellation is asynchronous, that the job settles with close_reason 'cancelled', and that finished/deleted jobs are unaffected. It stops short of stating permission/auth requirements or whether the call is idempotent on repeated invocation, which keeps it out of the top band.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, then the async caveat, then the no-op boundary. Every sentence carries distinct information and none is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation tool with no output schema, the description covers everything an agent needs: what it does, the async result shape (close_reason), the verification path, and the no-op case. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'job' parameter is documented in the schema with its project/spider/job format and pattern. The description adds no parameter-level detail beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Stop a pending or running Scrapy Cloud job.' The scope qualifier ('pending or running') and the explicit no-op condition ('Does nothing for finished or deleted jobs') separate it cleanly from siblings like scrapy_cloud_get_job and scrapy_cloud_delete_periodic_job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives both sides of the when-to-use boundary: applicable to pending/running jobs, ineffective for finished or deleted jobs. It also names the follow-up tool ('poll scrapy_cloud_get_job to confirm'), so an agent knows the workflow rather than just the call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_create_periodic_jobAInspect
Create a recurring schedule that runs one or more spiders (or scripts) on a cron expression, in UTC. Cron fields must be numbers or '*': no ranges, steps or lists; use several schedules for several times. This changes what runs on the account on an ongoing basis; confirm spider and schedule with the user first. Returns the new periodic job with its id.
| Name | Required | Description | Default |
|---|---|---|---|
| cron | Yes | Schedule in UTC as 5 cron fields, each a number or '*'. Example: '0 9 * * 1' is Mondays 09:00 UTC. No ranges, steps or lists. | |
| tasks | Yes | Spiders or scripts to run on each tick, in order. | |
| addtags | No | Tags added to every job this schedule starts. | |
| disabled | No | Create the schedule paused. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| description | No | Human-readable purpose of the schedule. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose the important traits: the effect is ongoing/persistent on the account, UTC is the timezone, cron syntax is restricted (no ranges/steps/lists), and it returns the new periodic job with its id. It omits permission/auth requirements and what happens on cron conflicts, so it is not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, each earning its place: purpose, cron constraint, mutation warning, return value. The most decision-relevant content (what it creates and that it changes ongoing behavior) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by stating the return ('the new periodic job with its id'). Combined with the mutation warning and the cron-format constraint, an agent has what it needs to call it correctly; only secondary details like permissions are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (cron, tasks, project_id, disabled, addtags, description) is already documented inline with examples and defaults. The description restates the cron constraint and the spider-or-script notion already present in the schema, adding no new parameter semantics beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Create a recurring schedule that runs one or more spiders (or scripts) on a cron expression, in UTC.' The word 'recurring' plus the cron mechanism separates it from one-off siblings like scrapy_cloud_run_spider and from the update/delete periodic-job siblings, so an agent can place it without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditions and a caution: 'This changes what runs on the account on an ongoing basis; confirm spider and schedule with the user first.' It also explains how to express multiple run times ('use several schedules for several times'), which is a usage workaround for the cron limitation. It does not name a sibling alternative (e.g. run_spider for one-off runs), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_delete_collectionAInspect
Permanently delete a whole Scrapy Cloud collection with all its records. Irreversible, and it removes every collection of that name whatever its type ('s', 'cs', 'vs', 'vcs'), since storage treats those as one. Confirm with the user first. To remove some records only, use scrapy_cloud_delete_collection_records.
| Name | Required | Description | Default |
|---|---|---|---|
| collection | Yes | Collection name, as listed by scrapy_cloud_list_collections. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses irreversibility, the surprising cross-type deletion scope, and a user-confirmation requirement. It does not mention permissions or what the tool returns, so it falls just short of exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each earning its place: core action, irreversibility + scope caveat, safety prerequisite, and sibling routing. The most consequential warning (permanent deletion) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive two-parameter tool with no annotations and no output schema, the description supplies the context an agent needs: permanence, the counterintuitive all-types deletion, user confirmation, and the exact alternative for partial removal. Nothing critical is left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented with format cues (collection name as listed by list_collections; numeric project id). The description adds no parameter-level syntax or constraint beyond that baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Permanently delete') and resource ('Scrapy Cloud collection'), and immediately names the sibling tool for partial deletion to distinguish scope. The type-collapsing nuance ('s', 'cs', 'vs', 'vcs') further sharpens what uniquely gets deleted.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to confirm with the user first and names the alternative (scrapy_cloud_delete_collection_records) with the condition that selects it ('To remove some records only'). This is clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_delete_collection_recordsAInspect
Delete records from a Scrapy Cloud collection by key. Irreversible. To remove a whole collection use scrapy_cloud_delete_collection.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | Yes | Record keys to delete. | |
| type | No | Collection type: 's' store (default), 'cs' cached store (records expire after a month), 'vs' versioned store, 'vcs' versioned cached store. | s |
| collection | Yes | Collection name, as listed by scrapy_cloud_list_collections. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It correctly and critically warns that the operation is 'Irreversible,' which is important for a destructive tool. However, it does not mention required permissions, behavior for missing keys, or what the response contains, leaving notable gaps for a mutation operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no waste. It front-loads the core action, then immediately flags the irreversible nature, then routes to the correct alternative tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description covers the essential purpose, danger, and alternative tool. It is largely complete given that the schema documents parameters thoroughly, though it could add a note about authorization or partial-failure behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the input schema already documents all four parameters, including the enum values and default for collection type. The description adds the phrase 'by key,' which maps to the keys parameter but does not extend meaning beyond what the schema provides. Baseline 3 is appropriate when the schema fully documents the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete records from a Scrapy Cloud collection by key.' It also explicitly distinguishes this from the sibling tool that deletes entire collections, naming scrapy_cloud_delete_collection. An agent can identify exactly what this tool does and how it differs from nearby tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states both when to use this tool (delete individual records by key) and when to use a specific alternative (scrapy_cloud_delete_collection for removing a whole collection). The condition for the alternative is explicit, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_delete_periodic_jobAInspect
Permanently delete a periodic job. Irreversible; prefer scrapy_cloud_update_periodic_job with disabled=true to pause a schedule. Past jobs it started are not affected.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| periodic_job_id | Yes | Periodic job id, as listed by scrapy_cloud_list_periodic_jobs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose the key traits: the operation is irreversible and deleting the schedule does not remove jobs it already started. It stops short of stating permission/auth requirements or what the response returns, but the destructive semantics and blast radius are clearly communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with zero waste; the destructive nature and irreversibility are front-loaded before the alternative and the side-effect note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter delete with no output schema, the description covers purpose, danger, and the preferred alternative, which is enough to call it correctly. Only the absence of explicit auth/permission context keeps it from being fully exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters already document themselves (project id format, job id source). The description adds no parameter-level meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently delete a periodic job') and distinguishes the operation from its sibling by naming scrapy_cloud_update_periodic_job as the non-destructive alternative. An agent can differentiate it from the list/get/update periodic-job tools immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the when-not case: prefer scrapy_cloud_update_periodic_job with disabled=true to pause a schedule instead of deleting. It names the concrete alternative and the condition that selects it, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_collection_recordsAInspect
Read records of a Scrapy Cloud collection, each with its '_key' and '_ts' write time in milliseconds. Select by exact keys, by key prefix (indexed, fast) or by write-time window; otherwise reads from the start and also reports the record count ('total', or 'total_at_least' for a very large collection). Long strings are truncated to 'max_chars'. Output passes through the conversation, so beyond a small sample use the storage REST API instead.
| Name | Required | Description | Default |
|---|---|---|---|
| key | No | Only these record keys. | |
| type | No | Collection type: 's' store (default), 'cs' cached store (records expire after a month), 'vs' versioned store, 'vcs' versioned cached store. | s |
| count | No | Maximum number of records to return. | |
| endts | No | Only records written before this UNIX timestamp, in milliseconds. | |
| prefix | No | Only records whose key starts with this prefix; uses an index, prefer it over scanning. | |
| startts | No | Only records written at or after this UNIX timestamp, in milliseconds. | |
| max_chars | No | Truncate string values longer than this; 0 disables truncation. | |
| collection | Yes | Collection name, as listed by scrapy_cloud_list_collections. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the per-record fields, the 'total'/'total_at_least' count semantics, truncation behavior, and that output passes through the conversation (a token-cost warning). It omits auth/permission requirements, but the read-only nature is unambiguous.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but well front-loaded: purpose first, then selection modes, then the caveat about output size. Every sentence carries content, though the single long paragraph packs several distinct concerns together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must cover returns — and it does (_key, _ts, total, total_at_least, truncation). For a 9-parameter read tool this is nearly complete; only the absence of any permission/auth note keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents every parameter and the baseline is 3. The description reinforces the selection strategy ('indexed, fast' prefix, truncation to max_chars) but adds little syntax or meaning beyond what the schema descriptions already state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (read records of a Scrapy Cloud collection) and immediately characterizes the returned data ('_key' and '_ts' write time). It is trivially distinguishable from siblings like scrapy_cloud_delete_collection_records and scrapy_cloud_put_collection_record, which mutate the same resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the selection modes (exact keys, prefix, write-time window) and, crucially, warns that beyond a small sample the caller should use the storage REST API instead — an explicit when-not-to-use rule. It doesn't route to any sibling MCP tool by name, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_jobBInspect
Get one Scrapy Cloud job: state, close_reason, spider, tags, who scheduled and cancelled it, pending/running/finished times (ms), spider_args and job_settings (credential-like values masked) and 'scrapystats', Scrapy's counters such as item_scraped_count, downloader/response_status_count/, log_count/ERROR, retry/count, memusage/max and scrapy-zyte-api/*. A healthy finished job has close_reason 'finished'.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. It discloses useful specifics like masked credential-like values in spider_args, job_settings, and the meaning of close_reason 'finished', but does not cover permissions, error behavior, or the exact format of the response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It's a single, dense sentence that packs a lot of field-level detail without unnecessary repetition. Front-loaded with the core action and resource, though the list of fields makes it somewhat long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description provides substantial value by enumerating the returned fields, including scrapystats and masked settings. It still lacks details on error handling, rate limits, or authentication requirements, which would be needed for a fully complete picture of a network-dependent retrieval tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the single 'job' parameter with a regex pattern and example. The description adds no parameter-specific information, which is acceptable given the high schema coverage baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get one Scrapy Cloud job') and enumerates the rich fields returned. It's clear the tool retrieves a single job, distinguishing it from list_jobs, but does not explicitly name siblings like scrapy_cloud_get_job_items or scrapy_cloud_get_job_log.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It doesn't explain the difference between this and scrapy_cloud_get_job_items, scrapy_cloud_get_job_log, or scrapy_cloud_list_jobs, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_job_itemsAInspect
Read a sample of a Scrapy Cloud job's scraped items, in scrape order: at most 'count' items from 'offset', long string fields truncated to 'max_chars', optionally filtered server-side with 'where' (e.g. price above 100, title containing a word). Output passes through the conversation, so beyond a small sample use the shub CLI (shub items) instead.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. | |
| count | No | Maximum number of items to return. | |
| where | No | Server-side filters, all of which an item must match, e.g. [{"field": "price", "op": ">", "value": 100}]. | |
| offset | No | Number of items to skip, for paging. | |
| max_chars | No | Truncate string field values longer than this; 0 disables truncation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it warns that output passes through the conversation (context-window cost), discloses truncation of long strings, server-side filtering, and ordering. It omits auth/permission requirements and error behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the core capability plus its main parameters are front-loaded before the fallback-tool caveat. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter read tool with full schema coverage and no output schema, the description is nearly complete: paging, truncation, filtering, ordering, and the practical limitation are all covered. Remaining gaps (auth needs, item shape) are minor for a passthrough sample reader.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is already 100%, so baseline is 3; the description adds genuine meaning beyond the schema by stating that results come in scrape order, that count items start from offset, that where is applied server-side with a concrete example ('price above 100, title containing a word'), and that max_chars truncates long string fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read a sample of a Scrapy Cloud job's scraped items') plus a meaningful scope qualifier ('in scrape order'). An agent can distinguish this from siblings like scrapy_cloud_get_job_item_stats or scrapy_cloud_get_collection_records without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit context ('small sample') and a when-not rule: for anything beyond a small sample, use the shub CLI (shub items). This is strong routing guidance, though it points to an external CLI rather than naming the closest MCP sibling tool, and does not state prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_job_item_statsAInspect
Item count and per-field coverage of a Scrapy Cloud job without downloading items: how many items were scraped and, for each field, how many items have it (count and percentage). A field at 0% coverage usually means a broken extractor. Use it to validate a finished job's results before reading samples.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral load, and it does disclose that no items are downloaded (a cheap metadata read) plus an interpretation rule that 0% coverage signals a broken extractor. It does not mention auth/permission requirements or what happens if the job has not finished, which are the remaining behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with what is returned, then the diagnostic caveat, then the intended use. Nothing is padding; each sentence adds either output semantics, interpretation, or routing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain the return shape, and it does: total item count plus per-field count and percentage. For a one-parameter read-only stats tool this is sufficient to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single 'job' parameter with 100% schema description coverage, including a pattern and an example, so the schema does the work. The description only refers to 'a Scrapy Cloud job' generically and adds no format details beyond what is already documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource (a Scrapy Cloud job) and the exact computation (item count plus per-field coverage), and the phrase 'without downloading items' functionally separates it from the sibling scrapy_cloud_get_job_items. An agent can tell what it produces without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit use case — 'validate a finished job's results before reading samples' — which implies the ordering against get_job_items, but it never names the alternative tool or states when not to use this one (e.g., for a still-running job).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_job_logAInspect
Read the log of a Scrapy Cloud job. By default returns the last 'count' entries. Pass level='ERROR' to get only errors, 'contains' to search messages for a text (both scan from the start of the log), or 'offset' to page from the start. Entries have time (ms), level and message. Output passes through the conversation, so beyond a small sample use the shub CLI (shub log) instead.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. | |
| tail | No | Return the last entries of the log instead of the first. Ignored when 'offset', 'level' or 'contains' is given. | |
| count | No | Maximum number of log entries to return. | |
| level | No | Only entries at this level or above. | |
| offset | No | Number of entries to skip from the start of the log. | |
| contains | No | Only entries whose message contains this text, filtered server-side. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses the entry shape (time/level/message), the default tail behavior, the non-obvious scan semantics (level/contains scan from the start of the log, not the tail), and a cost warning that output passes through the conversation. It stops short of naming auth requirements or hard truncation limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and efficient overall, though the middle sentence packs several modes together and the CLI recommendation could be tightened. Every sentence still earns its place by covering distinct behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only log tool with no output schema, the description supplies the return-entry format, the filtering/paging model, and a throughput caveat, which is enough for an agent to invoke it correctly and know when it is a bad fit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: how level/contains/offset interact with tail and that filtering scans from the log start rather than the tail. It reinforces rather than repeats the schema's parameter docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the log of a Scrapy Cloud job'), which cleanly separates it from siblings like scrapy_cloud_get_job_items or scrapy_cloud_get_job_requests. An agent can identify the target artifact (job log) without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes mode selection (tail default, level='ERROR', contains, offset) and routes the agent to an alternative (shub CLI / 'shub log') for anything beyond a small sample. This is actionable when-to-use and when-to-use-something-else guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_job_requestsAInspect
The HTTP requests a Scrapy Cloud job made, in order: url, status, method, response size (rs), duration (ms), time and the parent request index. Pass min_status=400 to see only failed requests, or 'url_contains' to find requests to one path or host. Also returns the job's total request count. Paged with 'count' and 'offset'. Output passes through the conversation, so beyond a small sample use the shub CLI (shub requests) instead.
| Name | Required | Description | Default |
|---|---|---|---|
| job | Yes | Job key in project/spider/job form, for example 123/1/4. | |
| count | No | Maximum number of requests to return. | |
| offset | No | Number of requests to skip, for paging. | |
| min_status | No | Only requests whose response status is this or higher; 400 gives every failed request. | |
| url_contains | No | Only requests whose URL contains this text. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses output ordering, that the total request count is included, that results are paged via count/offset, and that output passes through the conversation (a token/size warning). It stops short of stating auth requirements or rate limits, so it is not fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what is returned, then filters, then pagination, then the volume caveat. Dense but every sentence earns its place; only the field enumeration is slightly list-like.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description takes responsibility for return values by listing the fields and noting the total count, plus pagination and the conversation-passthrough caveat. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantic guidance about how to use the filters together (400 catches all failures; url_contains matches a path or host) and confirms paging semantics for count/offset, going beyond raw field documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('the HTTP requests a Scrapy Cloud job made, in order') and enumerates the returned fields, so an agent can distinguish it from siblings like get_job_items or get_job_log without opening a schema. The scope (requests of a single job) is precise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance for the filters ('Pass min_status=400 to see only failed requests, or url_contains to find requests to one path or host') and names an alternative for large pulls ('beyond a small sample use the shub CLI (shub requests) instead'). This is a clear when-to-use and when-not-to-use statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_periodic_jobBInspect
Get one periodic job of a Scrapy Cloud project by id: cron, tasks, addtags, description, disabled.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| periodic_job_id | Yes | Periodic job id, as listed by scrapy_cloud_list_periodic_jobs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Get' implies a read-only fetch, and the description usefully names the fields it returns (cron, tasks, addtags, description, disabled), but it says nothing about permissions, error behavior for a missing/invalid id, or whether the result is a snapshot.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence identifying the verb, resource, and id-based scope, with the returned field list appended compactly. No wasted words, though the trailing field list reads slightly like a dump.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description's field list partially compensates by hinting at the return shape, which is helpful for a 2-parameter fetch tool. However, it stops short of describing the response structure or failure modes, leaving some gaps for a tool with no annotations and no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (project_id, periodic_job_id) are already fully documented in the schema, including the hint that periodic_job_id comes from scrapy_cloud_list_periodic_jobs. The description adds no parameter detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Get) and resource (one periodic job of a Scrapy Cloud project) scoped by id, so an agent can identify the operation. The word 'one' implicitly distinguishes it from scrapy_cloud_list_periodic_jobs, though the sibling is not named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'by id' and the schema note that the id comes from scrapy_cloud_list_periodic_jobs, but the description itself gives no explicit when-to-use guidance or exclusions relative to list_periodic_jobs or update_periodic_job.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_projectAInspect
Details of one Scrapy Cloud project, including its organization id (needed by zyte_api_usage_stats), dashboard URL and current job counts.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. The 'get' framing and the enumeration of returned fields signal a safe read, and disclosing that org id feeds another tool is genuinely useful. But it says nothing about error behavior for a missing project, auth requirements, or rate limits, leaving gaps for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every clause earns its place by naming the resource and its most useful return fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema, describing the key returned fields (org id, dashboard URL, job counts) usefully compensates for the absent output schema. However, with no annotations and no mention of failure modes, it is only adequately complete rather than thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single project_id parameter is fully documented in the schema, including the 'numeric id in the dashboard URL' hint. The description adds no parameter syntax or format detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Details of one Scrapy Cloud project') and enumerates what is returned (organization id, dashboard URL, job counts). The singular 'one project' implicitly distinguishes it from siblings like scrapy_cloud_list_projects and scrapy_cloud_get_project_settings, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one useful cross-tool hint — the organization id is 'needed by zyte_api_usage_stats' — which implies a chained workflow. However, it never states when to use this tool versus get_project_settings, get_project_activity, or list_projects, so the routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_project_activityAInspect
The activity log of a Scrapy Cloud project, newest first: events such as jobs started, completed or cancelled and deployments, each with the user or system that caused it. Use it to answer who did what and when in a project.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Maximum number of events to return, newest first. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full behavioral burden. It usefully discloses ordering (newest first), the event taxonomy, and that each entry identifies the causing actor, but it says nothing about the operation being read-only, pagination/truncation behavior, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no waste; the resource and its ordering are front-loaded, followed by the intended use.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the shape of returned entries (event, actor, timestamp ordering), enough for an agent to know what it gets back. It leaves return size/pagination behavior unaddressed, but the schema covers the count cap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so project_id and count are already fully documented in the schema, and the description's 'newest first' merely echoes the count parameter's own wording. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (a Scrapy Cloud project's activity log) and enumerates the event types it contains (jobs started/completed/cancelled, deployments), which no sibling offers. The verb is only implied by the noun phrase and tool name, so it falls just short of a crisp verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it to answer who did what and when in a project' gives a clear scenario for selecting this tool over e.g. get_job_log or list_jobs. It does not name alternatives or state when not to use it, so it is context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_get_project_settingsAInspect
The Scrapy settings stored on a Scrapy Cloud project (project-wide defaults such as LOG_LEVEL or DOWNLOAD_DELAY), its enabled add-ons and default job units. Settings not listed use Scrapy Cloud defaults; spiders and jobs can override them. Values of credential-like settings are masked. Explains, for example, why a job logs no DEBUG lines.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so well: it discloses that missing settings fall back to Scrapy Cloud defaults, that spiders/jobs can override, and that credential-like values are masked. It does not describe return shape or error behavior, but the masking detail is genuinely useful beyond structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each front-loading the key noun (settings, overrides, masking) with no filler. The final sentence about DEBUG logs is a helpful concrete example rather than bloat; structure is tight though slightly sentence-heavy for its size.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param read tool with no annotations or output schema, the description covers scope, defaults, overrides, and a key data caveat (masked credentials), which is enough for correct invocation. It stops short of describing the response format, but since no output schema is declared that gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and its schema description coverage is 100%, so the schema already fully documents project_id. The description adds no parameter-level meaning, matching the baseline 3 for a fully covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Scrapy settings stored on a Scrapy Cloud project') and enumerates what's included (project-wide defaults, enabled add-ons, default job units). It is clearly distinguishable from the sibling update_project_settings (a write) and from settings on jobs/spiders, which the text mentions as overriders.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly targets the read-settings use case and notes that it explains job behavior like missing DEBUG logs, which is a clear context cue. However, it never explicitly states when to prefer scrapy_cloud_get_project or scrapy_cloud_get_job, nor does it exclude any use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_list_collectionsAInspect
List a Scrapy Cloud project's collections: key-value stores that spiders and scripts use to keep state across jobs, such as seen URLs, checkpoints or configuration. Each has a name and a type ('s' store, 'cs' cached, 'vs' versioned, 'vcs' versioned cached).
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully describes return fields and type codes, and 'List' implies a read operation, but it omits permissions, rate limits, pagination, and explicit read-only safety information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and resource, followed by a compact explanation of the domain object and its type codes. Every sentence carries useful information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter list tool with no output schema, the description explains what collections are and what each returned item contains (name and type). It is missing only secondary details like pagination, authorization, or ordering behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single required parameter project_id is already documented in the schema. The description adds no parameter syntax, format, or semantic detail beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('Scrapy Cloud project's collections'), then defines what collections are in this context. It is clearly distinguishable from siblings such as scrapy_cloud_get_collection_records or scrapy_cloud_delete_collection by resource scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what collections are and what a result contains, but gives no when-to-use, when-not-to-use, or alternatives. It does not say whether to call this before get_collection_records, delete_collection, or similar tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_list_jobsAInspect
List Scrapy Cloud jobs of a project, newest first. Returns only finished jobs unless 'state' is given. Each job summary has key, spider, state, close_reason, items, errors, pages, logs, elapsed, ts, pending/running/finished_time (ms), tags and version; ask for more fields via 'meta'. Use 'start' with 'count' to page.
| Name | Required | Description | Default |
|---|---|---|---|
| job | No | Only these job keys, in project/spider/job form, for example 123/1/4. | |
| meta | No | Extra job fields to include, for example spider_args, scheduled_by, units. | |
| count | No | Maximum number of jobs to return. | |
| endts | No | Only jobs updated before this UNIX timestamp, in milliseconds. | |
| start | No | Number of jobs to skip, for paging. | |
| state | No | Only jobs in these states. Default is finished jobs only. | |
| spider | No | Only jobs of this spider name. | |
| has_tag | No | Only jobs having at least one of these tags. | |
| startts | No | Only jobs updated at or after this UNIX timestamp, in milliseconds. | |
| lacks_tag | No | Only jobs having none of these tags. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses the default state filter, sort order, paging mechanism, and the full set of returned summary fields plus the 'meta' escape hatch. Auth requirements and rate limits are unaddressed, but core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, all front-loaded with the scope, default behavior, and result shape. No filler; every clause conveys operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description enumerates the returned fields and explains filtering defaults and pagination, giving the agent everything needed to call and interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds cross-parameter semantics the schema does not: the finished-only default, the meta field-expansion pattern, and start/count paging. This is meaningful added meaning beyond per-field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb plus resource ('List Scrapy Cloud jobs of a project') and adds scope details (newest first, per-project). This cleanly distinguishes it from singular siblings like scrapy_cloud_get_job and filtered variants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains a crucial default ('Returns only finished jobs unless state is given') and how to page ('Use start with count'). It stops short of naming sibling alternatives, but the when-to-use context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_list_periodic_jobsAInspect
List the periodic (cron-scheduled) jobs of a Scrapy Cloud project: id, cron, tasks (spiders or scripts with priority and arguments), addtags, description and whether the schedule is disabled. Includes the UTC hour Scrapy Cloud suggests for new schedules.
| Name | Required | Description | Default |
|---|---|---|---|
| count | No | Maximum number of periodic jobs to return. | |
| offset | No | Number of periodic jobs to skip, for paging. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It implies a read-only listing and discloses that results include a 'disabled' flag and the suggested UTC hour, which is useful context. However, it says nothing about authentication, paging behavior, or result ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence opens with the verb and resource, then enumerates returned fields. It is dense but every clause conveys real information, with only minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with fully documented parameters and no output schema, the description compensates by enumerating the returned fields (id, cron, tasks, addtags, description, disabled, suggested UTC hour). Only paging expectations and auth needs are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with count, offset, and project_id all documented in the schema. The description adds no syntax or format detail beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and a precise resource (periodic/cron-scheduled jobs of a Scrapy Cloud project), then enumerates the fields returned. This clearly separates it from siblings like scrapy_cloud_list_jobs (regular jobs) and scrapy_cloud_list_spiders.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the name and description; there is no explicit guidance on when to choose this over scrapy_cloud_get_periodic_job (single job) or scrapy_cloud_list_jobs. No prerequisites or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_list_projectsAInspect
List the Scrapy Cloud projects the configured account can access, with job counts (pending/running/finished) and, when available, name and organization id. Use it to find a project id, or the organization id that zyte_api_usage_stats needs.
| Name | Required | Description | Default |
|---|---|---|---|
| search | No | Only projects whose name contains this text, case-insensitive. | |
| organizationId | No | Only projects of this organization (the id zyte_api_usage_stats takes). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load. It does disclose the return shape (pending/running/finished job counts, name and org id 'when available'), which is genuinely useful and hints the account scope is enforced server-side. It says nothing about pagination, result limits, authentication requirements, or whether the listing is truncated, which are the remaining behavioral gaps for a list endpoint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the payload description is front-loaded before the routing guidance. Every clause carries information an agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description usefully enumerates what comes back, and both optional filters are covered by the schema. The only shortfall is the absence of any note on result volume/pagination for a listing whose size depends on the account.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'search' (case-insensitive name substring) and 'organizationId' (numeric id, same one zyte_api_usage_stats takes) are fully documented in the schema. The description adds only the cross-tool note about the organization id, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the Scrapy Cloud projects the configured account can access') and even describes the payload (job counts, name, organization id). It does not explicitly name or contrast with the nearest siblings (scrapy_cloud_get_project for a single project, scrapy_cloud_list_jobs), so it falls short of a 5 by the stated rubric.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case and downstream link: 'Use it to find a project id, or the organization id that zyte_api_usage_stats needs.' That routes the agent to the right follow-up tool. There is no statement of when NOT to use it or a named alternative lister, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_list_spidersAInspect
List the spiders deployed in a Scrapy Cloud project: name ('id'), tags, type and deployed version. Use it to find the spider name before scrapy_cloud_run_spider or scrapy_cloud_create_periodic_job.
| Name | Required | Description | Default |
|---|---|---|---|
| search | No | Only spiders whose name contains this text, case-insensitive. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the returned fields, which implies a safe read-only listing, but says nothing about pagination, project scoping, or auth requirements for a tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the primary action and return shape, then the usage routing. No wasted language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter list tool with no output schema, the description covers the action, return fields, and usage context adequately. Minor gaps remain around pagination and whether results are scoped or limited.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'project_id' and 'search' (including case-insensitive matching). The description's note that the spider name maps to 'id' is useful but concerns output, not input parameters, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the spiders deployed in a Scrapy Cloud project') and enumerates the returned fields (name/'id', tags, type, deployed version). This clearly differentiates it from siblings like scrapy_cloud_list_jobs or scrapy_cloud_list_collections.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: use it to find the spider name before scrapy_cloud_run_spider or scrapy_cloud_create_periodic_job. Clear context and downstream alternatives, though it offers no explicit 'when not to use' or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_put_collection_recordAInspect
Write one record to a Scrapy Cloud collection under a key, replacing any record with that key. Creates the collection if it does not exist. Use it to store configuration or state that spiders read.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | Record key. Writing an existing key replaces the record. | |
| type | No | Collection type: 's' store (default), 'cs' cached store (records expire after a month), 'vs' versioned store, 'vcs' versioned cached store. | s |
| value | Yes | The record, a JSON object of up to 1 MB. | |
| collection | Yes | Collection name, as listed by scrapy_cloud_list_collections. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It usefully discloses overwrite semantics and collection auto-creation, but does not mention required permissions, rate limits, or what the operation returns. For a mutation tool with zero annotation coverage, this is adequate but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and behavior, with no filler. The use-case sentence is short and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, overwrite behavior, collection creation, and a primary use case. With a rich schema and no output schema, this is mostly complete, though it omits auth/permission and error context that an agent might need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all five parameters are documented in the schema. The description reinforces the key semantics ('replacing any record with that key') but adds little beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Write') and resource ('record to a Scrapy Cloud collection under a key'), and clarifies overwrite and auto-create semantics. It does not explicitly name a sibling alternative or contrast with read/delete siblings, so it falls short of the top tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear use case ('store configuration or state that spiders read'), giving context for when to use it. It does not state when not to use it or name alternatives like get_collection_records, so no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_run_scriptAInspect
Start a Scrapy Cloud job that runs a standalone Python script deployed with the project (declared under 'scripts' in its setup.py), instead of a spider. Pass the script's command-line arguments as one string in 'args'. Returns the job key and dashboard URL; the job starts in state 'pending'. Check progress with scrapy_cloud_get_job and read its output with scrapy_cloud_get_job_log.
| Name | Required | Description | Default |
|---|---|---|---|
| args | No | Command-line arguments for the script, as one string, e.g. '--limit 100 --dry-run'. | |
| tags | No | Tags to add to the job. | |
| units | No | Scrapy Cloud units for the job. Default is the project's setting. | |
| script | Yes | Script file name as deployed, e.g. hello.py; the 'py:' prefix Scrapy Cloud uses is optional. | |
| priority | No | Queue priority, 0 (lowest) to 4 (highest). Default 2. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| job_settings | No | Scrapy settings overriding the project's, readable by the script via sh_scrapy.utils.get_project_settings. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses the initial job state ('pending'), the returned artifacts (job key and dashboard URL), and the deployment requirement for the script. It omits auth/permission requirements and quota/unit consumption behavior, which for a job-launching mutation would be valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the primary action, then the argument convention, then the return value and follow-up tools. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description explains the return payload (job key, dashboard URL) and the follow-up tools needed to observe completion, which is what an agent needs for a fire-and-poll job launch. Minor gap: no mention of required credentials or what happens if the named script is not deployed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter already carries its own documentation and the baseline is 3. The description only restates the 'args' semantics already given in the schema and adds nothing about units, priority, tags, or job_settings beyond what the structured fields state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start a Scrapy Cloud job that runs a standalone Python script') and immediately scopes it against the sibling by saying it runs a script 'instead of a spider', which distinguishes it from scrapy_cloud_run_spider. It also clarifies the deployment precondition (declared under 'scripts' in setup.py).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context and routes the agent forward: 'Check progress with scrapy_cloud_get_job and read its output with scrapy_cloud_get_job_log.' It implies the script-vs-spider selection condition, but never states explicitly when NOT to use this tool (e.g. use run_spider for spiders), and gives no prerequisites around project deployment state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_run_spiderAInspect
Start a new Scrapy Cloud job for a spider now. Returns the job key and dashboard URL; the job starts in state 'pending' and is picked up by the queue. Fails with 'already scheduled' if an identical job is pending or running. Check progress with scrapy_cloud_get_job. For a recurring schedule use scrapy_cloud_create_periodic_job instead; for a standalone script use scrapy_cloud_run_script.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tags to add to the job. | |
| units | No | Scrapy Cloud units for the job. Default is the project's setting. | |
| spider | Yes | Spider name, as listed by scrapy_cloud_list_spiders. | |
| job_args | No | Spider arguments, passed as -a name=value; numbers and booleans are sent as strings ('100', 'true'). | |
| priority | No | Queue priority, 0 (lowest) to 4 (highest). Default 2. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| job_settings | No | Scrapy settings overriding the project's, for example {"CLOSESPIDER_ITEMCOUNT": 100}. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses the async lifecycle (starts 'pending', picked up by queue), the return payload (job key and dashboard URL), and an important dedup failure mode ('already scheduled'). It omits auth/permission requirements and any expectation about how long a job waits in queue, which is the only meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the action and return value, then failure behavior, then sibling routing. Every sentence carries distinct, actionable information with no repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no output schema and no annotations, the description fills the critical gaps: it says what is returned (job key, dashboard URL), that execution is asynchronous, that duplicate submissions are rejected, and which tool to use instead for adjacent use cases. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (spider, project_id, job_args, units, priority, tags, job_settings) is already documented in the schema, including the job_args stringification rule. The description adds no parameter-level detail beyond 'for a spider', so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a new Scrapy Cloud job for a spider now') and immediately distinguishes itself from the two nearest siblings, scrapy_cloud_create_periodic_job and scrapy_cloud_run_script, by name. An agent can route to it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules for all three cases: recurring schedules go to create_periodic_job, standalone scripts go to run_script, and progress checks go to get_job. It also states the condition under which this tool itself fails ('already scheduled' when an identical job is pending or running).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_update_job_tagsAInspect
Add and/or remove tags on a Scrapy Cloud job, for example to mark it 'consumed' after processing its items. Returns the job's tags after the change. Tags can be used to filter scrapy_cloud_list_jobs (has_tag, lacks_tag).
| Name | Required | Description | Default |
|---|---|---|---|
| add | No | Tags to add. | |
| job | Yes | Job key in project/spider/job form, for example 123/1/4. | |
| remove | No | Tags to remove. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return value ('returns the job's tags after the change'), which is genuinely useful given there is no output schema, but says nothing about permissions required, whether re-adding an existing tag is idempotent, or how a missing tag behaves on remove.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero waste, and the operation is front-loaded. Each sentence adds distinct value: the operation, the return value, and the cross-tool linkage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description steps in to explain the return value and the tag-filtering use case. It is nearly complete for a 3-parameter mutation, though auth/permission expectations remain unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with add/job/remove all documented, including the job key pattern. The description adds no syntax or format detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb pair (add and/or remove) and resource (tags on a Scrapy Cloud job), plus the return value. An agent can distinguish it from every sibling in the list, since no other tool manipulates job tags.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete motivating scenario ('mark it consumed after processing its items') and links to the consumer of tags, scrapy_cloud_list_jobs with has_tag/lacks_tag. It does not state any when-not conditions or prerequisites, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_update_periodic_jobAInspect
Change a periodic job: reschedule (cron), pause or resume (disabled), or replace its tasks, tags or description. Only the given fields change. Pausing with disabled=true is the reversible alternative to deleting.
| Name | Required | Description | Default |
|---|---|---|---|
| cron | No | Schedule in UTC as 5 cron fields, each a number or '*'. Example: '0 9 * * 1' is Mondays 09:00 UTC. No ranges, steps or lists. | |
| tasks | No | Replaces the whole task list when given. | |
| addtags | No | Replaces the whole tag list when given. | |
| disabled | No | true pauses the schedule without deleting it, false resumes it. | |
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| description | No | ||
| periodic_job_id | Yes | Periodic job id, as listed by scrapy_cloud_list_periodic_jobs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load, and it delivers two genuinely valuable traits: 'Only the given fields change' establishes partial-update semantics, and the pause/delete note establishes reversibility. It omits permission requirements, error behavior, and whether changes take effect immediately, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the operation and its field-level effects, ending on the one strategic decision an agent faces. No filler and nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with no annotations and no output schema, the description covers the essential semantics: partial updates and the replace-wholesale behavior of tasks/tags (via 'replace'). It is arguably thin on auth/permission prerequisites and return behavior, which keeps it just below fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents cron, tasks, tags, disabled, and the ids. The description maps fields to intents (cron=reschedule, disabled=pause/resume) but adds little syntax or meaning beyond the already-rich schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Change) and resource (a periodic job) and then enumerates the exact mutations possible: reschedule via cron, pause/resume via disabled, or replace tasks/tags/description. It also distinguishes itself from the sibling scrapy_cloud_delete_periodic_job by naming pause as the reversible alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the pause case as the reversible alternative to deleting, which is real decision guidance against a sibling. However it gives no guidance relative to the other close siblings (create, get, list periodic jobs), so it is clear context rather than a complete when/when-not map.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrapy_cloud_update_project_settingsAInspect
Change a Scrapy Cloud project's stored Scrapy settings and/or default job units. Settings are merged: only the given keys change, and a null value removes a setting. Applies to jobs started afterwards. This changes how every spider in the project runs; confirm with the user first.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Scrapy Cloud project id (the numeric id in the dashboard URL). | |
| scrapy_settings | No | Scrapy settings to set on the project, merged into the existing ones; a null value removes that setting. | |
| default_job_units | No | Default Scrapy Cloud units for the project's jobs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden, and it does well: merge semantics, null-value removal, forward-only application, and project-wide blast radius are all disclosed. It omits permission/auth requirements and whether changes are reversible, which are the remaining behavioral gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: what changes, how it merges, and the caution. The most decision-relevant caveat (project-wide effect, confirm first) is placed last for emphasis without burying the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no annotations and no output schema, the description covers mechanism (merge/null), timing (future jobs), and risk (confirm with user). It could still note that only settings/units are mutable and what a successful call returns, but nothing required for a correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the three parameters are already documented; baseline is 3. The description restates the merge/null-removal contract already present in the scrapy_settings schema, adding reinforcement but no new syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Change a Scrapy Cloud project's stored Scrapy settings and/or default job units'), naming the exact scope of mutation. An agent can immediately distinguish it from the read-side sibling scrapy_cloud_get_project_settings and from job-level tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives real context: changes apply only to jobs started afterwards, they affect every spider in the project, and the user should be consulted first. It stops short of explicitly naming the alternative (e.g. altering a single job's settings), so it is clear but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchAInspect
Search the web through Zyte API: submit a keyword query to a supported search engine domain (currently Google domains, e.g. "google.com") and get structured results back. Use this instead of pointing fetch_page or fetch_http at a search-engine URL — SERP HTML is huge and bot-protected, and this tool returns parsed results directly; use extract_from_browser or extract_from_http for structured data from a specific known page, and the fetch tools for raw pages. 'include' picks the response components: "organic" — structured organic results (rank, url, title, snippet, sitelinks), usually what you want; "aiOverview" — the SERP's AI-generated answer with citations, returned only when the search engine shows one (absence is not an error); "html" — the raw SERP page (large). maxResults asks the engine for up to that many organic results (10-100 in steps of 10, default 10). queryParameters tunes the search in exactly one style: "generic" is portable — geolocation (a country code like "US" or a canonical geotarget name like "New York,New York,United States") and locale (e.g. "en-US") — while "engineSpecific" passes native Google parameters (uule, gl, hl, cr, lr, safe, nfpr) straight through. Zero organic results is a success, not an error. Returns a JSON metadata block (canonical SERP url, status, fetchedAt, totals; on status "partial" an error naming the failed component), then a JSON block with organicResults/aiOverview when requested, then the raw HTML when requested. An unsupported engine domain is rejected with an error naming the supported engines — do not retry it unchanged.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The keyword or phrase to search for (1-2048 chars). Zyte API constructs the canonical SERP URL from it; do not URL-encode. | |
| domain | Yes | Search engine domain in registrable form, lowercase (e.g. 'google.com', 'google.co.uk'). Only engines supported by Zyte API are accepted — currently Google domains (any google.* country domain); an unsupported domain is rejected with an error naming the supported set (do not retry it unchanged). | |
| include | Yes | Response components to return, each enabling one output field: 'organic' -> structured organic results (rank, url, title, snippet, sitelinks) — usually what you want; 'aiOverview' -> the SERP's AI-generated answer with citations, returned only when the SERP shows one; 'html' -> the raw SERP HTML (large). Example: ["organic"]. | |
| maxResults | No | How many organic results to ask the engine for (Google's 'num'; default 10). Also affects the fetched SERP page when 'html' is included. | |
| organizationId | Yes | Required. The Zyte organization to attribute this call to (max 100 characters, printable ASCII without spaces). If you do not already have an id, call the user_info tool: it lists the organizations your credential belongs to. IMPORTANT: if it lists more than one, ask the user which to use and wait for their answer — this call is billed to whichever organization you name here, so it is the user's choice to make, not yours. Never guess an id, and never fall back to a default. Once the user has chosen, reuse that id across the session unless they ask for a different organization. | |
| queryParameters | No | Search-tuning parameters in exactly one style: 'generic' is portable and translated per engine; 'engineSpecific' passes native Google parameters through. The styles cannot be mixed in one request. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses that absent aiOverview is not an error, that zero results is a success, the return-block order (metadata, then organicResults/aiOverview, then raw HTML), the 'partial' status error behavior, and the non-retryable unsupported-domain error. This is unusually thorough behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and alternative routing, then dense per-parameter semantics. It is long, but nearly every clause earns its place given no annotations and no output schema; only minor trimming would be possible without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description explains the return structure, error/partial behavior, and success conditions, and covers the nested queryParameters styles. An agent has everything needed to invoke it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description nevertheless adds real meaning beyond the schema — 'include' component-by-component intent, maxResults as '10-100 in steps of 10, default 10', and the two mutually exclusive queryParameters styles with portable-vs-native framing. It clarifies rather than merely repeats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search the web through Zyte API: submit a keyword query... get structured results back') and explicitly separates itself from four siblings by name. An agent can identify the tool's role without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use this instead of pointing fetch_page/fetch_http at a search-engine URL, and routes to extract_from_browser/extract_from_http for known pages and the fetch tools for raw pages. It also states that zero organic results is success and that an unsupported engine domain should not be retried unchanged.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
user_infoAInspect
List the Zyte organizations your credential belongs to — for "who am I / what account am I on", for an id to pass as another tool's organizationId, or to check whether the credential covers more than one. Takes no arguments; the answer is always about the credential's own user. IMPORTANT: when more than one comes back, ask the user which to use and wait for their answer before calling any tool that takes an organizationId. Do not choose for them, and treat nothing as an obvious default — not the first listed, not the only one with a plan flag, not the one whose name resembles the work. Every Zyte API call is billed to the organization it names, so guessing spends money on an account the user did not pick; offer them by name and id. With exactly one organization there is nothing to ask: use it, and reuse it all session. Needs an OAuth access token in the Authorization header. Returns 'organizations', each with id, name, website and the flags 'useZapi' and 'useScrapycloud' for a live Zyte API or Scrapy Cloud plan (one switched off upstream counts as neither). Only organizations with at least one such plan are listed, and the listing stops at 1000. Zero organizations is a success, not an error.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses the OAuth access token requirement in the Authorization header, the filtering rule (only orgs with at least one live plan are listed), the 1000-item listing cap, and that zero organizations is a success rather than an error. It also explains the billing consequence that motivates the user-confirmation rule.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and the sentence order (what it does, why, how to behave, auth, return shape, edge cases) is logical. It is on the long side and the triple negative example ('not the first listed, not the only one with a plan flag, not the one whose name resembles the work') is slightly repetitive, though it does add concrete clarification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description still documents the return shape ('organizations' with id, name, website, useZapi and useScrapycloud flags) and semantics of the flags. Combined with the auth and edge-case notes, an agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which sets a baseline of 4. The description reinforces this with 'Takes no arguments; the answer is always about the credential's own user,' adding meaning beyond the empty schema without needing to document any parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('List the Zyte organizations your credential belongs to') and enumerates the three concrete reasons to call it: identity check, obtaining an organizationId, and detecting multi-org credentials. No sibling tool overlaps this identity-listing role, so an agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly covers the when: it names the follow-on use (passing the result as another tool's organizationId) and gives concrete decision rules for the ambiguous case — ask the user when multiple orgs come back, never default to the first/plan-flagged/similar-named one, and reuse the single org all session when exactly one exists. This is when/how guidance with an explicit exclusion ('do not choose for them').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zyte_api_domain_pricingAInspect
Whether Zyte API supports a website and what a request to it costs: pay-as-you-go USD per 1000 requests for httpResponseBody and browserHtml, with tier names. 'behind_login' means supported but priced only after sign-in; 'unavailable' means not supported; 'no_permanent_tier' means no price yet. A subdomain is priced as its registered domain, returned as 'pricing_domain'. Public data, no account needed. Not for plan prices, limits or discounts, and not for what you already spent (use zyte_api_usage_stats).
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Website to look up, e.g. amazon.com. A full URL or 'www.' prefix is accepted and reduced to the bare domain. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does so well: it discloses no-account/public-data access, decodes the return statuses ('behind_login', 'unavailable', 'no_permanent_tier'), and explains subdomain normalization. This is exactly the behavioral context the agent needs and cannot get from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then behavior, then exclusions in a logical order, with every sentence carrying information. Slightly dense across three long sentences and could be tightened, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only one fully-documented parameter, the description supplies everything an agent needs: purpose, cost semantics, status meanings, normalization, auth posture, and exclusions. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds semantics beyond the schema: 'A subdomain is priced as its registered domain, returned as pricing_domain', clarifying normalization behavior the schema only hints at. Modest but genuine added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: reports whether Zyte API supports a website and its pay-as-you-go cost, with the exact unit (USD per 1000 requests) and fields (httpResponseBody, browserHtml). It differentiates from siblings by naming zyte_api_usage_stats as the tool for spend data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly carves out what this tool is NOT for: plan prices, limits, discounts, and past spend, routing the latter to zyte_api_usage_stats. The condition 'public data, no account needed' also tells the agent this requires no auth, which is actionable usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zyte_api_usage_statsAInspect
Recorded Zyte API usage of the organization: requests made, cost, response times and status codes for a time window (default: the last 7 days), as one summary or broken down with group_by into one row per hour/day/month/year, per target domain (sorted by cost), per requested feature (browserHtml, httpResponseBody, screenshot, actions, ...) or per automatic extraction type (product, article, ...). Filter by domain, API key label, status code, feature, extraction type or request tags ('tag:value'). Costs are in the currency's main unit (e.g. USD), one entry per currency. Only recorded usage: it cannot estimate a future crawl (use zyte_api_domain_pricing for rates) and knows nothing about Scrapy Cloud jobs. Rate limit: 20 Stats API calls per minute; group_by 'feature' spends 8 of them and 'extraction_type' 11, so do not repeat those within a minute.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Only requests carrying all of these tags; each entry is 'tag' or 'tag:value'. | |
| domains | No | Only requests to these domains. | |
| feature | No | Only requests that used this feature. | |
| end_time | No | End of the window, UTC, same formats. A bare date is included whole. Default: now. | |
| group_by | No | 'none' for one total; 'hour'/'day'/'month'/'year' for one row per time bucket; 'domain' for one row per target domain; 'feature' or 'extraction_type' for one row per Zyte API feature or automatic extraction type used. | none |
| start_time | No | Start of the window, UTC, as an ISO 8601 date (2026-08-01) or date-time (2026-08-01T12:00:00Z). Default: 7 days ago. | |
| status_codes | No | Only requests that got these HTTP status codes. | |
| domain_health | No | Add Zyte's health status per domain (top 100 domains of the last 7 days). Implies group_by=domain. | |
| api_key_labels | No | Only requests made with API keys having these labels. | |
| organizationId | Yes | Required. The Zyte organization whose usage to report: the number in https://app.zyte.com/o/<id>, the 'organizationId' of a Scrapy Cloud project, or one of the organizations the user_info tool lists. Never guess an id; once chosen, reuse it across the session unless the user asks for another. | |
| extraction_from | No | Only extraction requests from this source. | |
| extraction_type | No | Only requests with this automatic extraction type. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the 20 calls/minute rate limit, that group_by 'feature' spends 8 and 'extraction_type' 11 of those, that costs are in the currency's main unit with one entry per currency, the default 7-day window, and that only recorded usage is covered. This is unusually rich behavioral context for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and grouping/filter capabilities are front-loaded in the first sentence, followed by limitations and rate limits. It is dense and reads as two very long sentences, which slightly hurts scanability, but essentially every clause carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter analytics tool with no output schema, the description covers the window, grouping modes, filter dimensions, currency semantics, limitations, and rate-limit cost of expensive group_by values. The row-per-bucket semantics are also spelled out in the schema, so nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond it: the expected 'tag:value' filter syntax, that domain grouping is sorted by cost, and the conceptual mapping of each filter dimension. It stops short of explaining defaults for every filter, keeping it at 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Recorded Zyte API usage of the organization') and enumerates the exact metrics returned (requests, cost, response times, status codes). It explicitly distinguishes itself from siblings by naming zyte_api_domain_pricing and ruling out Scrapy Cloud jobs, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (historical usage over a time window with grouping/filtering) and when-not ('cannot estimate a future crawl – use zyte_api_domain_pricing'; 'knows nothing about Scrapy Cloud jobs'). It also warns not to repeat expensive group_by calls within a minute, which is actionable routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Publisher details
- Operator
- Not applicable
- Operator website
- https://zyte.com · Publisher source
- Vendor relationship
- First-party
- Documentation
- https://docs.zyte.com/zyte-web-data/mcp.html · Publisher source
- Trust center
- Not applicable
- Restrictions
- Requires Zyte account · Publisher source
Related MCP Connectors
Zyte API MCP — unified web fetch + AI extraction (zyte.com)
Your agent needs the open web — searched by more than one engine, and read as clean markdown rather than raw HTML. **What you can ask for** • "Search this question with two providers and tell me where they disagree." • "Scrape these 40 URLs into markdown, in one batch." • "Crawl this documentation site and give me every page." • "Do deep research on this topic and cite the sources." • "Find the academic papers behind this claim." **How to use it** Point any MCP client at https://mcp.aisa.one/search/mcp and sign in with OAuth — there is no key to create or paste. 30 tools across several independent providers: Tavily and Exa search, answers, contents and agent runs; Firecrawl scrape, batch scrape, crawl, map and search; Perplexity Sonar, Sonar Pro, reasoning and deep research; Oxylabs AI search and LLM jobs; OpenAI and Anthropic web search; and scholarly search. **Why this rather than the source** Several independent indexes behind one account, because one engine's blind spot is not visible from inside it. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Find the page here, then ask the same agent who links to it or how much traffic it gets — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/seo-serp/mcp for the Google results page itself, https://mcp.aisa.one/seo-serp-other-engines/mcp for Bing, Baidu and Naver.
Your agent needs marketplace data — what a product costs on Amazon and Google Shopping, who the sellers are, what reviewers actually complain about. **What you can ask for** • "What is this ASIN's price history, rating and seller list?" • "Who else sells this product, and at what price?" • "Pull the reviews for this product and group the complaints." • "What comes up on Google Shopping for this query in the UK?" • "Compare these products across both marketplaces." **How to use it** Point any MCP client at https://mcp.aisa.one/seo-merchant/mcp and sign in with OAuth — there is no key to create or paste. 22 tools: Amazon products, ASIN detail and sellers; Google Shopping products, product info, sellers and reviews; live and queued forms, with raw HTML where you need it. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Price the product here, then ask the same agent what the brand's site traffic or ad spend looks like — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/seo/mcp for all of it at once — rankings, keywords, backlinks, site health and AI-answer visibility across DataForSEO, Semrush and Ahrefs.
Your agent needs live data — a competitor's traffic, who to contact there, what people are saying, what Google and ChatGPT answer about you, a company's filings. Normally that is six vendor accounts, six sets of keys and six SDKs. This is one URL. **What you can ask for** • "How much traffic does stripe.com get, where does it come from, and who competes for the same keywords?" • "Find 20 Series-B fintech companies in Germany and the heads of marketing there, with emails." • "Does ChatGPT mention our brand when someone asks for the best CRM — and what does it cite?" • "What is X saying about $NVDA today, and what did the stock actually do?" • "Search the web for this, then scrape the three best pages into markdown." **How to use it** Point any MCP client at https://mcp.aisa.one/mcp and sign in with OAuth — there is no key to create or paste. Then just ask: the agent calls search to find the right operation and use to run it. **Why this rather than the source** 26 sources behind one account and one bill — DataForSEO, Semrush, Ahrefs, Similarweb, Apollo, X/Twitter, Instagram, Reddit, Pinterest, YouTube, Tavily, Exa, Perplexity, Firecrawl, CoinGecko, Kalshi, Polymarket, AgentMail and more, 580+ operations. tools/list returns five tools, not 580, so the introduction does not eat your context window. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** One slice at a time: https://mcp.aisa.one/seo/mcp · /finance/mcp · /social/mcp · /search/mcp · /sales/mcp · /mail/mcp · /gtm/mcp, or a single provider like /twitter-api/mcp. Same account, fewer tools listed, and search still reaches everything. Full list at https://mcp.aisa.one/servers
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceWeb scraping MCP server for Al agents. 6 tools: extract clean text/markdown from any URL, structured scraping with CSS selectors, full-page screenshots via Playwright, link extraction with regex filtering, metadata extraction (OG tags, Twitter cards), and Google search. Free tier: 50 requests/IP/day.8MIT

DLBrowserofficial
AlicenseNot gradedqualityBmaintenanceMCP server that gives AI agents reliable web access with self-healing, anti-bot bypass and clean content extraction. It provides 11 tools like fetch, scrape, and search, metered by credit with cost transparency.21 PyPI1MIT- AlicenseNot gradedqualityDmaintenanceMCP tool server that gives any AI agent the ability to search, scrape, and analyze content across the internet.42MIT
- AlicenseBqualityDmaintenanceMCP server for the Spider web crawling and scraping API, enabling AI agents to crawl, scrape, search, and extract web data.1347 npm2MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.