Skip to main content
Glama
paulet4a-commits

WebDataTools Developer, app & research data MCP server

shopify_products_scraper

Extract product data from any Shopify store's public products.json feed, returning title, price, discount, stock, SKU and images per product or variant.

Instructions

Shopify Store Products Scraper reads any Shopify store's public products.json feed — title, price, compare-at price, discount %, stock, SKU and images, one row per product or variant. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
storesYesStore URLs — Enter the Shopify store URLs to scrape, e.g. https://www.allbirds.com. A bare domain such as allbirds.com also works. Every store is checked against its public products.json feed; non-Shopify or password-protected stores get one error row instead of crashing the run. Example: ["https://www.allbirds.com"].
onlyAvailableNoOnly available (in-stock) items — Keep this off to include out-of-stock variants and products. Turn it on to drop any variant (or, in product mode, any product) that is not currently available for sale.
includeVariantsNoOne row per variant — Keep this on to get one dataset row per product variant (size/color), with its own price, SKU and stock. Turn it off to get one row per product with aggregated minPrice, maxPrice and anyAvailable instead.
collectionHandlesNoCollection handles (optional) — Optionally enter collection handles to scrape instead of the whole catalog, e.g. mens-shoes. Each handle is fetched from /collections/<handle>/products.json. Leave empty to scrape every product in the store via /products.json.
maxProductsPerStoreNoMax products per store — Enter the maximum number of products to fetch per store, e.g. 250. Products are paged 250 at a time from products.json, so 250 is a single page; raise it for a full catalog export, up to 20000.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does add real value: it discloses cost (~$0.001 per result, billed to the caller's own Apify account, lower on paid plans) and that reads come from a public feed. It does not state rate limits or how non-Shopify stores fail, though that error behavior is covered in the parameter schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the capability and followed by the pricing note; there is no filler. The response is dense but readable and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description correctly enumerates what comes back (title, price, compare-at price, discount %, stock, SKU, images). Combined with the fully covered parameter schema and disclosed billing, an agent has enough to invoke it correctly, though failure modes and pagination behavior are only implied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already thoroughly documented in the schema (store URLs, onlyAvailable, includeVariants, collectionHandles, maxProductsPerStore). The description adds only the 'one row per product or variant' framing and does not add syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('reads any Shopify store's public products.json feed') and enumerates the returned fields, cleanly distinguishing it from the other scraper siblings (GitHub, Hacker News, app stores). An agent immediately knows what data source this targets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes the context of use (any public Shopify store, one row per product or variant) and the billing model, but gives no explicit when-to-use/when-not guidance or named alternatives. The store-mode vs collection-mode choice, if anything, lives in the schema rather than the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.