shopify-admin-mcp-server
Provides tools for interacting with the Shopify Admin GraphQL API, enabling management of orders, products, variants, media, inventory and per-location fulfilment by SKU, customers, collections, discounts, abandoned checkouts, blog content, pages, metafields, bulk variant updates, and analytics.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@shopify-admin-mcp-servercheck where SKU TSHIRT-BLK-M is stocked"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Shopify MCP Server v4.0.1
Shopify Admin GraphQL API 2026-01 MCP server. 54 tools aligned to the OAuth scopes actually granted on the token — orders, products, variants, media, inventory & per-location fulfilment, customers, collections, discounts, abandoned checkouts, blog content, pages, metafields, bulk variant updates, and analytics.
Every GraphQL operation is validated against the 2026-01 schema, and all 24 read tools plus the inventory write tools are verified against a live store.
npx -y shopify-admin-mcp-serverRequires SHOPIFY_STORE_DOMAIN and SHOPIFY_ACCESS_TOKEN. Full setup in Setup.
Location fulfilment by SKU
The headline capability: turn a location's ability to stock and fulfil a variant on or off, addressed by SKU rather than by GID.
shopify_get_variant_locations { sku: "TSHIRT-BLK-M" }
→ active_locations: where it IS stocked, with available/on_hand/committed/incoming,
plus can_deactivate and any deactivation_blocked_reason
→ inactive_locations: where it is NOT stocked
→ inventory_item_id: for the stock tools below
shopify_set_variant_locations {
sku: "TSHIRT-BLK-M",
locations: [
{ location: "Melbourne Warehouse", activate: true }, // by name or GID
{ location: "Pop-up Store", activate: false }
]
}Backed by inventoryBulkToggleActivation, so multiple locations are applied in one call. Activating starts the location at 0 available — follow with shopify_set_inventory.
Deactivation is destructive: it discards that location's stock for the variant. Shopify refuses when the variant has committed stock or pending fulfilments there. Check can_deactivate first.
Guardrails: a SKU that matches zero or more than one variant is rejected rather than guessed at, and an unknown location name lists the valid ones.
Related MCP server: Shopify MCP Server
Tools (54)
Shop & Orders
Tool | Scopes | Description |
| — | Store info: name, domain, plan, currency, timezone |
|
| List orders in a date range with financials and line items |
|
| Single order by GID: addresses, fulfillments, refunds |
|
| Auto-paginated aggregate: revenue, AOV, product breakdown |
|
| Lightweight count for a date range |
Inventory & Locations
Tool | Scopes | Description |
|
| All store locations with addresses |
|
| By SKU: which locations stock/fulfil a variant, and which don't |
|
| By SKU: activate/deactivate fulfilment at locations |
|
| Quantities per location for a variant GID |
|
| Change stock by a delta ("we sold 3") |
|
| Set stock to an exact value, with compare-and-set |
|
| Move stock between locations transactionally |
|
| Per-SKU settings: tracked, unit cost, HS/HTS code, country of origin |
Customs / HTS data
shopify_update_inventory_item sets the customs fields Shopify needs before it will generate international
shipping labels — harmonizedSystemCode (6–13 digits; 6110.20.20 and 61102020 are both accepted),
countryCodeOfOrigin, provinceCodeOfOrigin, and per-destination countryHarmonizedSystemCodes overrides.
Country codes are case-insensitive. Read the current values back under customs in
shopify_get_variant_locations.
shopify_set_inventory defaults to compare-and-set: pass compareQuantity and Shopify rejects the write if
someone changed the value underneath you. ignoreCompareQuantity: true forces it, at the risk of clobbering a
concurrent update.
Bulk variant updates
Tool | Scope | What it does |
|
| Apply customs, cost, pricing, identifier and weight changes to many variants at once |
|
| Report which variants are missing HS codes or country of origin |
Select variants with skus, productIds, query, onlyMissingHsCode or onlyMissingOrigin,
combined with AND. The common customs backfill is one call:
onlyMissingHsCode: true, harmonizedSystemCode: "611030", countryCodeOfOrigin: "AU"dryRun defaults to true when price or compareAtPrice is set and false otherwise, so
pricing changes need an explicit dryRun: false. maxVariants (default 500) refuses rather than
truncating. sku cannot be changed in bulk.
Analytics & marketing
Tool | Scope | What it does |
|
| Run a ShopifyQL query and get a table back |
|
| Campaign/post events with UTM parameters |
|
| One customer's activity timeline |
Marketing events are only populated when a marketing integration publishes to Shopify; an empty list means none is connected.
Products, Variants & Media
Tool | Scopes | Description |
|
| List/search products with variants, inventory, pricing (paginated) |
|
| Add tags to every product matching a search query, paging server-side; skips already-tagged products, supports |
|
| Full detail: variants, options, media, collections, SEO |
|
| Create a product, with options and images |
|
| Update fields and variant prices/SKUs |
|
| Permanent delete (requires |
|
| Copy a product; the copy defaults to DRAFT |
|
| Attach images/video from public URLs |
|
| Add variants against the product's options |
|
| Permanent delete (requires |
Customers
Tool | Scopes | Description |
|
| Search by name, email, or any Shopify filter |
|
| Full profile with addresses and recent orders |
|
| Add tags to any Shopify resource |
|
| Remove tags from any Shopify resource |
Collections
Tool | Scopes | Description |
|
| List/search collections with SEO |
|
| Create a manual or smart (rule-based) collection |
|
| Update title, description, and SEO |
|
| Add products to a manual collection (async) |
|
| Remove products from a manual collection (async) |
Pages, Blogs & Articles
Tool | Scopes | Description |
|
| List/search store pages |
|
| Full page content and publish status |
|
| Update title, HTML body, handle, publish status |
|
| List/search blogs |
|
| Blog with 10 most recent articles |
|
| Update title, handle, comment policy |
|
| Articles for a blog |
|
| Full article with HTML, author, tags |
|
| Create a blog article |
|
| Update content, tags, author, publish status |
Search, Metafields & Commerce
Tool | Scopes | Description |
|
| Unified search across products, articles, blogs, pages |
| varies by resource | Get metafields for any resource |
| varies by resource | Create or update a metafield |
| varies by resource | Delete a metafield by GID |
|
| List draft orders with line items and totals |
|
| List all discounts (code + automatic) via |
|
| Abandoned checkouts with recovery URL and line items |
Not included (scopes not granted)
These are deliberately absent — the tools would 403 at runtime. Add the scopes to the custom app and reinstall to get a new token, then they can be built:
Capability | Scopes needed |
Create/cancel fulfilments, tracking numbers, move fulfilment between locations |
|
Order edits, refunds, cancellations, order tagging |
|
Returns |
|
Create/complete draft orders |
|
Create discount codes |
|
Publish products to sales channels |
|
Create/edit locations |
|
Architecture
src/
index.ts # entry point — wires the tool modules together
shopify-client.ts # config, GraphQL client (429 retry), errors, MCP response helpers
resolvers.ts # SKU → variant, location name → GID
variant-fields.ts # shared variant field validation/normalisation (HS codes, weight, etc.)
variant-select.ts # selector (skus/productIds/query/missing-customs) → concrete variant list
bulk.ts # field-agnostic bulk variant writer: grouping, chunking, per-product results
tools/
orders.ts # shop, orders, summary, count
products.ts # products, variants, media
inventory.ts # locations, variant⇄location activation, stock
customers.ts # customers, resource tagging
content.ts # pages, blogs, articles, metafields, search
commerce.ts # collections, draft orders, discounts, abandoned checkouts
bulk-variants.ts # bulk variant field updates, customs coverage audit
analytics.ts # ShopifyQL query tool
marketing.ts # marketing events, customer event timelines
tests/
*.test.ts # Vitest unit tests — one file per module under testEach module exports a single registerXTools(server) function. No file exceeds 800 lines.
Setup
1. Create a Shopify Custom App
Shopify Admin → Settings → Apps and sales channels → Develop apps
Click Create an app
Under Configure Admin API scopes, enable:
read_all_orders,read_ordersread_analytics,read_reports,read_customer_eventsread_checkoutsread_customersread_price_rules,read_discountsread_draft_ordersread_inventory,write_inventory,read_inventory_transfers,write_inventory_transfersread_locationsread_marketing_integrated_campaigns,read_marketing_eventsread_online_store_pages,write_online_store_pagesread_content,write_contentread_products,write_products
Click Install app
Copy the Admin API access token
To confirm what a token actually has, query { currentAppInstallation { accessScopes { handle } } }.
2. Install
Run straight from npm — no clone, no build:
npx -y shopify-admin-mcp-serverOr install it globally:
npm install -g shopify-admin-mcp-servergit clone <repo> shopify-admin-mcp-server
cd shopify-admin-mcp-server
npm install
npm run build3. Configure the MCP client
{
"mcpServers": {
"shopify": {
"command": "npx",
"args": ["-y", "shopify-admin-mcp-server"],
"env": {
"SHOPIFY_STORE_DOMAIN": "mystore.myshopify.com",
"SHOPIFY_ACCESS_TOKEN": "shpat_xxxxxxxxxxxxxxxxxxxxx"
}
}
}
}{
"mcpServers": {
"shopify": {
"command": "node",
"args": ["/full/path/to/shopify-admin-mcp-server/dist/index.js"],
"env": {
"SHOPIFY_STORE_DOMAIN": "mystore.myshopify.com",
"SHOPIFY_ACCESS_TOKEN": "shpat_xxxxxxxxxxxxxxxxxxxxx"
}
}
}
}4. Restart the client
Environment Variables
Variable | Required | Example |
| Yes |
|
| Yes |
|
Notes
All GIDs use the format
gid://shopify/ResourceType/12345Rate limits are handled automatically with retry-after backoff (up to 3 retries)
Responses over 100,000 characters are truncated with a notice
API version: 2026-01
discountNodesreplacescodeDiscountNodes(removed in 2026-01)Destructive tools (
shopify_delete_product,shopify_delete_variants) require an explicitconfirm: trueinventoryAdjustQuantities/inventorySetQuantitiesidempotency keys are optional until 2026-04, when they become required
Fixed in v4.0.0
Five tools were silently broken against 2026-01 and failed on every call. All are now fixed and verified live:
Tool | Bug |
| Used non-existent |
| Selected |
| Selected |
| Selected |
| Used four money fields that don't exist on |
productCreate / productUpdate were also migrated off the deprecated input: ProductInput! argument to product: ProductCreateInput / ProductUpdateInput.
Available Tools
54 toolsshopify_add_product_mediaAdd Media to Shopify ProductA
Attach images or video to a product from publicly reachable URLs.
Shopify fetches each URL asynchronously, so media comes back with status UPLOADED/PROCESSING and becomes READY shortly after. A URL Shopify cannot reach fails silently into a FAILED status — re-read the product with shopify_get_product_by_id to confirm.
| Name | Required | Description | Default |
|---|---|---|---|
| media | Yes | ||
| productId | Yes | Product GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the generic safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=true). The description goes well beyond that, disclosing asynchronous fetching, the UPLOADED/PROCESSING to READY lifecycle, and the important failure mode where unreachable URLs fail silently with a FAILED status.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and the operational caveat (async lifecycle and silent failure) following. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by describing the status progression and the FAILED case, plus a verification path. Nothing essential is missing for correct invocation or interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; url and alt are documented in the schema while type has an enum but no description. The description reinforces the crucial 'publicly reachable URL' constraint, which adds real value, but it says nothing about the media array structure or the type field, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (attach) and resource (images or video to a product) with the key constraint that sources are publicly reachable URLs. It is clearly distinguishable from siblings like shopify_update_product or shopify_get_product_by_id without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the context of use and explicitly routes the agent to shopify_get_product_by_id to verify a suspected silent failure. It does not state exclusions (e.g., when to prefer shopify_update_product for media edits), so it falls just short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_add_tagsAdd Tags to Shopify ResourceAIdempotent
Add tags to any Shopify resource you have write access to (products, articles, collections). Tags are additive — existing tags are preserved.
Note: requires write access on the target resource type (e.g. write_products for products).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Resource GID (product, collection, article, etc.) | |
| tags | Yes | Tags to add |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly=false, idempotent=true, and destructive=false. The description adds genuinely useful context beyond them: that tags are additive with existing tags preserved, and the exact write scope required per resource type.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loading the action and scope, with the permission caveat deferred to a note. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, only two well-documented params, and annotations cover the safety profile. The description is complete for correct invocation, missing only sibling disambiguation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (id as Resource GID, tags as string array), so the schema carries the semantics. The description adds no format or syntax detail beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (add tags) and resource, and enumerates the target types (products, articles, collections). It doesn't explicitly distinguish itself from siblings like shopify_remove_tags or shopify_bulk_tag_products, but the single-resource add scope is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition (write access on the target resource type) and scope of applicability. It stops short of naming when to prefer it over a sibling like shopify_bulk_tag_products or excluding cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_adjust_inventoryAdjust Shopify Inventory (by delta)A
Change stock by a RELATIVE amount at one location — e.g. delta: -3 to remove three units.
Accepts a SKU (preferred) or an explicit inventoryItemId. Location may be a GID or a name. To set an exact number instead of a delta, use shopify_set_inventory.
The variant must already be stocked at the location — activate it first with shopify_set_variant_locations.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact variant SKU | |
| name | No | Which quantity to adjust | available |
| delta | Yes | Change: positive adds stock, negative removes it | |
| reason | No | correction | |
| location | Yes | Location GID or location name | |
| inventoryItemId | No | InventoryItem GID — alternative to sku |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readOnly, non-destructive, non-idempotent, open-world behavior, so the safety profile is covered. The description adds real behavioral context beyond that: the relative-delta mutation model, the requirement that the item be stocked at the location first, and identifier flexibility (SKU preferred, GID or name for location). It does not mention permission requirements or response shape, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight paragraphs, front-loaded with the core semantic distinction, then the alternative tool, then the prerequisite. No filler sentences and no repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations-level detail on returns, yet the description covers everything needed to invoke the tool correctly: mutation model, identifier choices, location format, and the activation prerequisite. Only the return/error behavior and the enum-driven fields (reason, name) are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 83% so the baseline is 3, but the description adds genuine semantics: delta sign convention with an example, SKU as the preferred identifier with inventoryItemId as the alternative, and that location accepts either a GID or a name. It does not explain the 'name' (available vs on_hand) or 'reason' enums, which the schema already labels.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource plus the key modifier: 'Change stock by a RELATIVE amount at one location,' with a concrete example (delta: -3). It explicitly distinguishes itself from the sibling shopify_set_inventory, which adjusts to an exact number.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative tool and the condition for choosing it ('To set an exact number instead of a delta, use shopify_set_inventory'), and states a prerequisite with the fix ('The variant must already be stocked at the location — activate it first with shopify_set_variant_locations'). Both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_analytics_queryQuery Store Analytics (ShopifyQL)ARead-onlyIdempotent
Run a ShopifyQL query against the store's analytics and get back a table.
ShopifyQL shape: FROM SHOW [GROUP BY ] [SINCE ] [UNTIL ] [ORDER BY ] [LIMIT n]
Datasets include sales, orders, products, customers.
Examples: FROM sales SHOW total_sales GROUP BY month SINCE -12m ORDER BY month FROM sales SHOW total_sales, orders GROUP BY product_title SINCE -30d ORDER BY total_sales DESC LIMIT 10 FROM orders SHOW average_order_value SINCE -90d
A syntax error is reported as an error, not as an empty table — an empty result genuinely means no matching data.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ShopifyQL query, e.g. 'FROM sales SHOW total_sales SINCE -30d' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld/non-destructive, so safety is covered. The description adds genuinely non-structured behavior: a syntax error surfaces as an error rather than an empty table, so an agent can distinguish 'broken query' from 'no data'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary, then grammar, datasets, examples, and an error caveat. Each example demonstrates a different capability (time grouping, multi-metric + limit, simple scalar), so no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description tells the agent the return shape (a table) and the error/empty-result distinction. For a single-parameter query tool with full annotation coverage, nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema only says 'a ShopifyQL query', so the description carries the real semantics: the clause grammar, allowed datasets, and three concrete query templates. That is well beyond the single-line schema hint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Run/Query) and resource (ShopifyQL analytics) and a concrete result (a table). It is unmistakably distinct from the CRUD-oriented siblings like shopify_list_orders or shopify_get_products.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The grammar shape, the enumerated datasets (sales, orders, products, customers) and three worked examples make clear this is the tool for aggregate/analytics questions. It never explicitly says when NOT to use it or which sibling to prefer for raw record listing, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_audit_variant_customsAudit Variant Customs CoverageARead-onlyIdempotent
Report which variants are missing customs data — the HS/HTS code and country of origin Shopify needs before it will generate international shipping labels.
Returns totals, the distinct HS codes already in use with their variant counts, and the list of variants missing data. Use this to aim a backfill before running shopify_bulk_update_variants.
Read-only — it never writes.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Cap on listed variants; totals always cover the whole catalogue | |
| includeVariants | No | Include the per-variant list of what is missing |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered; the closing 'Read-only — it never writes' is largely redundant. However, the description usefully discloses the return contents (totals, distinct HS codes with variant counts, list of missing variants) and the business consequence of missing data (labels won't generate), which adds real context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then the return shape, then the usage routing in three tight sentences. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing what is returned (totals, distinct HS code counts, missing-variant list), which is what an agent needs to chain into a backfill. Annotations cover the safety profile, so nothing critical is missing, though pagination/scale behavior of the returned list is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (limit, includeVariants) are fully documented in the schema, including the important nuance that totals always cover the whole catalogue. The description adds nothing about these parameters, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report/audit) and resource (variant customs data), and defines what 'missing customs data' means concretely (HS/HTS code and country of origin). An agent can distinguish this from sibling variant tools like shopify_get_variant_locations or shopify_bulk_update_variants without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the use case — 'aim a backfill before running shopify_bulk_update_variants' — naming the downstream sibling tool and the condition that selects it. The agent knows both when to call this and what to call next.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_bulk_tag_productsBulk Add Tags to Matching ProductsAIdempotent
Add tags to every product matching a Shopify search query, paging through all matches server-side in one call — no manual pagination or per-product tagging round trips.
Example: query: "title:jersey OR title:bib OR title:t-shirt", tags: ["reviewsizing"]
Skips products that already carry every tag given. Set dryRun:true to preview the match set (and see which products would be skipped as already-tagged) before writing anything.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | Yes | Tags to add to every matching product | |
| query | Yes | Shopify product search query, e.g. "title:*jersey* OR title:*bib*" | |
| dryRun | No | If true, report matches without writing any tags | |
| maxProducts | No | Safety cap on how many products a single call will touch |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=false and idempotentHint=true already declared, the description still adds real value: server-side paging across all matches, skipping of already-fully-tagged products (explaining the idempotency), and a dryRun preview mode. It doesn't discuss auth scope or rate limits, but covers the mutation's practical behavior well.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with the core capability, followed by example, then skip/dryRun behavior. Every sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation with no output schema, the description covers what happens, what's skipped, and how to preview safely. The maxProducts safety cap is left entirely to the schema, a minor gap in an otherwise complete picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds a worked query-syntax example ('title:*jersey* OR title:*bib* OR title:*t-shirt*') and clarifies dryRun's preview-and-skip reporting semantics beyond the schema's terse wording.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Add tags to every product matching a Shopify search query', with the server-side paging behavior called out. This clearly distinguishes it from single-product tagging siblings like shopify_add_tags.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage example and an explicit preview workflow ('Set dryRun:true to preview the match set ... before writing anything'), which tells the agent when to use it. It stops short of naming the alternative sibling (shopify_add_tags) or stating when not to use bulk tagging.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_bulk_update_variantsBulk Update VariantsAIdempotent
Apply the same field changes to many variants in one go — customs data, cost, pricing, identifiers and shipping weight.
Select the variants with any combination of skus, productIds, query, onlyMissingHsCode and onlyMissingOrigin (combined with AND). At least one selector is required.
The common customs backfill is: onlyMissingHsCode: true, harmonizedSystemCode: '611030', countryCodeOfOrigin: 'AU'
dryRun defaults to TRUE when price or compareAtPrice is set, and FALSE otherwise — so pricing changes require an explicit dryRun: false, while customs and cost writes apply immediately. A dry run reports exactly which variants would change and writes nothing.
maxVariants (default 500) is a hard refusal, not a truncation: if the selector matches more, nothing is written and you are told the match count.
SKU cannot be changed here — bulk-rewriting SKUs would break the selectors used to address variants. Use shopify_update_inventory_item for single-SKU changes.
| Name | Required | Description | Default |
|---|---|---|---|
| cost | No | Unit cost e.g. '12.50' | |
| skus | No | Exact variant SKUs | |
| price | No | Price e.g. '29.99' | |
| query | No | Shopify variant search, e.g. 'product_type:Socks' | |
| dryRun | No | Preview without writing. Defaults true when price fields are set. | |
| barcode | No | Barcode / GTIN | |
| taxCode | No | ||
| taxable | No | ||
| tracked | No | Whether Shopify tracks stock for this SKU | |
| productIds | No | Product GIDs — targets every variant of each | |
| weightUnit | No | ||
| maxVariants | No | Refuse if the selector matches more than this | |
| weightValue | No | Shipping weight, e.g. 0.25 | |
| compareAtPrice | No | Compare-at price e.g. '39.99' | |
| inventoryPolicy | No | Whether to keep selling when out of stock | |
| requiresShipping | No | false for digital/service items | |
| onlyMissingHsCode | No | Restrict to variants with no HS code | |
| onlyMissingOrigin | No | Restrict to variants with no country of origin | |
| countryCodeOfOrigin | No | 2-letter ISO country code e.g. 'AU' | |
| harmonizedSystemCode | No | HS/HTS code, 6-13 digits e.g. '611030' or '6110.30' | |
| provinceCodeOfOrigin | No | Province/state code e.g. 'BC' | |
| countryHarmonizedSystemCodes | No | Destination-specific HTS overrides — replaces the existing list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly/destructive/idempotent/openWorld; the description adds far more: dryRun defaults to true only when price or compareAtPrice is set, maxVariants is a hard refusal rather than a truncation, dry runs write nothing and report the affected variants, and countryHarmonizedSystemCodes replaces the existing list. These are exactly the mutation-safety details an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the tool's effect, then groups selectors, dry-run behavior, limits, and exclusions into short scannable blocks. Length is justified by a 22-parameter mutation surface; no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries return-value burden; it describes dry-run and refusal outputs but not the shape of a successful write response. For selection, safety, and limits it is otherwise thorough given the parameter count.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 86%, so the baseline is 3, but the description adds real meaning the schema lacks: the AND-combination and required-selector semantics, the default-dryRun logic keyed to price fields, and the refusal behavior of maxVariants. It stops short of explaining a few params such as tracking/inventoryPolicy interactions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (bulk update variants) plus the exact field families touched (customs data, cost, pricing, identifiers, shipping weight). An agent can distinguish it from shopify_create_variants, shopify_delete_variants, and shopify_update_inventory_item without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states at least one selector is required, that selectors combine with AND, gives the canonical customs-backfill invocation, and names the alternative for the excluded case (SKU changes → shopify_update_inventory_item). When-to-use and when-not-to-use are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_collection_add_productsAdd Products to CollectionAIdempotent
Add products to a MANUAL collection.
This fails on smart (rule-based) collections — their membership is decided by the rules, not by hand. Shopify processes the change asynchronously and returns a job; job.done: false just means it's still running, not that it failed.
| Name | Required | Description | Default |
|---|---|---|---|
| productIds | Yes | Product GIDs to add | |
| collectionId | Yes | Collection GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly=false, idempotent=true, destructive=false, openWorld=true). The description adds genuinely new behavior: the operation is asynchronous and returns a job whose job.done=false means still-running rather than failed, plus a hard failure mode on smart collections. That is exactly the contextual disclosure annotations cannot provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, and the MANUAL scoping constraint is front-loaded before the failure mode and the async caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the key operational gotchas (smart-collection failure, async job interpretation) for a 2-param mutation tool with annotations and no output schema. It stops short of describing the job's shape or how/whether to poll for completion, which an agent invoking an async tool would benefit from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (Product GIDs, Collection GID), so the schema already carries the semantics. The description adds no format, batching, or limit detail beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (add) and resource (products) and qualifies the target precisely as a MANUAL collection. The manual-vs-smart distinction also implicitly separates it from rule-based collection mutation paths and from shopify_collection_remove_products.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the when-not case: it fails on smart/rule-based collections because membership is rule-driven. It does not name an alternative tool (e.g., update_collection) for that case, so it stops just short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_collection_remove_productsRemove Products from CollectionADestructiveIdempotent
Remove products from a MANUAL collection. The products themselves are not deleted — only their membership of this collection.
Processed asynchronously, same as adding.
| Name | Required | Description | Default |
|---|---|---|---|
| productIds | Yes | Product GIDs to remove | |
| collectionId | Yes | Collection GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds genuinely new context: the products themselves survive and only membership is removed, plus that the operation is processed asynchronously — both are behavior the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no waste, with the destructive-scope clarification and the MANUAL constraint front-loaded before the async note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description reasonably does not explain return values, and it discloses both the async nature and the non-destructive scope. It stops short of saying how to confirm completion or whether progress can be polled, which is a minor gap for an async job.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters (productIds, collectionId), so the schema carries full semantics. The description adds no format, cardinality, or constraint detail beyond it, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Remove) and resource (products from a collection) with a meaningful scope qualifier: only MANUAL collections are supported. An agent can distinguish this from shopify_collection_add_products and from product-deletion tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The MANUAL-collection constraint is an implicit rule about when this tool applies, but the description never states what to do for non-manual (smart/automated) collections or names an alternative tool. Usage is implied rather than guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_create_articleCreate Shopify ArticleC
Create a new blog article with HTML body, author, summary, and tags.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | HTML body content | |
| tags | No | ||
| title | Yes | ||
| author | Yes | ||
| blogId | Yes | Blog GID | |
| summary | No | ||
| published | No | Publish immediately (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-readonly, non-idempotent, non-destructive open-world write, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say whether a duplicate call creates a second article, what the default publish state is, or what happens on missing blogId — all meaningful for a create tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the action front-loaded and no filler. It is efficient, though its brevity is also the source of the content gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation with a nested author object, no output schema, and only 43% schema coverage, the description leaves too much unstated — required blogId, the nested author.name shape, and the published default are all left to the schema or guesswork.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 43% (below 50%), so the description is expected to compensate, but it only echoes body/author/summary/tags. It omits blogId (required, the most error-prone parameter), title (required), and published, leaving the required-vs-optional split and the blog scoping unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Create a new blog article") and enumerates the main content fields, so an agent can distinguish it from shopify_update_article and the other create_* tools without opening a schema. It stops short of explicitly naming the update/read siblings it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus shopify_update_article, no mention that an existing blog (blogId) must already exist, and no note about draft-vs-published workflow. The agent must infer all invocation context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_create_collectionCreate Shopify CollectionA
Create a collection.
Two kinds:
Manual (default): you choose the products, via shopify_collection_add_products.
Smart: pass ruleSet and Shopify keeps membership up to date automatically. e.g. { appliedDisjunctively: false, rules: [{ column: "TAG", relation: "EQUALS", condition: "sale" }] } appliedDisjunctively: false = products must match ALL rules; true = ANY rule. Common columns: TAG, TITLE, TYPE, VENDOR, VARIANT_PRICE, VARIANT_INVENTORY.
A collection's rule set cannot be added later — decide smart vs manual now.
| Name | Required | Description | Default |
|---|---|---|---|
| seo | No | ||
| title | Yes | ||
| handle | No | ||
| ruleSet | No | Provide to create a smart (automated) collection | |
| sortOrder | No | ||
| descriptionHtml | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write/non-idempotent/non-destructive profile, so the description adds real value by disclosing the irreversible constraint that a rule set cannot be added after creation. It does not discuss rate limits, failure modes, or return payload, but the immutability warning is the critical behavioral fact for this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the manual/smart fork, uses a concrete ruleSet example, and avoids padding. The bulleted structure is scannable, though the trailing note about rule-set immutability is slightly buried at the end where it arguably belongs up front.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter, nested-object, no-output-schema tool, the description covers the hardest concept (ruleSet) and the key constraint well, but several parameters with no schema documentation are left silent. An agent can call the smart path confidently but gets little help for seo, handle, or sortOrder.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17%, so the description must carry more weight; it does explain ruleSet, appliedDisjunctively, and common column values (going beyond the schema). However, it leaves seo, handle, sortOrder (an 9-value enum), and descriptionHtml entirely unexplained, so the compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a collection') and immediately splits it into two meaningful modes (Manual vs Smart), which is the key distinction an agent needs. It clearly differentiates itself from siblings like shopify_update_collection and shopify_collection_add_products by naming the latter as the follow-up for manual membership.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit guidance on which mode to pick and routes the agent to shopify_collection_add_products for manual product assignment. It also warns 'decide smart vs manual now', which is actionable, though it does not spell out when NOT to create a collection or prerequisites (permissions, naming collisions).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_create_productCreate Shopify ProductA
Create a product. Defaults to DRAFT status — pass status: "ACTIVE" to publish immediately.
Options and variants:
Give productOptions (e.g. Size → S/M/L) to create a product with real variants.
Then call shopify_create_variants to add the priced variants against those options.
With no productOptions, Shopify creates a single default variant; set its price with shopify_update_product.
Images can be attached at creation via imageUrls (each must be a publicly reachable URL).
| Name | Required | Description | Default |
|---|---|---|---|
| seo | No | ||
| tags | No | ||
| title | Yes | ||
| handle | No | ||
| status | No | DRAFT | |
| vendor | No | ||
| imageUrls | No | Publicly reachable image URLs to attach | |
| productType | No | ||
| productOptions | No | Product options — required if the product has real variants | |
| descriptionHtml | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this as a non-read-only, non-destructive, non-idempotent, open-world write. The description adds real behavioral context beyond them: the tool defaults to DRAFT (so nothing publishes by accident), a default variant is silently created when productOptions is omitted, and images require publicly reachable URLs. It does not disclose what the response contains, which matters since the downstream create_variants step needs the product's identity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and default status, then organized under clear "Options and variants" and "Images" headings. Tight and skimmable, with only a slight restatement of the imageUrls constraint already present in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter create tool with nested objects and no output schema, the workflow guidance is solid, but the description never says what comes back. Since it explicitly instructs the agent to call shopify_create_variants afterward, the missing note about the returned product ID (the handle needed for that next call) leaves a real gap for the documented workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20%, so the description carries extra weight, and it does explain status's publish semantics, the productOptions shape, and the imageUrls public-URL constraint. However six of ten parameters (handle, vendor, productType, tags, seo, descriptionHtml) get no explanation beyond their self-evident names, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Create a product") and immediately scopes the variant/no-variant cases, which is what separates it from shopify_create_variants and shopify_update_product. An agent can tell exactly what this tool produces without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the follow-on tools explicitly (shopify_create_variants for real variants, shopify_update_product to price a default variant) and gives the condition that selects each path. It stops short of a true when-not (e.g. "to copy an existing product use shopify_duplicate_product"), so it lands at 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_create_variantsCreate Product VariantsA
Add new variants to an existing product.
Each variant must supply optionValues matching the product's options — e.g. for a product with options Size and Colour: optionValues: [{ optionName: "Size", name: "M" }, { optionName: "Colour", name: "Black" }]. Read the product's options first with shopify_get_product_by_id.
strategy REMOVE_STANDALONE_VARIANT deletes the auto-created default variant — use it when adding the first real variants to a product that was created without options.
| Name | Required | Description | Default |
|---|---|---|---|
| strategy | No | DEFAULT | |
| variants | Yes | ||
| productId | Yes | Product GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true). The description adds real behavior beyond them: the REMOVE_STANDALONE_VARIANT strategy deletes the auto-created default variant, which is a conditional destructive side effect not reflected in destructiveHint. Minor tension, but it is a mode-specific delete, not a contradiction of the tool's declared behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded paragraphs: purpose, the tricky optionValues contract with an example, then the strategy caveat. No filler; every sentence carries information an agent needs before calling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and 33% schema coverage, the definition covers the two error-prone areas (optionValues matching, standalone-variant removal). It omits permissions requirements, whether new variants land on inventory locations, and the response shape, which are modest gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33%, so the description must compensate — and it does, explaining that optionValues must match the product's existing options with a concrete example and clarifying the semantics of the strategy enum's non-default value. It adds little on price/sku/barcode, but the high-risk parameters are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Add new variants to an existing product.' An agent can immediately distinguish this from shopify_create_product, shopify_update_product, and shopify_bulk_update_variants without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit procedural guidance ('Read the product's options first with shopify_get_product_by_id') and a conditional rule for strategy REMOVE_STANDALONE_VARIANT. It does not, however, say when to prefer this over siblings like shopify_bulk_update_variants or shopify_delete_variants.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_delete_metafieldDelete Shopify MetafieldADestructiveIdempotent
Delete a metafield by GID. Get the ID first with shopify_get_metafields.
| Name | Required | Description | Default |
|---|---|---|---|
| metafieldId | Yes | Metafield GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is fully covered elsewhere. The description adds only the GID lookup prerequisite, not what deletion destroys, whether it cascades to variants/collections, or whether it can be undone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the destructive action front-loaded and the prerequisite sequenced after it. Nothing redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter delete whose annotations already flag destructiveness and idempotency, and with no output schema to explain, the definition is close to sufficient. It still omits whether the deletion is permanent or recoverable, a meaningful gap for a destructive operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with a single documented parameter ('Metafield GID'), so the schema carries the semantics. The description's 'by GID' merely echoes the schema rather than adding format examples or constraints, which is the baseline-3 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a metafield') plus the identifier form ('by GID'), which distinguishes it from the sibling shopify_get_metafields and shopify_set_metafield without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent the prerequisite step ('Get the ID first with shopify_get_metafields'), which routes it to the right sibling before calling this tool. It does not state when-not-to-use or note irreversibility, so it falls short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_delete_productDelete Shopify ProductADestructiveIdempotent
PERMANENTLY delete a product and all its variants, media, inventory items, and collection memberships. This CANNOT be undone.
Prefer setting status to ARCHIVED via shopify_update_product if there is any chance you'll want it back. Completed orders containing the product are unaffected.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | Yes | Must be true — confirms this permanent deletion is intended | |
| productId | Yes | Product GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint and non-readonly, but the description goes well beyond by disclosing the irreversible cascade of destroyed entities, the 'cannot be undone' warning, and the compensating effect on orders. This is exactly the extra context a destructive tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly-scoped sentences, front-loaded with the destruction scope and irreversibility, then the safer alternative, then the order caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema is present or needed for a delete. The description covers scope, irreversibility, the recommended alternative, and the order side-effect — everything an agent needs to call this safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so productId (GID) and confirm are already documented in the schema. The description adds no parameter-specific syntax or format detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('PERMANENTLY delete') and resource ('a product'), then enumerates the cascade (variants, media, inventory items, collection memberships). This clearly distinguishes it from sibling mutations like shopify_update_product and shopify_delete_variants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Prefer setting status to ARCHIVED via shopify_update_product if there is any chance you'll want it back.' Names the alternative tool and the condition that selects it, and clarifies a non-obvious boundary (completed orders unaffected).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_delete_variantsDelete Product VariantsADestructiveIdempotent
PERMANENTLY delete variants from a product, along with their inventory items. This CANNOT be undone.
A product must keep at least one variant — deleting all of them fails.
| Name | Required | Description | Default |
|---|---|---|---|
| confirm | Yes | Must be true — confirms this permanent deletion is intended | |
| productId | Yes | Product GID | |
| variantIds | Yes | Variant GIDs to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but the description adds real context beyond them: the deletion is permanent and irreversible, and it cascades to associated inventory items. The 'must keep at least one variant' business rule is behavior no annotation or schema encodes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the most important fact (permanent, irreversible) is front-loaded in caps. The failure constraint follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers irreversibility, cascading effect, and the minimum-variant rule. It leaves open only edge behavior on partial batch failure, which is a minor omission given annotations already carry the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so productId, variantIds, and the confirm guard are already fully documented in the schema. The description adds no argument-level detail (e.g., ID format, batch size limits), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('PERMANENTLY delete variants from a product') plus the cascade scope ('along with their inventory items'). This distinguishes it cleanly from shopify_delete_product and shopify_bulk_update_variants in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition for failure — a product must retain at least one variant — which tells the agent when this call is invalid. It does not name alternative tools (e.g., bulk_update_variants for non-destructive edits), so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_duplicate_productDuplicate Shopify ProductA
Copy a product, including its variants and options, under a new title. The copy defaults to DRAFT so it can't accidentally go live.
| Name | Required | Description | Default |
|---|---|---|---|
| newTitle | Yes | Title for the copy | |
| newStatus | No | DRAFT | |
| productId | Yes | Product GID to copy | |
| includeImages | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-destructive write operation, and the description adds useful behavioral detail: the copy defaults to DRAFT status so it cannot accidentally go live. It does not contradict any annotation and adds safety-relevant context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and scope, followed immediately by the default-status safety note. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter duplication tool with no output schema, the description covers the core operation, what is copied, and the default status. The only notable gap is the unmentioned includeImages parameter, but annotations already carry the mutation safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%. The description maps newTitle ('under a new title') and newStatus ('defaults to DRAFT') and implies productId, but it never mentions includeImages, leaving one parameter with no explanatory text in either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: copy a product including its variants and options under a new title. This clearly distinguishes it from generic create/update product tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: duplicating an existing product with its variants and options, rather than creating a blank product. It does not explicitly name alternatives or exclusions, but the core condition is implicit and strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_article_by_idGet Shopify Article by IDBRead-onlyIdempotent
Full article with HTML body, SEO, author, tags, image, and parent blog reference.
| Name | Required | Description | Default |
|---|---|---|---|
| articleId | Yes | Article GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered. The description adds useful return-shape context (HTML body, SEO, parent blog reference), but with no output schema it could say more about completeness/nulls for missing fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded clause that lists the payload contents without filler. It is appropriately short for a simple one-parameter read, though it omits an explicit verb like 'retrieve'.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-record read with one trivially-documented param and full safety annotations, the description covers the essentials. With no output schema, however, it could better characterize the return (e.g., what happens when the article is absent), leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (articleId) with 100% schema coverage describing it as an 'Article GID'. The description adds nothing about the ID beyond what the schema states, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name and description together state a specific verb+resource (get a single article by ID) and enumerate what comes back (HTML body, SEO, author, tags, image, parent blog). It implicitly separates itself from the list-oriented sibling shopify_get_articles, but never names that sibling or explicitly contrasts 'single article' vs 'list of articles'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternatives are named. An agent must infer from the name that this is the single-record retrieval versus shopify_get_articles for listing, and nothing addresses what to do if the ID is unknown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_articlesGet Shopify Blog ArticlesARead-onlyIdempotent
Get articles from a specific blog. Use shopify_get_blogs first to find the blog GID.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| blogId | Yes | Blog GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered and the bar is lower. The description adds the dependency on shopify_get_blogs, but says nothing about pagination, whether limit truncates results, or the volume/shape of what comes back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the core action is front-loaded and the prerequisite follows immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description bears the burden of describing the return, yet it says nothing about result shape or pagination despite exposing a limit parameter. The prerequisite note is helpful, but an agent still lacks enough to call this confidently at scale.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: blogId is documented as 'Blog GID' in the schema, while limit (default 10, max 250) is undocumented anywhere. The description reinforces how to obtain the GID but adds no meaning about limit's effect, so it does not fully compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get articles from a specific blog'), and the plural 'articles' plus the 'specific blog' scoping implicitly separates it from the singular shopify_get_article_by_id sibling. It is clear what the tool returns, though it does not explicitly name the sibling it differs from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite workflow: 'Use shopify_get_blogs first to find the blog GID,' which is exactly the guidance an agent needs before it can populate blogId. It stops short of stating when to prefer this over search-oriented siblings or how it relates to shopify_get_article_by_id.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_blog_by_idGet Shopify Blog by IDARead-onlyIdempotent
Get a specific blog with its 10 most recent article summaries.
| Name | Required | Description | Default |
|---|---|---|---|
| blogId | Yes | Blog GID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnlyHint, idempotentHint, destructiveHint=false) and open-world access, so the description's job is to add beyond that. It does disclose a non-obvious behavioral trait: the response embeds the 10 most recent article summaries rather than just blog metadata. Pagination and truncation behavior for that article list remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; every clause (the resource, the ID restriction, and the embedded article summaries) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the essential work of stating what the response contains via the '10 most recent article summaries' clause. For a simple read-by-id tool this is nearly complete, with only the article-list ordering/pagination cut-off unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter with 100% schema description coverage; the schema already documents blogId as a Blog GID. The description adds nothing about the ID format, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (blog by ID) plus the scope of what comes back (10 most recent article summaries). This clearly separates it from the list-style sibling shopify_get_blogs, though it does not name that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the required blogId makes it obvious this is for fetching a known blog, but the description never says when to prefer this over shopify_get_blogs or shopify_get_articles, nor does it name any alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_blogsGet Shopify BlogsCRead-onlyIdempotent
List all blogs or search by title.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| searchTitle | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is fully covered. The description adds no behavioral context beyond that — it says nothing about pagination behavior, default result size, or result shape, despite a limit parameter existing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Its brevity is appropriate in form, though it is arguably too short rather than padded — the gap is under-specification, not verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-param read tool with rich annotations and no output schema, the description is minimally viable. It omits the pagination/default-limit semantics an agent needs to call it predictably, so it is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with 2 parameters, so the description must compensate. It maps 'search by title' onto searchTitle, but says nothing about limit (default 10, max 250), leaving half the parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (list/search) and resource (blogs), so the agent knows it enumerates or filters blogs rather than fetching one. It does not explicitly differentiate itself from the sibling shopify_get_blog_by_id, but the plural resource and 'list all' phrasing make the distinction inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use context, no prerequisites, and no mention of alternatives such as shopify_get_blog_by_id (single blog) or shopify_get_articles (child resources). The 'or search by title' clause hints at the filtering mode but gives no condition for choosing it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_collectionsGet Shopify CollectionsBRead-onlyIdempotent
List all collections or search by title. Returns SEO fields, sort order, and template suffix.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| searchTitle | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so safety is covered. The description adds the return-field profile (SEO fields, sort order, template suffix), which is genuinely useful, but it says nothing about pagination behavior or how the default limit of 10 interacts with the maximum of 250.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with no filler. The listing behavior comes first and the return payload second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Absent an output schema, the description reasonably covers the return fields, which is the right use of that space. However, for a list tool with a pagination-style limit parameter and no schema descriptions, the missing information about truncation and result count leaves an agent unable to predict the response size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains searchTitle semantically ('search by title') but says nothing about 'limit' – its default of 10, its ceiling of 250, or whether it caps the returned list. Half the parameters remain undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List all collections') plus the search variant, which is enough to distinguish it from mutation siblings like shopify_create_collection and shopify_update_collection. It also names the returned fields. It stops short of explicitly naming an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'List all collections or search by title' implies the branch condition for using searchTitle, so usage is inferable. But there is no explicit when-to-use/when-not guidance and no pointer to sibling collection tools such as shopify_update_collection or shopify_collection_add_products.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_customerGet Shopify Customer by IDARead-onlyIdempotent
Full customer profile: all addresses, tags, order history (last 5 orders), and lifetime spend.
| Name | Required | Description | Default |
|---|---|---|---|
| customerId | Yes | Customer GID e.g. gid://shopify/Customer/1234567890 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, openWorld behavior, so the safety profile is covered. The description adds return-scope context beyond the annotations, notably that order history is limited to the last 5 orders and includes lifetime spend, which helps an agent anticipate payload shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with the content list front-loaded and zero filler. Every clause earns its place by describing the returned profile.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the job of conveying return content and does so reasonably (addresses, tags, orders, spend). It lacks error/not-found behavior and pagination notes, but for a single-record read with full annotation coverage it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single customerId parameter with 100% schema description coverage including the GID format example. The description adds no additional parameter meaning, so the schema carries the burden and baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific retrieval action ('Full customer profile') for the customer resource and enumerates the returned content (addresses, tags, last 5 orders, lifetime spend). This clearly distinguishes it from siblings like shopify_search_customers (listing/lookup) and shopify_get_customer_events (events).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the ID-based title and the 'Full customer profile' phrasing; there is no explicit when-to-use guidance and no mention of alternatives such as shopify_search_customers for finding a customer ID first. Adequate minimum viable context but no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_customer_eventsGet Customer Activity TimelineARead-onlyIdempotent
The activity timeline for one customer — the same events Shopify shows on the customer's admin page, newest first.
Many customers have no recorded events; an empty list is a valid answer.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| customerId | Yes | Customer GID e.g. gid://shopify/Customer/123 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, so safety is covered. The description adds genuine behavioral context beyond that: events are returned newest first and an empty list is a legitimate outcome for many customers, which prevents the agent from treating an empty result as a failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the resource definition, with the empty-result caveat placed last where it belongs. Zero wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should carry more of the return-value burden than it does; it covers ordering and the empty case but says nothing about the shape of an event, the limit/pagination semantics, or error conditions. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only 50% of parameters are documented in the schema (limit has no description), and the description adds nothing about either parameter. It never mentions the limit cap of 250 or pagination behavior, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('activity timeline for one customer') and scopes it by analogy to the Shopify admin page. It doesn't explicitly distinguish itself from sibling shopify_get_customer or shopify_search_customers, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the analogy to the admin page timeline, but there is no explicit when-to-use guidance and no mention of the obvious alternative shopify_get_customer for profile data. The agent must infer that this is the historical-events view versus a profile lookup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_inventory_levelsGet Inventory Levels for VariantARead-onlyIdempotent
Inventory quantities at each location for a specific variant: available, on_hand, committed, incoming.
Prefer shopify_get_variant_locations — it accepts a SKU and also shows which locations are NOT stocking the variant.
| Name | Required | Description | Default |
|---|---|---|---|
| variantId | Yes | Variant GID e.g. gid://shopify/ProductVariant/1234567890 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds value by naming the exact fields returned (available, on_hand, committed, incoming), which matters acutely because there is no output schema. It stops short of pagination or rate-limit behavior, keeping it just below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded: the capability first, the routing advice second. Zero filler, and the field enumeration does real work rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing the returned quantities, and it handles sibling disambiguation for a 50-tool namespace. Error behavior and location-scoping mechanics (e.g., whether all locations always appear) are left unstated, so it is strong but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single variantId parameter is documented in the schema with a GID example. The description adds nothing about parameter format, though the sibling comparison implicitly clarifies that this tool needs a variant ID rather than a SKU. Baseline 3 is correct when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (inventory quantities per location for a variant) and enumerates the exact quantities returned. It explicitly differentiates itself from the closest sibling, shopify_get_variant_locations, so an agent can choose between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the preferred alternative and the conditions that select it: that tool accepts a SKU and also reveals locations NOT stocking the variant. This is a clear when-to-use-this-vs-that signal, not vague context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_locationsGet Shopify LocationsCRead-onlyIdempotent
All store locations: name, address, active status, and whether each fulfills online orders.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and non-destructive behavior, so safety is covered. The description adds what data is returned (name, address, active status, fulfillment capability), but says nothing about pagination, rate limits, or the limit parameter's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficiently worded fragment that front-loads the resource. It wastes no words, though as a sentence fragment it leans terse rather than fully structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list operation with safety annotations already present, the description gives enough to understand the return shape. However, it omits any mention of the limit parameter or how results are constrained, leaving an agent to infer from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the limit parameter at all. With a single input parameter that has a default, minimum, and maximum, the description should explain its role or at least acknowledge pagination, but it does neither.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (store locations) and enumerates key attributes returned, making the tool's output clear. It distinguishes from siblings like variant locations or inventory levels, though the verb 'get' is only in the tool name, not the description text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool, when not to use it, or which alternatives exist. The description is a static field list with no contextual routing information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_metafieldsGet Shopify MetafieldsARead-onlyIdempotent
Get metafields for any Shopify resource. Requires read access on the owner resource type (e.g. read_products for product metafields).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| ownerId | Yes | Owner GID (product, collection, customer, order, etc.) | |
| namespace | No | Filter by namespace |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds genuine value with the per-resource-type scope requirement for read access, but says nothing about pagination, limits, or what an empty result means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the core action front-loaded and the permission requirement immediately after. Every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter read tool with rich annotations and no output schema, the description covers purpose, scope, and access requirements adequately. The remaining gap — return shape and pagination behavior given the limit parameter — is minor since annotations carry the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67%: ownerId is well documented as a GID with resource examples and namespace as a filter, while limit is only constrained numerically with no description. The description adds no parameter-level meaning beyond the vague 'any Shopify resource' hint, so the schema does the heavy lifting — a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get metafields') and adds useful scope ('for any Shopify resource'), so an agent knows it retrieves rather than writes. It does not, however, differentiate itself from the nearby shopify_set_metafield / shopify_delete_metafield siblings, leaving that routing to inference from the verb alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a real prerequisite ('Requires read access on the owner resource type, e.g. read_products for product metafields'), which is actionable context. It gives no explicit when-to-use versus the set/delete metafield siblings and no exclusions, so guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_orderGet Shopify Order by IDARead-onlyIdempotent
Get a single order by GID with full detail: line items, shipping/billing address, fulfillments, and refunds.
| Name | Required | Description | Default |
|---|---|---|---|
| orderId | Yes | Order GID e.g. gid://shopify/Order/1234567890 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the payload scope, which is useful, but says nothing about error behavior for a missing/invalid GID or whether returned data is rate-limited or paginated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and resource, followed by the returned fields. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only single-record fetch with complete annotations and full schema coverage, the description is nearly sufficient; enumerating the returned sections partly compensates for the absent output schema. Only the error/missing-record case is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, and the schema already documents the GID format with an example. The description's 'by GID' merely echoes this, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a single order by GID') and enumerates the returned detail (line items, addresses, fulfillments, refunds). The 'single' framing cleanly separates it from sibling listing tools such as shopify_list_orders or shopify_order_count.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the tool clearly retrieves one order, and the GID requirement hints at the access pattern, but there is no explicit statement of when to prefer it over shopify_list_orders or shopify_search, nor any prerequisite/limitation noted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_page_by_idGet Shopify Page by IDARead-onlyIdempotent
Full page content including HTML body, SEO, and publish status. Use after shopify_get_pages to get the full body of a specific page.
| Name | Required | Description | Default |
|---|---|---|---|
| pageId | Yes | Page GID e.g. gid://shopify/Page/1234567890 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds what the payload contains (HTML body, SEO, publish status), which is modest value beyond the annotations, but says nothing about auth scope, errors, or how large the body may be.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the return-content summary and followed by the usage hint. Nothing is redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A single-parameter read tool with annotations and no output schema; the description covers what a caller wants to know (payload contents and when to reach for it). Only minor gaps remain around pagination/error behavior, which are largely irrelevant here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents the single pageId parameter including the GID format example. The description adds no parameter-level detail such as format expectations or ID sourcing, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (get a specific page) and enumerates what the response contains: HTML body, SEO, publish status. It differentiates from the sibling list tool by naming shopify_get_pages as the entry point rather than a competitor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit sequencing guidance: 'Use after shopify_get_pages to get the full body of a specific page.' This tells the agent when this tool is the right next step. It stops short of naming when NOT to use it or listing edge cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_pagesGet Shopify PagesCRead-onlyIdempotent
List all store pages or search by title.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| searchTitle | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and open-world behavior, so the safety profile is covered. The description adds only a scope statement, and 'List all store pages' is arguably misleading since the schema's default limit of 10 means not all pages are returned; pagination behavior is never mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. It is appropriately compact for a simple list tool, though it is arguably too terse to be fully useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool, the description still omits key invocation context: no pagination guidance despite a default limit of 10, no indication of what fields a page contains, and no output-schema support since none exists. An agent could call it, but not confidently predict the result set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It covers the searchTitle concept qualitatively ('search by title') but says nothing about the limit parameter, its default of 10, or its maximum of 250.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all store pages') and names the second mode ('search by title'). It implicitly distinguishes itself from the sibling shopify_get_page_by_id (single page) and shopify_update_page (mutation), though it never names those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'or search by title' clause hints at two modes, but there is no guidance on when to list vs. search, when to prefer shopify_get_page_by_id, or what pagination/limit assumptions apply. Usage must be inferred entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_product_by_idGet Shopify Product by IDBRead-onlyIdempotent
Full product detail: all variants with inventory item IDs, product options, media, collections, and SEO fields.
| Name | Required | Description | Default |
|---|---|---|---|
| productId | Yes | Product GID e.g. gid://shopify/Product/1234567890 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered structurally. The description adds only that a broad payload (variants, media, collections, SEO) is returned; it says nothing about not-found behavior, rate limits, or auth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tightly packed sentence with no filler, and the most valuable fact (full detail including nested resources) is front-loaded. It is appropriately sized for a one-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by enumerating the included fields. It is nearly complete for a single-record getter, missing only error behavior on an invalid or missing ID.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single productId parameter is fully documented in the schema with a GID example. The description adds no parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The returned-content phrase 'Full product detail' plus the title makes it clear this fetches a single product's complete record, which implicitly contrasts with the plural shopify_get_products list tool. However, the description itself never states the action or the ID-based lookup, leaving that to the title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no alternative tools named. The agent might infer usage from the enumerated fields, but nothing routes it between this and shopify_get_products.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_productsGet Shopify ProductsARead-onlyIdempotent
List or search products. Supports cursor pagination.
Args:
searchTitle: Partial title match
query: Raw Shopify filter e.g. "status:active vendor:Nike tag:sale"
limit: Max results (default 10)
after/before: Pagination cursors
detail: "full" (variants, inventory, pricing, SKUs, images — default) or "summary" (id/title/handle/status/vendor/productType/tags/totalInventory only — much lighter, use this for bulk scans, filtering, or deciding what to tag)
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Cursor for next page | |
| limit | No | ||
| query | No | Raw Shopify query e.g. 'status:active vendor:Nike' | |
| before | No | Cursor for previous page | |
| detail | No | summary = id/title/status/vendor/productType/tags/totalInventory only, no variants/images/pricing | full |
| reverse | No | Reverse sort order | |
| searchTitle | No | Filter by title (partial match) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: cursor pagination support, a default limit of 10, and the payload-weight tradeoff between full and summary ("much lighter"). It does not describe rate limits or return envelope shape, keeping it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence plus pagination note, then a tight args list — every line is short and scannable. The args list partially duplicates the schema, which is the only mild waste, but nothing is bloated or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, zero-required list tool with no output schema, the description covers the decisions that matter: how to filter (searchTitle vs raw query), how to page, and which detail level to request. It leaves the relationship between searchTitle and query, and the return shape, implicit, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the baseline is 3, and the description does add value on top: it spells out what "full" includes (variants, inventory, pricing, SKUs, images) and why summary exists. The remaining args (searchTitle, query, limit, after/before) largely restate the schema, so it is a modest improvement rather than a transformation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ("List or search products") that an agent can act on immediately. It implicitly separates itself from shopify_get_product_by_id (single-product fetch) and write tools like shopify_create_product, but never names a sibling or states the boundary explicitly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The detail parameter carries genuine usage guidance: "use this for bulk scans, filtering, or deciding what to tag" tells the agent when to pick summary over full. There is no explicit when-not or alternative-tool routing (e.g. vs shopify_search), so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_shopGet Shopify Shop InfoARead-onlyIdempotent
Get store details: name, email, domain, Shopify plan, currency, timezone, billing address.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered. The description adds genuine value by naming exactly which attributes the shop resource exposes, which no annotation or schema conveys (there is no output schema). It stops short of describing response shape, error cases, or rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a verb, resource and field list; there is no filler, hedging, or redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters, full annotation coverage and no output schema, the field enumeration is effectively the only missing structured information, and the description supplies it. It is nearly complete, only lacking notes on error handling or whether a single shop object is always returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to document and the baseline of 4 applies. The description does not need to add parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Get') plus a specific resource ('store details') with an explicit enumeration of the returned fields (name, email, domain, plan, currency, timezone, billing address). This cleanly separates it from siblings like shopify_get_locations or shopify_get_products, which return other resource classes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when store-level details are needed) but never states it explicitly, names no alternatives, and gives no exclusions or prerequisites. For a zero-parameter getter with an obvious distinct scope, this is minimally viable rather than helpful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_get_variant_locationsGet Variant Location Fulfilment StateARead-onlyIdempotent
Look up a variant by SKU (or variant GID) and see exactly which locations stock and fulfil it.
Returns:
active_locations: locations that CAN stock & fulfil this variant, with available/on_hand/committed/incoming quantities, plus can_deactivate and any deactivation_blocked_reason
inactive_locations: locations that CANNOT — activate one with shopify_set_variant_locations
inventory_item_id: needed by shopify_adjust_inventory / shopify_set_inventory
Call this before shopify_set_variant_locations to see the current state.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact variant SKU e.g. 'TSHIRT-BLK-M' | |
| variantId | No | Variant GID — use instead of sku when the SKU is ambiguous or blank |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered. The description adds real value by enumerating the return structure (active vs inactive locations, per-location quantities, can_deactivate, deactivation_blocked_reason) and showing the mutation path via the sibling tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence followed by a bulleted breakdown of the return shape and the sequencing instruction. Every line earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the burden of describing returns—and it does so thoroughly (active/inactive buckets, quantity fields, inventory_item_id). Combined with the sequencing guidance, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (sku, variantId) are already documented in the schema, including the ambiguity/blank guidance that the description only echoes as '(or variant GID)'. Baseline 3 is appropriate; the description adds no syntax detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up) and resource (variant by SKU/GID) and the exact dimension it reveals (which locations stock and fulfil it). This clearly distinguishes it from siblings like shopify_get_locations and shopify_get_inventory_levels.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Call this before shopify_set_variant_locations to see the current state,' and names shopify_set_variant_locations as the way to activate an inactive location. It also names the downstream consumers (shopify_adjust_inventory / shopify_set_inventory) of inventory_item_id, giving the agent clear routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_list_abandoned_checkoutsList Abandoned CheckoutsARead-onlyIdempotent
List abandoned checkouts with customer details, line items, pricing, and the recovery URL. Useful for identifying recovery opportunities and lost revenue.
Only returns checkouts that have not been completed (completedAt is null).
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor | |
| limit | No | ||
| query | No | Filter e.g. 'created_at:>2025-01-01' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world. The description adds a genuinely non-obvious filter constraint — only checkouts where completedAt is null — which is real behavioral context an agent cannot infer from the schema. It stops short of describing pagination behavior or result volume.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The payload summary comes first and the scope constraint second, both front-loaded and information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-shape burden and does so adequately by naming the fields returned. For a simple read-only listing with optional filters, nothing essential is missing beyond pagination guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (limit has no description text but carries default/min/max; after and query are documented). The description adds no parameter meaning at all, so it lands at the baseline for a partially documented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (list abandoned checkouts) and enumerates the payload fields returned (customer details, line items, pricing, recovery URL). It doesn't name a sibling alternative, but no sibling performs a comparable listing, so confusion risk is low.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'useful for identifying recovery opportunities and lost revenue' implies the business context for calling it, but there is no explicit when-to-use vs. when-not guidance and no named alternative (e.g., shopify_list_orders) for cases where completed checkouts are also wanted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_list_discountsList Shopify DiscountsARead-onlyIdempotent
List all discounts — code-based (basic, free shipping, BXGY) and automatic. Returns codes, usage counts, and discount values.
Filter by query string. Examples:
"status:ACTIVE"
"discount_type:percentage"
"discount_type:fixed_amount"
"title:SUMMER"
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor | |
| limit | No | ||
| query | No | Filter e.g. 'status:ACTIVE' or 'discount_type:percentage' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description usefully adds what the response contains (codes, usage counts, discount values), but says nothing about pagination behavior despite exposing an 'after' cursor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence followed by a tight bulleted example list; every line is doing work. Slightly longer than strictly necessary but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly covers return contents, and the query examples cover the main filtering parameter. Pagination semantics remain undocumented, a minor but real gap for a list endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67%, with 'limit' carrying no description. The description compensates by showing real query syntax (status:ACTIVE, discount_type:percentage, title:SUMMER), which adds genuine meaning beyond the schema's minimal 'Filter e.g.' note; it does not, however, explain the pagination cursor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (discounts), and enumerates the sub-types it covers (code-based basic/free shipping/BXGY and automatic). It also names the return payload (codes, usage counts, values), so an agent knows exactly what class of object it is dealing with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says the tool supports filtering via a query string and gives concrete query examples, which implies when to reach for it. However, it never states prerequisites, when not to use it, or routes to any alternative sibling, so usage is inferred rather than prescribed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_list_draft_ordersList Shopify Draft OrdersBRead-onlyIdempotent
List draft orders with status, customer info, line items, and totals.
Filter examples: "status:open", "status:completed", "customer_id:gid://shopify/Customer/123"
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor | |
| limit | No | ||
| query | No | Shopify query filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds the set of fields that will be returned, but does not disclose pagination behavior or rate-limit considerations; with annotations carrying the behavioral burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with zero waste, and the core purpose is front-loaded before the filter examples. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool with three optional parameters and annotations covering safety, the definition is nearly complete: it states the returned fields and gives query filter examples. It could mention pagination briefly or distinguish itself from shopify_list_orders, but the schema and annotations cover the main structural gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%; 'after' and 'query' are documented in the schema, while 'limit' is not. The description adds concrete filter syntax examples for the query parameter, which is helpful beyond the schema's 'Shopify query filter' text, but it does not explain the cursor or limit semantics. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List draft orders,' and adds the returned fields (status, customer info, line items, totals). It is clear but does not explicitly distinguish itself from the sibling shopify_list_orders tool, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides useful query filter examples, but gives no guidance on when to use this tool instead of alternatives such as shopify_list_orders. There is no mention of prerequisites, exclusions, or context for choosing draft orders over regular orders.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_list_marketing_eventsList Marketing EventsARead-onlyIdempotent
Marketing events recorded against the store — ad campaigns, posts and other attributed activity, with their UTM parameters and run dates.
Only populated when a marketing app or integration publishes events to Shopify. An empty list means no such integration is connected, not that the query failed.
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor from a previous response's pageInfo.endCursor | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so safety is covered. The description adds a genuinely non-obvious behavioral fact not derivable from annotations or schema: an empty list means no marketing integration is connected rather than a failed query. It stops short of covering pagination behavior or result ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The payload description comes first and the critical empty-result caveat follows where it will actually be read, before the agent inspects the result.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully sketches what a record contains (UTM parameters, run dates) and how to interpret an empty result. Pagination mechanics are left to the schema's 'after' description, which is a minor gap for a two-parameter list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%: 'after' is fully documented in the schema, while 'limit' carries only default/min/max with no prose. The description contributes nothing to either parameter, so it neither compensates for the gap nor restates the schema. Since both are simple, self-explanatory pagination controls, this is an acceptable baseline rather than a serious omission.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Marketing events recorded against the store') and enumerates the payload contents (ad campaigns, posts, attributed activity, UTM parameters, run dates). No sibling in the list covers marketing events, so differentiation is implicit rather than stated, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the resource name, but there is no explicit when-to-use statement or routing to an alternative tool. The strongest guidance is interpretive rather than selective: it tells the agent what an empty result means, not when to reach for this tool over another.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_list_ordersList Shopify OrdersBRead-onlyIdempotent
List orders within a date range with financial details, line items, and customer info.
Args:
date_from: Start date YYYY-MM-DD (inclusive)
date_to: End date YYYY-MM-DD (inclusive)
limit: Max orders 1–250 (default 50)
status_filter: Extra Shopify query filter e.g. "financial_status:paid fulfillment_status:unfulfilled"
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max orders (default 50) | |
| date_to | Yes | End date YYYY-MM-DD | |
| date_from | Yes | Start date YYYY-MM-DD | |
| status_filter | No | Extra filter e.g. financial_status:paid |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered structurally. The description contributes the returned-data shape ('financial details, line items, and customer info'), which is genuinely useful, but says nothing about pagination, the 250-order cap behavior, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and the argument list is compact with one line each. Nothing is padded, though the Args block partly restates the schema at the cost of a little redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does take on the job of signaling the return content, and it does so. The main remaining gap is pagination/limit interaction with the date range, which an agent listing large ranges would want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value beyond the schema by marking both date bounds as inclusive and giving a concrete composite example for status_filter ('financial_status:paid fulfillment_status:unfulfilled'), which clarifies the query-filter syntax the schema only hints at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List orders') plus the scope ('within a date range') and what the payload contains. It does not, however, distinguish itself from close siblings like shopify_get_order, shopify_order_count, or shopify_list_draft_orders, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no mention of alternatives. The required date range implies a bulk listing context, but nothing tells the agent why to pick this over shopify_order_count or shopify_get_order.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_move_inventoryMove Inventory Between LocationsA
Move stock of one variant from one location to another in a single transactional call — the origin is decremented and the destination incremented together.
Both locations must already stock the variant (activate with shopify_set_variant_locations first).
quantityName pairs the ledger states being moved between, e.g. moving 'available' at the origin to 'available' at the destination. For most stock transfers, leave both as 'available'.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact variant SKU | |
| reason | No | correction | |
| toName | No | available | |
| fromName | No | available | |
| quantity | Yes | Units to move | |
| toLocation | Yes | Destination location GID or name | |
| fromLocation | Yes | Origin location GID or name | |
| inventoryItemId | No | InventoryItem GID — alternative to sku | |
| referenceDocumentUri | No | Traceability URI for this movement |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-readOnly, non-idempotent, open-world, non-destructive mutation. The description adds meaningful behavior beyond that: it is a single transactional call that atomically decrements origin and increments destination, and both locations must already stock the variant. It does not cover auth needs, rate limits, or failure behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded: the first sentence states the action and atomicity, the second states the prerequisite, and the third gives parameter guidance. Each sentence carries a distinct piece of information, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no output schema, the description covers the core behavior, prerequisite, and key ledger-state semantics. Annotations already carry the safety/idempotency profile, so the remaining gaps around sku alternates and reason/referenceDocumentUri usage are tolerable, though not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%, so the description needs to add value, and it does for ledger-state parameters: it explains pairing origin/destination states and recommends 'available' for most transfers. However, it calls that pair 'quantityName' while the schema fields are fromName and toName, and it does not clarify sku-vs-inventoryItemId, reason, or referenceDocumentUri. The added semantics are useful but imperfect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: moving stock of one variant from one location to another. It adds the atomic behavior ('origin is decremented and the destination incremented together') and references shopify_set_variant_locations as a prerequisite, which helps distinguish this from a location-activation tool. The purpose is unambiguous without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: both locations must already stock the variant, activate with shopify_set_variant_locations first, and for most stock transfers leave ledger states as 'available'. It does not explicitly contrast this tool with sibling inventory-mutating tools such as shopify_adjust_inventory or shopify_set_inventory, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_order_countShopify Order CountARead-onlyIdempotent
Lightweight order count for a date range without fetching order details.
| Name | Required | Description | Default |
|---|---|---|---|
| date_to | Yes | End date YYYY-MM-DD | |
| date_from | Yes | Start date YYYY-MM-DD | |
| status_filter | No | Extra filter e.g. financial_status:paid |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered. The description adds that it avoids fetching order details (an efficiency trait), but says nothing about count limits, status scoping defaults, or return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the key distinction ('lightweight', 'count') front-loaded and zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple counting tool with annotations and full schema coverage this is adequate, but with no output schema it never states what is returned (a bare integer?) or whether the count respects a default status filter, leaving a small but real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so date_from/date_to/status_filter are already documented in the schema. The description echoes the date-range concept but adds no syntax, default, or semantics for status_filter beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('order count') scoped to a date range, and implicitly distinguishes itself from detail-fetching siblings like shopify_list_orders via 'without fetching order details'. No sibling is named explicitly, so it misses the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Lightweight ... without fetching order details' implies you should use it when only a number is needed, but no alternative tool is named and there is no stated when-not guidance beyond the efficiency hint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_remove_tagsRemove Tags from Shopify ResourceAIdempotent
Remove specific tags from any Shopify resource you have write access to (products, articles, collections).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Resource GID | |
| tags | Yes | Tags to remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds only the write-access prerequisite and the resource scope; it says nothing about behavior when a tag is absent or whether partial failures are possible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero padding; the scope and the resource list are stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with a complete schema and rich annotations, the definition covers what is needed to call it. It lacks any statement of what the call returns or how removed tags are confirmed, which is a minor gap given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with two fully documented parameters, so the schema carries the load. The mention of products, articles and collections loosely signals that 'id' accepts GIDs of varying resource types, but no syntax or array semantics are added beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove specific tags') and scopes it to 'any Shopify resource you have write access to (products, articles, collections)'. The opposite verb on the sibling shopify_add_tags makes the distinction clear, though siblings are never named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context — it applies to any resource the caller has write access to, which is a real prerequisite. It provides no exclusions or explicit routing away from alternatives such as shopify_add_tags or shopify_bulk_tag_products.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_searchSearch Shopify StoreARead-onlyIdempotent
Unified search across products, articles, blogs, and pages. Runs parallel queries for each type and returns results grouped by type.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Results per type | |
| query | Yes | Search query | |
| types | No | Types to search (default: all) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuine behavior beyond that: it runs parallel queries per type and returns results grouped by type, which tells the agent how output is structured for parsing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the search scope front-loaded before the execution/return detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully states that results are grouped by type, which is the key return-shape fact. For a read-only search tool with full schema coverage and safety annotations, this is nearly complete, though pagination behavior is unmentioned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so query, types, and limit are all documented in the schema. The description names the searchable types, loosely mapping to the types enum, but adds no format or default information beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and names the exact resource set (products, articles, blogs, pages), so the agent can tell this is a cross-type search tool. It stops short of explicitly differentiating from the similarly-named sibling shopify_search_customers, which is excluded only implicitly by the type list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'unified' implies use when a single cross-type lookup is wanted, but there is no explicit when-to-use versus alternatives like shopify_search_customers or the per-type list/get tools, and no exclusions stated. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_search_customersSearch Shopify CustomersARead-onlyIdempotent
Search customers by name, email, tags, or any Shopify customer query filter.
Example queries: "Jane Smith", "email:bob@example.com", "tag:vip", "state:enabled"
| Name | Required | Description | Default |
|---|---|---|---|
| after | No | Pagination cursor | |
| limit | No | ||
| query | No | Search query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the query-filter syntax, which is genuine value, but says nothing about pagination behavior or result ordering beyond what the annotations and schema provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core capability and followed by high-value examples. No filler; every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With readOnly/idempotent annotations covering safety and no output schema to explain, the description only needs to convey search semantics, which it does via the filter examples. It is slightly thin on pagination ('after'/'limit') expectations, but that is documented in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (query and after described, limit undocumented). The description compensates by demonstrating concrete query syntax strings ('email:bob@example.com', 'tag:vip', 'state:enabled'), which adds real meaning beyond the schema's bare 'Search query' label.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (search customers) and enumerates the searchable fields (name, email, tags, or any customer query filter). It clearly separates this from get_customer by framing it as a search, though it does not explicitly name the sibling tools it competes with (shopify_get_customer, shopify_search).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the query examples showing filter syntax, but there is no explicit guidance on when to use this versus shopify_get_customer (single lookup) or the generic shopify_search. No prerequisites or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_set_inventorySet Shopify Inventory (absolute)AIdempotent
Set stock to an EXACT number at one or more locations — e.g. quantity: 42 means "there are now 42".
Use this for stocktakes and when syncing from a system that is the source of truth. For relative changes ("we sold 3"), use shopify_adjust_inventory instead.
Concurrency: by default this uses compare-and-set — pass compareQuantity (the quantity you believe is currently there) and Shopify rejects the write if someone changed it underneath you. Set ignoreCompareQuantity: true to force the write regardless, which risks clobbering a concurrent update.
The variant must already be stocked at the location — activate it first with shopify_set_variant_locations.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact variant SKU | |
| name | No | Which quantity to set | available |
| reason | No | correction | |
| quantities | Yes | ||
| inventoryItemId | No | InventoryItem GID — alternative to sku | |
| referenceDocumentUri | No | Traceability URI e.g. 'logistics://warehouse/stocktake/2026-01-14' | |
| ignoreCompareQuantity | No | Skip the compare-and-set safety check |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (idempotent, non-destructive, open-world). The description goes well beyond them by explaining the compare-and-set concurrency model, the role of compareQuantity, and the data-loss risk of ignoreCompareQuantity: true clobbering a concurrent update, plus an activation prerequisite.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core semantics and the alternative tool in the first two sentences, then the concurrency caveat, then the prerequisite. Slightly long with a few adjacent restatements of compareQuantity, but every section carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema needed and 7 parameters, the description covers the mutation semantics, the safety/override switch, the sibling routing, and the activation prerequisite — everything required to invoke it correctly, including the risky path.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 71% and the description compensates by explaining what compareQuantity means (the quantity you believe is currently there), what ignoreCompareQuantity does, and that quantity is the exact target. It does not clarify the name enum (available vs on_hand) or reason values, which the schema carries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (set stock to an exact number) with a concrete example showing that quantity: 42 means 'there are now 42', and explicitly contrasts the absolute semantics with the relative sibling shopify_adjust_inventory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('stocktakes', 'syncing from a source of truth') and when-not ('we sold 3' → use shopify_adjust_inventory), and adds a prerequisite: the variant must already be stocked at the location, activated via shopify_set_variant_locations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_set_metafieldCreate/Update Shopify MetafieldAIdempotent
Upsert a metafield on any resource you have write access to (products, collections, pages, articles).
Common types: single_line_text_field, multi_line_text_field, integer, boolean, json, url, color, date, file_reference
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| type | Yes | Metafield type e.g. single_line_text_field, json | |
| value | Yes | ||
| ownerId | Yes | Owner GID | |
| namespace | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-destructive, idempotent, open-world write, and 'Upsert' is consistent with idempotentHint=true. The description reinforces the write-access requirement but adds no detail on overwrite semantics, error behavior for invalid types, or what happens to an existing metafield value. Some value, but thin beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the operation front-loaded and the type list separated. Efficient and readable, though the second sentence could be flagged as loosely formatted rather than fully earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-required-param mutation with no output schema, the description covers the core action and the type vocabulary but omits key/namespace/value meaning and response behavior. Annotations cover the safety profile, so the remaining gaps are moderate rather than severe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, with key, namespace, and value having no descriptions. The enumerated types list partially compensates for the `type` parameter and the resource list hints at ownerId, but namespace/key/value semantics are left entirely undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Upsert) and resource (metafield), plus the scope of applicable owner resources (products, collections, pages, articles). It cleanly separates itself from shopify_get_metafields and shopify_delete_metafield by naming the write operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Prerequisite is implied with 'any resource you have write access to', which is useful context. However, it never states when to use this versus shopify_get_metafields or shopify_delete_metafield, nor any ordering/quoting guidance. Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_set_variant_locationsActivate/Deactivate Variant Fulfilment at LocationsADestructiveIdempotent
Turn location fulfilment ON or OFF for a variant, identified by SKU or variant GID.
activate: true → the location can stock and fulfil this variant (starts at 0 available; follow with shopify_set_inventory to set stock)
activate: false → the location no longer stocks the variant, and its stock there is discarded
Locations may be given as a GID or by name (e.g. "Melbourne Warehouse"), matched case-insensitively. Multiple locations are applied in a single call.
DEACTIVATION IS DESTRUCTIVE: it discards that location's inventory for the variant. Shopify refuses when the variant has committed stock or pending fulfilments there — check can_deactivate via shopify_get_variant_locations first.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact variant SKU | |
| locations | Yes | Locations to activate or deactivate | |
| variantId | No | Variant GID — use instead of sku when the SKU is ambiguous |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and idempotentHint=true, and the description goes further, warning that deactivation discards the location's inventory and that Shopify refuses when committed stock or pending fulfilments exist. These are concrete behavioral consequences beyond what the structured fields convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action, then uses tight bullet points for the two activate states, followed by location-matching rules and a highlighted destructive warning. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, multi-location mutation with no output schema, the description covers identification, matching rules, batching, side effects, and failure conditions. An agent has everything needed to invoke it correctly and safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: locations accept a GID or a name (e.g. "Melbourne Warehouse") matched case-insensitively, and multiple locations can be applied in one call. The activate flag semantics (starts at 0 available) also enrich the boolean beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource — turning location fulfilment ON/OFF for a variant — and specifies identification by SKU or variant GID. It is clearly distinguishable from siblings like shopify_get_variant_locations (read) and shopify_set_inventory (stock levels), not fulfilment enablement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly sequences usage: use shopify_get_variant_locations to check can_deactivate first, then follow activation with shopify_set_inventory to set stock. It names the alternative tools and the conditions that select them, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_update_articleUpdate Shopify ArticleCIdempotent
Update an article's title, body, summary, tags, author, or publish status.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | HTML body | |
| tags | No | ||
| title | No | ||
| author | No | ||
| summary | No | ||
| articleId | Yes | Article GID | |
| published | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, idempotent=true, destructive=false, and openWorld=true, so the safety profile is covered. The description adds almost nothing behavioral — it does not say whether the update is partial or full-replace, whether omitted fields are cleared, or what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the verb and field list front-loaded and no filler. Efficient, though it is arguably too terse for a mutation tool with a nested author object.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, low schema coverage, and a nested author object, the description should clarify partial-update semantics, the required articleId, and the response shape. None of that is present, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 29% (7 params, 1 required), so the description does carry weight: it names title, body, summary, tags, author, and publish status, covering six of the seven parameters. However it omits articleId (the required one), omits that body is HTML, and gives no format/content constraints, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (Shopify article) and enumerates the mutable fields, so the agent can tell it apart from update_page/update_blog/update_product by the resource name. No explicit sibling comparison, but purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of alternatives such as shopify_create_article or shopify_update_blog, and no prerequisites (e.g. that an existing articleId is required). The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_update_blogUpdate Shopify BlogCIdempotent
Update a blog's title, handle, template suffix, or comment policy.
Comment policies: MODERATED, AUTO_PUBLISHED, CLOSED
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| blogId | Yes | Blog GID | |
| handle | No | ||
| commentPolicy | No | ||
| templateSuffix | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation is a non-read, idempotent, non-destructive mutation, so the safety profile is covered. The description adds nothing beyond annotations that carries behavioral meaning: it does not say whether the update is a partial merge or full replace, what permissions are needed, or how unpublished/unset fields are treated. The comment-policy list simply restates the schema enum.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight segments, front-loaded with the action and fields, followed by the enum values. Every sentence is short and there is no filler, though the enum listing partially duplicates the schema's enum declaration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation with no output schema and annotations covering the safety profile, the definition is minimally adequate. However it omits partial-update semantics and permission requirements, which are the main remaining gaps an agent would want before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 20% (just blogId = 'Blog GID'), so the description has to compensate. It usefully names the four optional fields (title, handle, template suffix, comment policy), giving the agent an inventory of what is updatable, but it adds no format or syntax detail and does not clarify the relationship between handle and templateSuffix.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (blog) and enumerates the four mutable fields, so an agent can immediately tell this apart from read siblings like shopify_get_blogs and shopify_get_blog_by_id. It does not explicitly name any alternative tool, but the write verb plus the 'blog' resource makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus other update tools, no prerequisites (e.g. a valid blogId), and no statement that only supplied fields change. The description is purely a field inventory with no contextual routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_update_collectionUpdate Shopify CollectionCIdempotent
Update a collection's title, description HTML, and SEO metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| seo | No | ||
| title | No | ||
| description | No | ||
| collectionId | Yes | Collection GID | |
| descriptionHtml | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is covered. The description adds only implicit partial-update scope by enumerating fields; it does not state whether omitted fields are left unchanged, whether SEO can be cleared, or what authorization is needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the verb and affected fields front-loaded and zero padding. It is efficient, though the brevity contributes to the missing usage and parameter detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with a required ID, a nested SEO object, and 20% schema coverage but no output schema, the description is too thin. It never explains partial-update behavior, the required collectionId, or the nested seo structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20%, so the description must compensate, but it names just title, description HTML, and SEO while omitting the plain 'description' parameter and the required collectionId entirely. The nested 'seo' object's inner title/description semantics are left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (collection) plus the fields affected (title, description HTML, SEO metadata). An agent can distinguish it from shopify_create_collection and the add/remove-products siblings, but the description does not explicitly name or exclude those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use context, no prerequisites (e.g. existing collection ID), and no pointer to competing collection tools such as shopify_collection_add_products. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_update_inventory_itemUpdate Inventory Item SettingsAIdempotent
Update a variant's inventory-item settings — the per-SKU attributes that aren't pricing:
tracked: whether Shopify tracks stock for this SKU at all (untracked SKUs always sell)
cost: unit cost, used for profit reporting and COGS
requiresShipping: false for digital or service items
Customs / HTS fields, needed before Shopify will generate international shipping labels:
harmonizedSystemCode: the HS/HTS tariff code, 6 to 13 digits. Dots and spaces are stripped, so '6110.20.20' and '61102020' are both accepted.
countryCodeOfOrigin: 2-letter ISO country where the item was made, e.g. 'AU', 'CN'
provinceCodeOfOrigin: province/state code, only used by a few destinations (e.g. 'BC' for Canada)
countryHarmonizedSystemCodes: destination-specific HTS overrides, for when a country wants a longer national code than your 6-digit base. Replaces the whole override list.
Read the current values back with shopify_get_variant_locations. Accepts a SKU or an explicit inventoryItemId. Only provided fields change.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact variant SKU | |
| cost | No | Unit cost e.g. '12.50' | |
| tracked | No | Whether Shopify tracks stock for this SKU | |
| inventoryItemId | No | InventoryItem GID — alternative to sku | |
| requiresShipping | No | false for digital/service items | |
| countryCodeOfOrigin | No | 2-letter ISO country code e.g. 'AU', 'CN' — case-insensitive | |
| harmonizedSystemCode | No | HS/HTS tariff code, 6-13 digits e.g. '611020' or '6110.20.20' | |
| provinceCodeOfOrigin | No | Province/state code e.g. 'BC' — only needed for some destinations | |
| countryHarmonizedSystemCodes | No | Destination-specific HTS overrides — replaces the existing list |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover idempotency, non-destructiveness, and open-world behavior, so the bar is lower. The description nonetheless adds real behavioral detail: 'Only provided fields change' (partial-update semantics), HS codes have dots/spaces stripped, and countryHarmonizedSystemCodes 'replaces the whole override list' — a field-level data-loss behavior that the destructiveHint=false annotation does not convey. It is the reason an agent must re-read before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The front-loaded opening and bulleted grouping are readable, and the closing lines (read-back tool, SKU-or-id, partial update) are the highest-value content. However, most bullet entries restate the schema descriptions nearly verbatim ('requiresShipping: false for digital or service items', 'Province/state code e.g. BC'), so a meaningful share of the text does not earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, zero-required, no-output-schema tool, the description covers every parameter with meaning plus normalization and list-replacement semantics, and names a sibling for verification. It is nearly complete; what is missing is the relationship to adjacent inventory tools (bulk update, inventory levels), which matters for choosing this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage the baseline is 3; the description earns above that by adding normalization rules the schema lacks ('6110.20.20' and '61102020' both accepted), the purpose of each field (cost for COGS, requiresShipping false for digital), and the replace-not-merge behavior of the override array.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (a variant's inventory-item settings) and explicitly scopes it as the per-SKU attributes that are not pricing. The field list makes the boundary concrete. It does not, however, distinguish itself from the quantity-oriented siblings (shopify_adjust_inventory, shopify_set_inventory, shopify_bulk_update_variants), which an agent could plausibly pick for a value like 'tracked'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Points at shopify_get_variant_locations for reading current values and notes the tool accepts either a SKU or an explicit inventoryItemId, which is useful routing. But there is no when-to-use/when-not guidance relative to shopify_bulk_update_variants, shopify_update_product, or the inventory-level tools, and no preconditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_update_pageUpdate Shopify PageAIdempotent
Update a page's title, HTML body, handle, template suffix, and publish status. Only provided fields are changed.
| Name | Required | Description | Default |
|---|---|---|---|
| body | No | HTML body content | |
| title | No | ||
| handle | No | ||
| pageId | Yes | Page GID | |
| isPublished | No | Publish or unpublish the page | |
| templateSuffix | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, so safety is covered. The description adds real value beyond that by disclosing PATCH-style partial-update behavior ('Only provided fields are changed'), telling the agent that unspecified fields are preserved. It stops short of noting permission requirements or side effects such as a handle change altering the public URL.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the field scope is front-loaded ahead of the partial-update caveat. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with annotations present and no output schema, the definition covers the mutation surface and the partial-update contract adequately. It omits error behavior for an invalid pageId and any permission or publication-timing nuances, but nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, so the description must compensate, and it does: it names title, HTML body, handle, template suffix, and publish status, covering five of the six parameters semantically. Only pageId is left to the schema, and no format or validation detail (e.g., handle slug rules) is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (a page) and enumerates the mutable fields, making the scope unambiguous. The resource naturally separates it from sibling updaters like shopify_update_blog, shopify_update_article, and shopify_update_collection, though it never explicitly names an alternative to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisite (e.g., obtaining pageId from shopify_get_page_by_id), and no exclusions. The only usage-adjacent statement is 'Only provided fields are changed,' which describes update semantics rather than when this tool should be selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_update_productUpdate Shopify ProductAIdempotent
Update a product's title, description, SEO, status, vendor, tags, or its variants' prices/SKUs. Only provided fields change.
Variants passed here are UPDATED (each needs its variant GID). To add new variants use shopify_create_variants; to remove them use shopify_delete_variants.
| Name | Required | Description | Default |
|---|---|---|---|
| seo | No | ||
| tags | No | ||
| title | No | ||
| handle | No | ||
| status | No | ||
| vendor | No | ||
| variants | No | ||
| productId | Yes | Product GID | |
| productType | No | ||
| descriptionHtml | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, so safety is covered. The description adds genuine non-redundant behavior: partial-patch semantics ('Only provided fields change') and the GID requirement for variant updates. It stops short of describing error/return behavior, which keeps it at a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the mutable field list, followed by the patch-semantics constraint and then the sibling routing. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, nested variant objects, no output schema, and near-zero schema descriptions, the definition covers the majority of the mutation surface and the add/update/delete variant distinction well, but omits several scalar fields (handle, productType, descriptionHtml) and variant sub-fields an agent may need to populate correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 10%, so the description must carry the load. It maps the main parameter groups (title, description, SEO, status, vendor, tags, variants) and clarifies the variant id semantics, but handle, productType, descriptionHtml, and the variant-level barcode/compareAtPrice/inventoryPolicy fields are left undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) plus resource (Shopify product) and enumerates the mutable surface: title, description, SEO, status, vendor, tags, variant prices/SKUs. It also distinguishes itself from shopify_create_variants and shopify_delete_variants by declaring that variants passed here are updated in place.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: update existing variants here, add via shopify_create_variants, remove via shopify_delete_variants, and each variant requires its GID. It does not, however, address when to prefer shopify_bulk_update_variants over this tool for multi-product variant edits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shopify_weekly_summaryShopify Weekly Financial SummaryARead-onlyIdempotent
Aggregate financial summary — fetches ALL orders in the period (auto-paginated), calculates revenue totals, AOV, items per order, and product breakdown sorted by revenue.
| Name | Required | Description | Default |
|---|---|---|---|
| date_to | Yes | End date YYYY-MM-DD | |
| date_from | Yes | Start date YYYY-MM-DD | |
| status_filter | No | Extra filter e.g. financial_status:paid |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive). The description adds genuinely useful behavior beyond that: it fetches ALL orders in the period and auto-paginates, which warns the agent that this is a heavy, full-scan read rather than a bounded query. It stops short of 5 by not noting rate limits, cost, or result size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the aggregation purpose first and the computed outputs after. No filler, no restatement of the title or parameter list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description appropriately enumerates the return metrics (revenue totals, AOV, items per order, product breakdown), which is the key missing structured information. It is slightly incomplete in that it never addresses the naming mismatch between 'weekly' and the caller-supplied arbitrary date range, nor how status_filter interacts with the aggregates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so date_from, date_to, and status_filter are already documented in the schema. The description only implies date-range and status scoping ('in the period', 'sorted by revenue') without adding format or filtering semantics beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb set (fetches, calculates) and resource (orders/financial summary) and enumerates the computed metrics (revenue totals, AOV, items per order, product breakdown), which separates it from raw-listing siblings like shopify_list_orders. It does not explicitly name the alternative it is not, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. With siblings such as shopify_list_orders, shopify_order_count, and shopify_analytics_query in the same namespace, the agent gets no signal about when this aggregation is preferable to a raw order list or a general analytics query.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
54 tool updates
v4.0.1- First observed
shopify_add_product_media - First observed
shopify_add_tags - First observed
shopify_adjust_inventory - First observed
shopify_analytics_query - First observed
shopify_audit_variant_customs - First observed
shopify_bulk_tag_products - First observed
shopify_bulk_update_variants - First observed
shopify_collection_add_products - First observed
shopify_collection_remove_products - First observed
shopify_create_article - First observed
shopify_create_collection - First observed
shopify_create_product - First observed
shopify_create_variants - First observed
shopify_delete_metafield - First observed
shopify_delete_product - First observed
shopify_delete_variants - First observed
shopify_duplicate_product - First observed
shopify_get_article_by_id - First observed
shopify_get_articles - First observed
shopify_get_blog_by_id - First observed
shopify_get_blogs - First observed
shopify_get_collections - First observed
shopify_get_customer - First observed
shopify_get_customer_events - First observed
shopify_get_inventory_levels - First observed
shopify_get_locations - First observed
shopify_get_metafields - First observed
shopify_get_order - First observed
shopify_get_page_by_id - First observed
shopify_get_pages - First observed
shopify_get_product_by_id - First observed
shopify_get_products - First observed
shopify_get_shop - First observed
shopify_get_variant_locations - First observed
shopify_list_abandoned_checkouts - First observed
shopify_list_discounts - First observed
shopify_list_draft_orders - First observed
shopify_list_marketing_events - First observed
shopify_list_orders - First observed
shopify_move_inventory - First observed
shopify_order_count - First observed
shopify_remove_tags - First observed
shopify_search - First observed
shopify_search_customers - First observed
shopify_set_inventory - First observed
shopify_set_metafield - First observed
shopify_set_variant_locations - First observed
shopify_update_article - First observed
shopify_update_blog - First observed
shopify_update_collection - First observed
shopify_update_inventory_item - First observed
shopify_update_page - First observed
shopify_update_product - First observed
shopify_weekly_summary
TDQS
Scored across 54 tools
Most tools target a distinct resource and action, and descriptions actively clarify overlaps such as get_variant_locations vs get_inventory_levels and list vs by-id retrieval. The main ambiguity is in the inventory/tagging area, where several tools have related but overlapping responsibilities.
All tools share the shopify_ prefix and use snake_case throughout, with a mostly predictable verb_noun structure. A few descriptive names like weekly_summary, order_count, and search do not follow the exact verb pattern, but the overall convention is highly consistent.
54 tools is far beyond the typical well-scoped MCP range, even for a broad Shopify admin server. The domain justifies breadth, but many operations could be consolidated, making the surface heavy for an agent to navigate reliably.
The surface is strong for products, variants, inventory, collections, content, metafields, tags, and analytics. However, major admin workflows are missing write operations for orders, customers, draft orders, and discounts, and there are no create/delete paths for several content types.
Maintenance
Related MCP Connectors
Shopify MCP Pack — wraps the Shopify Admin REST API (2024-01)
Connect AI to store orders, products and inventory with scoped access and human approvals.
Manage a Shopify store's Sternify bundles, gifts, upsells, banners and reviews via OAuth.
200+ read/write tools for GA4, Search Console, Google Ads, Shopify, WooCommerce, Shopware & more.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables interaction with Shopify stores through GraphQL API, providing tools for managing products, customers, orders, and more.125 npm10MIT
- AlicenseCqualityDmaintenanceEnables interaction with Shopify store data through GraphQL API, providing tools for managing products, customers, orders, blogs, and articles.152,107 npm4MIT
- AlicenseNot gradedqualityDmaintenanceProvides full access to the Shopify Admin GraphQL API through 75 tools for managing products, orders, customers, inventory, and analytics. It supports both static token and OAuth authentication while featuring cost-aware rate limiting for efficient store management.2,107 npmMIT
- AlicenseDqualityDmaintenanceEnables interaction with Shopify store data through the GraphQL API, providing tools for managing products, customers, orders, and collections.212,107 npm9MIT