Skip to main content
Glama

Shopify Store Ops

Low-stock reorders, where-is-my-order replies, daily sales digests and catalogue imports.

An MCP server with 12 workflows across Shopify, Google Sheets, Gmail, Slack, Firecrawl, Linear and Stripe. Each workflow is a prompt your agent runs as a slash command, over the 26 tools it needs and no others.

uv tool install https://github.com/r28ai/shopify-ops-mcp/releases/download/v0.1.0/shopify_ops_mcp-0.1.0-py3-none-any.whl
claude mcp add shop -- shopify-ops-mcp

It installs with uv from this repository's release, with no git and nothing to build; nothing but Charter and the libraries it uses comes from PyPI. To update, run the install line from the latest release. If a desktop app cannot find shopify-ops-mcp, give it the full path from which shopify-ops-mcp (where shopify-ops-mcp on Windows).

Then ask your agent to connect your apps, or run /mcp__shop__setup.

Connect your apps

Ask the agent to connect one ("connect Linear"). It tells you where to get that app's key and the command that stores it, and the next call works, with no restart. The agent never asks for a key in the chat.

Or connect everything this server uses from a terminal:

shopify-ops-mcp login            # each app in turn
shopify-ops-mcp login shopify    # just one
shopify-ops-mcp status           # what is connected

Tokens and keys go to your operating system's keychain (macOS Keychain, Windows Credential Manager, the Secret Service on Linux), and are checked with one read-only call to the app's own API before they are kept. Every key, token and OAuth client is yours: we register no app with any of these services, and nothing passes through a server of ours, because there isn't one.

App

How it connects

Or set

Shopify

Your own key (get one), entered once.

SHOPIFY_SHOP, SHOPIFY_CLIENT_ID, SHOPIFY_CLIENT_SECRET

Google

Browser sign-in, over your own OAuth client (make one).

GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET

Slack

Your own key (get one), entered once. A bot token from your own Slack app, which the guide sets up in about three minutes.

SLACK_BOT_TOKEN

Firecrawl

Your own key (get one), entered once.

FIRECRAWL_API_KEY

Linear

Your own key (get one), entered once.

LINEAR_API_KEY

Stripe

Your own key (get one), entered once.

STRIPE_API_KEY

A variable set in your client's config always wins over the keychain.

Related MCP server: Commerce-MCP

Workflows

Workflow

What you get

Apps

Low stock → supplier PO draft low_stock_to_supplier_po_draft

Reorders drafted before a bestseller goes out of stock.

Shopify, Google Sheets, Gmail

Daily sales digest daily_sales_digest

Yesterday's orders, revenue and top products, posted and logged.

Shopify, Google Sheets, Slack

Where-is-my-order replies where_is_my_order_replies

The commonest support email answered with the actual fulfillment status.

Gmail, Shopify

Competitor product price watch competitor_product_price_watch

A competitor undercuts a SKU and you see both prices side by side.

Firecrawl, Shopify, Slack

Supplier sheet → new products supplier_sheet_to_new_products

A season's catalogue goes from spreadsheet to store draft in one run.

Google Sheets, Shopify, Slack

Product copy from the supplier's page product_copy_from_the_supplier_s_page

Specs and materials pulled from the manufacturer, written as your listing.

Firecrawl, Shopify

Wholesale order by email → draft order wholesale_order_by_email_to_draft_order

B2B buyers email a list and get an invoice link back.

Gmail, Shopify

Cancellation request handled cancellation_request_handled

Unfulfilled orders cancelled on request with the confirmation drafted.

Gmail, Shopify

VIP customer outreach vip_customer_outreach

Top spenders get a personal note before a launch, not a blast.

Shopify, Google Sheets, Gmail

Fulfillment exceptions fulfillment_exceptions

Orders unfulfilled after 48 hours become an ops issue with the order attached.

Shopify, Linear, Slack

Stock count from a sheet stock_count_from_a_sheet

The warehouse's count becomes the store's inventory without manual edits.

Google Sheets, Shopify

Shopify + Stripe revenue in one sheet shopify_and_stripe_revenue_in_one_sheet

Stores selling in two places see one revenue number.

Shopify, Stripe, Google Sheets

Every prompt takes one optional argument, details: the repo, team, channel, customer or date range you mean, so the agent does not have to ask. In Claude Code, put it in quotes, or only its first word arrives:

/mcp__shop__low_stock_to_supplier_po_draft "anything under 10 units, supplier sheet 'POs'"

Reads run without asking. Before anything that creates, sends, changes or deletes, the prompt tells the agent to show you the call and wait.

3 of the 12 workflows need no Google or Granola credential.

Other clients

Claude Desktop: install uv if you have not, since Claude Desktop starts the server with it, then open the .mcpb from the latest release. Claude asks for any keys in its own settings and keeps them in your keychain. The first start takes a few seconds longer, while uv installs it.

VS Code (.vscode/mcp.json): VS Code asks for each key the first time the server starts and stores it securely. Leave out any you stored with login.

{
  "inputs": [
    {
      "type": "promptString",
      "id": "shopify-client-secret",
      "description": "Shopify: App client secret",
      "password": true
    },
    {
      "type": "promptString",
      "id": "google-client-secret",
      "description": "Google: OAuth client secret",
      "password": true
    },
    {
      "type": "promptString",
      "id": "slack-bot-token",
      "description": "Slack: Bot token (xoxb-\u2026)",
      "password": true
    },
    {
      "type": "promptString",
      "id": "firecrawl-api-key",
      "description": "Firecrawl: API key",
      "password": true
    },
    {
      "type": "promptString",
      "id": "linear-api-key",
      "description": "Linear: Personal API key",
      "password": true
    },
    {
      "type": "promptString",
      "id": "stripe-api-key",
      "description": "Stripe: Secret or restricted key",
      "password": true
    }
  ],
  "servers": {
    "shop": {
      "type": "stdio",
      "command": "shopify-ops-mcp",
      "env": {
        "SHOPIFY_CLIENT_SECRET": "${input:shopify-client-secret}",
        "GOOGLE_CLIENT_SECRET": "${input:google-client-secret}",
        "SLACK_BOT_TOKEN": "${input:slack-bot-token}",
        "FIRECRAWL_API_KEY": "${input:firecrawl-api-key}",
        "LINEAR_API_KEY": "${input:linear-api-key}",
        "STRIPE_API_KEY": "${input:stripe-api-key}",
        "SHOPIFY_SHOP": "",
        "SHOPIFY_CLIENT_ID": "",
        "GOOGLE_CLIENT_ID": ""
      }
    }
  }
}

Cursor (.cursor/mcp.json) starts it the same way:

{
  "mcpServers": {
    "shop": {
      "command": "shopify-ops-mcp"
    }
  }
}

Codex (~/.codex/config.toml) starts a turn without waiting for a server unless it is required, and then the agent has none of its tools. required = true makes the session wait for it, and startup_readiness = "catalog" waits for its tool list rather than just its connection:

[mcp_servers.shop]
command = "shopify-ops-mcp"
required = true
startup_readiness = "catalog"
startup_timeout_sec = 30

Name the server shop. A host builds each tool's name from that key, and a longer one can push a tool past the 64 characters a function name allows.

Built with Charter

Every tool here is a Charter declaration: a Pydantic schema saying where each field goes on the wire. Charter's runtime builds the request, attaches and refreshes the credential, and trims the response before the model reads it. It runs in your process, with no proxy and no telemetry.

The 26 tool schemas come to 25,389 tokens.

The same tools work in your own agent, without MCP:

from charter.adapters.openai import to_openai_tools
from charter_packs_mcp import FAMILIES

tools = FAMILIES["commerce"].tools()
definitions = to_openai_tools(tools)   # or charter.adapters.langchain

Need an API that isn't here? Write a pack: your coding agent writes the declarations, and Charter's conformance suite checks them.

  • Shopify: shopify_products_list, shopify_variant_inventory_level, shopify_orders_list, shopify_order_fulfillment_orders, shopify_product_get, shopify_product_create, shopify_product_variants_bulk_create, shopify_product_update, shopify_customers_list, shopify_draft_order_create, shopify_order_get, shopify_order_cancel, shopify_locations_list, shopify_inventory_adjust_quantities

  • Google Sheets: gsheets_spreadsheets_values_append, gsheets_spreadsheets_values_get, gsheets_spreadsheets_values_update

  • Gmail: gmail_drafts_create, gmail_threads_list, gmail_threads_get

  • Slack: slack_chat_post_message

  • Firecrawl: firecrawl_monitor_create, firecrawl_extract, firecrawl_scrape

  • Linear: linear_issue_create

  • Stripe: stripe_balance_transactions_list

License

Apache 2.0.

Available Tools

28 tools
connectA

Connect one app this server uses. For an app that issues keys, says where to get one and the terminal command that stores it. For Google, once the user's own OAuth client is set, starts the browser sign-in and returns at once: the user approves in the browser and the next call works. To see which apps are connected, call connection_status. Never ask the user for a key in the chat.

ParametersJSON Schema
NameRequiredDescriptionDefault
appYesThe app to connect.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=false, so the description carries most of the behavioral load. It discloses meaningful traits: key-apps explain where to obtain a key and the terminal command that stores it, Google's flow starts a browser sign-in and returns immediately so the next call works, and the agent must not request keys in chat. Auth/permission nuances beyond this aren't covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by compact per-case behavior and a memorable closing rule; every sentence contributes. It is slightly dense/multi-clause in the middle but avoids redundancy with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter connect tool with no output schema and minimal annotations, the description covers the critical decisions: how different app types behave, the immediate return on Google sign-in, how to check status, and the key-handling safety rule. Return shape is left implicit but is not required here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single enum parameter, setting baseline 3. The description adds real semantic value by branching the 'app' values into behavioral categories (key-issuing apps vs Google), which tells the agent what a given enum value will trigger beyond the schema's flat 'The app to connect.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a clear verb (Connect) and resource (one app this server uses), and the body differentiates this tool from the sibling connection_status by routing status checks there. The scope is specific enough that an agent can pick it without opening the schema, though the declarative 'what it does' could be sharper.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit per-app conditions (key-issuing apps vs Google) and names the alternative tool (`connection_status`) for seeing which apps are connected. It also gives a hard usage rule: never ask for a key in chat. It stops short of stating when not to use the tool beyond the one routing example.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connection_statusA
Read-only

See which apps this server is connected to, and how to connect each one that is not. Changes nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered by structured data. 'Changes nothing' restates the readOnly hint rather than adding new behavior; the only incremental value is noting that connect instructions are returned for unconnected apps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence that front-loads the primary purpose and appends the secondary benefit with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, read-only status tool with no output schema, the description covers both what is inspected and the shape of the useful payload (connect guidance). Return format details are absent but minimal given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly implies no input is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('See') and resource ('which apps this server is connected to'), and adds the secondary payload of connect instructions for missing apps. This distinguishes it from the sibling 'connect' tool, which performs the connection rather than reporting status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'how to connect each one that is not' implies this tool is the discovery step before using 'connect', but the sibling is never named and there is no explicit when-to-use/when-not statement. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_extractA

Extract structured data from one or more URLs using an LLM. Poll results with extract_status.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYesThe URLs to extract data from. URLs should be in glob format.
promptNoPrompt to guide the extraction process.
schemaNoSchema to define the structure of the extracted data. Must conform to JSON Schema.
showSourcesNoWhen true, the sources used to extract the data will be included in the response as `sources`.
ignoreSitemapNoWhen true, sitemap.xml files will be ignored during website scanning.
scrapeOptionsNoOptions applied when scraping pages for extraction.
enableWebSearchNoWhen true, the extraction will use web search to find additional data.
threatProtectionNoPer-request threat protection override. Enterprise feature.
ignoreInvalidURLsNoIf invalid URLs are specified, they are ignored and returned in invalidURLs instead of failing the request. The server applies true when this is absent.
includeSubdomainsNoWhen true, subdomains of the provided URLs will also be scanned. The server applies true when this is absent.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, openWorldHint=true) are minimal and do not reveal the async job lifecycle, so the description correctly surfaces the most important behavioral trait: results must be polled via extract_status. It stops there, omitting what the initial call returns (a job id?), failure behavior, and whether web-search/credit costing applies. Decent added context over annotations, but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste. The core purpose is front-loaded and the polling instruction follows immediately. Nothing to trim and nothing misplaced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, no output schema, and an async job model, the description covers the poll step but not the full lifecycle: it does not explain what the initial invocation returns or how to correlate extract_status results back to it. For a complex async tool with no output schema to fall back on, this is a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 10 parameters (urls, prompt, schema, scrapeOptions, enableWebSearch, etc.) are already documented in the schema. The description adds no parameter-level detail (e.g., the glob URL format or JSON Schema requirement) beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Extract structured data from one or more URLs') and adds the key mechanism ('using an LLM'), so the agent understands this is LLM-driven structured extraction rather than raw scraping. It does not, however, distinguish itself from the very close sibling firecrawl_scrape, leaving the boundary between the two to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence ('Poll results with extract_status') is genuine usage guidance: it tells the agent this is an async operation that must be polled. But it gives no guidance on when to choose this tool over firecrawl_scrape or other extraction options, and does not name any exclusion criteria. Usage is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_monitor_createB

Create a scheduled monitor for scrape, crawl, or search targets.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalNoPlain-language goal used to judge whether changed pages are meaningful.
nameYesMonitor name.
targetsYesTargets to run on each check.
webhookNoWebhook destination for monitor events.
scheduleYesSchedule for monitor checks.
judgeEnabledNoWhether to judge changed pages against goal.
notificationNoNotification destinations.
retentionDaysNoHow long to retain monitor history. The server applies 30 when this is absent.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the mutation and external-reach profile is already covered; the description's 'create' is consistent with these and adds no contradiction. It adds only the notion that the created object is scheduled/recurring, but says nothing about lifecycle, cost, or persistence beyond that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, correctly leading with the verb and the resource. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex creation tool with required schedule/targets, nested target union types, webhooks, notifications, retention, and an optional judge/goal mechanism, and no output schema to fall back on. The one-line description does not explain what a monitor does over time or what happens after creation, leaving significant gaps for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 8 parameters (schedule, targets, webhook, judgeEnabled, notification, retentionDays, goal, name) are already documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Create a scheduled monitor') and scopes it to the three target kinds the schema supports (scrape, crawl, search). It implicitly separates this from the one-off firecrawl_scrape/firecrawl_crawl siblings via 'scheduled,' but never names them, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites (e.g. that monitors run repeatedly and incur ongoing credit usage), and no pointer to alternatives like firecrawl_crawl or firecrawl_crawl_status for one-off jobs. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

firecrawl_scrapeB

Scrape a single URL and optionally extract information. Use when the user wants to read or summarize a specific webpage. Supports markdown, HTML, screenshots, and structured JSON extraction.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe URL to scrape
proxyNoSpecifies the type of proxy to use.
maxAgeNoReturns a cached version of the page if it is younger than this age in milliseconds. The server applies 172800000 (2 days) when this is absent.
minAgeNoWhen set, the request only checks the cache and never triggers a fresh scrape.
mobileNoEmulate scraping from a mobile device.
actionsNoActions to perform on the page before grabbing the content.
formatsNoOutput formats to include in the response. Strings or objects. The server applies markdown when this is absent.
headersNoHeaders to send with the request.
parsersNoControls how files are processed during scraping.
profileNoPersistent browser storage across scrape and interact sessions.
timeoutNoTimeout in milliseconds. The server applies 60000 when this is absent.
waitForNoSpecify a delay in milliseconds before fetching the content. The server applies 0 when this is absent.
blockAdsNoEnables ad-blocking and cookie popup blocking.
locationNoLocation settings for the request.
lockdownNoServe from cache only and never make an outbound request. On miss, returns 404 SCRAPE_LOCKDOWN_CACHE_MISS.
redactPIINoRedact personally identifiable information from returned markdown. Pass true for defaults, or an object to tune it.
excludeTagsNoTags to exclude from the output.
includeTagsNoTags to include in the output.
storeInCacheNoIf true, the page will be stored in the Firecrawl index and cache.
auditMetadataNoUser attribution included with SIEM logging events when SIEM is enabled.
onlyMainContentNoOnly return the main content of the page excluding headers, navs, footers, etc. The server applies true when this is absent.
onlyCleanContentNoBeta. LLM pass over markdown to remove residual boilerplate that onlyMainContent can miss.
threatProtectionNoPer-request threat protection override. Enterprise feature.
zeroDataRetentionNoIf true, this will enable zero data retention for this scrape. To enable this feature, please contact help@firecrawl.dev
removeBase64ImagesNoRemoves all base64 images from the markdown output.
skipTlsVerificationNoSkip TLS certificate verification when making requests.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the agent knows this touches external state. The description adds the supported output formats, which is useful, but it omits behavior implied by the schema — that actions (click/write/executeJavascript) mutate the page, that storeInCache writes to an external index, and that some features cost credits. Nothing contradicts the annotations, but the added behavioral detail is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core purpose front-loaded and no filler. The trailing format list is somewhat redundant with the schema's formats enum, which keeps it short of ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a 26-parameter tool with no output schema and only minimal annotations, yet the description is three sentences long. It omits cost/credit implications, caching/lockdown semantics, the relationship to firecrawl_extract, and any hint of what the response looks like, so an agent invoking it correctly still depends almost entirely on reading the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 26 parameters are already documented in-schema and the baseline is 3. The description echoes the formats dimension ('markdown, HTML, screenshots, structured JSON extraction') but adds no format syntax, precedence, or interaction detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Scrape a single URL') plus an optional outgrowth ('optionally extract information'), so an agent can tell it is a per-URL content fetcher. The phrase 'a single URL' gestures at the multi-URL alternative but never names firecrawl_extract, so sibling differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one clear trigger ('Use when the user wants to read or summarize a specific webpage'), which is real usage guidance. But it never states when NOT to use it, nor does it point to firecrawl_extract for bulk/structured extraction, so the routing decision against the closest sibling is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_drafts_createC

Save an email draft to Gmail.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesThe draft to create.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/non-read nature is covered structurally. The description adds nothing beyond that: it does not say the draft is not sent, whether creation is idempotent, or what auth is required, so the behavioral burden is largely unmet.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is well structured. It is arguably under-specified rather than bloated, but as a size/structure judgment it is tight and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with only minimal annotations and no output schema, the description should at least clarify that it saves rather than sends and hint at the returned draft. The rich input schema compensates for parameters, but the core behavioral distinction is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the nested Message/Draft/EmailContent fields are richly documented (threadId rules, bodyHtml multipart behavior, in_reply_to threading). The description adds no parameter meaning at all, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Save an email draft') and destination (Gmail), which is enough to know it creates a draft rather than sending. However, it does not distinguish itself from the adjacent gmail_messages_send sibling, so an agent gets no explicit routing cue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of the alternative tool (gmail_messages_send) for actually delivering mail. The agent must infer that this only persists a draft.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_threads_getC
Read-onlyIdempotent

Read a Gmail thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe unique ID of the Gmail thread to retrieve.
formatNoThe format to return the messages in.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me
metadataHeadersNoWhen format is 'METADATA', only include these headers in the response.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds nothing beyond that—no note on auth requirements, what a 'thread' contains, or how format affects the response—so it does not enrich the behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero waste. It is appropriately sized, though the brevity reflects under-specification rather than tight editing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool whose annotations cover safety and whose schema documents all four parameters, the description is minimally adequate. It would be stronger if it hinted at the format options or return structure, but nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with every parameter (id, format, userId, metadataHeaders) documented in the schema itself. The description adds no parameter meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read') and resource ('a Gmail thread'), which cleanly separates it from the list-oriented sibling gmail_threads_list. It does not, however, explicitly contrast itself with gmail_threads_list or gmail_messages_list, so sibling differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus gmail_threads_list or gmail_messages_list, and no prerequisites or context are given. The agent must infer usage purely from the name and the required 'id' parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_threads_listC
Read-onlyIdempotent

List Gmail threads.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoOnly return threads matching this Gmail search query string.
userIdNoThe user's email address. The special value 'me' can be used to indicate the authenticated user.me
labelIdsNoReturn only threads with all of these label IDs.
pageTokenNoPage token to retrieve a specific page of results in the list.
maxResultsNoMaximum number of threads to return (default 100, max 500).
includeSpamTrashNoInclude threads from SPAM and TRASH in the results.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered by structured data. The description adds nothing beyond that — no note on pagination behavior, result ordering, or the fact that a full mailbox scan may be needed without filters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the key information front-loaded. It is efficient, though the extreme brevity shades into under-specification rather than pure conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter read tool with no output schema, the description is minimally viable: the schema covers all inputs, but the description omits pagination semantics and what a thread result contains. Nothing is misleading, but an agent gets no help beyond the structured fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (q, userId, labelIds, pageToken, maxResults, includeSpamTrash) is already documented in the schema. The description contributes no additional meaning, which is the baseline 3 case when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('Gmail threads'), which is unambiguous and distinguishable from write-oriented siblings like gmail_messages_send and gmail_drafts_create. It stops short of scope details (e.g. mailbox-wide vs. label-filtered), so it is clear but not maximally informative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives, no mention of prerequisites or the conditions under which a caller should prefer it. The sibling set contains other Gmail operations but the description offers no routing signal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_appendC

Appends values to a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation of a range to search for a logical table of data. Values are appended after the last row of the table.
valueRangeYesThe request body contains an instance of ValueRange.
spreadsheetIdYesThe ID of the spreadsheet to update.
insertDataOptionNoHow the input data should be inserted.
valueInputOptionYesHow the input data should be interpreted.
includeValuesInResponseNoDetermines if the update response should include the values of the cells that were appended. By default, responses do not include the updated values.
responseValueRenderOptionNoDetermines how values in the response should be rendered. The default render option is FORMATTED_VALUE.
responseDateTimeRenderOptionNoDetermines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false and openWorldHint=true, indicating a mutating, external operation. The description adds nothing beyond this – it doesn't disclose how appended data interacts with existing tables, whether headers are auto-detected, or rate-limit/permission requirements. For a mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single, front-loaded sentence that wastes no words. However, the extreme brevity contributes to gaps elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, nested value types, mutation behavior), the description is critically underspecified. It omits key behavioral details like how the append range is determined, interaction with existing data, and available options (insertDataOption, valueInputOption). No output schema exists to compensate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds no parameter-level detail beyond what's already provided. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (appends) and resource (values to a spreadsheet), which is clear but does not differentiate from the sibling gsheets_spreadsheets_values_update or clarify the 'logical table' append behavior. It's clear but lacks sibling differentiation within the gsheets tool family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool vs alternatives like gsheets_spreadsheets_values_update, and no mention of required preconditions or idempotency considerations. The description is silent on usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_getC
Read-onlyIdempotent

Returns a range of values from a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation or R1C1 notation of the range to retrieve values from.
spreadsheetIdYesThe ID of the spreadsheet to retrieve data from.
majorDimensionNoThe major dimension that results should use. For example, if the spreadsheet data in Sheet1 is: A1=1,B1=2,A2=3,B2=4, then requesting range=Sheet1!A1:B2?majorDimension=ROWS returns [[1,2],[3,4]], whereas requesting range=Sheet1!A1:B2?majorDimension=COLUMNS returns [[1,3],[2,4]].
valueRenderOptionNoHow values should be represented in the output. The default render option is FORMATTED_VALUE.
dateTimeRenderOptionNoHow dates, times, and durations should be represented in the output. This is ignored if valueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered structurally. The description adds nothing beyond that — no mention of behavior for missing/empty ranges, error cases, or whether the range must already exist — so it contributes essentially no behavioral context of its own.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero padding. It is efficient, though arguably terse given the tool takes five parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover the safety profile and the schema fully documents inputs, so the description's burden is lighter. Still, with no output schema, it says only that 'values' are returned without hinting at the row/column shape that majorDimension controls, which is the main thing an agent must reason about.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the parameters (range, spreadsheetId, majorDimension, valueRenderOption, dateTimeRenderOption) are already richly documented in the schema, including an example for majorDimension. The description adds no parameter meaning at all, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Returns a range of values from a spreadsheet'), which is clearly readable. However, it does not differentiate this from siblings such as gsheets_spreadsheets_values_update or gdocs_documents_get, leaving the agent to infer the read-vs-write distinction from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of when NOT to use it, and no reference to the sibling update tool. The agent gets an implied read-only purpose from the verb but nothing explicit about context or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gsheets_spreadsheets_values_updateC

Sets values in a range of a spreadsheet.

ParametersJSON Schema
NameRequiredDescriptionDefault
rangeYesThe A1 notation of the values to update.
valueRangeYesThe request body contains an instance of ValueRange.
spreadsheetIdYesThe ID of the spreadsheet to update.
valueInputOptionYesHow the input data should be interpreted.
includeValuesInResponseNoDetermines if the update response should include the values of the cells that were updated. By default, responses do not include the updated values. If the range to write was larger than the range actually written, the response includes all values in the requested range (excluding trailing empty rows and columns).
responseValueRenderOptionNoDetermines how values in the response should be rendered. The default render option is FORMATTED_VALUE.
responseDateTimeRenderOptionNoDetermines how dates, times, and durations in the response should be rendered. This is ignored if responseValueRenderOption is FORMATTED_VALUE. The default dateTime render option is SERIAL_NUMBER.

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the write nature is known. But the description adds nothing beyond the name – it does not disclose that existing cell values are overwritten, how valueInputOption affects interpretation, or what the update returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, which is tight. But for a 7-parameter mutation tool it is under-specified rather than appropriately sized; conciseness here is closer to omission.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with 7 params and no output schema, the description omits overwrite semantics, auth requirements, and response behavior. An agent could call it, but not safely without reading the schema closely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with 7 well-documented parameters, so the schema carries the meaning. The description adds no parameter detail beyond 'in a range', which is baseline 3 for fully documented schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (sets values in a range of a spreadsheet), which an agent can distinguish from gsheets_spreadsheets_values_get by direction of data flow. However it offers no explicit sibling differentiation and largely restates the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the sibling gsheets_spreadsheets_values_get or when reading vs writing applies. The agent must infer context entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

linear_issue_createA

Create an issue. team_id and title are required; everything else is optional. The UUIDs for team, assignee, state and labels come from teams_list, users_list and workflow_states_list — Linear does not accept names here. Set parent_id to create a sub-issue.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe issue to create.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the mutation/external-scope profile is covered. The description adds genuinely useful behavior: Linear rejects names and only accepts UUIDs resolved via teams_list/users_list/workflow_states_list, and parent_id turns this into a sub-issue creation. It stops short of noting side effects like notifications or returned identity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and the required fields, followed by the ID-resolution constraint and the sub-issue tip. No filler, though the UUID guidance partially duplicates the schema text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool whose one parameter is a deeply nested object fully documented by the schema, plus annotations covering the safety profile, the description covers the essentials an agent needs. It lacks any mention of what creation returns, which is a minor gap given there is no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every field including the resolver-tool hints and the sub-issue semantics. The description's parameter notes (team_id/title required, everything else optional) largely restate the schema rather than adding new meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an issue'), which is instantly distinguishable from the list-oriented sibling linear_issues_list. It does not explicitly name a sibling it is not, so it falls short of a 5, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives practical invocation guidance (required vs optional fields, which resolver tools supply UUIDs, how to make a sub-issue) but never states when to reach for this tool versus alternatives such as linear_issues_list, nor any exclusion or prerequisite conditions. Usage is implied rather than framed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_customers_listB
Read-onlyIdempotent

List customers. Narrow with a Shopify search string in query, e.g. 'email:ada@example.com' or 'orders_count:>5'.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesNoPaging and filtering.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered elsewhere. The description itself adds no behavioral context beyond filtering — nothing on pagination behavior, result size, or consistency (that nuance lives only in the schema's parameter text).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, with the core purpose front-loaded and the filtering hint immediately after. The examples earn their place by being copy-pasteable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema and 100% schema coverage, the description is adequate but thin. Paging and consistency behavior are documented in the schema, but the description offers no routing guidance among the many sibling list tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter including qualifier lists, page size limits, cursor semantics and eventual-consistency warnings. The description's query examples duplicate rather than extend that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ("List customers") and immediately scopes it with the filtering mechanism, so the agent knows exactly what the tool returns. It does not, however, distinguish itself from the other Shopify list siblings (shopify_orders_list, shopify_products_list), relying on the resource name alone to do that work.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete narrowing syntax with two worked examples ('email:ada@example.com', 'orders_count:>5'), which is genuine usage guidance. It stops short of stating when to reach for this tool versus alternatives, and there are no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_draft_order_createA

Create a draft order, the way to build an order by hand before it is placed. A draft holds no inventory until it is completed.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe draft order to create.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, and the description adds a genuine behavioral trait beyond them: 'A draft holds no inventory until it is completed,' which tells the agent that creating a draft does not reserve stock. It does not cover permissions, error behavior, or what completion does, but the key consequence is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, no filler, with the purpose front-loaded and the inventory caveat following as supporting context. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation with no output schema, the description supplies the essential conceptual and behavioral context (what a draft is, that inventory is not held). Adding permission requirements or what happens on completion would make it fully complete, but an agent can call this correctly as written.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with a single parameter, and the schema description already enumerates the useful input keys (lineItems, customerId, email, note, tags, shippingAddress). The description adds nothing about input semantics, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Create) and resource (draft order) and adds a conceptual gloss — building an order by hand before placement — that clarifies what a 'draft order' is. It does not, however, explicitly distinguish itself from the sibling create tools (shopify_product_create, linear_issue_create, gmail_drafts_create), so sibling differentiation is only implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'the way to build an order by hand before it is placed' implies the usage context (assembling an order prior to checkout), but there is no explicit when-to-use, when-not-to-use, or named alternative. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_inventory_adjust_quantitiesA

Move stock by a relative amount: delta of 5 adds five, -5 removes five. This is not a way to set an absolute count. Every change must also carry change_from_quantity, the count you expect to be overwriting — read it with variant_inventory_level immediately before — so two adjustments racing each other cannot both win. committed is never writable, and on_hand cannot be named here although adjusting available moves it by the same delta. The result reports the deltas applied and leaves quantity_after_change null, so read the level back for the new count. Since 2026-04 adjusting an item that is not stocked at the location succeeds and creates an inactive level, so a success does not prove the stock is sellable.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe adjustment to apply.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries most of the load and does it well: `committed` is never writable, `on_hand` cannot be named, the result reports deltas with a null `quantity_after_change` so a read-back is required, and the 2026-04 inactive-level behavior is disclosed. The one gap is it never states permission/auth requirements for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core delta semantics, then layers constraints and edge cases; every sentence earns its place with no filler. It is dense and long for a single-parameter wrapper, which keeps it just short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the write behavior, the concurrency requirement, the return shape (deltas reported, count null), and the surprising success semantics. Nothing an agent needs to call it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: it clarifies delta sign semantics with an example, explains why `change_from_quantity` exists (concurrency guard), and documents which states can and cannot be named (`committed` never writable, `on_hand` excluded). This genuinely extends the schema rather than restating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (adjust inventory quantities) and immediately scopes it as a relative delta operation ('Move stock by a relative amount'), explicitly distinguishing it from an absolute set ('This is not a way to set an absolute count'). An agent can tell it apart from the read sibling `shopify_variant_inventory_level` without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions and prerequisites: every change must carry `change_from_quantity`, and it must be read with `variant_inventory_level` immediately before to avoid racing adjustments. It also names a surprising edge case (non-stocked item now succeeds and creates an inactive level) so the agent knows success does not imply sellability.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_locations_listB
Read-onlyIdempotent

List the store's locations. Inactive ones are excluded unless asked for.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesNoPaging and filters.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description's only added claim (inactive locations excluded unless requested) merely restates the includeInactive schema field, and it omits pagination limits, cost behavior, or return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and immediately followed by the scoping caveat. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with a fully documented schema and covering annotations, the definition is mostly adequate. With no output schema, however, nothing tells the agent what a returned location contains or how paging results are shaped, leaving a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter (after, first, query, reverse, includeInactive) is documented in the schema. The description adds no syntax or format detail beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('List the store's locations') and states a filter scope, so the agent knows exactly what it does. It doesn't distinguish itself from siblings beyond the distinct resource name, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Inactive ones are excluded unless asked for' implies when the default behavior applies and hints at the includeInactive path. However, there is no explicit when-to-use versus alternatives or any prerequisite guidance, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_order_cancelA

Cancel an order. restock has to be stated: true returns the items to inventory, false leaves stock as it is. Shopify does this in the background and answers with a job, so no errors means accepted rather than finished. Cancelling cannot be undone.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe order to cancel.

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true; the description adds substantial non-obvious behavior: the operation is asynchronous (returns a job, success means queued not completed), the restock choice materially affects inventory, and the action is irreversible. These are exactly the traits an agent needs beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the operation stated first, then the parameter semantics, then the async/irreversibility caveat. No filler and the most important warning (cannot be undone) lands last for emphasis.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no annotations covering danger and no output schema, the description covers irreversibility, async job semantics, and inventory effects well. It omits preconditions (which order states allow cancellation) and what the returned job looks like, but nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and every parameter is documented, including the restock consequence and the notifyCustomer default, so the description's restock sentence largely duplicates the schema. Baseline 3 applies because the schema does the heavy lifting, though 'has to be stated' usefully stresses that restock is mandatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Cancel an order') with no ambiguity, and cancellation is intrinsically distinguishable from the sibling tools (orders_list, order_get, order_fulfillment_orders). An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clarifies required semantics of `restock` and warns that no errors means 'accepted' not 'finished', which is useful invocation context. However it offers no when-to-use guidance versus alternatives and no prerequisites such as the order state that permits cancellation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_order_fulfillment_ordersA
Read-onlyIdempotent

List an order's fulfillment orders, which is the first half of fulfilling it. Shopify creates these itself, and fulfillment_create names them. The location is at assignedLocation.location.id.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe order to resolve.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so safety/idempotency are covered. The description adds genuinely useful context: that Shopify auto-creates these records and that fulfillment_create operates on them, plus where the location id lives. It doesn't add pagination or rate-limit behavior, but the annotations carry the main burden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, with no filler. The pointer to assignedLocation.location.id is a compact, useful addition. Slightly terse on usage, but structurally sound.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with full annotation safety coverage and 100% schema coverage, the description is largely complete. With no output schema, pointing to assignedLocation.location.id is a helpful compensating hint about the return structure, though more on the relationship to fulfillment_create's inputs could round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the id and first parameters are fully documented in the schema; baseline 3 applies. The description's mention of assignedLocation.location.id refers to the response shape rather than adding syntax or meaning to the input parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List an order's fulfillment orders') and clarifies what fulfillment orders are ('the first half of fulfilling it'). It distinguishes the concept from ordinary orders, though it doesn't explicitly differentiate why to call this instead of the sibling shopify_order_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'which is the first half of fulfilling it' and the reference to fulfillment_create naming them, suggesting it's the prerequisite step before fulfilling. However, there is no explicit when-to-use vs when-not or named alternative selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_order_getA
Read-onlyIdempotent

Get one order with its line items, shipping address and totals. Takes a global ID (gid://shopify/Order/...), which orders_list returns.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesWhich order to fetch.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint and idempotentHint, so the safety profile is covered externally. The description adds useful shape-of-result information (line items, address, totals) and constrains the ID format, but says nothing about auth requirements, rate limits, or error behavior. With annotations carrying the safety burden, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the purpose is front-loaded and the ID-format detail is deferred to the second sentence where it belongs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with full annotation coverage, the description does the work of standing in for an absent output schema by naming what gets returned. Nothing essential is missing, though error/not-found handling is unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'id' parameter is already fully documented in the schema with the same gid:// example. The description's mention of the global ID format duplicates rather than extends the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Get') plus a precisely scoped resource ('one order') and it enumerates the payload contents (line items, shipping address, totals). An agent can tell this is a single-record fetch, distinct from any list/collection operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description supplies concrete workflow context by noting the required global ID is the one 'orders_list' returns, which tells the agent when this tool fits (after a listing). It stops short of stating exclusions or what to do if the ID is unknown, so it is not a full when/when-not treatment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_orders_listA
Read-onlyIdempotent

List orders. Narrow with a Shopify search string in query, e.g. 'financial_status:paid fulfillment_status:unshipped' to find orders waiting to ship. Only the last 60 days of orders are visible unless this app holds Shopify's read_all_orders scope, and older orders are absent rather than reported — do not read an empty result as an empty history.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesNoPaging and filtering.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/idempotent annotations by disclosing a critical data-visibility boundary: only the last 60 days are visible unless the app holds read_all_orders, and older orders are silently absent. The explicit warning not to interpret an empty result as an empty history is exactly the kind of failure-mode context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the purpose, then a usage example, then the visibility caveat. Every sentence earns its place and the most important constraint (silent 60-day truncation) is placed prominently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description covers purpose, filtering usage, and the key visibility caveat well enough to call the tool correctly. It stops short of describing the returned order fields or pagination flow, but paging is documented in the schema and return shape is a minor gap for a list endpoint.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema's `query` field already documents qualifiers, the typo-ignored behavior, and the created_at caveat in detail. The description's example largely duplicates the schema's own example, so it adds only marginal meaning over the structured field. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Starts with a specific verb+resource ('List orders'), which cleanly separates it from the singular shopify_order_get sibling. The rest of the description elaborates filtering rather than the core purpose, so the differentiation is implicit in the name rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete usage pattern with a worked search example ('financial_status:paid fulfillment_status:unshipped to find orders waiting to ship'), which tells an agent how and when to apply filtering. It does not name alternatives or state when-not to use this versus shopify_order_get, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_product_createA

Create a product. Only title is required. Note that Shopify creates products unpublished — a new product is not on the storefront until it is published to a sales channel, which is a separate operation.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe product to create.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true. The description adds a genuinely useful non-obvious trait: new products are created unpublished and won't appear on the storefront until published to a sales channel. That is real behavioral context beyond the structured fields, though it omits any note on auth/permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and the required-parameter fact, followed by the one non-obvious caveat. No filler; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with no output schema, the description covers the required input and the key publishing side effect. It doesn't say what the response contains (e.g., the new product's id), which would help chaining, but the essentials for correct invocation are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every field (title, handle, status, tags, vendor, productType, descriptionHtml) is already documented. The description's 'Only title is required' merely restates the schema's required list, adding no format or interaction detail. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a product') that an agent can immediately distinguish from the create/get/update/list siblings. It never names a sibling explicitly, so it stops short of a 5, but the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clarifies that title is required and flags that publishing is a separate operation, which steers the agent away from assuming storefront visibility. However, it never names the follow-up tool or states when to prefer create vs. update/bulk-create, so the guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_product_getA
Read-onlyIdempotent

Get one product with its description and its first 50 variants, including each variant's SKU, price and inventory. Takes a global ID (gid://shopify/Product/...), not a bare number.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesWhich product to fetch.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, openWorld), so the bar is lower; the description adds real value by disclosing that only the first 50 variants are returned, a scope limit the annotations cannot express. It stops short of saying how to retrieve the remaining variants.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what is returned and followed by the input constraint; nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of describing the return payload (description plus variant SKU/price/inventory), which is enough to call the tool correctly. The absence of any hint on paging past the first 50 variants is the only minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the id parameter is already documented in the schema, so the baseline is 3. The 'not a bare number' emphasis reinforces the format but adds no new syntax or semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('one product'), and enumerates the payload (description plus first 50 variants with SKU, price, inventory), which cleanly distinguishes it from shopify_products_list and shopify_product_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The single-record scope implies it is for fetching a known product rather than browsing, but the description never names an alternative (e.g. shopify_products_list) or states when not to use it. The ID-format warning is parameter guidance, not usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_products_listB
Read-onlyIdempotent

List products. Narrow with a Shopify search string in query, e.g. 'status:active vendor:Acme' or 'inventory_total:<5'.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesNoPaging and filtering.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered by structured data. The description itself adds no behavioral context beyond that — no mention of pagination, query-cost implications, or the eventual-consistency caveat that actually lives in the schema. That richer behavioral detail exists but not in the description, which is what this dimension evaluates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the purpose and zero filler. The second sentence restates a `query` example that the schema already provides at greater length, which is mild redundancy rather than waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Annotations cover safety and the schema documents all five paging/filter params at 100% coverage, and there is no output schema to explain. An agent has enough to invoke this correctly; only cursor/pagination behavior is unstated in the description, and that is recoverable from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the `query` property in the schema is far richer than the description (qualifier list, AND/OR semantics, eventual-consistency warning). The description's `query` example is a strict subset of what the schema already documents, so it adds no meaning beyond the schema. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('List products'), which is unambiguous on its own. It does not, however, explicitly differentiate itself from the singular sibling `shopify_product_get` or the adjacent `shopify_orders_list`/`shopify_customers_list`, leaving that distinction to be inferred from the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description shows *how* to narrow results with a search string but never states *when* this tool is the right choice versus `shopify_product_get` for a known ID or another list tool. Usage is implied by the examples rather than stated as guidance or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_product_updateA

Update a product. id is required; only the other fields provided are changed. Note that tags replaces the product's tags rather than adding to them.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe product changes.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, so the mutation nature is already known. The description adds genuine behavioral context beyond that: partial-update semantics ('only the other fields provided are changed') and the destructive-surprise caveat that `tags` replaces rather than appends.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and the required-field/patch rule, followed by the one non-obvious caveat. Nothing is wasted, though the tags warning is duplicated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with annotations covering its write/safety profile and a fully documented schema, the description covers the key semantics an agent needs: required `id`, partial-update behavior, and the tag-replacement gotcha. No output schema exists, so return values need not be described; only return-shape expectations are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including the `tags` replacement note and the `id` format is already documented in the schema. The description's parameter information is therefore a restatement rather than an addition, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Update a product'), which is unambiguous and distinguishable from shopify_product_create and shopify_product_get. It stops short of explicitly naming those siblings or the boundary between create/update, so it is clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the update context by stating `id` is required and that only supplied fields change, which tells the agent how to call it. It offers no explicit when-to-use/when-not guidance or routing to siblings like shopify_product_create for new products, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_product_variants_bulk_createA

Add variants to a product. The SKU goes inside inventory_item, not on the variant. When the product still carries its default placeholder variant, pass the strategy that removes it.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe variants to add.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the safety profile is covered. Beyond that, the description discloses a genuine, non-obvious behavioral quirk (the default placeholder variant and the strategy that removes it) that the schema alone does not surface. It stops short of covering return or partial-failure behavior for a bulk mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core purpose front-loaded and zero filler. The SKU-placement note is a slight non-sequitur in ordering but each sentence carries useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a bulk mutation with no output schema and full schema coverage, the description covers purpose, a key gotcha, and conditional strategy guidance. It omits what happens on partial failure across a bulk array, a minor gap given the otherwise complete coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents every field. The description still adds structural meaning beyond the schema by clarifying that the SKU belongs in `inventory_item` rather than on the variant, and by explaining when the `strategy` value is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Add variants to a product," which is clearly distinct from siblings such as shopify_product_create, shopify_product_update, and shopify_variant_inventory_level. It does not explicitly name or contrast with those siblings, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives targeted conditional guidance for one scenario ("When the product still carries its default placeholder variant..."), which is useful context, but it never states when to choose this bulk-create over alternatives or what prerequisites apply. Usage is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

shopify_variant_inventory_levelA
Read-onlyIdempotent

Read a variant's stock at one location. Returns null where the variant is not stocked there, which is not an error.

ParametersJSON Schema
NameRequiredDescriptionDefault
variablesYesThe variant and location to read.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, openWorldHint), so the bar is lower. The description adds genuinely useful non-annotation behavior: a null return means the variant is not stocked at that location and is explicitly not an error, which prevents an agent from misinterpreting the response.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the single most important edge case. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining the return value, and it does so for the key edge case (null = not stocked, not an error). The shape of a successful non-null response (which stock states come back) is left implicit, but the critical ambiguity is resolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema documents the variant ID, location ID, and the stock-state names array in detail. The description's 'at one location' reinforces the locationId semantics but adds no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: reading a variant's stock level at a single location. This distinguishes it clearly from write-oriented siblings like shopify_inventory_adjust_quantities, though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'Read' implies a retrieval use case, and 'at one location' hints at the per-location granularity, but there is no explicit when-to-use guidance, no prerequisites, and no routing to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

slack_chat_post_messageA

Send a message to a Slack channel, private group, or DM. Provide text for a plain message; set thread_ts to reply inside an existing thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
textNoThe main body text of the message. Required unless blocks or attachments are provided. Used as the fallback string for notifications when blocks are provided, so it is worth setting even then.
parseNoChange how messages are treated. Accepts 'none' or 'full'.
blocksNoA JSON-based array of structured Block Kit blocks.
mrkdwnNoDisable Slack markup parsing by setting to false. Defaults to true.
channelYesAn encoded ID or channel name that represents a channel, private group, or IM channel to send the message to. Prefer the encoded ID (e.g. 'C123ABC456').
iconUrlNoURL to an image to use as the icon for this message. Requires the chat:write.customize scope.
metadataNoApplication-specific metadata to attach to the message.
threadTsNoProvide another message's 'ts' value to make this message a reply in that thread. Avoid using a reply's ts value; use the parent's.
usernameNoSet the bot's user name. Requires the chat:write.customize scope.
iconEmojiNoEmoji to use as the icon for this message, e.g. ':chart_with_upwards_trend:'. Requires the chat:write.customize scope.
linkNamesNoFind and link user groups.
attachmentsNoA JSON-based array of structured attachments.
unfurlLinksNoPass true to enable unfurling of primarily text-based content.
unfurlMediaNoPass false to disable unfurling of media content.
markdownTextNoAccepts message text formatted in markdown. Limit this field to 12,000 characters. Cannot be used together with blocks or text.
replyBroadcastNoUsed in conjunction with thread_ts and indicates whether the reply should be made visible to everyone in the channel. Defaults to false.
unfurlAppLinksNoPass true to enable unfurling of links to installed apps.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the write/external nature is covered. The description adds the threading behavior, but does not disclose required scopes, rate limits, message-size limits, or what a successful send returns — and most scope info already lives in the schema parameter descriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and then the two most important parameter behaviors. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with no output schema, the description is thin but the schema carries full parameter documentation, so an agent can call it correctly. Missing behavioral context (rate limits, required scopes, response shape) keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 17 parameters are documented in the schema itself. The description's notes on `text` and `thread_ts` largely restate what the schema already provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Send') and resource ('a message to a Slack channel, private group, or DM'), making the action and destination unambiguous. An agent can distinguish this from siblings like slack_conversations_create without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence gives implied usage for `text` and `thread_ts`, but there is no explicit when-to-use vs. when-not, no mention of prerequisites (e.g. chat:write scope), and no reference to alternative messaging tools such as gmail_messages_send. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stripe_balance_transactions_listA
Read-onlyIdempotent

List every movement across the Stripe balance: charges, refunds, fees and payouts. For accounting, reporting_category on each result groups them better than type does.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoOnly return transactions of this type.
limitNoA limit on the number of objects to be returned, between 1 and 100. Defaults to 10.
payoutNoOnly return transactions paid out in this payout. Automatic payouts only.
sourceNoOnly return transactions for this object ID.
createdNoOnly return records created in this window.
currencyNoThree-letter lowercase ISO currency code.
endingBeforeNoA cursor for use in pagination: an object ID that defines your place in the list. Returns the page before the named object. Mutually exclusive with starting_after.
startingAfterNoA cursor for use in pagination: an object ID that defines your place in the list. To get the next page, pass the id of the last object in the current page.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint, so the safety profile is covered. The description adds the useful detail that each result carries a reporting_category field, but says nothing about pagination behavior, volume limits, or return shape beyond that. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, and the scope statement is front-loaded ahead of the accounting tip. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, all-optional list endpoint with no output schema, the description gives the resource scope and a key field-level hint. It stops short of mentioning pagination or the shape of a result, which an agent would have to infer from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema carries the parameter burden and the baseline is 3. The description earns a point above baseline by explaining the relationship between the `type` filter and the `reporting_category` output field, telling the agent which is more useful for accounting grouping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (balance transactions) and enumerates what those movements are: charges, refunds, fees and payouts. This scope is precise enough to separate it from stripe_balance_retrieve (totals) and from the narrower list endpoints like stripe_payouts_list or stripe_charges_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The accounting framing ('for accounting, reporting_category groups them better than type') implies the reconciliation use case, but there is no explicit when-to-use or when-not-to-use guidance and no routing to alternative list endpoints such as stripe_payouts_list or stripe_charges_list. Usage is inferable, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 28 tool updatesv0.1.0
    • First observedconnect
    • First observedconnection_status
    • First observedfirecrawl_extract
    • First observedfirecrawl_monitor_create
    • First observedfirecrawl_scrape
    • First observedgmail_drafts_create
    • First observedgmail_threads_get
    • First observedgmail_threads_list
    • First observedgsheets_spreadsheets_values_append
    • First observedgsheets_spreadsheets_values_get
    • First observedgsheets_spreadsheets_values_update
    • First observedlinear_issue_create
    • First observedshopify_customers_list
    • First observedshopify_draft_order_create
    • First observedshopify_inventory_adjust_quantities
    • First observedshopify_locations_list
    • First observedshopify_order_cancel
    • First observedshopify_order_fulfillment_orders
    • First observedshopify_order_get
    • First observedshopify_orders_list
    • First observedshopify_product_create
    • First observedshopify_product_get
    • First observedshopify_product_update
    • First observedshopify_product_variants_bulk_create
    • First observedshopify_products_list
    • First observedshopify_variant_inventory_level
    • First observedslack_chat_post_message
    • First observedstripe_balance_transactions_list

TDQS

B3.3/5.0

Scored across 28 tools

Disambiguation4/5

Tools are grouped by resource+action and mostly distinct: Shopify product/order/customer/inventory/location tools don't overlap, and read vs write inventory tools are clearly separated. Minor overlap between firecrawl_scrape and firecrawl_extract (both can extract structured JSON) and the broad multi-domain span makes cross-app selection occasionally ambiguous, but descriptions disambiguate well.

Naming Consistency4/5

A predictable domain_resource_action pattern runs throughout (shopify_products_list, shopify_product_get, gmail_threads_get, stripe_balance_transactions_list). The only real deviations are the un-prefixed connect and connection_status utilities, which break the otherwise consistent scheme.

Tool Count3/5

28 tools is heavy, and they are split across eight unrelated services (Shopify, Sheets, Gmail, Slack, Firecrawl, Linear, Stripe, connection management), leaving each app with only a few operations. It is defensible for a multi-app integration server but sits above the comfortable range for a Shopify-centric 'Store Ops' surface.

Completeness3/5

Core Shopify read/write flows (products, orders, inventory, locations) are covered, but there are notable gaps: no product delete, no order fulfillment creation even though shopify_order_fulfillment_orders' description references a nonexistent fulfillment_create, Gmail cannot send (draft only), and Slack only posts. These leave some obvious dead ends an agent would hit.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A production-grade MCP server and CLI tool that enables AI agents to manage Shopify stores through 49 built-in tools across products, orders, inventory, and analytics. It supports natural language workflows for tasks like inventory tracking, customer support, and sales reporting.
    153 npm
    18
    MIT
  • A
    license
    D
    quality
    D
    maintenance
    Enables natural-language control of e-commerce operations including product management, order processing, inventory tracking, customer service, content generation, and advertising analytics through 15 integrated MCP tools. Provides a local-first commerce automation solution with SQLite storage and extensible channel adapters for end-to-end online store workflows.
    15
    3
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to manage Shopify stores via natural language, with tools for products, orders, inventory, customers, and analytics.
    8 npm
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Runs 20 slash-command workflows across Google Calendar, Gmail, Linear, Slack, Granola, Google Docs, GitHub, Stripe, Notion, Google Forms, Sheets and Drive to produce morning briefs, meeting prep, action items assigned to owners and weekly updates. Reads proceed without asking, while anything that creates, sends, changes or deletes is shown for approval first.
    47
    Apache 2.0