Skip to main content
Glama
trhonpavel
by trhonpavel

medusa-mcp

πŸ‡¨πŸ‡Ώ Česky

An MCP server for the Medusa v2 Admin API. It gives Claude (or any MCP client) access to orders, customers, products and inventory, computes sales reports, and performs a small set of carefully scoped write actions.

It runs in two modes:

  • stdio – locally for Claude Desktop, Claude Code and other MCP clients

  • Streamable HTTP + OAuth 2.1 – as a remote custom connector you can use from Claude on the web, desktop and mobile

Tools

Tool

What it does

Kind

get_store_info

regions and currencies, sales channels, stock locations

read

list_orders

orders – full-text, date range, customer, order/payment/fulfillment status

read

get_order

full order detail by ID or order number (1042, #1042)

read

list_customers / get_customer

customers, order history, total spent

read

list_products / get_product

products, variants, prices, linked inventory items

read

list_inventory

stock per location, low_stock_threshold to find what's running out

read

sales_report

revenue, AOV, units, unique customers, day/week/month series, top products

report

create_fulfillment

fulfill an order (defaults: all remaining items, the only stock location)

write

create_shipment

mark as shipped with a tracking number

write

complete_order

mark an order as completed

write

cancel_order

cancel an order (destructiveHint)

write

update_product

title, description, status, handle, metadata

write

delete_product

delete a product and its variants, plus their unreserved inventory items; requires confirm_title (destructiveHint)

write

set_variant_price

set a variant's base price in one currency – all other prices, including ones with price rules, are preserved

write

set_stock_level

restock by SKU, absolute or relative (adjust_by: +10)

write

Amounts are in major currency units (Medusa v2 does not store minor units). Plain dates in filters (2026-09-01) are interpreted in REPORT_TIMEZONE (default UTC). With MEDUSA_READ_ONLY=true the write tools are not registered at all.

Related MCP server: managed-agent-control-mcp

1. Create a Medusa API key

In the Medusa Admin go to Settings β†’ Developer β†’ Secret API Keys β†’ Create. The key (sk_…) acts with the permissions of the user who created it, so consider a dedicated admin user that you can revoke independently.

2. Local use (stdio)

Claude Desktop – claude_desktop_config.json:

{
  "mcpServers": {
    "medusa": {
      "command": "npx",
      "args": ["-y", "medusa-mcp", "stdio"],
      "env": {
        "MEDUSA_BACKEND_URL": "https://api.example.com",
        "MEDUSA_API_KEY": "sk_...",
        "MEDUSA_READ_ONLY": "true"
      }
    }
  }
}

Claude Code:

claude mcp add medusa \
  -e MEDUSA_BACKEND_URL=https://api.example.com -e MEDUSA_API_KEY=sk_... \
  -- npx -y medusa-mcp stdio

3. Remote connector (HTTP + OAuth)

docker run -d --name medusa-mcp -p 127.0.0.1:3000:3000 -v medusa-mcp-data:/data \
  -e MEDUSA_BACKEND_URL=https://api.example.com \
  -e MEDUSA_API_KEY=sk_... \
  -e PUBLIC_URL=https://mcp.example.com \
  -e OWNER_PASSWORD="$(openssl rand -base64 24)" \
  ghcr.io/trhonpavel/medusa-mcp:latest

Or clone the repo, copy .env.example to .env and run docker compose up -d --build.

The server listens on 127.0.0.1:3000; expose it through a reverse proxy with TLS. Claude connects to remote connectors from Anthropic's servers, so the endpoint must be publicly reachable over HTTPS. Caddy example:

mcp.example.com {
    reverse_proxy 127.0.0.1:3000
}

Then add a custom connector in Claude with the URL https://mcp.example.com/mcp. Claude registers itself (Dynamic Client Registration), opens the consent page, you enter OWNER_PASSWORD and click Allow.

Claude Code can use the same OAuth flow, or a static token if you set MCP_STATIC_TOKEN:

claude mcp add --transport http medusa https://mcp.example.com/mcp \
  --header "Authorization: Bearer <MCP_STATIC_TOKEN>"

Configuration

Variable

Required

Default

Description

MEDUSA_BACKEND_URL

yes

Medusa backend URL

MEDUSA_API_KEY

yes

Secret API key (sk_…)

MEDUSA_READ_ONLY

false

Register read and report tools only

REPORT_TIMEZONE

UTC

IANA timezone for date filters and report buckets

MEDUSA_TIMEOUT_MS

20000

Timeout for Medusa requests

PUBLIC_URL

HTTP

Public HTTPS origin of this server (without /mcp)

OWNER_PASSWORD

HTTP

Password required on the consent page

MCP_STATIC_TOKEN

Optional static bearer token

ALLOWED_REDIRECT_HOSTS

claude.ai,claude.com,localhost,127.0.0.1

Hosts OAuth clients may use as redirect targets

TRUST_PROXY

1

Express trust proxy – number of proxies in front

PORT / HOST

3000 / 0.0.0.0

Listen address

DATA_DIR

./data

Where OAuth clients and token hashes are stored

ACCESS_TOKEN_TTL / REFRESH_TOKEN_TTL

3600 / 2592000

Token lifetimes in seconds

Endpoints

Path

Purpose

POST /mcp

MCP over Streamable HTTP (stateless), requires a bearer token

/.well-known/oauth-protected-resource/mcp

RFC 9728 protected resource metadata

/.well-known/oauth-authorization-server

RFC 8414 authorization server metadata

/register, /authorize, /token, /revoke

OAuth 2.1 (DCR, PKCE S256)

POST /oauth/login

consent form (rate limited: 10 attempts / 15 min / IP)

GET /healthz

health check

Security model

  • The Medusa API key never leaves the server. Clients get their own short-lived tokens (1 h access, 30-day refresh with rotation).

  • Only SHA-256 hashes of tokens are stored, in DATA_DIR/oauth-state.json (mode 600). Delete the file to sign out every client.

  • Dynamic Client Registration only accepts redirect URIs on ALLOWED_REDIRECT_HOSTS, so an arbitrary app cannot register its own callback and phish a token.

  • Authorization codes are single-use, expire after 5 minutes, and PKCE S256 is mandatory.

  • The consent page sends Content-Security-Policy: default-src 'none' and X-Frame-Options: DENY, and compares the password in constant time.

  • Write tools are not marked readOnlyHint and cancel_order / delete_product carry destructiveHint, so clients like Claude ask for approval before running them.

  • Set TRUST_PROXY to the number of reverse proxies in front of the server, otherwise rate limiting only sees the proxy's IP.

See SECURITY.md for reporting vulnerabilities.

Development

npm ci
npm test        # build + tests against a mock Medusa (tools and the full OAuth flow)
npm run smoke   # read-only check against a real Medusa – prints response shapes only, no data
npm run dev     # HTTP mode via tsx

npm run smoke needs MEDUSA_BACKEND_URL and MEDUSA_API_KEY. Its output contains only keys and types, so it is safe to paste into an issue.

License

MIT

Available Tools

17 tools
cancel_orderCancel orderA
DestructiveIdempotent

CANCELS the order. Irreversible – get explicit confirmation from the user before calling. The order must not have active fulfillments.

ParametersJSON Schema
NameRequiredDescriptionDefault
orderYesOrder ID or order number

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds real value beyond that: irreversibility and the active-fulfillment precondition. It does not describe the failure mode when the precondition is violated or what the caller gets back.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the operation, then the confirmation requirement, then the precondition. Every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter destructive mutation, the annotations cover the safety profile and the description covers irreversibility, confirmation, and preconditions. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single 'order' parameter is documented as 'Order ID or order number'. The description adds no syntax or format detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('CANCELS the order') with no ambiguity, and no sibling tool competes for the cancel operation, so the agent can route to it directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit call-time requirement ('get explicit confirmation from the user before calling') and a precondition ('must not have active fulfillments'), which is strong contextual guidance. It stops short of naming alternatives, but none exist among the siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

complete_orderComplete orderC
Idempotent

Marks the order as completed.

ParametersJSON Schema
NameRequiredDescriptionDefault
orderYesOrder ID or order number

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds nothing beyond restating the name: no state-transition rules (can a cancelled order be completed? does it require payment?), no side effects, no permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no wasted words. It is appropriately sized, though its brevity reflects under-specification rather than disciplined conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation with a single well-documented parameter and annotations covering idempotency and safety, the description is minimally adequate. It omits the state-transition semantics an agent would need to call it correctly in edge cases (already-completed or cancelled orders).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter ('order') exists and the schema documents it fully at 100% coverage ('Order ID or order number'). Per the rubric, full schema coverage yields a baseline of 3, and the description adds no extra parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ('Marks the order as completed'), so the operation is unambiguous. However, it does nothing to distinguish this from sibling mutations like cancel_order or create_fulfillment beyond the obvious wording of the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when this tool should be used versus alternatives such as cancel_order, nor any prerequisite (e.g., order must be paid/shipped). The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_fulfillmentCreate fulfillmentA

Creates a fulfillment for an order. Without 'items' it fulfills all remaining unfulfilled quantities. Without 'location_id' it uses the only stock location, if there is exactly one.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsNoSpecific line items; defaults to everything remaining
orderYesOrder ID or order number
location_idNoStock location to ship from
notify_customerNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a non-idempotent write (readOnlyHint=false, idempotentHint=false, non-destructive). The description adds genuinely useful behavioral context beyond them: the implicit bulk action of fulfilling all remaining quantities when items is omitted, and the single-location fallback. It does not disclose the customer-visible effect of notify_customer, which is the one notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place, with the core action front-loaded and the two conditional defaults following immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-idempotent mutation with no output schema, the description covers the essential default behaviors an agent must know before calling. Missing only the notify_customer side effect and the failure mode when no location_id is given and multiple locations exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% and the description adds cross-parameter semantics the schema cannot express: the default expansion of items to 'everything remaining' and the conditional resolution of location_id. Only notify_customer's customer-facing effect remains unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Creates a fulfillment for an order.' An agent can distinguish it from order-mutating siblings like complete_order and cancel_order, though it doesn't explicitly separate itself from the closest sibling create_shipment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives conditional default behavior ('Without items...', 'Without location_id...') which implies when the optional parameters can be omitted, but it never states when to prefer this tool over create_shipment or gives any prerequisites/exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_shipmentMark as shippedB

Marks a fulfillment as shipped and attaches a tracking number. Without 'fulfillment_id' it uses the only unshipped fulfillment.

ParametersJSON Schema
NameRequiredDescriptionDefault
orderYesOrder ID or order number
tracking_urlNo
fulfillment_idNo
notify_customerNo
tracking_numberNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true). The description adds the non-obvious fallback behavior when fulfillment_id is omitted, which is genuine extra context, but it says nothing about notify_customer's effect, validation failures, or what happens when multiple unshipped fulfillments exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the primary purpose is front-loaded and the fallback rule follows immediately. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with 20% schema coverage and no output schema, the description covers purpose and one default rule but omits the role of the other three optional parameters and any error/precondition behavior. It is adequate to identify the tool but thin for calling it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 20% and 5 parameters exist, so the description must compensate and largely does not. Only fulfillment_id gets any explanation; tracking_number, tracking_url, and notify_customer (default true) are left entirely undocumented, so the agent must guess whether they are required or what they control.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Marks a fulfillment as shipped') plus the side effect (attaches a tracking number), which is enough to distinguish it from siblings like create_fulfillment and complete_order. It stops short of explicitly naming those siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description adds one useful conditional rule ('Without fulfillment_id it uses the only unshipped fulfillment'), which is closer to behavioral context than usage guidance. It never says when to prefer this over create_fulfillment or complete_order, nor any prerequisite for shipping, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_productDelete productA
DestructiveIdempotent

DELETES the product with all its variants. Irreversible – get explicit confirmation from the user before calling. 'confirm_title' must match the product title exactly. By default the inventory items of its variants are deleted too (only those with nothing reserved).

ParametersJSON Schema
NameRequiredDescriptionDefault
product_idYes
confirm_titleYesThe exact product title, as a safeguard against deleting the wrong product
delete_inventory_itemsNo

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, but the description adds substantially: the irreversible nature, the cascade to all variants, and the non-obvious caveat that inventory items are only deleted when nothing is reserved. That is exactly the behavioral context an agent cannot get from the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the destructive action and the confirmation requirement; every clause carries information with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no output schema, the description covers confirmation, irreversibility, and cascade semantics. It omits what happens when confirm_title does not match (error vs. silent refusal) and whether the inventory caveat can be overridden, but the essentials are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, so the description must compensate. It does explain the confirm_title safeguard and the default behavior of delete_inventory_items, but product_id is undocumented in both places and the exactness requirement for confirm_title largely repeats the schema text. Adequate but with a real gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (DELETES) and resource (product with all its variants), and the scope of the cascade is included. This clearly separates it from update_product among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs the agent to obtain user confirmation before calling and explains the confirm_title safeguard condition, which is real usage guidance. It does not name update_product as the non-destructive alternative, so it stops short of full when-not-to-use routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_customerGet customerA
Read-only

Customer detail with addresses, groups and order history (count, total spent, recent orders).

ParametersJSON Schema
NameRequiredDescriptionDefault
customer_idYesCustomer ID (cus_…)

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds real value by disclosing what is returned (addresses, groups, counts, total spent, recent orders), which is needed since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence with no filler, and the resource is front-loaded. It is appropriately sized for a simple single-parameter read, though the parenthetical enumeration is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the return contents, and annotations cover the safety profile. For a one-parameter read tool it is nearly complete, missing only a note about error/not-found behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the single customer_id parameter (with the 'cus_…' format hint) is fully documented in the schema. The description adds no additional parameter semantics beyond that, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (get customer) and enumerates the detail returned: addresses, groups, order history. It implicitly distinguishes itself from the sibling list_customers by being a single-customer detail call, but never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'detail' vs the sibling 'list_customers' suggests this is for one customer rather than enumeration. No explicit when-to-use, when-not-to-use, or alternative is stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_orderGet orderA
Read-only

Full order detail – line items, addresses, payments, fulfillments and tracking numbers. Accepts an order ID (order_…) or the order number.

ParametersJSON Schema
NameRequiredDescriptionDefault
orderYesOrder ID (order_…) or order number, e.g. 1042

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real value by disclosing the breadth and nesting of the returned payload (line items, addresses, payments, fulfillments, tracking), which an agent cannot infer from a bare 'get' name. It stops short of describing auth requirements or behavior for missing orders.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the payload scope is front-loaded and the accepted identifier formats follow. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully compensates by naming the fields returned. For a one-parameter read tool this is close to complete, though it omits error/not-found behavior and whether the response is paginated or truncated for large orders.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is fully documented in the schema, including the same order_/order-number formats the description mentions. The description therefore adds no meaning beyond the structured field, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get order') and immediately enumerates what the response covers (line items, addresses, payments, fulfillments, tracking). This clearly separates it from sibling list_orders, which would return many orders rather than full detail for one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call it when you have an order identifier. However, it never states when to prefer it over list_orders or get_customer, nor any prerequisite such as needing a valid order ID. No exclusions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_productGet productB
Read-only

Product detail – variants, prices in all currencies, linked inventory items, categories, collection, tags.

ParametersJSON Schema
NameRequiredDescriptionDefault
product_idYesProduct ID (prod_…)

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is clear. The description adds useful context by listing the returned data fields (variants, prices, inventory, categories, collection, tags), which helps since no output schema exists. It does not cover auth, rate limits, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact phrase that front-loads the resource and efficiently enumerates the returned details. Every word earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read-only lookup with a fully documented parameter and clear annotations. The description compensates for the missing output schema by listing return fields, though it omits usage guidance relative to siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the parameter already documents the 'prod_…' ID format. The description adds no further meaning about the product_id parameter, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and return content ('Product detail – variants, prices in all currencies, linked inventory items, categories, collection, tags'), making it clear this is a single-product retrieval tool. However, it does not differentiate from siblings like list_products or update_product.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided. It does not mention alternatives such as using list_products to browse or update_product to modify, leaving the agent to infer the tool's appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_store_infoStore overviewA
Read-only

Returns regions (currencies), sales channels and stock locations. A good first call to learn location IDs and currencies.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety and scope are covered. The description adds that this is a discovery call and enumerates returned categories, but nothing about pagination, auth, or failure modes, which is minor for a static lookup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first lists the returned data, the second gives the actionable cue. Nothing is wasted and the important framing is not buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, read-only overview tool with no output schema, describing the returned categories is sufficient for an agent to call it correctly. Slightly more detail on the response shape would help, but the essentials are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to disambiguate; baseline for a parameterless tool is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (returns) plus the exact resources it exposes: regions/currencies, sales channels, and stock locations. It is distinguishable from list-oriented siblings, though it does not name them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says it is 'a good first call to learn location IDs and currencies,' which tells an agent when to prefer it β€” bootstrap/discovery. No when-not conditions or named alternatives, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_customersList customersC
Read-only

Searches customers by name, email or company.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoFull-text search – name, email, company
emailNo
limitNoNumber of records (max 200)
offsetNoRecords to skip (pagination)
has_accountNotrue = registered, false = guests

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds nothing beyond them: no mention of pagination limits, default result count, whether results are capped, or what happens with an empty query.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with no waste, but it is so brief that it omits information the agent would find useful. Conciseness is good; completeness suffers as a result.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, the description should at least note pagination and the guest/registered filter. It gestures at searchable fields but leaves the tool's result and filtering behavior underspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80%, so the schema already documents q, limit, offset, and has_account. The description's 'name, email or company' merely restates q's schema text and omits any mention of the has_account filter or pagination controls, adding no new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (searches) and resource (customers) plus the facets it queries, which distinguishes it from the singular get_customer sibling. The verb 'searches' is slightly at odds with the name 'list_customers', but intent is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this vs get_customer or list_orders, and no statement of prerequisites or default behavior. The agent must infer that this is the browse/search entry point from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_inventoryInventory levelsB
Read-only

Inventory items with stock per location (stocked, reserved, available). With low_stock_threshold returns only items at or below the threshold.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoFull-text search – title, SKU
skuNoExact SKU
limitNoNumber of records (max 200)
offsetNoRecords to skip (pagination)
location_idNoOnly this stock location
low_stock_thresholdNoOnly return items whose available quantity is less than or equal to this value

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds useful behavioral detail by enumerating the returned quantity fields (stocked, reserved, available) and clarifying that the threshold filters rather than sorts, but says nothing about pagination defaults or result limits beyond what the schema states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with what the tool returns before the optional filter modifier. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema and fully documented parameters, this covers what an agent needs: the payload shape and the one non-obvious filter. Missing only pagination/ordering guidance, which the schema partially implies via limit/offset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented, including that low_stock_threshold compares against available quantity β€” which the description simply restates. Baseline 3 for high coverage with no meaningful added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (inventory items with per-location stocked/reserved/available quantities) and a filtering behavior, so an agent knows exactly what data comes back. It stops short of an explicit 'list' verb and never distinguishes itself from list_products, a close sibling in the same catalog.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus list_products or get_product, and no prerequisites or exclusions. The only usage signal is implicit in the low_stock_threshold sentence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_ordersList ordersA
Read-only

Lists orders, newest first. Filters: full-text, date range (YYYY-MM-DD in the reporting timezone), customer, order/payment/fulfillment status.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNoFull-text search – order number, email, name…
limitNoNumber of records (max 200)
offsetNoRecords to skip (pagination)
statusNoOrder status
created_toNoTo date (inclusive), e.g. 2026-09-30
customer_idNo
created_fromNoFrom date, e.g. 2026-09-01
payment_statusNo
fulfillment_statusNoE.g. ['not_fulfilled'] = waiting to be fulfilled

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely new behavioral context: results are ordered newest-first, and date filter inputs are YYYY-MM-DD interpreted in the reporting timezone β€” neither of which the schema states (created_from/created_to say only 'From date, e.g. ...'). It stops short of describing result size or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The core behavior and sort order come first, then the filter list β€” exactly the order an agent needs them.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, zero-required read tool with no output schema, the description covers ordering, date format/timezone, and filter categories. It leaves default behavior when no filters are supplied and result-set/pagination implications of limit/offset to the schema, a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 78%, and the description fills the most important gaps: it supplies the timezone semantics for the date range parameters and confirms that 'customer' maps to the otherwise undocumented customer_id field. It also groups status vs payment_status vs fulfillment_status for the reader, though it adds no enum values beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists orders') plus the sort order ('newest first') and the full set of filter dimensions. The list-vs-get distinction against the sibling get_order is implied rather than named, and no sibling is explicitly referenced, so it falls short of the top band.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The enumeration of available filters implies the browsing/search use case, but there is no explicit when-to-use guidance and no mention of the closest alternative, get_order, for retrieving a single order. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_productsList productsA
Read-only

Lists products with their variants (SKUs). Filter by full-text, status, collection or category.

ParametersJSON Schema
NameRequiredDescriptionDefault
qNo
limitNoNumber of records (max 200)
offsetNoRecords to skip (pagination)
statusNo
category_idNo
collection_idNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that results include variants (SKUs), which is useful return-shape context, but says nothing about pagination behavior, default result counts, or ordering.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action front-loaded before the filter list. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, the definition covers scope and available filters adequately. It omits pagination/default-limit behavior that an agent might need, but the schema supplies limit and offset documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (limit and offset), yet the description maps meaningful filter semantics: full-text via q, and filtering by status, collection, and category. That covers four of the six params conceptually, compensating well for the sparse schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Lists products with their variants (SKUs)') and even clarifies return granularity (variants/SKUs). It does not explicitly name sibling tools like get_product, but the singular/plural split makes the distinction obvious to an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by naming the available filters, but never states when to prefer this over get_product or update_product, nor any prerequisites. Usage is inferable rather than guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sales_reportSales reportA
Read-only

Computes sales for a period: order count, revenue, average order value, units sold, unique customers, a time series (day/week/month) and top products. Amounts are per currency. Canceled and draft orders are excluded.

ParametersJSON Schema
NameRequiredDescriptionDefault
toYesTo date (inclusive), e.g. 2026-09-30
fromYesFrom date, e.g. 2026-09-01
top_nNoHow many top products to return
group_byNoday
timezoneNoIANA timezone used for date boundaries and bucketsUTC
only_paidNoOnly count paid orders (captured / partially_refunded / partially_captured)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=false, and the description adds real behavioral context: canceled and draft orders are excluded, amounts are per-currency (so no cross-currency rollup), and a time series is produced. It stops short of disclosing performance, limits, or empty-range behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the metric list and followed by the important scoping caveats. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned metrics and the exclusion rules that determine the numbers, which is the key information an agent needs. It omits how currency grouping interacts with the totals and what the top-products ordering is based on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents from/to/top_n/timezone/only_paid. The description only indirectly echoes parameters by mentioning day/week/month buckets and top products, and does not clarify only_paid or timezone semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (computes) and resource (sales report) and enumerates the exact metrics returned: order count, revenue, AOV, units, unique customers, time series, and top products. This makes it clearly distinguishable from the sibling list/get tools, which return raw entities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implicitly signals it is an aggregation tool rather than a record lister, and notes that canceled and draft orders are excluded. However, it never states when to prefer this over list_orders plus manual aggregation, nor any prerequisite or exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_stock_levelSet stock levelA

Sets the stocked quantity of an item at a location. Provide either an absolute 'stocked_quantity' or a relative 'adjust_by' (+/-). Identify the item by inventory_item_id or SKU. Without location_id the item's only location is used.

ParametersJSON Schema
NameRequiredDescriptionDefault
skuNo
adjust_byNoE.g. +10 when restocking, -2 when writing off
location_idNo
stocked_quantityNo
inventory_item_idNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the write/idempotency/destructiveness profile, so the description's added value is the mutually exclusive quantity modes and the default behavior when location_id is omitted ("the item's only location is used"). It does not disclose permission requirements or what happens if both quantity fields are supplied. Notably it implicitly clarifies why idempotentHint=false: adjust_by is relative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the core action, then input modes, identification, and the location default. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param mutation tool with no output schema and partial schema coverage, the description covers the operation, the input alternatives, and location resolution. It stops short of error behavior (e.g., both quantity params provided) and permission prerequisites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20% (just adjust_by), so the description must compensate, and it does: it defines stocked_quantity as absolute, adjust_by as relative with +/- semantics, names the two identifier options, and explains location_id's default resolution.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (sets the stocked quantity of an item at a location) and is clearly distinguishable from inventory read tools like list_inventory and product-mutation tools like update_product.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains how to supply the quantity (absolute stocked_quantity vs relative adjust_by) and how to identify the item (inventory_item_id or SKU), plus the location fallback. No guidance on when-not-to-use or how this relates to sibling tools, but the invocation context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_variant_priceSet variant priceA
Idempotent

Sets the base price (no price rules) of a variant in one currency. All other prices of the variant are preserved. The amount is in major currency units (e.g. 49.99 = 49.99 EUR).

ParametersJSON Schema
NameRequiredDescriptionDefault
amountYes
product_idYes
variant_idYes
currency_codeYesE.g. eur, usd

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotent=true, destructive=false, openWorld=true, so the safety profile is covered. The description adds real behavioral value beyond them: other prices are preserved (no collateral mutation) and the amount is in major units, preventing the classic 4999-vs-49.99 error.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the action, the preservation guarantee, and the unit convention. The critical unit warning is front-loaded after the action rather than buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema, the description covers the action, side effects, and the most error-prone parameter convention. It is slightly thin on the identifier parameters and gives no error/permission expectations, but is otherwise adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only currency_code documented), so the description must compensate, and it does for the highest-risk parameter by explaining amount units with a concrete example. It says nothing, however, about product_id or variant_id semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource ('Sets the base price ... of a variant') and narrows scope with '(no price rules)' and 'in one currency', so an agent knows exactly what is being modified. No sibling tool covers pricing, so this is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the scope note ('base price, no price rules'), which tells the agent this is not for rule-based pricing, but there is no explicit when-to-use or exclusion statement. With no pricing sibling to route between, the gap is modest.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_productUpdate productB
Idempotent

Updates basic product fields (title, description, status, handle, metadata). Send only the fields that should change.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNo
handleNo
statusNo
metadataNoMerged into existing metadata; an empty string deletes a key
subtitleNo
product_idYes
descriptionNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the mutation safety profile is covered. The description's 'send only fields that should change' adds useful partial-update semantics beyond the annotations, but it does not say whether omitted fields are preserved or how metadata merges (that detail lives only in the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the operation and field scope, then the usage rule. No filler, though the sentence is slightly compressed and could not be called richly structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no output schema and low parameter description coverage, the description leaves meaningful gaps: what happens to fields not sent, whether metadata keys are merged or replaced, and what the response returns. Annotations cover safety hints, but the partial-update contract is only loosely stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% across 7 parameters, so the description must carry more of the load. It lists title, description, status, handle, and metadata, but omits subtitle and product_id entirely and provides no enum or format detail for status beyond what the schema already shows. It adds partial value over the bare schema, not enough to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb+resource and enumerates the updatable field groups, so the agent knows this modifies basic product fields. It does not explicitly distinguish itself from siblings like delete_product or set_variant_price, but the scope is clear enough to route correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Send only the fields that should change' gives a partial-update usage rule, which is genuinely helpful. However, there is no guidance on when to prefer this over delete_product, set_variant_price, or set_stock_level, nor any prerequisite/auth context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 17 tool updatesv0.1.0
    • First observedcancel_order
    • First observedcomplete_order
    • First observedcreate_fulfillment
    • First observedcreate_shipment
    • First observeddelete_product
    • First observedget_customer
    • First observedget_order
    • First observedget_product
    • First observedget_store_info
    • First observedlist_customers
    • First observedlist_inventory
    • First observedlist_orders
    • First observedlist_products
    • First observedsales_report
    • First observedset_stock_level
    • First observedset_variant_price
    • First observedupdate_product

TDQS

A3.6/5.0

Scored across 17 tools

Disambiguation5/5

Each tool targets a distinct resource and action: list/get pairs for orders, customers, products and inventory; mutating tools like set_stock_level, complete_order, cancel_order, create_fulfillment and create_shipment are clearly separated. There is no meaningful overlap that would cause an agent to misselect between tools.

Naming Consistency4/5

The set uses a consistent snake_case verb_noun pattern for the vast majority of tools (list_orders, get_order, set_stock_level, complete_order, create_fulfillment, cancel_order, update_product, delete_product, set_variant_price). The only notable deviation is sales_report, which is a noun phrase rather than a verb_noun, but overall the convention is predictable.

Tool Count4/5

17 tools is slightly on the heavy side but appropriate for an e-commerce admin surface spanning orders, customers, products, inventory, fulfillments and reporting. Each tool appears to earn its place, though the set could be tightened marginally.

Completeness3/5

The surface covers reads and several lifecycle actions well but has notable gaps: there is no create_product despite update/delete, no create_customer or update/delete_customer, and no order creation. These are common admin operations that an agent may reasonably expect, causing dead ends in otherwise well-covered domains.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A minimal local MCP server that wraps any Claude Messages API-compatible upstream into a unified ask_model tool. It enables MCP clients to interact with these models through a standard tool interface using stdio transport.
    1
    8
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Start, observe, and interact with Claude Managed Agents from any MCP client β€” launch an agent, watch its events, reply, approve the tools it wants to run, and stop it. Runs over stdio, HTTP, or AWS Lambda with pluggable auth.
    17
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Turns any CLI tool or REST API into an MCP server for Claude, enabling Claude to use git, databases, or any API through natural language.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server that connects Claude to Shopify stores, enabling natural language queries and actions on products, orders, customers, inventory, and sales analytics. Includes a demo mode with bundled fixtures for trying tools without credentials.
    97 npm
    MIT