Skip to main content
Glama

commerce-mcp-sandbox

A zero-dependency mock commerce MCP server for building and testing shopping agents.

给购物类 AI agent 用的模拟商城后端:搜索、下单、模拟支付、物流自动推进、退款——全部内存模拟,没有任何真实商品与真实资金。让你在写购物 agent 时不用等真实电商接口,先把 agent 逻辑跑通。

English below | 中文说明见下半部分


Why

Every developer building a shopping agent hits the same wall: "I need a commerce backend to test against." Real APIs need credentials, contracts, and money. This sandbox gives you the full shopping lifecycle — search → order → pay → ship → deliver → refund — as an MCP server your agent can talk to, with deterministic in-memory state.

Related MCP server: Agentic Shopping MCP

Install & Run

Zero dependencies. Node ≥ 18.

npx commerce-mcp-sandbox          # stdio mode (for MCP hosts)
npx commerce-mcp-sandbox --http   # not supported yet; use node src/http.js 3300
# or clone & run:
node src/index.js                 # stdio
node src/http.js 3300             # HTTP mode: POST http://localhost:3300/mcp

MCP host configuration

{
  "mcpServers": {
    "commerce-sandbox": {
      "command": "npx",
      "args": ["commerce-mcp-sandbox"]
    }
  }
}

Streamable-HTTP hosts (Dify / Coze / web agents):

url: http://localhost:3300/mcp

Tools (7)

tool

说明

search_products

Search mock catalog (query/category/page). Prices in CNY cents.

get_product

Product detail with specs.

create_order

Create order. Requires receiver {name, phone, address}.

pay_order

Simulate payment. Shipping auto-advances in ~5s, delivery in ~15s — so agents can exercise the full lifecycle in seconds.

get_order

Status + event timeline (created/paid/shipped/delivered/refunded).

list_orders

List orders, filter by user_ref.

request_refund

Simulate refund (no real money).

Example flow

agent: search_products {query: "umbrella"}     → 1 item, ¥39.00
agent: create_order {items:[{sku:"UMB-001",qty:1}], receiver:{...}}
user:  pay_order {order_no:"MOCK100001"}       → paid
(wait ~5s) get_order → shipped (waybill SF...)
(wait ~15s) get_order → completed
agent: request_refund {order_no:"MOCK100001"}  → refunded (simulated)

Honest limits

  • In-memory only — restart wipes state. This is a test double, not a store.

  • No real payments — pay_order flips state; nothing is charged.

  • Single-instance — no persistence layer by design.

Building a real shopping agent?

This sandbox is built by 源头淘 (Yuantoutao) — an open product shelf + fulfillment network for AI agents in China (MCP + REST, WeChat Pay checkout, merchant drop-shipping via ERP, automated aftersales). When your agent is ready for real goods: same tool surface, swap the endpoint, done.

Docs: https://mcp.yuantoutao.com/docs/ · llms.txt: https://mcp.yuantoutao.com/llms.txt

License

MIT

Available Tools

7 tools
create_orderB

Create an order. Requires receiver {name, phone, address}. Idempotency is the caller's job in this sandbox (orders are in-memory).

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYes
receiverYes
user_refNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two genuinely useful traits: orders are in-memory (sandbox, non-durable) and idempotency is the caller's responsibility. It still omits auth/permission requirements, whether the created order must be paid via pay_order, and what identifier is returned for later use.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and immediately followed by the required shape and the key sandbox caveat. No filler, no repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and nested object parameters, the description is thin. It never explains the items payload or user_ref, nor what the caller gets back (an order id is presumably needed to call pay_order), leaving real gaps an agent must guess at.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only restates the receiver fields, which the schema already marks required, and says nothing about the required 'items' array (sku/qty shape, whether qty defaults) or the optional 'user_ref' meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create an order'), which is unambiguous and distinguishable from read-side siblings like get_order and list_orders. It does not, however, explicitly contrast itself with pay_order, which is the nearest write-side sibling and the natural next step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use or when-not-to-use guidance and names no alternative. An agent must infer from the sibling list that create_order precedes pay_order and that list_orders/get_order are read paths; nothing in the text routes it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_orderC

Order status + event timeline (created/paid/shipped/delivered/refunded).

ParametersJSON Schema
NameRequiredDescriptionDefault
order_noYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the return contents (status values and event timeline), which is genuinely useful, but says nothing about read-only behavior, auth requirements, or what happens for a nonexistent or unauthorized order.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the parenthetical enumeration of statuses is dense but informative rather than redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description rightly sketches the return shape (status + timeline), which is the minimum an agent needs. It nonetheless omits error/empty-case behavior and permission context for a lookup tool with zero annotation coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter. The description never explains order_no — its format, source, or whether it is the internal ID versus a human-facing number — so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (order) and the specific payload (status plus event timeline with enumerated lifecycle stages), which lets an agent distinguish it from siblings like list_orders or pay_order. It is a noun phrase rather than a verb+resource, but the retrieval intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this versus list_orders, get_product, or other siblings, and no prerequisites (e.g. 'use after create_order'). Usage is only inferable from the tool name and the single-order parameter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_productB

Product detail by sku, incl. specs. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
skuYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden, and it does disclose the single most important trait for an agent ('Read-only'), which is genuinely useful here. Beyond that it says nothing about behavior for an unknown SKU, auth requirements, or rate limits, so the disclosure is thin relative to the unannotated baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Extremely compact and front-loaded: the resource and key come first, the scope hint second, the safety trait last. The telegraphic style costs nothing in clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter lookup with no output schema, the description is minimally adequate, and 'incl. specs' hints at the return shape. It falls short on usage routing and on what happens when the SKU is not found.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It mentions 'by sku' but adds no meaning beyond restating the parameter name already present in the schema, leaving format/validation expectations (case sensitivity, exactness, partial matches) unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (fetch product detail), the lookup key (sku), and even the scope of the payload ('incl. specs'). It does not, however, differentiate itself from the sibling search_products, which an agent must disambiguate on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no routing to alternatives. The presence of search_products as a sibling makes the omission consequential: nothing tells the agent to use this for exact-SKU lookup versus search for discovery.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_ordersC

List orders, optionally filtered by user_ref.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_refNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read operation via "List" but never states read-only semantics, auth/permission requirements, pagination behavior, or what the return contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the purpose front-loaded. It is efficient, though its brevity leaves behavioral and usage gaps rather than earning a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter list tool with no output schema, the description covers purpose and the one filter, but omits read-only confirmation, pagination, and ordering behavior that an agent would need to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is one parameter. The description names user_ref as an optional filter, confirming its role and optionality, but adds no format, syntax, or expected-value detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("List") and resource ("orders"), so the agent knows this retrieves order collections. It does not explicitly distinguish itself from siblings like get_order or search_products, so it stops short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no when-to-use guidance, no conditions selecting it over get_order or create_order, and no prerequisites. It only states the optional filter, which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pay_orderA

Simulate a successful payment for an order. Shipping auto-advances (~5s) and delivery (~15s) so agents can exercise full flows.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_noYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does disclose genuinely useful side effects: shipping auto-advances (~5s) and delivery (~15s). It still omits prerequisites (must the order be unpaid?), idempotency/error behavior, and confirmation that this is a mock that never touches real money, so it is strong but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action stated first and the side-effect timings immediately after. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema and no annotations, the description covers what happens and roughly how long it takes, which is the most important part. It is still thin on preconditions, failure modes, and what state the order is left in, which an agent needs before invoking a state-changing simulation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'order_no' has no description in the schema. The description only implies 'an order' generically and adds no format, source, or example for the identifier, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Simulate a successful payment for an order') and the word 'simulate' distinguishes it from real-payment siblings like create_order or request_refund. However, it never explicitly contrasts itself with those siblings, so the differentiation is inferred rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'so agents can exercise full flows' implies the tool is meant for end-to-end test scenarios, which is a usable hint about context. But there is no explicit when-to-use/when-not guidance, no prerequisites, and no mention of what to call instead for a real (non-simulated) payment flow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_refundB

Simulate a refund for a paid/shipped/completed order (no real money).

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNo
order_noYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that this is a simulation with no real monetary effect, which is important context. However, it says nothing about effects on order state, reversibility, permissions, or idempotency, leaving significant gaps for a mutation-style tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the core action and the key non-destructive caveat front-loaded. Every element earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no annotations and no output schema, the description conveys purpose and a usage precondition but leaves parameter meaning and behavioral consequences uncovered. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for both parameters, and the description adds no meaning for 'order_no' or 'reason' beyond implying an order reference. With two undocumented parameters it should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Simulate a refund ... for an order'), which is clearly distinct from siblings like pay_order or get_order. The parenthetical '(no real money)' reinforces the simulate semantics. No explicit naming of siblings, but the action is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for a paid/shipped/completed order' implies an eligibility precondition that guides usage, but there is no explicit when-not-to-use or comparison to alternatives among the siblings. The condition is implied rather than framed as guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_productsC

Search the mock catalog. Prices in CNY cents (int). Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
queryNo
categoryNo
page_sizeNo

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose the important facts that the tool is read-only and that prices are in CNY cents (int), which aids interpretation of results, but it says nothing about pagination behavior, default page size, or the shape of returned results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three terse sentences with no filler, and the core action is front-loaded. Nothing redundant or bloated for a definition of this scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A four-parameter search tool with no annotations, no output schema, and 0% schema coverage needs the description to explain filtering and pagination, yet it covers none of that. The result is too thin for the tool's complexity, leaving the agent to guess at page/page_size semantics and result structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and none of the four parameters (page, query, category, page_size) are described. The description's only data note, CNY cents pricing, concerns return values rather than inputs, so it adds no parameter meaning whatsoever.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ("Search") and resource ("mock catalog"), which an agent can distinguish from get_product's single-item retrieval. It does not explicitly name or contrast with any sibling, so it falls short of the 5-level sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no mention of alternatives such as get_product, and no prerequisites. Usage is only implied by the verb "Search".

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.0
    • First observedcreate_order
    • First observedget_order
    • First observedget_product
    • First observedlist_orders
    • First observedpay_order
    • First observedrequest_refund
    • First observedsearch_products

TDQS

A3.5/5.0

Scored across 7 tools

Disambiguation5/5

Each tool has a distinct resource-action pair: product search/detail, order create/pay/get/list, and refund. No overlaps or ambiguous boundaries; an agent can easily select the right tool.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (search_products, get_product, create_order, pay_order, get_order, list_orders, request_refund). The convention is uniform and readable.

Tool Count5/5

Seven tools is well-scoped for a commerce sandbox, covering the essential product and order operations without excess. Each tool earns its place.

Completeness4/5

The surface covers the core commerce lifecycle: product search/detail, order creation, payment, status, listing, and refund. Minor gaps exist (e.g., no order cancellation or update), but agents can work around these for typical sandbox flows.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to interact with Dynamics 365 Commerce systems through 125+ tools covering customer management, sales orders, cart operations, product searches, inventory tracking, and store operations. Provides comprehensive mock data for development and testing purposes.
    3
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to perform e-commerce operations including product search, budget-constrained shopping recommendations, and sustainability analysis. Includes a secure HTTP bridge with OAuth integration and observability features for production deployment.
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    A local-first Model Context Protocol server that provides LLM agents with structured access to e-commerce data (products, inventory, orders, sales analytics) using a realistic mock dataset, zero configuration, and a swappable DataProvider interface for live APIs.
    MIT