commerce-mcp-sandbox
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@commerce-mcp-sandboxsearch products for 'umbrella' and place an order for the cheapest one"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
commerce-mcp-sandbox
A zero-dependency mock commerce MCP server for building and testing shopping agents.
给购物类 AI agent 用的模拟商城后端:搜索、下单、模拟支付、物流自动推进、退款——全部内存模拟,没有任何真实商品与真实资金。让你在写购物 agent 时不用等真实电商接口,先把 agent 逻辑跑通。
English below | 中文说明见下半部分
Why
Every developer building a shopping agent hits the same wall: "I need a commerce backend to test against." Real APIs need credentials, contracts, and money. This sandbox gives you the full shopping lifecycle — search → order → pay → ship → deliver → refund — as an MCP server your agent can talk to, with deterministic in-memory state.
Related MCP server: Agentic Shopping MCP
Install & Run
Zero dependencies. Node ≥ 18.
npx commerce-mcp-sandbox # stdio mode (for MCP hosts)
npx commerce-mcp-sandbox --http # not supported yet; use node src/http.js 3300
# or clone & run:
node src/index.js # stdio
node src/http.js 3300 # HTTP mode: POST http://localhost:3300/mcpMCP host configuration
{
"mcpServers": {
"commerce-sandbox": {
"command": "npx",
"args": ["commerce-mcp-sandbox"]
}
}
}Streamable-HTTP hosts (Dify / Coze / web agents):
url: http://localhost:3300/mcpTools (7)
tool | 说明 |
| Search mock catalog (query/category/page). Prices in CNY cents. |
| Product detail with specs. |
| Create order. Requires |
| Simulate payment. Shipping auto-advances in ~5s, delivery in ~15s — so agents can exercise the full lifecycle in seconds. |
| Status + event timeline ( |
| List orders, filter by |
| Simulate refund (no real money). |
Example flow
agent: search_products {query: "umbrella"} → 1 item, ¥39.00
agent: create_order {items:[{sku:"UMB-001",qty:1}], receiver:{...}}
user: pay_order {order_no:"MOCK100001"} → paid
(wait ~5s) get_order → shipped (waybill SF...)
(wait ~15s) get_order → completed
agent: request_refund {order_no:"MOCK100001"} → refunded (simulated)Honest limits
In-memory only — restart wipes state. This is a test double, not a store.
No real payments —
pay_orderflips state; nothing is charged.Single-instance — no persistence layer by design.
Building a real shopping agent?
This sandbox is built by 源头淘 (Yuantoutao) — an open product shelf + fulfillment network for AI agents in China (MCP + REST, WeChat Pay checkout, merchant drop-shipping via ERP, automated aftersales). When your agent is ready for real goods: same tool surface, swap the endpoint, done.
Docs: https://mcp.yuantoutao.com/docs/ · llms.txt: https://mcp.yuantoutao.com/llms.txt
License
MIT
Available Tools
7 toolscreate_orderB
Create an order. Requires receiver {name, phone, address}. Idempotency is the caller's job in this sandbox (orders are in-memory).
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| receiver | Yes | ||
| user_ref | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two genuinely useful traits: orders are in-memory (sandbox, non-durable) and idempotency is the caller's responsibility. It still omits auth/permission requirements, whether the created order must be paid via pay_order, and what identifier is returned for later use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and immediately followed by the required shape and the key sandbox caveat. No filler, no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and nested object parameters, the description is thin. It never explains the items payload or user_ref, nor what the caller gets back (an order id is presumably needed to call pay_order), leaving real gaps an agent must guess at.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only restates the receiver fields, which the schema already marks required, and says nothing about the required 'items' array (sku/qty shape, whether qty defaults) or the optional 'user_ref' meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create an order'), which is unambiguous and distinguishable from read-side siblings like get_order and list_orders. It does not, however, explicitly contrast itself with pay_order, which is the nearest write-side sibling and the natural next step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use or when-not-to-use guidance and names no alternative. An agent must infer from the sibling list that create_order precedes pay_order and that list_orders/get_order are read paths; nothing in the text routes it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_orderC
Order status + event timeline (created/paid/shipped/delivered/refunded).
| Name | Required | Description | Default |
|---|---|---|---|
| order_no | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the return contents (status values and event timeline), which is genuinely useful, but says nothing about read-only behavior, auth requirements, or what happens for a nonexistent or unauthorized order.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the parenthetical enumeration of statuses is dense but informative rather than redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description rightly sketches the return shape (status + timeline), which is the minimum an agent needs. It nonetheless omits error/empty-case behavior and permission context for a lookup tool with zero annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one required parameter. The description never explains order_no — its format, source, or whether it is the internal ID versus a human-facing number — so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (order) and the specific payload (status plus event timeline with enumerated lifecycle stages), which lets an agent distinguish it from siblings like list_orders or pay_order. It is a noun phrase rather than a verb+resource, but the retrieval intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to call this versus list_orders, get_product, or other siblings, and no prerequisites (e.g. 'use after create_order'). Usage is only inferable from the tool name and the single-order parameter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_productB
Product detail by sku, incl. specs. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden, and it does disclose the single most important trait for an agent ('Read-only'), which is genuinely useful here. Beyond that it says nothing about behavior for an unknown SKU, auth requirements, or rate limits, so the disclosure is thin relative to the unannotated baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely compact and front-loaded: the resource and key come first, the scope hint second, the safety trait last. The telegraphic style costs nothing in clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup with no output schema, the description is minimally adequate, and 'incl. specs' hints at the return shape. It falls short on usage routing and on what happens when the SKU is not found.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'by sku' but adds no meaning beyond restating the parameter name already present in the schema, leaving format/validation expectations (case sensitivity, exactness, partial matches) unstated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (fetch product detail), the lookup key (sku), and even the scope of the payload ('incl. specs'). It does not, however, differentiate itself from the sibling search_products, which an agent must disambiguate on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no routing to alternatives. The presence of search_products as a sibling makes the omission consequential: nothing tells the agent to use this for exact-SKU lookup versus search for discovery.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_ordersC
List orders, optionally filtered by user_ref.
| Name | Required | Description | Default |
|---|---|---|---|
| user_ref | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read operation via "List" but never states read-only semantics, auth/permission requirements, pagination behavior, or what the return contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler and the purpose front-loaded. It is efficient, though its brevity leaves behavioral and usage gaps rather than earning a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter list tool with no output schema, the description covers purpose and the one filter, but omits read-only confirmation, pagination, and ordering behavior that an agent would need to call it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there is one parameter. The description names user_ref as an optional filter, confirming its role and optionality, but adds no format, syntax, or expected-value detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("List") and resource ("orders"), so the agent knows this retrieves order collections. It does not explicitly distinguish itself from siblings like get_order or search_products, so it stops short of the top tier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no conditions selecting it over get_order or create_order, and no prerequisites. It only states the optional filter, which is not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pay_orderA
Simulate a successful payment for an order. Shipping auto-advances (~5s) and delivery (~15s) so agents can exercise full flows.
| Name | Required | Description | Default |
|---|---|---|---|
| order_no | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose genuinely useful side effects: shipping auto-advances (~5s) and delivery (~15s). It still omits prerequisites (must the order be unpaid?), idempotency/error behavior, and confirmation that this is a mock that never touches real money, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action stated first and the side-effect timings immediately after. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema and no annotations, the description covers what happens and roughly how long it takes, which is the most important part. It is still thin on preconditions, failure modes, and what state the order is left in, which an agent needs before invoking a state-changing simulation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter 'order_no' has no description in the schema. The description only implies 'an order' generically and adds no format, source, or example for the identifier, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Simulate a successful payment for an order') and the word 'simulate' distinguishes it from real-payment siblings like create_order or request_refund. However, it never explicitly contrasts itself with those siblings, so the differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'so agents can exercise full flows' implies the tool is meant for end-to-end test scenarios, which is a usable hint about context. But there is no explicit when-to-use/when-not guidance, no prerequisites, and no mention of what to call instead for a real (non-simulated) payment flow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_refundB
Simulate a refund for a paid/shipped/completed order (no real money).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| order_no | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that this is a simulation with no real monetary effect, which is important context. However, it says nothing about effects on order state, reversibility, permissions, or idempotency, leaving significant gaps for a mutation-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core action and the key non-destructive caveat front-loaded. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the description conveys purpose and a usage precondition but leaves parameter meaning and behavioral consequences uncovered. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, and the description adds no meaning for 'order_no' or 'reason' beyond implying an order reference. With two undocumented parameters it should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Simulate a refund ... for an order'), which is clearly distinct from siblings like pay_order or get_order. The parenthetical '(no real money)' reinforces the simulate semantics. No explicit naming of siblings, but the action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for a paid/shipped/completed order' implies an eligibility precondition that guides usage, but there is no explicit when-not-to-use or comparison to alternatives among the siblings. The condition is implied rather than framed as guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_productsC
Search the mock catalog. Prices in CNY cents (int). Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| query | No | ||
| category | No | ||
| page_size | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose the important facts that the tool is read-only and that prices are in CNY cents (int), which aids interpretation of results, but it says nothing about pagination behavior, default page size, or the shape of returned results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse sentences with no filler, and the core action is front-loaded. Nothing redundant or bloated for a definition of this scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A four-parameter search tool with no annotations, no output schema, and 0% schema coverage needs the description to explain filtering and pagination, yet it covers none of that. The result is too thin for the tool's complexity, leaving the agent to guess at page/page_size semantics and result structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and none of the four parameters (page, query, category, page_size) are described. The description's only data note, CNY cents pricing, concerns return values rather than inputs, so it adds no parameter meaning whatsoever.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ("Search") and resource ("mock catalog"), which an agent can distinguish from get_product's single-item retrieval. It does not explicitly name or contrast with any sibling, so it falls short of the 5-level sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no mention of alternatives such as get_product, and no prerequisites. Usage is only implied by the verb "Search".
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
create_order - First observed
get_order - First observed
get_product - First observed
list_orders - First observed
pay_order - First observed
request_refund - First observed
search_products
TDQS
Scored across 7 tools
Each tool has a distinct resource-action pair: product search/detail, order create/pay/get/list, and refund. No overlaps or ambiguous boundaries; an agent can easily select the right tool.
All tools follow a consistent snake_case verb_noun pattern (search_products, get_product, create_order, pay_order, get_order, list_orders, request_refund). The convention is uniform and readable.
Seven tools is well-scoped for a commerce sandbox, covering the essential product and order operations without excess. Each tool earns its place.
The surface covers the core commerce lifecycle: product search/detail, order creation, payment, status, listing, and refund. Minor gaps exist (e.g., no order cancellation or update), but agents can work around these for typical sandbox flows.
Maintenance
Related MCP Connectors
Hosted MCP for e-commerce: live product catalog, stock, and pricing for AI agents.
Agentic commerce gateway: discovery, search, checkout across Shopify/Woo/Odoo/PrestaShop.
AI-agent product catalog: search, lookup & purchase routing over verified merchant data.
Product search for AI agents: Amazon + Shopify, cart-to-checkout buy path. Pay-per-call, no API key.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to interact with Dynamics 365 Commerce systems through 125+ tools covering customer management, sales orders, cart operations, product searches, inventory tracking, and store operations. Provides comprehensive mock data for development and testing purposes.3-
- FlicenseNot gradedqualityDmaintenanceEnables AI agents to perform e-commerce operations including product search, budget-constrained shopping recommendations, and sustainability analysis. Includes a secure HTTP bridge with OAuth integration and observability features for production deployment.-
- AlicenseAqualityDmaintenanceProvides seamless access to the Fake Store API for AI assistants with 18 CRUD tools for managing e-commerce data including products, carts, and users. Perfect for e-commerce demos, testing, and learning MCP development with zero configuration required.1810 npmMIT
- AlicenseNot gradedqualityCmaintenanceA local-first Model Context Protocol server that provides LLM agents with structured access to e-commerce data (products, inventory, orders, sales analytics) using a realistic mock dataset, zero configuration, and a swappable DataProvider interface for live APIs.MIT