Zoho Inventory Connector
Connects to Zoho Inventory (read-only) to expose a merchant's stock, sales orders and delivery evidence. Provides availability checks for up to 25 SKUs (optionally per warehouse, with alternative-location stock), order and catalogue listing/search/get primitives, and fulfilment evidence lookup returning carrier, AWB, ship/delivery dates and a quotable verdict for dispute handling, with OAuth 2.0 auth, rate limiting, PII masking and audit logging.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Zoho Inventory ConnectorA shopper abandoned ST-TEE-OLV-M and ST-TEE-BLK-M. What can we still pitch?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Zoho Inventory connector for Razorpay Agent Studio
A private, read-only connector that gives Agent Studio agents a merchant's stock, orders and proof of delivery from Zoho Inventory. It is an MCP server with OAuth 2.0, list/get/search primitives, three-layer rate limiting and a published tool spec. It runs fully offline in demo mode, so you can try it in two minutes without a Zoho account.
The merchant problem
Saffron Threads is a D2C apparel and home brand. It takes payments on Razorpay and runs stock and fulfilment in Zoho Inventory, with warehouses in Bengaluru and Delhi NCR. It has switched on Agent Studio.
The cart recovery agent keeps calling shoppers about an olive tee that has been out of stock for a week.
The dispute responder can't answer a "never received" chargeback. The carrier, AWB and delivery date live in Zoho, and someone on the ops team copies them out by hand.
The ask was "let the agent read inventory and orders." The real need is "let agents check fulfilment facts before they talk to a customer or answer a dispute."
So on top of the list/get/search primitives, the connector has two tools built for those jobs:
check_availability: up to 25 SKUs in one call, optionally for one warehouse. Answers sellable: yes/no with a reason.get_fulfilment_evidence: looks up an order by its Razorpay order ID. Returns carrier, AWB, ship and delivery dates, a verdict, a sentence the agent can quote, and the gaps that weaken the evidence.
Related MCP server: Korral StoreLink MCP
Try it in 2 minutes (no Zoho account)
Needs Python 3.11+ and uv.
git clone <this repo> && cd zoho-inventory-connector
uv sync
uv run python scripts/demo.py # narrated walkthrough through a real MCP client session
uv run pytest -q # 66 tests, ~10 s, no networkscripts/demo.py connects through the OAuth code path, then runs the cart and dispute scenarios. It also shows PII masking, a prompt-injection record, recovery from Zoho 429s, and the structured error an agent gets when the budget runs out.
Use it from an MCP host (Claude Desktop, Claude Code, MCP Inspector) with fictional demo data:
uv run zoho-connector serve --demo # stdio
npx -y @modelcontextprotocol/inspector uv run zoho-connector serve --demo # click through every toolClaude Code picks up the bundled .mcp.json automatically. For Claude Desktop, add this to claude_desktop_config.json:
{
"mcpServers": {
"zoho-inventory": {
"command": "uv",
"args": ["--directory", "/absolute/path/to/zoho-inventory-connector", "run", "zoho-connector", "serve", "--demo"]
}
}
}Then ask, for example: "A shopper abandoned ST-TEE-OLV-M and ST-TEE-BLK-M. What can we still pitch?" or "Customer disputes Razorpay order order_DemoQ7Hk2pXa01 as not received. What's our evidence?"
Connect a real Zoho Inventory org
Everything below is free: Zoho Inventory's free plan includes API access at 1,000 calls/day.
Org. Sign up at zoho.com/in/inventory (free plan).
OAuth client. At api-console.zoho.in, choose Add Client → Server-based Applications and set:
Homepage URL:
http://localhostAuthorized Redirect URI:
http://localhost:8765/callback
Configure.
cp .env.example .envand setZOHO_CLIENT_ID,ZOHO_CLIENT_SECRETandZOHO_DC(infor zoho.in accounts).Connect.
uv run zoho-connector connect # opens Zoho's consent screen (5 read-only scopes) → pick org → tokens go to the OS keyring uv run zoho-connector status # verifies with one API call, shows budgetOn a server without a browser, create a Self Client in the API console, put its client ID/secret in
.env, generate a code with the same read scopes, and runzoho-connector connect --code 1000.xxxx.Demo data (optional). Seed the same fictional records into your org with
scripts/seed_zoho.py. It uses a separate, short-lived write grant and revokes it at the end, because the connector itself never gets write access.uv run python scripts/seed_zoho.py --dry-run # see the ~55 requests first uv run python scripts/seed_zoho.py --code 1000.xxxx --client-id <self client id> --client-secret <self client secret>Serve.
uv run zoho-connector serve, or use the Claude Desktop config above without--demo. Add--transport http --port 8000to serve over streamable HTTP.Disconnect.
uv run zoho-connector disconnectrevokes the grant at Zoho and deletes the local tokens.
Command | What it does |
| OAuth consent, bind to one organization |
| Connection, scopes, token expiry, per-minute and daily budget |
| Run the MCP server |
| Revoke and forget |
| Mock Zoho on :8900, to try the browser OAuth flow without an account |
| Regenerate |
--connection <name> (or ZOHO_CONNECTION) keeps one named connection per merchant org.
Tools
The full JSON Schemas, including output schemas, are in docs/tools.json. A test keeps that file in sync with the code.
Tool | Use it when | Zoho calls |
| Before pitching or promising a product. ≤25 SKUs, optional warehouse | 1 per SKU, or 1 catalogue page for ≥6 SKUs |
| Disputes, "where is my order". By Razorpay ref, SO number or id | 2 + 1 per package |
| Catalogue, low stock, per-warehouse stock | 1–2 |
| Orders by status, date, ref, customer | 1–2 |
| Which org, scopes, budget left | 0 |
All tools are annotated readOnlyHint: true, destructiveHint: false. Pagination uses opaque next_cursor tokens bound to the original query. Errors are JSON with a stable code, an optional retry_after_s and a hint telling the agent what to do next.
docs/CAPABILITIES.md is the one-page summary of what an agent can and cannot do.
How it handles the hard parts
OAuth. Authorization-code flow with offline access, a loopback redirect, a
statecheck, and validation of Zoho's multi-DCaccounts-server. Token refresh is single-flight, and a 401 triggers one refresh and retry. Tokens are stored in the OS keyring and revoked on disconnect.Rate limits. Zoho enforces 100 requests/min, 5–10 concurrent and 1,000–10,000/day per org. The connector:
keeps a sliding 90/min window and a concurrency semaphore;
persists a daily quota counter;
honours
Retry-Afterand otherwise backs off with jitter, within a bounded total wait;fails fast with
RATE_LIMITEDorDAILY_QUOTA_EXHAUSTEDinstead of hanging an agent;caches reads (items 60s, orders 30s) to save quota.
Safety.
Read-only, enforced by OAuth scope, by the tool surface and by MCP annotations.
Customer email, phone and street are masked unless the operator sets
ZOHO_PII_MODE=full.Merchant free text is returned as
UntrustedText.Every call writes a JSON audit line with arguments (names redacted), status, latency, Zoho calls spent and quota left.
The reasoning behind these choices is in docs/DESIGN.md. How to tell whether the connector is worth deploying is in docs/IMPACT.md.
Tests and evals
uv run pytest -q # 66 tests
uv run python evals/run_eval.py # LLM eval; needs EVAL_API_KEY (free Groq key works; Anthropic/Ollama via env)The tests cover:
the full browser OAuth flow on real sockets;
statemismatch and untrustedaccounts-server;expired-token refresh, including concurrent callers sharing one refresh;
revoked grants;
429 with
Retry-After, the 1070 concurrency error and 5xx retries;a 150-call burst throttled client-side with zero upstream 429s;
the daily quota persisting across restarts and rolling over at IST midnight;
every fulfilment verdict;
cursor misuse, PII masking and the prompt-injection record;
the MCP schema, plus a real stdio subprocess launched the way Claude Desktop launches it.
Eval result (gpt-oss-120b on Groq, demo data): 12/12 scenarios, with tool choice 12/12 and grounded answers 12/12. Groq rejected one malformed tool call (tool_use_failed); the harness resampled it once and counts that in the report. The scenarios cover cart recovery, disputes, order tracking, inventory, prompt injection and write refusal.
The first run scored 10/12, and the misses were useful:
A real product gap. For "is the mug available in Bengaluru?" the tool returned only "out of stock", so the model couldn't offer Delhi.
check_availabilitynow returnsother_locations, and the model answers "not from Bengaluru; Delhi NCR WH has 26".Two harness bugs. The model wrote
SO‑00104with a non-breaking hyphen andcan’twith a curly apostrophe, so the scorer now normalises punctuation. A provider error used to crash the whole run.
On the injection record, the model described the candle and noted that "the description field contains internal notes; no action is required". It listed no customer data and attempted no refund.
Re-run with make eval (free Groq key) after changing any tool description. A scripted oracle in the tests shows every scenario is answerable from the data.
Assumptions
The Razorpay order or payment ID is stored in the Zoho sales order's
reference_number. This is common for storefront syncs. If a merchant uses a custom field, that becomes a per-connection setting (see DESIGN "next steps").One connection = one Zoho organization. Multi-org merchants create one connection each.
Delivery facts are what the merchant recorded in Zoho. The connector does not query courier APIs.
The daily quota resets at midnight IST. Zoho resets per org day, and
ZOHO_QUOTA_UTC_OFFSET_MINUTESadjusts this.
Limitations
Verified against a live Zoho Inventory org (free plan, India DC) seeded with
scripts/seed_zoho.py. Runscripts/verify_live.pyto check yours. That run turned up four differences from Zoho's docs, now handled and covered by tests:contact details live on
contact_person_details;the ship date is
shipping_dateon package and shipment records;customer_nameis exact-match, so search usescustomer_name_contains;the best stock figure is
actual_available_for_sale_stock.
Delivery dates can be missing. When a shipment is marked delivered through the API (or without a date), Zoho leaves
shipment_delivered_dateempty. The connector reports this as an evidence gap rather than guessing a date.Single-warehouse orgs (Zoho's free plan) return no per-warehouse stock, so
location_namechecks explain that and point to total stock.Stock can be stale by up to 60s because of the read cache. Order data can be up to 30s old.
Pull only. No webhooks.
Search is limited to Zoho's exact, contains and starts-with matching.
The daily quota counter is per process host. Two servers on different machines for the same org would each count separately. A shared store such as Redis would fix that in production.
The demo mode is a test double, not a Zoho emulator. It implements only what this connector uses.
Project layout
src/zoho_connector/
config.py settings, data centres, read-only scopes, plan quotas
auth.py OAuth: consent URL, loopback callback, code exchange, single-flight refresh, revoke
store.py token storage: OS keyring / file / memory
ratelimit.py sliding window, daily quota, backoff
client.py Zoho HTTP client: auth, retries, error mapping, cache
models.py agent-facing output models, PII masking, untrusted text
tools.py the 9 capabilities (transport-agnostic)
server.py MCP adapter, tool descriptions, audit log
cli.py connect / status / serve / disconnect / spec / mock-server
mock/ mock Zoho accounts + Inventory API, fictional dataset
docs/ CAPABILITIES.md · DESIGN.md · IMPACT.md · tools.json
evals/ scenarios.json + run_eval.py
scripts/ demo.py · seed_zoho.py
tests/ 66 testsNo real customer data, passwords or keys are included. All demo records are fictional, and emails use the reserved example.com domain.
Available Tools
9 toolscheck_availabilityARead-onlyIdempotent
Can these SKUs be sold right now? Returns sellable yes/no with a reason per SKU (out of stock, insufficient quantity, inactive, unknown SKU). Call this before an agent pitches, recovers or promises any product to a customer.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | SKUs and quantities to check, e.g. [{"sku": "ST-TEE-BLK-M", "quantity": 2}]. | |
| location_name | No | Only count stock in this warehouse/location (exact name, e.g. 'Bengaluru WH'). Omit for all locations. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| as_of | Yes | |
| results | Yes | |
| all_sellable | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds value beyond that by disclosing the exact response shape (yes/no plus enumerated reasons: out of stock, insufficient quantity, inactive, unknown SKU), which an agent needs to interpret results. It omits any auth or rate-limit context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core question, then the return contract, then the usage trigger. No filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be re-explained, yet the description still summarizes the reason codes usefully. Annotations cover safety and the schema covers parameters, so the definition is nearly complete; the only minor omission is any hint about latency, batch limits, or behavior when a location name is invalid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – both 'lines' and 'location_name' are documented in the schema, including the 25-line cap and the location scoping behavior. The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Can these SKUs be sold right now?') and explicitly enumerates the return semantics (sellable yes/no with a per-SKU reason). It is clearly distinguishable from generic item tools like list_items or get_item because its purpose is sellability, not retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: 'Call this before an agent pitches, recovers or promises any product to a customer.' That is strong when-to-use guidance. It does not name when not to use it or point to a sibling for the alternative case, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
connector_statusARead-onlyIdempotent
Which Zoho organization is connected, granted scopes, and how much of the per-minute and daily API budget is left. Costs no Zoho API calls.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description adds the valuable behavioral trait that this operation costs no Zoho API calls, which is not in the annotations. It does not go into auth details beyond granted scopes, but that is acceptable given the annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, perfectly sized. The first sentence front-loads the returned information; the second adds a key behavioral benefit. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only status tool that has an output schema, the description is complete enough. It tells the agent what it will learn and that the call is free, and the annotations cover the safety profile. No critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the schema description coverage is trivially 100%. Per the rules, a zero-parameter tool receives a baseline of 4 for parameter semantics. There is nothing further the description needs to add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool reports: the connected Zoho organization, granted scopes, and remaining per-minute/daily API budget. It is clearly distinct from sibling tools that retrieve business objects (items, sales orders, availability). The only minor gap is the lack of an explicit verb like 'Returns', but the resource and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool is a diagnostic status check, so its use is somewhat implied. The sentence 'Costs no Zoho API calls' hints that it can be called freely, but the description does not explicitly say when to use it versus alternatives, nor does it state any prerequisites such as needing an existing connection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fulfilment_evidenceARead-onlyIdempotent
Proof-of-fulfilment pack for one order: every shipment with carrier, tracking (AWB), ship and delivery dates and items, a verdict (DELIVERED, PARTIALLY_DELIVERED, SHIPPED_NOT_DELIVERED, NOT_SHIPPED, CANCELLED), a one-line factual summary and the gaps that weaken it. Use for chargebacks/disputes and 'where is my order' questions.
| Name | Required | Description | Default |
|---|---|---|---|
| salesorder_id | No | Zoho salesorder_id. | |
| reference_number | No | Storefront / Razorpay order or payment ID. | |
| salesorder_number | No | e.g. SO-00101. |
Output Schema
| Name | Required | Description |
|---|---|---|
| source | No | |
| summary | Yes | One factual sentence suitable for a dispute response draft. |
| verdict | Yes | |
| shipments | Yes | |
| order_date | Yes | |
| order_total | Yes | |
| order_status | Yes | |
| customer_name | Yes | |
| evidence_gaps | Yes | What is missing before this is strong dispute evidence. |
| salesorder_id | Yes | |
| quantity_ordered | Yes | |
| quantity_shipped | Yes | |
| reference_number | Yes | |
| salesorder_number | Yes | |
| quantity_delivered | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered. The description's extra content is mostly about the return payload (verdicts, summary, gaps), which is output-schema territory rather than behavioral disclosure, and it says nothing about aggregation cost, freshness, or how identifiers resolve.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resource and its contents, with the usage cue last. The verdict enumeration is long but genuinely informative, so there is little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With rich annotations and an output schema, the description need not explain return format, and it sensibly focuses on scope and use cases. It falls just short of full completeness by not noting that all three identifier parameters are optional and at least one must be supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three identifiers (salesorder_id, reference_number, salesorder_number) are already documented, making the baseline 3 appropriate. The description adds no guidance on which identifier to supply or whether one is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Proof-of-fulfilment pack for one order') and enumerates exactly what the pack contains (shipments, carrier, AWB, dates, items, verdict, summary, gaps). This clearly distinguishes it from siblings like get_sales_order or list_sales_orders, which return order data rather than an evidence pack.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for chargebacks/disputes and “where is my order” questions' gives explicit and concrete when-to-use context. It does not name alternatives or when-not conditions, but the intended scenarios are unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_itemARead-onlyIdempotent
Full detail for one item, including stock per warehouse/location. Give exactly one of item_id or sku.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | Exact SKU (case-insensitive). | |
| item_id | No | Zoho item_id. |
Output Schema
| Name | Required | Description |
|---|---|---|
| sku | Yes | |
| name | Yes | |
| rate | Yes | |
| unit | Yes | |
| status | Yes | |
| item_id | Yes | |
| locations | Yes | |
| low_stock | Yes | |
| description | Yes | |
| last_modified | Yes | |
| reorder_level | Yes | |
| stock_on_hand | Yes | |
| available_for_sale | No | Physical stock not committed to other orders. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered. The description adds genuine content context by stating the response includes per-warehouse stock, but says nothing about auth, rate limits, or behavior when an identifier matches nothing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what is returned and followed by the identifier rule. No filler, no repetition of the name or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With rich annotations and an output schema present, the description need not explain return values or safety. It covers the identifier contract and the payload shape, which is enough to call the tool correctly; only the sibling-routing guidance is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a real constraint the schema does not express: the schema makes both params nullable with no required list, while the description mandates exactly one of item_id or sku. That is meaningful guidance the structured fields lack.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource: returns full detail for one item, plus the scope of the payload (stock per warehouse/location). This distinguishes it from list_items and search_items, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Give exactly one of item_id or sku" is an invocation constraint, not a when-to-use rule. The description never says when to pick this over check_availability or search_items, leaving the routing decision entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sales_orderARead-onlyIdempotent
One sales order with line items, shipped quantities, packages and (masked) customer contact. Give exactly one identifier.
| Name | Required | Description | Default |
|---|---|---|---|
| salesorder_id | No | Zoho salesorder_id. | |
| reference_number | No | Storefront / Razorpay order or payment ID. | |
| salesorder_number | No | e.g. SO-00101. |
Output Schema
| Name | Required | Description |
|---|---|---|
| date | Yes | |
| notes | Yes | |
| total | Yes | |
| status | Yes | |
| customer | Yes | |
| packages | Yes | |
| line_items | Yes | |
| paid_status | Yes | |
| customer_name | Yes | |
| last_modified | Yes | |
| salesorder_id | Yes | |
| shipped_status | Yes | |
| delivery_method | Yes | |
| invoiced_status | Yes | |
| reference_number | No | Often the storefront / Razorpay order or payment ID. |
| shipping_address | Yes | |
| salesorder_number | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, covering the safety profile. The description adds useful context: the response includes line items and masked customer contact. Beyond that, it doesn't disclose auth requirements, error behavior for failed lookups, or rate limits, so it adds moderate value on top of solid annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what is returned and followed by the input constraint. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation isn't required, and annotations carry the safety profile. Combined with a summary of returned fields and the one-identifier constraint, this is nearly complete; the only gap is an explicit sibling routing statement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents the three identifier parameters. The description's 'Give exactly one identifier' reinforces exclusivity but adds no format or syntax detail beyond what the schema provides; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (sales order) and enumerates the returned content (line items, shipped quantities, packages, masked customer contact), clearly distinguishing this single-record fetch from sibling list_sales_orders and search_sales_orders.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction 'Give exactly one identifier' implies usage as a single-record lookup, and the contrast with list/search siblings is discernible from the name. It does not explicitly state when not to use it or name the human-readable alternative, but the constraint is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_itemsARead-onlyIdempotent
Browse the merchant's catalogue with stock levels (all warehouses combined). Use for overviews such as 'what is running low?'. For specific SKUs use check_availability.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | next_cursor from the previous page; omit for the first page. | |
| status | No | Item status to include. | active |
| sort_by | No | name | |
| page_size | No | Results per page (1-100). | |
| low_stock_only | No | Only items at or below their reorder level. |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes | |
| has_more | Yes | |
| returned | Yes | |
| next_cursor | No | Pass back as `cursor` to fetch the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the description's real contribution is the aggregation semantics: stock levels are combined across all warehouses, which the annotations cannot express. It does not restate destructive/read-only behavior, and pagination behavior is carried by the schema, so this is a genuine value-add without being exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what the tool returns and scoped immediately, then the alternative. No filler or restated parameter names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description covers scope, contents, use case, and the correct alternative. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents cursor, status, page_size, and low_stock_only. The description only implicitly gestures at low_stock_only through the 'running low' example and adds no syntax or defaulting detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Browse') plus resource ('the merchant's catalogue') with the added scope 'stock levels (all warehouses combined)'. It names the sibling check_availability as the non-overlapping alternative, so an agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use is given ('overviews such as "what is running low?"') along with the alternative for specific SKUs (check_availability), which is strong routing guidance. It does not differentiate from search_items, the other plausibly overlapping sibling, so it falls short of a full when/when-not map.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sales_ordersARead-onlyIdempotent
Browse sales orders, newest first, optionally by status and order-date range.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | next_cursor from the previous page; omit for the first page. | |
| status | No | Order status filter. | all |
| date_to | No | YYYY-MM-DD, inclusive. | |
| date_from | No | YYYY-MM-DD, inclusive. | |
| page_size | No | Results per page (1-100). |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes | |
| has_more | Yes | |
| returned | Yes | |
| next_cursor | No | Pass back as `cursor` to fetch the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world semantics, so the safety profile is covered. The description adds one genuinely useful behavioral fact, the newest-first ordering, but says nothing about result volume or how pagination terminates beyond what the cursor schema states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the default ordering constraint appears immediately and the optional filters follow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the filters are fully specified by the schema. The definition is nearly complete; only the relationship to the sibling search tool is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so cursor, status enum, date_from/date_to formats, and page_size bounds are all documented in the schema. The description only restates the status and date-range filter categories, adding no syntax or defaults beyond what is already structured. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Browse) and resource (sales orders) plus the default sort order, so the operation is unambiguous. It does not explicitly distinguish itself from the sibling search_sales_orders, leaving the list-vs-search boundary to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: browse when you want a full listing with optional status/date filters, as opposed to the sibling search tool. There is no explicit statement of when to prefer this over search_sales_orders or what prerequisites apply, so guidance is only adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_itemsARead-onlyIdempotent
Find items by partial name or SKU when you do not know the exact SKU.
| Name | Required | Description | Default |
|---|---|---|---|
| field | No | Which field to match. 'any' searches name and SKU. | any |
| match | No | Ignored when field='any'. | contains |
| query | Yes | Text to look for, e.g. 'kurta' or 'ST-TEE'. | |
| cursor | No | next_cursor from the previous page; omit for the first page. | |
| page_size | No | Results per page (1-100). |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes | |
| has_more | Yes | |
| returned | Yes | |
| next_cursor | No | Pass back as `cursor` to fetch the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds nothing about result ordering, case sensitivity, or empty-result behavior, and pagination is only covered by the schema's cursor/page_size docs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action first and the qualifying condition second. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and rich annotations, the description need not explain return values or pagination. It is complete enough to call correctly, though naming the exact-match alternative would close the last gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all five parameters (including the field/match enums and pagination) are fully documented in the schema. The description adds no semantics beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Find') and resource ('items') plus the match basis (partial name or SKU). It implicitly contrasts with exact-SKU lookup, but it never names the sibling (get_item) that handles the exact case, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger condition: use when you do not know the exact SKU. That implicitly routes exact lookups elsewhere, but no sibling tool is named explicitly and no exclusions (e.g. vs list_items) are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_sales_ordersARead-onlyIdempotent
Find sales orders by reference number, order number, customer or text. Give at least one criterion.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Free-text search across order number, reference, customer and item names. | |
| cursor | No | next_cursor from the previous page; omit for the first page. | |
| date_to | No | YYYY-MM-DD, inclusive. | |
| date_from | No | YYYY-MM-DD, inclusive. | |
| page_size | No | Results per page (1-100). | |
| customer_name | No | Part of the customer's name. | |
| reference_number | No | Exact reference number, usually the storefront or Razorpay order/payment ID. | |
| salesorder_number | No | Exact Zoho sales order number, e.g. SO-00101. |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes | |
| has_more | Yes | |
| returned | Yes | |
| next_cursor | No | Pass back as `cursor` to fetch the next page. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered by structured data. The description adds only the minimum-one-criterion requirement and nothing about result volume, ranking, or matching behavior; pagination is documented in the schema rather than the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the resource and searchable dimensions, then the constraint. No filler, though it is minimal enough that it leaves value on the table.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich output schema, full annotation coverage, and complete parameter documentation, the description does not need to explain return values or pagination. It is complete enough to call correctly, with only sibling differentiation missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with all 8 parameters documented individually (formats, ranges, defaults), so the schema carries the load. The description names the same searchable fields (reference number, order number, customer, text) without adding syntax or precedence detail, matching the baseline for high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find/search) and resource (sales orders) plus the searchable fields, so the agent knows this is a lookup tool. It does not distinguish itself from siblings like list_sales_orders or get_sales_order, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Give at least one criterion" is a useful invocation constraint, implying the search needs a filter and is not a bare list. However, there is no guidance on when to use this versus list_sales_orders or get_sales_order, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
check_availability - First observed
connector_status - First observed
get_fulfilment_evidence - First observed
get_item - First observed
get_sales_order - First observed
list_items - First observed
list_sales_orders - First observed
search_items - First observed
search_sales_orders
TDQS
Scored across 9 tools
Each tool targets a distinct resource+action: availability check, catalogue browsing, single-item detail, item search, order listing, order search, single-order fetch, fulfilment evidence, and connector status. The list/search and get/list pairs are cleanly separated by explicit guidance in the descriptions (e.g. 'For specific SKUs use check_availability').
Nearly all tools follow a consistent verb_noun snake_case pattern (check_availability, list_items, get_item, search_sales_orders). The only deviation is the noun-only connector_status, which is a minor and understandable exception for a meta/status tool.
Nine tools is well-scoped for a merchant inventory/sales-order connector, with each tool earning its place and no redundant padding.
Read/lookup coverage is strong (items, availability, orders, fulfilment evidence, status), but the surface is entirely read-only: there are no tools to create or update items, adjust stock, create sales orders, or record shipments. These are notable gaps for a connector whose agents may need to act on inventory or orders, though some may be intentional scoping.
Maintenance
Related MCP Connectors
Connect AI to store orders, products and inventory with scoped access and human approvals.
Agent-native travel platform: read-only flight, hotel, and brand tools over MCP. OAuth sign-in.
Real-time data APIs for 50+ commerce, grocery, social and maps platforms over one OAuth MCP server
MerchantFlow is a hosted, read-only ecommerce analytics MCP server for Shopify and WooCommerce. It gives AI assistants tenant-scoped access to revenue, profit and loss, product and SKU profitability, COGS coverage, fulfillment costs, advertising spend, ROAS, marketing performance, cohorts, LTV, and business valuation across connected commerce and marketing platforms. Connect with OAuth over Streamable HTTP. No local server installation is required.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides order lookup, customer lookup, and refund issuance tools with categorized errors to ensure accurate routing and distinguish access failures from valid empty results.-
- FlicenseNot gradedqualityCmaintenanceEnables stock assessment and replenishment by exposing three deterministic tools: inspect stock positions, raise replenishment orders, and check order status. Includes an auditable local client and follows a security-first design with limited API surface.-
- AlicenseNot gradedqualityBmaintenanceEnables an AI agent to query a warehouse management system for orders, stock, stalled orders, and audit logs, and to make narrowly scoped writes by changing order status with a required reason or appending notes, with all writes recorded in a full audit trail.MIT
- AlicenseAqualityBmaintenanceEnables AI agents to manage inventory conversationally: list and look up products, check current stock levels, record inbound and outbound movements idempotently, review movement history, and surface items below their minimum threshold. Every figure is grounded in tool calls rather than model estimation, and invalid operations return explicit error codes the agent can reason about.5MIT