oddly commerce corpus
Server Details
The world, on the record. AI infrastructure built on the observed web.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 3 tools
Each tool targets a clearly distinct step: benchmark_categories discovers valid paths, benchmark_lookup fetches the actual figures, and store_audit operates on an external domain. The descriptions explicitly sequence the two benchmark tools (call categories FIRST), eliminating any guesswork about which to use.
All names are consistent snake_case with domain-oriented prefixes (benchmark_*, store_*), which reads predictably. The only minor deviation is that the benchmark pair is noun-based while store_audit follows a noun_verb/verb_noun style, so it isn't a strict uniform verb_noun pattern.
Three tools is on the lean side but defensible for a narrow benchmark corpus plus one audit action, with no redundant or filler tools. A small gap remains (nothing to enumerate metrics or compare stores), so it is slightly under-scoped rather than bloated.
The surface covers the full workflow: discover categories, retrieve observed benchmarks, and audit a storefront against those norms. The scope limits are stated honestly (no conversion/revenue/AOV, empty answers instead of estimates), though there is no way to list metrics or query a specific product/SKU segment.
Available Tools
3 toolsbenchmark_categoriesList covered categoriesAInspect
List the product categories the oddly corpus currently covers, with how many observations and metrics back each one. Call this FIRST to discover valid category paths before calling benchmark_lookup, rather than guessing a path.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the burden and does disclose what comes back: category paths plus counts of observations and metrics backing each. It doesn't mention auth, rate limits, or freshness of the corpus, but for a zero-parameter enumeration the remaining risk surface is small.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the discovery-first instruction front-loaded before the alternative tool is mentioned. Every clause is load-bearing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by describing the returned categories and their backing counts, plus the ordering relationship to benchmark_lookup. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4; there is no parameter-level meaning for the description to add or omit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('product categories the oddly corpus currently covers') plus the return payload (observation and metric counts per category). An agent can distinguish it from benchmark_lookup immediately, since one enumerates coverage and the other retrieves data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly prescribes calling this FIRST to discover valid category paths before benchmark_lookup, and names the failure mode it prevents ('rather than guessing a path'). The when-to-use and the alternative are both stated outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
benchmark_lookupLook up observed category benchmarksAInspect
Observed price and assortment benchmarks for a product category: the median and the p25-p75 band across real public storefront listings, with the sample size, the currency, and the exact scope of each figure. Figures are observed market reference, not targets or advice. Returns an honest empty answer when a category has too little data to serve; it never estimates or fills a gap. This corpus does NOT hold conversion rates, revenue, or AOV, so do not ask it for those.
| Name | Required | Description | Default |
|---|---|---|---|
| market | No | Optional ISO 3166-1 alpha-2 market filter, e.g. 'NZ', 'US'. Omit for the all-markets aggregate. | |
| metric | No | Optional metric filter, e.g. 'price.romper'. Omit to get every servable metric for the category. | |
| category | Yes | Category path, e.g. 'commerce/baby/clothing'. Use benchmark_categories to list valid paths. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the honest-empty-answer behavior, that it never estimates or imputes gaps, the observed-not-target framing, and the currency/scope caveats attached to each figure. It omits operational traits like access/permission needs or any rate or paging behavior, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with what is returned before the caveats and the exclusion. Each sentence carries content, though the description is denser than strictly necessary and could compress the statistics list slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain the return shape, and it does: median, p25-p75 band, sample size, currency, and per-figure scope, plus the empty-result case and the data-domain limits. Nothing an agent needs to call this correctly and interpret the response appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (category, market, metric) are already documented with examples and omission semantics; the schema even points to benchmark_categories for valid paths. The description adds no parameter-level syntax or format detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (observed price and assortment benchmarks for a product category) and enumerates the returned statistics (median, p25-p75 band, sample size, currency, scope), so an agent knows exactly what this tool produces. It also implicitly separates itself from data tools that would return conversion/revenue/AOV, which no sibling provides.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit exclusions: the corpus does not hold conversion rates, revenue, or AOV, and it returns an empty answer rather than estimating when data is thin. That is strong when-not guidance, but it never routes the agent to the relevant siblings (benchmark_categories for valid paths, store_audit for other analyses) inside the description itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
store_auditAudit a public storefrontAInspect
Run oddly's free storefront audit against a public ecommerce domain and return the findings: detected platform, the checks that failed, and how the store's catalogue compares to observed category norms where the corpus can say. Reads the store's own public pages only, obeys robots.txt, and is rate limited. Use the bare domain (example.com). Never pass an internal host or an IP address.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The storefront domain or URL to audit, e.g. 'example.com'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does so well: it discloses reads of public pages only, robots.txt compliance, rate limiting, and the shape of the result. It stops short of the full profile (no mention of failure modes when the domain is unreachable, or whether results are cached).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences: purpose and output first, then behavioral guarantees, then input constraints. Every sentence earns its place except a bit of promotional framing ('oddly's free'), which is minor noise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no annotations and no output schema, the description covers the purpose, the return content, the safety/etiquette profile (robots.txt, rate limit, public-only), and the input format. Nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds real meaning: the accepted input must be a bare public domain and must not be an internal host or IP. That constraint goes beyond the schema's example and directly shapes how the agent formats the argument.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Run oddly's free storefront audit against a public ecommerce domain') and enumerates the returned findings (platform, failed checks, catalogue comparison). This is clearly distinguishable from the sibling benchmark_* tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit input-shape rules — 'Use the bare domain (example.com)' and 'Never pass an internal host or an IP address' — which is strong when-not guidance for the parameter. It does not, however, name the sibling tools or say when to prefer them over this audit, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
benchmark_categories - First observed
benchmark_lookup - First observed
store_audit
Related MCP Connectors
- mcpOAuthcom.sequentum
Turn the web into structured, reliable, actionable enterprise data for AI Agents
Cited reports on any company or person, credential checks and mention scans for your AI
Public AI web-readiness scanner and machine-facing observability discovery service.
Public data intelligence for AI agents — CVE, compliance, patents, contracts, domains.
Related MCP Servers
- AlicenseAqualityBmaintenanceProvenance-first web access for AI agents, delivering clean content with verifiable source metadata and SEC EDGAR financial data.843 npmMIT

Anakinofficial
AlicenseAqualityBmaintenanceWeb data for AI agents: scrape, crawl, search, deep research, site monitoring, browser automation2285 npm3Apache 2.0
siteglass-mcpofficial
AlicenseNot gradedqualityBmaintenanceEnables archival of any public URL as interactive snapshot (rrweb, PDF, PNG) and autonomous web QA for apps (register, scan, generate and run end-to-end flows).39 npmMIT- FlicenseNot gradedqualityBmaintenanceArchives live web pages for offline replay with full interactivity, enabling AI agents to analyze page structure and behavior.1-
Glama MCP Gateway
Add one secure layer between your agents and this server.