Skip to main content
Glama

price_benchmark

Summarise observed prices for products related to a capability, separated by delivery type (provider / mcp) and buyer tier (individual / pro / team_sme / enterprise) and never blended across incompatible pricing units. Supply the capability with query (natural language, e.g. 'AI code review'); task and the legacy niche are accepted aliases. The market is the semantic neighbourhood of the query (nearest providers by text embedding, no fixed category). Each cohort reports median, mean, stdev, p25/p75, min–max and n. A benchmark describes the observed comparable sample — it is not a quote and not evidence of willingness to pay. Supply provider_type and buyer_tier when the user makes them known; otherwise the populated per-type/per-tier matrix is returned. When nothing priced is semantically close it returns resolved:false with a note, never a fabricated figure.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
taskNoAlias of `query` (same text).
nicheNoLegacy alias of `query`, kept for older clients — prefer `query`.
queryNoThe capability to benchmark, in natural language, e.g. 'AI code review' or 'supplier invoice reconciliation'. Preferred input.
sectorNoSector name, e.g. 'legal' (ignored if a query is given)
buyer_tierNoBuyer tier being priced. Set when the user describes who is buying (an individual, a professional, a team/SME, or an enterprise). An individual licence must never be represented by the SME or enterprise price.
pricing_unitNoOptional pricing unit to hold constant (e.g. 'flat', 'per_seat', 'per_agent'). Incompatible units are never combined.
provider_typeNoDelivery type being priced. Set when the user says provider, MCP or API. Omit (or 'all') to get the per-type matrix instead of a blended figure. 'api' is recognised but not yet a separate commercial cohort (folded into provider).
response_modeNo'summary' (default): compact per-type/per-tier benchmark matrix. 'full': also returns the deprecated blended legacy block + AEPI index.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • changedInput schema / properties / provider_type / enum
      Previous value: -[
      -  "provider",
      -  "mcp",
      -  "api",
      -  "all"
      -]New value: +[
      +  "agent",
      +  "mcp",
      +  "api",
      +  "all"
      +]
    • addedInput schema / properties / query
      Added value: +{
      +  "description": "The capability to benchmark, in natural language, e.g. 'AI code review' or 'supplier invoice reconciliation'. Preferred input.",
      +  "type": "string"
      +}
    • addedInput schema / properties / task
      Added value: +{
      +  "description": "Alias of `query` (same text).",
      +  "type": "string"
      +}
  2. Changed2 schema fields changed
    • changedInput schema / properties / niche / description
      Previous value: -"Niche slug (or an approximate slug / natural-language name — it is resolved to the canonical niche), e.g. 'contract-review-automation'"New value: +"Legacy alias of `query`, kept for older clients — prefer `query`."
    • changedInput schema / properties / sector / description
      Previous value: -"Sector name, e.g. 'legal' (ignored if niche is given)"New value: +"Sector name, e.g. 'legal' (ignored if a query is given)"
  3. Changed4 schema fields changed
    • addedInput schema / properties / buyer_tier
      Added value: +{
      +  "description": "Buyer tier being priced. Set when the user describes who is buying (an individual, a professional, a team/SME, or an enterprise). An individual licence must never be represented by the SME or enterprise price.",
      +  "enum": [
      +    "individual",
      +    "pro",
      +    "team_sme",
      +    "enterprise",
      +    "all"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / pricing_unit
      Added value: +{
      +  "description": "Optional pricing unit to hold constant (e.g. 'flat', 'per_seat', 'per_agent'). Incompatible units are never combined.",
      +  "type": "string"
      +}
    • addedInput schema / properties / provider_type
      Added value: +{
      +  "description": "Delivery type being priced. Set when the user says provider, MCP or API. Omit (or 'all') to get the per-type matrix instead of a blended figure. 'api' is recognised but not yet a separate commercial cohort (folded into provider).",
      +  "enum": [
      +    "provider",
      +    "mcp",
      +    "api",
      +    "all"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / response_mode
      Added value: +{
      +  "description": "'summary' (default): compact per-type/per-tier benchmark matrix. 'full': also returns the deprecated blended legacy block + AEPI index.",
      +  "enum": [
      +    "summary",
      +    "full"
      +  ],
      +  "type": "string"
      +}
  4. Changed1 schema field changed
    • changedInput schema / properties / niche / description
      Previous value: -"Niche slug as returned by market_gaps/compare_agents, e.g. 'contract-review-automation'"New value: +"Niche slug (or an approximate slug / natural-language name — it is resolved to the canonical niche), e.g. 'contract-review-automation'"
  5. First observed

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses that incompatible pricing units are never blended, results come from a semantic neighbourhood, the benchmark is not a quote, and it returns resolved:false without fabricating figures when nothing is close. This is unusually transparent about edge cases and output semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, covering purpose, aliases, market semantics, output statistics, interpretation caveats, parameter conditions, and failure behavior. Every sentence adds value for a tool with no output schema, though the prose is a single dense paragraph and could be structured more cleanly with minimal loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 8 parameters and no output schema, the description is highly complete: it defines the query, aliases, optional filters, fallback behavior, statistical output per cohort, and the no-fabrication failure mode. An agent has enough context to call the tool correctly and interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful parameter context beyond the schema: `query` is natural language, `task` and `niche` are aliases, provider_type/buyer_tier are conditional, and pricing_unit must not be blended. It does not discuss sector or response_mode, but those are already well-described in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Summarise observed prices for products related to a capability', segmented by delivery type and buyer tier. It also clarifies what the tool is not ('not a quote and not evidence of willingness to pay'). However, it does not explicitly distinguish itself from siblings like get_price_index or compare_providers, so sibling differentiation is implicit rather than direct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit invocation guidance: supply the capability via `query`, use `task`/`niche` as aliases, and populate provider_type/buyer_tier only when the user makes them known. It also explains the fallback matrix behavior. It does not name alternatives or state when not to use this tool versus a sibling, so exclusion guidance is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources