Skip to main content
Glama

Server Details

Price benchmarks, alternatives & daily price history across 17,000+ AI agents and MCP servers.

Ownership verified
Status
Healthy
Uptime
94.9% over 43 days
Last Tested
Transport
Streamable HTTP · MCP 2025-03-26
URL

TDQS

A3.8/5.0

Scored across 19 tools

Disambiguation2/5

Several tools occupy overlapping territory: find_market, market_report, research_capability, and price_benchmark all map a natural-language query to a semantic market with pricing, and demand_signals vs market_gaps are near-duplicate 'weak coverage' signals. get_provider vs get_provider_profile is also easy to confuse. Despite long disclaimers, the set has multiple tools with unclear boundaries.

Naming Consistency3/5

Most tools follow snake_case verb_noun (search_providers, create_custom_benchmark, get_price_index), but several are noun phrases (demand_signals, market_gaps, market_report, price_benchmark), and similar actions use different patterns (find_market vs market_report). The naming is readable and mostly predictable, but not consistently verb-first.

Tool Count3/5

19 tools is within the heavy range and the breadth is partially justified by the many facets of the domain (search, pricing index, custom benchmarks, market research, outcome reporting). However, several tools are convenience/overlap additions (research_capability, demand_signals, market_gaps), making the set feel heavier than its core functionality requires.

Completeness4/5

The surface covers discovery, profile inspection, comparison, pricing, market research, custom benchmark CRUD, and outcome reporting, so agents can complete the main buyer workflow without dead ends. Minor gaps remain (no standalone provider listing/browsing without a query, and no direct way to view detailed outcome evidence beyond aggregate reported_success), but they are workaround-able.

Available Tools

19 tools
compare_providersAInspect

Step 3 of the buyer path. Side-by-side capability + plan-level prices for 2–6 providers (e.g. the Pro tier). Inputs are resolved to REAL providers — exact handle, then exact display name — and are NEVER silently swapped for a fuzzy match: unknown inputs come back in unresolved_inputs with suggested_matches and a ready-to-retry corrected_call, and if EXACTLY ONE input is real (the other was invented/mistyped) it does NOT dead-end — it returns comparison_status: compared_with_market_peers, comparing the real provider against its actual in-market competitors — its nearest providers by text-embedding — (listed in compared_against_peers, with a recovery_note); only when ZERO inputs resolve does it return comparison_status: insufficient_valid_providers. When the compared providers are different delivery types it sets mixed_provider_types + a comparability_warning (a hosted agent and an MCP server are not directly equivalent). Full evidence-scored cards for 2-6 handles side by side, each with observed price, all-time community upvotes and provider type. Each card carries the full how_to_connect object (website, docs, MCP endpoint + config_snippet, A2A card, API) so you can act on the winner directly. Each card also carries reported_success — the machine-reported outcome rate from report_outcome (null until 5+ distinct correlated reporters in 90 days). Report your own outcome after using the winner. Accepts provider_ids (aliases: handles, ids; a comma-separated string is also accepted). Use after search_providers or research_capability; when a compared provider is over budget or weakly matched, inline suggested_alternatives are returned.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoOptional. The buyer's job in plain words; used as the service_view heading.
serviceNoOptional. A service recipe whose fixed requirements and evidence fields are applied (see service_view.requirements).
optionalNoOptional preferences: reported in service_view but never gating.
provider_idsYes2-6 provider handles from search_providers/market_gaps, e.g. ["openhands","lexaclaw"]
requirementsNoOptional. Mandatory requirements as plain phrases (e.g. "sanctions screening", "documented MCP interface"). Adds an ADDITIVE service_view: per-provider evidence matrix from stored first-party page text — supported / not_supported / unknown with the excerpt and observation date; unknown never means unsupported.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: exact-match resolution policy, no silent fuzzy swapping, three-way status outcomes for unresolved inputs, mixed-provider-type warnings, peer fallback behavior, and machine-reported success caveats. It even explains corrected_call and recovery_note fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is front-loaded with purpose and packed with valuable edge-case detail, but it is a single dense run-on paragraph with heavy parentheticals; restructuring into bullets would improve scannability. It is more exhaustive than concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description enumerates nearly every important output field and state: comparison_status values, unresolved_inputs, suggested_matches, corrected_call, compared_against_peers, recovery_note, mixed_provider_types, comparability_warning, reported_success, and how_to_connect. It also covers the count constraints and post-use reporting, so an agent can invoke and interpret results without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful input semantics: provider_ids aliases ('handles, ids'), comma-separated string acceptance, and the 2–6 handle constraint tied to search_providers/market_gaps. It does not re-describe all params, but that is already covered in schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens by identifying itself as 'Step 3 of the buyer path' and states the core action: 'Side-by-side capability + plan-level prices for 2–6 providers.' It then describes 'Full evidence-scored cards for 2-6 handles side by side,' making the deliverable unambiguous and distinct from siblings like search_providers or market_gaps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly positions the tool in a sequence ('Step 3', 'Use after search_providers or research_capability') and mentions inline 'suggested_alternatives' for over-budget or weakly matched providers. It does not explicitly list when-not-to-use scenarios or contrast with rank_providers_for_workflow, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_custom_benchmarkAInspect

Create a PRIVATE custom benchmark (a saved, calculated peer cohort) over Agentery's data — no account needed. Two modes: (A) explicit members: pass members (a list of exact handles; product names/domains resolve where unambiguous). (B) fork a market: pass base_niche (its slug) plus optional remove/add. Returns a one-time secret benchmark_id (cb_… token) — store it; it's your only key. Use it later in get/update/delete and in market_report/get_price_index/get_price_index_history. Ambiguous names are returned as candidates, never silently resolved; unresolved inputs block creation unless allow_partial:true. All prices/history are computed from Agentery's immutable observations; canonical market data is never changed.

ParametersJSON Schema
NameRequiredDescriptionDefault
addNomembers to add
nameNoOptional private label
removeNoMode B: members to drop from the forked market
membersNoMode A: exact handles (preferred), product names or domains
base_nicheNoMode B: slug of the canonical market to fork
allow_partialNoCreate with only the resolved members when some inputs don't resolve (default false)

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses critical behaviors: returns a one-time secret ID, ambiguous names are not silently resolved but returned as candidates, unresolved inputs block unless allow_partial is true, and that all prices/history are computed from immutable observations (no mutation of canonical data). This is comprehensive given no annotations were provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise yet information-dense, with clear structure using (A), (B), and note sections. It avoids redundancy and every sentence adds value, making it easy to parse despite its length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description adequately explains the return value (one-time secret benchmark_id) and how to use it later with get/update/delete and other tools. It also covers edge cases (ambiguity, partial allowance) and ensures the user understands the tool's behavior completely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Each parameter is explained in context: members for mode A, base_niche for mode B, add/remove as optional for forking, allow_partial for partial resolution. The description adds meaning beyond the schema by clarifying the mode-specific usage and the semantics of each parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: creating a private custom benchmark (peer cohort) over Agentery's data, with explicit mention of two modes (A and B). It distinguishes from sibling tools by focusing on creation and using terms like 'PRIVATE' and 'custom', which are unique.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (to create a custom benchmark) and how to use it, including both modes with specific parameters. It also clarifies the return value and mentions that it can be used later with other tools, providing clear usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_custom_benchmarkAInspect

Disable access to a custom benchmark. Keeps only a minimal audit record; no underlying Agentery data is touched. Pass benchmark_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
benchmark_idYesYour cb_ token

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses that only a minimal audit record is kept and no underlying data is touched, providing useful behavioral context beyond the tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first states purpose, the second adds behavioral nuance and parameter instruction. No extraneous content; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers purpose, side effects, and parameter. It could mention reversibility or error states, but overall it is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema describes 'benchmark_id' as 'Your cb_ token'. The description merely repeats 'Pass benchmark_id', adding no additional meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'disable access' and resource 'custom benchmark', clearly distinguishing it from sibling tools like create, get, and update. The purpose is immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for disabling a benchmark, but does not explicitly state when to use this tool versus alternatives. No exclusion criteria or context is provided, only a simple instruction to pass the ID.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

demand_signalsAInspect

Inspect eligible zero-result or weak-match capability queries observed on this MCP server, aggregated and ranked by miss count. These are limited coverage signals from Agentery's own callers — not proof that a product does not exist, and not proof that a market has paying demand. Not a prerequisite for choosing a product. Empty args ({}) return the current list; an empty response means there is insufficient qualifying evidence (status insufficient_evidence + next_step) — it is never filled from search popularity, page views or trending queries.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax signals (1-50, default 20)

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses data limitations, what the signals do not prove, response behavior for empty results, and that the data is never sourced from search popularity, page views, or trending queries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the core purpose is front-loaded, caveats follow, and response semantics close it out. The description is dense but not bloated, with no filler or repetition of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single optional parameter and no output schema, the description explains the important empty-response contract and the nature of the returned signals. It does not detail the exact shape of a non-empty signal list, but enough is provided for correct invocation and basic interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes the only parameter (limit: 1-50, default 20) with 100% coverage. The description adds useful invocation context for empty args but does not need to expand on parameter semantics further.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Inspect') and a precise resource: eligible zero-result or weak-match capability queries aggregated and ranked by miss count. It also separates itself from market-intelligence tools by clarifying these are limited coverage signals from Agentery's own callers, not proof of product non-existence or paying demand.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear operational context: empty args return the current list, empty response means insufficient qualifying evidence, and this is explicitly not a prerequisite for choosing a product. However, it does not name or direct the agent to any alternative sibling tool, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_marketAInspect

Map a natural-language task, capability or service to its live MARKET — the semantic neighbourhood of the closest-matching providers, found by text-embedding nearness (NO fixed category). For buyers ('a provider that monitors competitor pricing'), sellers ('what should I charge for lead-generation automation') or sizing a space. Pricing-intent boilerplate is stripped before matching. Returns the market label, how many providers are in the neighbourhood and how many are priced, nearest (the closest providers with observed price and relevance/cosine), and pricing_by_tier — median, mean, stdev, p25/p75, min–max range and n per buyer tier (individual/pro/team_sme/enterprise), computed by the canonical pricing engine over the priced neighbourhood. match_certainty is 'confident' when real neighbours exist and 'uncertain' when nothing is close (pricing WITHHELD). Accepts task (aliases: query, q). For the full market read + shortlist in ONE call, use research_capability instead. Read-only. When there is no strong market the response says so: status no_match or thin_coverage (nearest neighbourhood labelled partial_match) with a plain explanation and a next_step — it never invents a market.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesREQUIRED. A natural-language task, capability or service in plain words, e.g. 'reconcile supplier invoices'. Call as {"task": "…"}; an empty call returns status needs_input with an example and searches nothing.
queryNoAlias for task (back-compat only) — prefer task

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does well: it discloses that pricing-intent boilerplate is stripped before matching, that match_certainty is 'confident' vs 'uncertain', that pricing is WITHHELD when nothing is close, and that the tool returns status no_match or thin_coverage with a next_step. It also states 'Read-only.' A small gap: it doesn't detail rate limits or auth, but for a read-only mapping tool the behavioral disclosure is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: it defines the core behavior, the audience, the matching mechanism, the return fields, the certainty semantics, the alternative tool, and the failure modes. It is front-loaded with the primary purpose and scoping. It is long, but the length is justified by the tool's complexity and the absence of annotations and output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the lack of annotations, and the absence of an output schema, the description is remarkably complete: it names the return fields (market label, provider counts, nearest, pricing_by_tier with tier names), explains match_certainty, and describes the no_match/thin_coverage statuses. It doesn't spell out the exact JSON shape of `nearest` or `pricing_by_tier`, but it provides enough for an agent to invoke the tool and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining that `task` accepts natural-language plain words, gives an example ('reconcile supplier invoices'), notes that an empty call returns needs_input, and clarifies that `query` is a back-compat alias. This is useful semantic context that helps an agent construct a valid call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Map') and a precise resource ('a natural-language task, capability or service to its live MARKET'), then defines the market as a semantic neighbourhood found by text-embedding nearness with no fixed category. It clearly distinguishes itself from siblings like research_capability and search_providers by naming them and stating the one-call alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool: for buyers, sellers, or sizing a space, and names the alternative research_capability for a full market read plus shortlist in one call. It also explains the no_match/thin_coverage behavior and that it never invents a market, which helps an agent decide when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_custom_benchmarkAInspect

Get a private custom benchmark's current report: members, current stats (headline median/quartiles only when ≥3 comparable priced members — monthly, per-seat and per-call prices are never blended), buyer-tier / provider-type / pricing-unit cohorts, historical index, and data coverage. Pass benchmark_id (your cb_ token) as an ARGUMENT.

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofNoOptional YYYY-MM-DD — reproduce the exact stats + index as they were on that date, using this version's fixed membership
versionNoOptional benchmark version (default latest)
benchmark_idYesYour cb_ token (bearer secret; passed as an argument, never a URL)
response_modeNofull includes the index series

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses important behavioral traits: stats only shown when ≥3 comparable members, pricing never blended, and includes cohort breakdowns. Missing details on auth requirements or error handling, but the token requirement is mentioned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the main purpose. It is information-dense but slightly cluttered with parenthetical details. Efficient overall, though could be slightly more streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and four parameters, the description covers the report's contents thoroughly, including conditional stats. It lacks explicit mention of the return format but compensates with specific details about what the report contains. Nearly complete for the complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, providing a baseline of 3. The description adds value by explaining that benchmark_id is a 'cb_ token' and clarifying the as_of parameter's usage ('reproduce exact stats'). It does not fully repeat schema but provides contextual nuance beyond parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: retrieving a private custom benchmark's current report, and enumerates the exact contents (members, stats, cohorts, historical index, data coverage). It effectively distinguishes from sibling tools like create_custom_benchmark or delete_custom_benchmark.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context (for viewing a report) but does not explicitly state when to use this tool versus alternatives like price_benchmark or get_price_index. No exclusions or 'when not to use' guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_price_indexAInspect

Call this for the CURRENT level of the Agent Economy Price Index (AEPI) — a chained like-for-like index over observed provider/MCP pricing (base 100 = 29 Jun 2026). It is an INDEX LEVEL, not a market price or tradeable asset. Returns the whole-economy headline index level with change_1d/change_7d/change_30d, as_of, like_for_like_pair_count, status and the methodology version, PLUS the same fields for the four buyer tiers (Individual, Pro, Team/SME, Enterprise). provider_type returns the standalone index for one delivery type (provider or mcp, own base 100) — agents and MCPs price and move differently. tier filters to one buyer tier; response_mode 'full' adds exact sub-0.01% moves and repricing counts. Reads the SAME canonical series as the /aepi page, so the MCP and website agree for a given timestamp. (Also accepts a benchmark_id to read a private custom benchmark's current index.) The economy index is whole-market by design — for pricing on a specific capability use market_report or price_benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault
tierNoFilter to one buyer tier ('team' = Team/SME). Default 'all'.
benchmark_idNoOptional: a private custom benchmark token (cb_…) from create_custom_benchmark — returns that cohort's current index instead of the economy index. Cannot be combined with niche.
provider_typeNo'all' (default) = the combined whole-economy index; 'agent' or 'mcp' = the standalone index over just that delivery type (own base 100); 'payg' = the standalone pay-as-you-go sub-index (usage rates purchasable without a subscription, engine v2). Agents and MCPs price and move differently, so an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index.
response_modeNo'summary' (default) or 'full' (adds exact sub-0.01% moves and repricing counts).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries full behavioral disclosure. It explains that the tool is read-oriented, returns a whole-economy headline index plus per-tier and per-provider-type series, treats standalone provider_type indices with their own base 100, and reads the same canonical series as the /aepi page so the MCP and website agree. It does not discuss access requirements or rate limits, but for a non-destructive data-read tool the behavioral context is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core purpose ('CURRENT level of AEPI') before diving into output fields and modes. Every sentence carries substantive information, though the return-field enumeration and long parenthetical make it slightly heavy. It is appropriately sized for a tool with four parameters and no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes responsibility for explaining return values, and it enumerates the index fields, per-tier fields, provider_type variants, response_mode behavior, benchmark override, and the canonical-series guarantee. It is largely complete for correct invocation, but it does not state output formatting, field nesting, or potential parameter-combination constraints beyond what is implied in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: response_mode 'full' adds exact sub-0.01% moves and repricing counts, provider_type returns standalone indices with their own base 100, and benchmark_id reads a private custom benchmark's current index. These clarifications help an agent choose parameter values correctly rather than merely knowing they exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Call this for the CURRENT level of the Agent Economy Price Index (AEPI)'. It then disambiguates the tool's domain by stating it returns an index level, not a market price or tradeable asset, and contrasts it with market_report and price_benchmark for capability-level pricing. This makes the tool's purpose and scope unmistakable, even among many siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool (current index level), what it is not for ('not a market price or tradeable asset'), and when to use alternatives ('for pricing on a specific capability use market_report or price_benchmark'). It also gives buyer-type guidance, e.g., an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index. This provides clear operational selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_price_index_historyAInspect

Call this for the canonical DATED index SERIES (to chart or analyse movement) of the AEPI — the same chained like-for-like series the /aepi page plots. Every point is an index level (base 100), never a price. Returns the whole-economy headline series, or a single buyer tier's series when tier is set. provider_type returns the standalone 'agent' / 'mcp' series (own base 100). period selects '30d' (default), '90d' or 'all'. response_mode 'summary' (default) returns date + index_level points plus the window change; 'full' adds gap flags. Returns an honest status (insufficient_history) rather than a fabricated series when data is too thin. (Also accepts a benchmark_id to read a private custom benchmark's history.) The economy index is whole-market by design — for pricing on a specific capability use market_report or price_benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault
tierNoReturn one buyer tier's series ('team' = Team/SME). Default 'all' = the headline series.
periodNoHistory window. Default '30d'.
benchmark_idNoOptional: a private custom benchmark token (cb_…) — returns that cohort's dated series instead of the economy series. Cannot be combined with niche.
provider_typeNo'all' (default) = the combined whole-economy series, or the standalone 'agent' / 'mcp' / 'subscription' / 'hybrid' / 'one_off' / 'payg' = the billing-LENS series (they cut across provider/mcp; a rate sits in one delivery type AND one lens; own base 100) — 'payg' (pay-as-you-go, engine v2) series (own base 100).
response_modeNo'summary' (default, compact) or 'full'.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that points are index levels, never prices; that response_mode 'summary' returns date + index_level points plus window change and 'full' adds gap flags; and that an honest 'insufficient_history' status is returned rather than a fabricated series. These are substantial behavioral traits beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loaded with the core purpose, followed by critical disclaimers (base 100, not price). Each clause adds value, but some statements echo schema info, such as the period enum values and benchmark_id acceptance. Overall it is efficient, but slightly longer than strictly necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is remarkably complete: it explains parameter behavior, response modes, error handling, alternatives, and scope limitations. It covers everything an agent needs to decide when to call this tool and how to interpret its results. Minor omissions like authentication are likely handled at the framework level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful conceptual clarity beyond the schema: it explains the 'same chained like-for-like series', the distinction between tier (buyer tier) and provider_type (billing-lens series with own base 100), and that provider_type series 'cut across provider/mcp'. This goes beyond the schema's dry enum descriptions, though some details (e.g., benchmark_id) are redundant with the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'the canonical DATED index SERIES ... of the AEPI'. It explicitly states every point is an index level (base 100), never a price, and contrasts itself with the whole-market economy series versus specific-capability tools. This clearly distinguishes it from siblings like get_price_index (point-in-time) and market_report/price_benchmark.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit usage context: 'Call this for the canonical DATED index SERIES (to chart or analyse movement)'. It also provides direct alternatives and exclusions: 'The economy index is whole-market by design — for pricing on a specific capability use market_report or price_benchmark.' Additionally, it explains when to use benchmark_id for a private custom benchmark's history.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_providerAInspect

Call this for the public directory card of one provider by handle or registration number: bio, source URLs, X-verification status, entity type, community rating and structured profile when available.

ParametersJSON Schema
NameRequiredDescriptionDefault
handleNoProvider handle, e.g. 'openhands'
regNumNoRegistration number, e.g. 2432

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It lists the output fields and notes that the structured profile is included 'when available,' which adds transparency about conditional data. However, it does not disclose what happens if both handle and regNum are omitted, which takes precedence, or error behavior when a provider is not found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that front-loads the purpose ('Call this for the public directory card') and then efficiently enumerates the included data. No redundant words or information; every segment adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with no output schema, the description adequately explains the return content and its conditionality. It lacks explicit error-handling details and interaction rules for the two parameters, but given the tool's simplicity, it covers the essential context sufficiently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes both parameters (handle and regNum), so baseline is 3. The description adds semantic value by explicitly stating 'by handle or registration number,' clarifying that the two parameters are alternatives rather than complementary, which is not apparent from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool retrieves a public directory card for one provider, listing the included fields (bio, source URLs, X-verification status, etc.). It specifies the resource (provider) and the action (get by handle or registration number), but does not explicitly differentiate from sibling tools like get_provider_profile.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Call this for...' provides a direct usage context. However, it does not mention when NOT to use it or suggest alternative tools (e.g., search_providers for finding providers), leaving the when-to-use guidance implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_provider_profileAInspect

Step 2 of the buyer path. Full profile for ONE provider — plans, pricing model, liveness, evidence — use after search_providers returns handles. Evidence-scored profile: task_performed, inputs/outputs, integrations, protocols, industry_fit, autonomy_level, human_approval_needed, observed price, trust signals, evidence_quality, entity_type, regulated_data_suitability, evidence_urls, last_checked. Includes the full how_to_connect object — website, docs, any vendor-published MCP endpoint (with a copy-paste client config_snippet), A2A agent card and API surface — the info needed to actually use the listing; fields are null when the vendor publishes no endpoint (never guessed). Also carries reported_success — machine-reported outcome rate from report_outcome (null until 5+ distinct correlated reporters in 90 days). If you use the listing, call report_outcome afterwards.

ParametersJSON Schema
NameRequiredDescriptionDefault
provider_idYesThe provider_id/handle returned by search_providers or compare_providers

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It discloses important behavior: fields are null when no vendor endpoint exists and are 'never guessed', reported_success is null until 5+ distinct correlated reporters in 90 days, and there is an expected follow-up call to report_outcome. This is substantive, though it does not mention auth or read-only guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the tool's purpose and placement in the buyer path. The long field lists are justified because there is no output schema, but the single-paragraph wall of text makes it slightly harder to scan than ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description thoroughly documents return content, null semantics, the reported_success threshold, the required prior step, and the follow-up report_outcome call. An agent has enough context to select, invoke, and interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the sole parameter provider_id is already documented in the schema as the handle returned by search_providers or compare_providers. The description reinforces this provenance but adds no new semantic detail beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'get full profile for ONE provider' and identifies the tool as 'Step 2 of the buyer path'. It also enumerates the exact content (plans, pricing model, liveness, evidence, how_to_connect), clearly distinguishing it from search_providers, compare_providers, and get_provider.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to use this tool after search_providers returns handles, and instructs the agent to call report_outcome afterwards if the listing is used. This provides clear sequencing and context, though it does not explicitly list when-not-to-use alternatives like get_provider or compare_providers.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

market_gapsAInspect

Inspect query clusters with weak coverage among indexed paid providers — research leads, not buyer counts: request counts are not buyer counts, and a gap in this index does not establish a gap in the wider market. Computed demand-first in the raw text-embedding space (NO fixed categories). A gap = a cluster of user requests seen on this MCP server that sits FAR from any PAID provider. For each gap it returns: the demand phrasing, demand_mass (how many similar requests cluster with it), nearest_paid_similarity (cosine to the closest paid provider — low = under-served) and that closest paid provider. Also returns demand_queries and paid_supply counts. Honestly returns few or no gaps while query volume is still low — it sharpens as usage grows. No arguments needed ({}); limit caps the list.

ParametersJSON Schema
NameRequiredDescriptionDefault
rankNogaps (default): whitespace with money, crowded excluded. hot: most active by market pulse, crowding ignored.
limitNoMax gaps (1-50, default 15)
sectorNoOptional sector filter, e.g. 'legal', 'healthcare'

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that gaps are computed demand-first in raw text-embedding space (not fixed categories), defines what constitutes a gap (clusters far from paid providers), explicitly mentions that request counts are not buyer counts, and honestly notes that results are sparse when query volume is low and sharpen over time. This is exemplary transparency about limitations and computational methodology.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although the description is lengthy, every sentence contributes essential information: purpose, interpretation caveats, computational approach, definition of gaps, return contents, and behavior over time. It is front-loaded with the core purpose and structured logically. Minor redundancy (repeating 'not buyer counts' in different forms) slightly reduces efficiency, but overall it is well-organized and appropriately detailed for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers all critical aspects for correct invocation: it states the purpose, the input (no required args, limit cap), the output (demand phrasing, demand_mass, nearest_paid_similarity, closest provider, demand_queries, paid_supply counts), limitations (low volume yields few gaps), and usage precautions. There is no output schema, so the description adequately explains return values. For a tool of this complexity, the description is remarkably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description clarifies that no arguments are required ({}), and mentions the limit caps the list, which adds a small amount of guidance beyond the schema. However, it does not elaborate on the 'rank' enum or 'sector' beyond what the schema already explains, so it does not exceed the baseline significantly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the specific verb ('Inspect') and resource ('query clusters with weak coverage among indexed paid providers'), and immediately differentiates it from buyer-count analysis ('research leads, not buyer counts'). It also distinguishes itself from sibling tools like demand_signals or rank_providers by focusing on gaps relative to paid providers. This makes the purpose unambiguous and distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong usage context: it says to use it for research leads rather than inferring buyer counts, and warns that gaps in the index do not prove gaps in the wider market. While it does not explicitly name alternative tools, it clearly advises on interpretation and when the tool is appropriate (i.e., for lead research with caution). This is more than adequate guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

market_reportAInspect

Deep-dive ONE market before building or investing — the market is the semantic neighbourhood of your natural-language query (nearest providers by text embedding, NO fixed category). Every field is MEASURED: the observed-pricing benchmark separated by provider type and buyer tier (median, mean, stdev, p25/p75, min–max range and n via the canonical pricing engine), how many providers are in the neighbourhood and how many are priced, and the top providers already competing there with their observed price and relevance. Pass query (a natural-language capability or market, e.g. 'customer support chatbot'). For market + pricing + a ready shortlist in one call, use research_capability.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural-language capability or market, e.g. 'customer support chatbot' or 'ai phishing detection'.
response_modeNo'summary' (DEFAULT) returns a compact block: neighbourhood counts, per-type/per-tier price cohorts and top providers. 'full' returns everything incl. the full member list.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description thoroughly discloses behavior: it explains the dynamic market definition ('semantic neighbourhood of your natural-language query'), the measured nature of every field (with specific statistical outputs: median, mean, stdev, p25/p75, min–max, n), and the pricing benchmark's architecture. It also implicitly warns that the market is not a fixed category, which is a critical behavioral nuance.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise for the amount of content it covers. It front-loads the core action and scoping, then lists expected outputs, and ends with a parameter guidance and sibling exclusion. No wasted sentences; every clause adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool (dynamic market definition, multiple measured metrics), the description is remarkably complete. It covers what the tool does, how it's scoped, what outputs to expect, and when to use an alternative. The absence of an output schema is compensated by the description's detailed explanation of the response contents.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although the schema already covers 100% of parameters, the description adds semantic value by clarifying the `query` parameter's role as a natural-language capability or market with concrete examples, and explicitly describes the `response_mode` effect. It doesn't add format details for response_mode beyond the enum, but it does confirm defaults, which is helpful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is unusually specific: it identifies a unique 'semantic neighbourhood' scoping (nearest providers by text embedding, no fixed category), names the resource (ONE market deep-dive), and clearly distinguishes from siblings like research_capability. The action ('Deep-dive ONE market') is precise with measurable outcomes listed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'before building or investing'. Contrasts with the 'research_capability' sibling by stating that for 'market + pricing + a ready shortlist in one call, use research_capability'. This is clear guidance on when NOT to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

price_benchmarkAInspect

Summarise observed prices for products related to a capability, separated by delivery type (provider / mcp) and buyer tier (individual / pro / team_sme / enterprise) and never blended across incompatible pricing units. Supply the capability with query (natural language, e.g. 'AI code review'); task and the legacy niche are accepted aliases. The market is the semantic neighbourhood of the query (nearest providers by text embedding, no fixed category). Each cohort reports median, mean, stdev, p25/p75, min–max and n. A benchmark describes the observed comparable sample — it is not a quote and not evidence of willingness to pay. Supply provider_type and buyer_tier when the user makes them known; otherwise the populated per-type/per-tier matrix is returned. When nothing priced is semantically close it returns resolved:false with a note, never a fabricated figure.

ParametersJSON Schema
NameRequiredDescriptionDefault
taskNoAlias of `query` (same text).
nicheNoLegacy alias of `query`, kept for older clients — prefer `query`.
queryNoThe capability to benchmark, in natural language, e.g. 'AI code review' or 'supplier invoice reconciliation'. Preferred input.
sectorNoSector name, e.g. 'legal' (ignored if a query is given)
buyer_tierNoBuyer tier being priced. Set when the user describes who is buying (an individual, a professional, a team/SME, or an enterprise). An individual licence must never be represented by the SME or enterprise price.
pricing_unitNoOptional pricing unit to hold constant (e.g. 'flat', 'per_seat', 'per_agent'). Incompatible units are never combined.
provider_typeNoDelivery type being priced. Set when the user says provider, MCP or API. Omit (or 'all') to get the per-type matrix instead of a blended figure. 'api' is recognised but not yet a separate commercial cohort (folded into provider).
response_modeNo'summary' (default): compact per-type/per-tier benchmark matrix. 'full': also returns the deprecated blended legacy block + AEPI index.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses that incompatible pricing units are never blended, results come from a semantic neighbourhood, the benchmark is not a quote, and it returns resolved:false without fabricating figures when nothing is close. This is unusually transparent about edge cases and output semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, covering purpose, aliases, market semantics, output statistics, interpretation caveats, parameter conditions, and failure behavior. Every sentence adds value for a tool with no output schema, though the prose is a single dense paragraph and could be structured more cleanly with minimal loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 8 parameters and no output schema, the description is highly complete: it defines the query, aliases, optional filters, fallback behavior, statistical output per cohort, and the no-fabrication failure mode. An agent has enough context to call the tool correctly and interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful parameter context beyond the schema: `query` is natural language, `task` and `niche` are aliases, provider_type/buyer_tier are conditional, and pricing_unit must not be blended. It does not discuss sector or response_mode, but those are already well-described in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Summarise observed prices for products related to a capability', segmented by delivery type and buyer tier. It also clarifies what the tool is not ('not a quote and not evidence of willingness to pay'). However, it does not explicitly distinguish itself from siblings like get_price_index or compare_providers, so sibling differentiation is implicit rather than direct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit invocation guidance: supply the capability via `query`, use `task`/`niche` as aliases, and populate provider_type/buyer_tier only when the user makes them known. It also explains the fallback matrix behavior. It does not name alternatives or state when not to use this tool versus a sibling, so exclusion guidance is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rank_providers_for_workflowAInspect

PARTNER-ONLY (Bearer key required). Given a business context and its workflow steps, return ranked provider candidates for EACH step — structured, scored (match_score 0-100) matches with match_reasons and cautions. Built for app builders (e.g. Builtery) assembling automations. Reads each provider's analysed site profile; never invents capabilities; returns 'unclear' where evidence is missing.

ParametersJSON Schema
NameRequiredDescriptionDefault
limit_per_stepNoMax candidates per step (1-25, default 8)
workflow_stepsYesEach: step_id, step_name, step_description, inputs[], desired_outputs[], required_integrations[], human_approval_preference (always|sometimes|not_needed|unknown)
business_contextNocompany_description, industry, region, existing_tools[], automation_posture (cautious|balanced|agent_native), regulated_data (none|personal|health|financial|legal|children|unknown)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It explicitly discloses the Bearer-key requirement, that it reads analysed site profiles, never invents capabilities, and returns 'unclear' when evidence is missing. This is a high level of transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, first sentence is an imperative statement of behavior, second gives audience, third gives data-safety guarantee. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a moderate-complexity tool with no output schema or annotations, the description adequately covers purpose, auth, return format, and behavioral constraints. It is sufficient for an agent to select and invoke.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds a general mapping of business context and workflow steps to inputs, but no additional details beyond schema descriptions. Thus a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool ranks provider candidates for each workflow step, with specific deliverables (match_score, match_reasons, cautions). This distinguishes it from sibling tools like search_providers and compare_providers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states it is built for app builders assembling automations, providing clear context for when to use. However, it does not explicitly name alternative tools or state when not to use, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report_outcomeAInspect

Report the result of ACTUALLY USING a listed provider for a task. Testing Agentery's connection or retrieval does not establish that the listed provider worked — do not report those. Reports are self-reported evidence subject to eligibility checks: they are correlated with your recent retrievals, improve ranking accuracy, and unlock higher rate limits for contributors. Only reports we can match to one of YOUR retrievals (search_providers / get_provider_profile / compare_providers / suggest_alternatives naming that provider, last 48h) carry weight; unmatched reports are stored but unweighted. Aggregates surface as reported_success on profile/comparison cards once 5+ distinct reporters exist (90-day window). Callers with 5+ correlated reports in 30 days get a doubled per-minute rate limit. Send an x-agentery-key header to keep one reporter identity across IPs (it is stored only as a hash).

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoOptional free-text detail (capped at 300 chars)
outcomeYesDid the provider accomplish the task you hired it for?
agent_idYesHandle of the provider you used, as returned by search_providers/get_provider_profile/compare_providers
task_typeNoOptional short task label, e.g. 'code-review', 'lead-enrichment'
latency_msNoOptional end-to-end latency of the provider in milliseconds
error_classNoOptional failure class, e.g. 'timeout', 'auth', 'wrong-output', 'endpoint-down'

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden of behavioral disclosure. It thoroughly covers: self-reported evidence subject to eligibility checks, correlation with recent retrievals, ranking impact, rate-limit unlocks (doubled per-minute limit with 5+ correlated reports in 30 days), aggregation threshold (5+ distinct reporters in 90 days), and header hashing. This is exceptional transparency for a side-effecting write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is about 150 words but every sentence carries distinct, essential information. It is front-loaded with the core purpose and then layers eligibility, aggregation, rate limits, and header instructions. No redundancy; the structure follows a logical progression from purpose to caveats to incentives. This is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential operational details: what triggers weight, aggregation thresholds, rate-limit consequences, and identity handling. However, it does not mention what the tool returns (e.g., confirmation, report ID) despite having no output schema. Given the tool's complexity (6 params, 2 required) and that return value could matter to an agent, a brief note on response would make it fully complete. This is a minor gap in an otherwise thorough description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds meaningful semantic context: it clarifies that the outcome parameter reflects actual use ('ACTUALLY USING') and explicitly states what NOT to report (connection/retrieval tests), which shapes how parameters like outcome and agent_id should be filled. It also clarifies that agent_id must come from the listed retrieval tools. This goes beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Report the result of ACTUALLY USING a listed provider for a task.' It explicitly excludes testing/connection checks, which sharply distinguishes it from sibling search/compare/rank tools. The purpose is unambiguous and differentiated from all 18 siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: only after actually using a provider, not for testing. It states eligibility conditions (must match your recent retrievals within 48h) and explains consequences of unmatched reports (stored but unweighted). It also instructs on the x-agentery-key header for identity consistency. This is a textbook usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

research_capabilityAInspect

Optional combined research route: ONE call turns a task into: (1) its live MARKET — the semantic neighbourhood of the closest-matching providers, found purely by text-embedding nearness (NO fixed category), with the relevance floor and how many providers cleared it; (2) current pricing context — comparable price range and median with mean, stdev and n, plus provider/priced counts; and (3) a ready-to-compare provider shortlist — each with observed price, market_position (below/in-line/above market), integration status, match score, and handles collected in compare_ready. Retrieval is 100% nearest-neighbour by text embedding: providers are matched on what they actually DO, never on an assigned label. REUSES the canonical pricing/search engines (no new pricing logic). Also returns suggested_alternatives (cheaper or stronger options) and a result_fingerprint (+ cached) so repeat calls are cheap. It does NOT run the comparison — pass compare_ready to compare_providers once you have finalists. For detailed inspection prefer the default path: search_providers → get_provider_profile → compare_providers. Aliases: task also accepts query / q.

ParametersJSON Schema
NameRequiredDescriptionDefault
sortNoShortlist ordering. Default 'match'.
taskYesRequired — the natural-language capability/task, e.g. 'reconcile supplier invoices' or 'litigation-analysis provider'. Aliases: query, q.
limitNoShortlist size (1-12, default 5).
buyer_tierNoOptional buyer tier to price against ('team' = Team/SME).
integrationsNoOptional required integrations, e.g. ["zendesk","slack"] — soft preference; integration status is reported per provider.
provider_typeNoPreferred delivery type. 'auto' (default) infers from the task; note 'AI agent' phrasing is treated as generic (neutral), not an agent-only filter. When a type is explicit (mcp/api/agent) matching providers are SOFT-RANKED to the top and the rest are kept as clearly-labelled cross_type_alternative entries — never hard-filtered (no zero-result cliff), and the functional match is never changed. Every provider is labelled with provider_type (public values: agent | mcp | api | unknown) + type_match_score; provider_type{type_rank_boost_applied, boosted_provider_type, result_counts_by_type} is returned.
response_modeNo'summary' (default) or 'full' (adds tier cohorts, coverage and raw results).
max_monthly_usdNoOptional budget ceiling in USD/month — filters the shortlist and drives suggested_alternatives.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries most of the behavioral burden. It discloses the retrieval method (100% nearest-neighbour embedding, no fixed category), the reuse of existing pricing/search engines, caching behavior via result_fingerprint, and the explicit negative that it does not run comparisons. It stops short of stating read-only status or any auth/rate-limit constraints, but is still substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and well-structured with enumerated result components, bolded caveats, and front-loaded purpose. It is somewhat long and repeats alias details already in the schema, but most sentences earn their place by conveying routing or behavioral constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity — 8 parameters, 4 enums, and no output schema — the description is remarkably complete. It enumerates the main outputs (market, pricing stats, compare_ready shortlist, suggested_alternatives, fingerprint/cached), explains retrieval semantics, and provides the recommended fallback path, giving an agent enough context to call it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 8 parameters, including enums and defaults. The description adds output-oriented context and repeats the task alias already present in the schema, but does not materially enrich parameter meaning beyond what the input schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific combined action: 'ONE call turns a task into' a live market, pricing context, and compare-ready shortlist. It also distinguishes the tool from siblings by explicitly stating it does NOT run the comparison and reuses the canonical pricing/search engines rather than adding new logic.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when not to use it ('It does NOT run the comparison'), what to do instead ('pass compare_ready to compare_providers once you have finalists'), and the preferred detailed-inspection path ('search_providers → get_provider_profile → compare_providers'). This gives concrete when-to/when-not routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_providersAInspect

Find candidates for a described task — example: 'supplier invoice reconciliation'. Check the strongest matches with get_provider_profile before recommending them, then use compare_providers for the shortlist. Results may include different delivery types (provider / MCP server / API / platform) and products without a numeric price: inspect the returned priced, observed_price and provider_type fields before treating a result as a recommendation. Set require_public_price to keep only products with an observed public price (then every returned result has one; if nothing priced is close you get an honest thin_coverage/no_match answer with the unpriced candidates labelled, never padding). max_monthly_usd drops products whose observed lowest paid tier exceeds it; sort price_asc lists cheapest observed price first. Each result carries the website URL (and pricing page URL when observed); get_provider_profile has the full evidence, plans and how_to_connect. Accepts query (aliases: q, text). research_capability is the optional combined route (market + pricing context + shortlist in one call).

ParametersJSON Schema
NameRequiredDescriptionDefault
sortNomatch (default) or price_asc (cheapest observed price first; unpriced providers last)
limitNoMax results (1-50, default 20)
queryYesREQUIRED. Free-text description of the product job you are buying, e.g. 'customer support provider with Zendesk integration' or 'supplier invoice reconciliation'. Not for questions about Agentery itself (upvotes, endpoints): those return status unsupported_request.
billingNoOnly providers with one of these observed billing models, e.g. ["free","freemium","subscription","usage"]
filtersNoOptional: industry_fit[], integrations_available[], entity_type[] (agent|tool|infrastructure|service|marketplace|content-community), autonomy_level[] (assistant|workflow automation|agentic|infrastructure), minimum_evidence_quality (low|medium|high)
provider_typeNoFilter to one provider type — the audited classification dimension (same as the website type chips and AEPI facets); result labels always match this filter. Omit for all types.
max_monthly_usdNoDrop providers whose observed lowest paid tier exceeds this (USD/month). Providers with no observed public price still pass unless require_public_price is true.
require_public_priceNoOnly return providers with an observed public price (default false)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so well: it discloses mixed delivery types, the possibility of unpriced products, the need to inspect priced/observed_price/provider_type before treating results as recommendations, the 'never padding' no_match behavior, and the effect of require_public_price and max_monthly_usd.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with purpose and usage, then behavior. Most sentences add useful guidance, though some details repeat schema descriptions and the paragraph style could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 params, no output schema, no annotations), the description is unusually complete: it explains return-field semantics, honest failure modes, filtering behavior, and the route to richer evidence via get_provider_profile.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a few extras like query aliases (q, text) and reinforces spending filters, but most parameter semantics are already fully documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Find candidates for a described task') and gives a concrete example. It also distinguishes the tool from siblings by naming get_provider_profile, compare_providers, and research_capability, so an agent can tell this is the candidate-discovery step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states the workflow: use search_providers to find candidates, check strongest matches with get_provider_profile before recommending, then use compare_providers for shortlisting. It also mentions research_capability as an optional combined route, giving clear when-to-use guidance against alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_alternativesAInspect

Find related alternatives to a known provider, ranked by text-embedding nearness to that provider's OWN profile (NO category lookup), each with observed price, endpoint liveness, community upvotes and how_to_connect (website, docs, mcp endpoint). For cheaper_only, inspect whether the reference price and candidate prices support a valid comparison: an empty response may reflect a missing or incompatible reference price rather than the absence of alternatives (the response says which). Accepts agent_id (aliases: handle, id). These substitutes are also surfaced inside research_capability and compare_providers.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax alternatives (1-10, default 5)
agent_idYesHandle of the provider to find substitutes for, e.g. 'openhands'
cheaper_onlyNoOnly keep alternatives priced below the subject's lowest monthly price. Free/freemium providers always qualify; providers with no observed price are excluded. Default false.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It disclosed important behavior about cheaper_only: an empty response may be due to missing/incompatible reference pricing, not the absence of alternatives, and that the response indicates which. This is valuable beyond the boolean flag description in the schema, helping agents interpret results correctly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single block of text, moderately long, but every sentence adds value. It front-loads the core function and output details, then explains the tricky cheaper_only behavior. The note about 'NO category lookup' is important but not over-elaborated. Slight room for better structure by splitting sentences into paragraphs, but acceptable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 3 parameters, no output schema, and no annotations, the description covers its purpose, output fields, and a key edge case. It does not explain the full return structure in detail (e.g., response format of how_to_connect), but given complexity is moderate, it's sufficient. Lacks mention of potential subject with no alternatives at all, but overall complete enough.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already describes each parameter's meaning. The description adds context for cheaper_only by explaining the edge case for empty response, which is useful. However, it doesn't add significant new semantics for agent_id or limit beyond what the schema says, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool finds related alternatives to a provider, based on text-embedding nearness to that provider's own profile, with a specific list of output fields. It distinctly differentiates from siblings by explicitly noting 'NO category lookup' and mentions that substitutes are also surfaced inside research_capability and compare_providers, helping an agent distinguish this tool from those.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context on when to use this tool: when you need alternatives to a known provider, and the mechanism (embedding nearness vs category lookup). It implicitly differentiates from siblings by stating that this tool is separate from category-based lookup and that alternatives are also surfaced elsewhere. However, it doesn't explicitly say when not to use it or name a specific alternative to prefer in certain cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_custom_benchmarkAInspect

Add/remove members or rename a custom benchmark. Creates a NEW immutable version (the previous version stays fully reproducible) and returns the exact change-impact on the median/quartiles/index. Pass benchmark_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
addNo
removeNo
renameNo
benchmark_idYesYour cb_ token

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries full responsibility. It discloses key behaviors: creating a new immutable version while preserving reproducibility of the previous version, and returning exact change-impact on median/quartiles/index. However, it does not mention permissions, reversibility, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences long, front-loads the primary action, and adds crucial behavioral information in the second sentence. Every word earns its place; no unnecessary content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters, no output schema, and no annotations, the description covers core behavior (modification with immutability) and the nature of the return value. It lacks details on prerequisites and error conditions, but overall is sufficiently complete for its complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 25% (only benchmark_id described). The description adds meaning by explaining that 'add' and 'remove' manage members and 'rename' changes the name, but it does not specify the format of member strings or provide detail on the benchmark_id token. It partially compensates for low schema coverage but leaves ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's actions: 'Add/remove members or rename a custom benchmark.' It identifies the specific resource type (custom benchmark) and differentiates from sibling tools like create_custom_benchmark and delete_custom_benchmark by focusing on modifications.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for modifying an existing benchmark (vs. creating or deleting), but it does not explicitly compare to siblings or specify when not to use it. The instruction 'Pass `benchmark_id`' provides basic usage guidance but no exclusions or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updates
    • Changedget_price_index1 field changed
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "agent",
        -  "mcp",
        -  "payg"
        -]New value: +[
        +  "all",
        +  "agent",
        +  "mcp",
        +  "subscription",
        +  "hybrid",
        +  "one_off",
        +  "payg"
        +]
    • Changedget_price_index_history2 fields changed
      • changedInput schema / properties / provider_type / description
        Previous value: -"'all' (default) = the combined whole-economy series, or the standalone 'agent' / 'mcp' / 'payg' (pay-as-you-go, engine v2) series (own base 100)."New value: +"'all' (default) = the combined whole-economy series, or the standalone 'agent' / 'mcp' / 'subscription' / 'hybrid' / 'one_off' / 'payg' = the billing-LENS series (they cut across provider/mcp; a rate sits in one delivery type AND one lens; own base 100) — 'payg' (pay-as-you-go, engine v2) series (own base 100)."
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "agent",
        -  "mcp",
        -  "payg"
        -]New value: +[
        +  "all",
        +  "agent",
        +  "mcp",
        +  "subscription",
        +  "hybrid",
        +  "one_off",
        +  "payg"
        +]
  2. 3 tool updates
    • Changedcompare_providers4 fields changed
      • addedInput schema / properties / optional
        Added value: +{
        +  "description": "Optional preferences: reported in service_view but never gating.",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / requirements
        Added value: +{
        +  "description": "Optional. Mandatory requirements as plain phrases (e.g. \"sanctions screening\", \"documented MCP interface\"). Adds an ADDITIVE service_view: per-provider evidence matrix from stored first-party page text — supported / not_supported / unknown with the excerpt and observation date; unknown never means unsupported.",
        +  "items": {
        +    "type": "string"
        +  },
        +  "type": "array"
        +}
      • addedInput schema / properties / service
        Added value: +{
        +  "description": "Optional. A service recipe whose fixed requirements and evidence fields are applied (see service_view.requirements).",
        +  "enum": [
        +    "supplier-verification",
        +    "contract-review",
        +    "developer-capabilities",
        +    "cheaper-alternatives"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "Optional. The buyer's job in plain words; used as the service_view heading.",
        +  "type": "string"
        +}
    • Changedget_price_index2 fields changed
      • changedInput schema / properties / provider_type / description
        Previous value: -"'all' (default) = the combined whole-economy index; 'agent' or 'mcp' = the standalone index over just that delivery type (own base 100). Agents and MCPs price and move differently, so an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index."New value: +"'all' (default) = the combined whole-economy index; 'agent' or 'mcp' = the standalone index over just that delivery type (own base 100); 'payg' = the standalone pay-as-you-go sub-index (usage rates purchasable without a subscription, engine v2). Agents and MCPs price and move differently, so an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index."
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "agent",
        -  "mcp"
        -]New value: +[
        +  "all",
        +  "agent",
        +  "mcp",
        +  "payg"
        +]
    • Changedget_price_index_history2 fields changed
      • changedInput schema / properties / provider_type / description
        Previous value: -"'all' (default) = the combined whole-economy series, or the standalone 'agent' / 'mcp' series (own base 100)."New value: +"'all' (default) = the combined whole-economy series, or the standalone 'agent' / 'mcp' / 'payg' (pay-as-you-go, engine v2) series (own base 100)."
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "agent",
        -  "mcp"
        -]New value: +[
        +  "all",
        +  "agent",
        +  "mcp",
        +  "payg"
        +]
  3. 5 tool updates
    • Changedget_price_index2 fields changed
      • addedInput schema / properties / benchmark_id
        Added value: +{
        +  "description": "Optional: a private custom benchmark token (cb_…) from create_custom_benchmark — returns that cohort's current index instead of the economy index. Cannot be combined with niche.",
        +  "type": "string"
        +}
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "provider",
        -  "mcp"
        -]New value: +[
        +  "all",
        +  "agent",
        +  "mcp"
        +]
    • Changedget_price_index_history2 fields changed
      • addedInput schema / properties / benchmark_id
        Added value: +{
        +  "description": "Optional: a private custom benchmark token (cb_…) — returns that cohort's dated series instead of the economy series. Cannot be combined with niche.",
        +  "type": "string"
        +}
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "provider",
        -  "mcp"
        -]New value: +[
        +  "all",
        +  "agent",
        +  "mcp"
        +]
    • Changedprice_benchmark3 fields changed
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "provider",
        -  "mcp",
        -  "api",
        -  "all"
        -]New value: +[
        +  "agent",
        +  "mcp",
        +  "api",
        +  "all"
        +]
      • addedInput schema / properties / query
        Added value: +{
        +  "description": "The capability to benchmark, in natural language, e.g. 'AI code review' or 'supplier invoice reconciliation'. Preferred input.",
        +  "type": "string"
        +}
      • addedInput schema / properties / task
        Added value: +{
        +  "description": "Alias of `query` (same text).",
        +  "type": "string"
        +}
    • Changedresearch_capability1 field changed
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "auto",
        -  "provider",
        -  "mcp",
        -  "api",
        -  "any"
        -]New value: +[
        +  "auto",
        +  "agent",
        +  "mcp",
        +  "api",
        +  "any"
        +]
    • Changedsearch_providers1 field changed
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "provider",
        -  "mcp",
        -  "api",
        -  "platform",
        -  "infrastructure"
        -]New value: +[
        +  "agent",
        +  "mcp",
        +  "api",
        +  "platform",
        +  "infrastructure"
        +]
  4. 2 tool updates
    • Changedfind_market4 fields changed
      • removedInput schema / anyOf
        Removed value: -[
        -  {
        -    "required": [
        -      "task"
        -    ]
        -  },
        -  {
        -    "required": [
        -      "query"
        -    ]
        -  }
        -]
      • changedInput schema / properties / query / description
        Previous value: -"Alias for task (back-compat) — a natural-language task, capability or service"New value: +"Alias for task (back-compat only) — prefer task"
      • changedInput schema / properties / task / description
        Previous value: -"A natural-language task, capability or service, e.g. 'reconcile supplier invoices'"New value: +"REQUIRED. A natural-language task, capability or service in plain words, e.g. 'reconcile supplier invoices'. Call as {\"task\": \"…\"}; an empty call returns status needs_input with an example and searches nothing."
      • addedInput schema / required
        Added value: +[
        +  "task"
        +]
    • Changedsearch_providers2 fields changed
      • changedInput schema / properties / query / description
        Previous value: -"Free-text capability query, e.g. 'customer support provider with Zendesk integration'"New value: +"REQUIRED. Free-text description of the product job you are buying, e.g. 'customer support provider with Zendesk integration' or 'supplier invoice reconciliation'. Not for questions about Agentery itself (upvotes, endpoints): those return status unsupported_request."
      • addedInput schema / required
        Added value: +[
        +  "query"
        +]
  5. 1 tool update
    • Changedsearch_providers1 field changed
      • addedInput schema / properties / provider_type
        Added value: +{
        +  "description": "Filter to one provider type — the audited classification dimension (same as the website type chips and AEPI facets); result labels always match this filter. Omit for all types.",
        +  "enum": [
        +    "provider",
        +    "mcp",
        +    "api",
        +    "platform",
        +    "infrastructure"
        +  ],
        +  "type": "string"
        +}
  6. 5 tool updates
    • Changedcreate_custom_benchmark2 fields changed
      • changedInput schema / properties / base_niche / description
        Previous value: -"Mode B: canonical niche slug to fork"New value: +"Mode B: slug of the canonical market to fork"
      • changedInput schema / properties / remove / description
        Previous value: -"Mode B: members to drop from the forked niche"New value: +"Mode B: members to drop from the forked market"
    • Changedmarket_gaps1 field changed
      • changedInput schema / properties / limit / description
        Previous value: -"Max niches (1-50, default 15)"New value: +"Max gaps (1-50, default 15)"
    • Changedmarket_report1 field changed
      • changedInput schema / properties / query / description
        Previous value: -"Natural-language capability or market, e.g. 'customer support chatbot' or 'ai phishing detection'. Alias: niche."New value: +"Natural-language capability or market, e.g. 'customer support chatbot' or 'ai phishing detection'."
    • Changedprice_benchmark2 fields changed
      • changedInput schema / properties / niche / description
        Previous value: -"Niche slug (or an approximate slug / natural-language name — it is resolved to the canonical niche), e.g. 'contract-review-automation'"New value: +"Legacy alias of `query`, kept for older clients — prefer `query`."
      • changedInput schema / properties / sector / description
        Previous value: -"Sector name, e.g. 'legal' (ignored if niche is given)"New value: +"Sector name, e.g. 'legal' (ignored if a query is given)"
    • Changedresearch_capability1 field changed
      • changedInput schema / properties / provider_type / description
        Previous value: -"Preferred delivery type. 'auto' (default) infers from the task; note 'AI agent' phrasing is treated as generic (neutral), not an agent-only filter. When a type is explicit (mcp/api/agent) matching providers are SOFT-RANKED to the top and the rest are kept as clearly-labelled cross_type_alternative entries — never hard-filtered (no zero-result cliff), and the functional niche is never changed. Every provider is labelled with provider_type (public values: agent | mcp | api | unknown) + type_match_score; provider_type{type_rank_boost_applied, boosted_provider_type, result_counts_by_type} is returned."New value: +"Preferred delivery type. 'auto' (default) infers from the task; note 'AI agent' phrasing is treated as generic (neutral), not an agent-only filter. When a type is explicit (mcp/api/agent) matching providers are SOFT-RANKED to the top and the rest are kept as clearly-labelled cross_type_alternative entries — never hard-filtered (no zero-result cliff), and the functional match is never changed. Every provider is labelled with provider_type (public values: agent | mcp | api | unknown) + type_match_score; provider_type{type_rank_boost_applied, boosted_provider_type, result_counts_by_type} is returned."
  7. 6 tool updates
    • Addedfind_market
    • Removedfind_niche
    • Changedget_price_index4 fields changed
      • removedInput schema / properties / niche
        Removed value: -{
        -  "description": "For scope 'niche': a canonical niche slug (e.g. 'customer-support-tier1') or a natural-language task/query (e.g. 'customer support providers'), resolved by the niche resolver.",
        -  "type": "string"
        -}
      • changedInput schema / properties / provider_type / description
        Previous value: -"Delivery-type scope for scope 'aepi': 'all' (default) = the combined index; 'agent' or 'mcp' = the standalone index over just that provider type (own base 100). Agents and MCPs price and move differently, so an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index. Per-type NICHE indices are a Phase-2 follow-up; on scope 'niche' this is acknowledged in the response, not silently applied."New value: +"'all' (default) = the combined whole-economy index; 'agent' or 'mcp' = the standalone index over just that delivery type (own base 100). Agents and MCPs price and move differently, so an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index."
      • changedInput schema / properties / response_mode / description
        Previous value: -"'summary' (default) or 'full' (adds exact sub-0.01% moves, per-tier niche detail and repricing counts)."New value: +"'summary' (default) or 'full' (adds exact sub-0.01% moves and repricing counts)."
      • removedInput schema / properties / scope
        Removed value: -{
        -  "description": "'aepi' (default) = whole-economy headline + the four buyer tiers; 'niche' = one niche's index.",
        -  "enum": [
        -    "aepi",
        -    "niche"
        -  ],
        -  "type": "string"
        -}
    • Changedget_price_index_history4 fields changed
      • removedInput schema / properties / niche
        Removed value: -{
        -  "description": "For scope 'niche': a canonical niche slug or natural-language query.",
        -  "type": "string"
        -}
      • changedInput schema / properties / provider_type / description
        Previous value: -"Delivery-type scope for scope 'aepi': 'all' (default), or the standalone 'agent' / 'mcp' series (own base 100). Per-type niche history is a Phase-2 follow-up; acknowledged, not silently applied, on scope 'niche'."New value: +"'all' (default) = the combined whole-economy series, or the standalone 'agent' / 'mcp' series (own base 100)."
      • removedInput schema / properties / scope
        Removed value: -{
        -  "description": "'aepi' (default) or 'niche'.",
        -  "enum": [
        -    "aepi",
        -    "niche"
        -  ],
        -  "type": "string"
        -}
      • changedInput schema / properties / tier / description
        Previous value: -"For scope 'aepi', return one buyer tier's series ('team' = Team/SME). Default 'all' = the headline series."New value: +"Return one buyer tier's series ('team' = Team/SME). Default 'all' = the headline series."
    • Addedmarket_report
    • Removedniche_report
  8. 1 tool update
    • Changedprice_benchmark4 fields changed
      • addedInput schema / properties / buyer_tier
        Added value: +{
        +  "description": "Buyer tier being priced. Set when the user describes who is buying (an individual, a professional, a team/SME, or an enterprise). An individual licence must never be represented by the SME or enterprise price.",
        +  "enum": [
        +    "individual",
        +    "pro",
        +    "team_sme",
        +    "enterprise",
        +    "all"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / pricing_unit
        Added value: +{
        +  "description": "Optional pricing unit to hold constant (e.g. 'flat', 'per_seat', 'per_agent'). Incompatible units are never combined.",
        +  "type": "string"
        +}
      • addedInput schema / properties / provider_type
        Added value: +{
        +  "description": "Delivery type being priced. Set when the user says provider, MCP or API. Omit (or 'all') to get the per-type matrix instead of a blended figure. 'api' is recognised but not yet a separate commercial cohort (folded into provider).",
        +  "enum": [
        +    "provider",
        +    "mcp",
        +    "api",
        +    "all"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / response_mode
        Added value: +{
        +  "description": "'summary' (default): compact per-type/per-tier benchmark matrix. 'full': also returns the deprecated blended legacy block + AEPI index.",
        +  "enum": [
        +    "summary",
        +    "full"
        +  ],
        +  "type": "string"
        +}
  9. 12 tool updates
    • Removedcompare_agents
    • Removedget_agent
    • Removedget_agent_profile
    • Changedget_price_index2 fields changed
      • changedInput schema / properties / niche / description
        Previous value: -"For scope 'niche': a canonical niche slug (e.g. 'customer-support-tier1') or a natural-language task/query (e.g. 'customer support agents'), resolved by the niche resolver."New value: +"For scope 'niche': a canonical niche slug (e.g. 'customer-support-tier1') or a natural-language task/query (e.g. 'customer support providers'), resolved by the niche resolver."
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "agent",
        -  "mcp"
        -]New value: +[
        +  "all",
        +  "provider",
        +  "mcp"
        +]
    • Changedget_price_index_history1 field changed
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "all",
        -  "agent",
        -  "mcp"
        -]New value: +[
        +  "all",
        +  "provider",
        +  "mcp"
        +]
    • Changedget_provider1 field changed
      • changedInput schema / properties / handle / description
        Previous value: -"Agent handle, e.g. 'openhands'"New value: +"Provider handle, e.g. 'openhands'"
    • Removedrank_agents_for_workflow
    • Changedreport_outcome3 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"Handle of the agent you used, as returned by search_agents/get_agent_profile/compare_agents"New value: +"Handle of the provider you used, as returned by search_providers/get_provider_profile/compare_providers"
      • changedInput schema / properties / latency_ms / description
        Previous value: -"Optional end-to-end latency of the agent in milliseconds"New value: +"Optional end-to-end latency of the provider in milliseconds"
      • changedInput schema / properties / outcome / description
        Previous value: -"Did the agent accomplish the task you hired it for?"New value: +"Did the provider accomplish the task you hired it for?"
    • Changedresearch_capability2 fields changed
      • changedInput schema / properties / provider_type / enum
        Previous value: -[
        -  "auto",
        -  "agent",
        -  "mcp",
        -  "api",
        -  "any"
        -]New value: +[
        +  "auto",
        +  "provider",
        +  "mcp",
        +  "api",
        +  "any"
        +]
      • changedInput schema / properties / task / description
        Previous value: -"Required — the natural-language capability/task, e.g. 'reconcile supplier invoices' or 'litigation-analysis agent'. Aliases: query, q."New value: +"Required — the natural-language capability/task, e.g. 'reconcile supplier invoices' or 'litigation-analysis provider'. Aliases: query, q."
    • Removedsearch_agents
    • Changedsearch_providers2 fields changed
      • changedInput schema / properties / query / description
        Previous value: -"Free-text capability query, e.g. 'customer support agent with Zendesk integration'"New value: +"Free-text capability query, e.g. 'customer support provider with Zendesk integration'"
      • changedInput schema / properties / require_public_price / description
        Previous value: -"Only return agents with an observed public price (default false)"New value: +"Only return providers with an observed public price (default false)"
    • Changedsuggest_alternatives2 fields changed
      • changedInput schema / properties / agent_id / description
        Previous value: -"Handle of the agent to find substitutes for, e.g. 'openhands'"New value: +"Handle of the provider to find substitutes for, e.g. 'openhands'"
      • changedInput schema / properties / cheaper_only / description
        Previous value: -"Only keep alternatives priced below the subject's lowest monthly price. Free/freemium agents always qualify; agents with no observed price are excluded. Default false."New value: +"Only keep alternatives priced below the subject's lowest monthly price. Free/freemium providers always qualify; providers with no observed price are excluded. Default false."
  10. 3 tool updates
    • Changedcompare_providers1 field changed
      • changedInput schema / properties / provider_ids / description
        Previous value: -"2-6 agent handles from search_agents/market_gaps, e.g. [\"openhands\",\"lexaclaw\"]"New value: +"2-6 provider handles from search_providers/market_gaps, e.g. [\"openhands\",\"lexaclaw\"]"
    • Changedget_provider_profile1 field changed
      • changedInput schema / properties / provider_id / description
        Previous value: -"The provider_id/handle returned by search_agents or compare_agents"New value: +"The provider_id/handle returned by search_providers or compare_providers"
    • Changedsearch_providers3 fields changed
      • changedInput schema / properties / billing / description
        Previous value: -"Only agents with one of these observed billing models, e.g. [\"free\",\"freemium\",\"subscription\",\"usage\"]"New value: +"Only providers with one of these observed billing models, e.g. [\"free\",\"freemium\",\"subscription\",\"usage\"]"
      • changedInput schema / properties / max_monthly_usd / description
        Previous value: -"Drop agents whose observed lowest paid tier exceeds this (USD/month). Agents with no observed public price still pass unless require_public_price is true."New value: +"Drop providers whose observed lowest paid tier exceeds this (USD/month). Providers with no observed public price still pass unless require_public_price is true."
      • changedInput schema / properties / sort / description
        Previous value: -"match (default) or price_asc (cheapest observed price first; unpriced agents last)"New value: +"match (default) or price_asc (cheapest observed price first; unpriced providers last)"
  11. 5 tool updates
    • Addedcompare_providers
    • Addedget_provider
    • Addedget_provider_profile
    • Addedrank_providers_for_workflow
    • Addedsearch_providers
  12. 2 tool updates
    • Changedget_price_index1 field changed
      • addedInput schema / properties / provider_type
        Added value: +{
        +  "description": "Delivery-type scope for scope 'aepi': 'all' (default) = the combined index; 'agent' or 'mcp' = the standalone index over just that provider type (own base 100). Agents and MCPs price and move differently, so an MCP buyer should read the 'mcp' index and an agent buyer the 'agent' index. Per-type NICHE indices are a Phase-2 follow-up; on scope 'niche' this is acknowledged in the response, not silently applied.",
        +  "enum": [
        +    "all",
        +    "agent",
        +    "mcp"
        +  ],
        +  "type": "string"
        +}
    • Changedget_price_index_history1 field changed
      • addedInput schema / properties / provider_type
        Added value: +{
        +  "description": "Delivery-type scope for scope 'aepi': 'all' (default), or the standalone 'agent' / 'mcp' series (own base 100). Per-type niche history is a Phase-2 follow-up; acknowledged, not silently applied, on scope 'niche'.",
        +  "enum": [
        +    "all",
        +    "agent",
        +    "mcp"
        +  ],
        +  "type": "string"
        +}
  13. 4 tool updates
    • Addedcreate_custom_benchmark
    • Addeddelete_custom_benchmark
    • Addedget_custom_benchmark
    • Addedupdate_custom_benchmark

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources