Skip to main content
Glama

compare_cohorts

INFLUENCE — two markets side by side on one rubric: "is US payments further along than UK banking?" Returns both distributions plus the deltas on score, agent readiness, every facet and every adoption rate. A question a written report cannot answer, because a report only ever covers one market. CHECK depth_confounded FIRST: when the higher-scoring cohort is also the one we enriched more deeply, the delta reflects our own coverage as much as the markets, and depth_note says by how much.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
aYesFirst cohort as <kind>:<slug>, e.g. "tag:payments".
bYesSecond cohort as <kind>:<slug>, e.g. "industry:banking".
contextNoOptional: why you are asking. One sentence — the task you are trying to complete, or what you expect to get back. Never included in the answer and never used to rank; it is read only when a result turns out to be wrong, which is when knowing the intent is what makes the report actionable.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden, and it does so exceptionally: it warns that when the higher-scoring cohort is also the more deeply enriched one, 'the delta reflects our own coverage as much as the markets,' and that depth_note quantifies the effect. It also discloses the return payload composition (distributions, deltas, depth_confounded). This prevents a serious misinterpretation an agent would otherwise make.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all earning their place: purpose+payload, usage context, and the critical interpretation warning. The 'CHECK depth_confounded FIRST' instruction draws attention despite being placed last rather than front-loaded, and the cryptic 'INFLUENCE' opener adds little functional value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the job of describing the return values: distributions, deltas across score/readiness/facets/adoption, and the depth_confounded/depth_note caveat fields. For a two-string-parameter comparison tool this is nearly sufficient; only a slight vagueness around 'every facet' and no explicit read-only statement keep it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the <kind>:<slug> format with examples and the context parameter's behavior. The description adds only a semantic frame — that the two parameters are 'markets' being compared — and an example question that mirrors the schema examples. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action — comparing 'two markets side by side on one rubric' — and specifies the output: 'both distributions plus the deltas on score, agent readiness, every facet and every adoption rate.' The concrete example question ('is US payments further along than UK banking?') anchors the purpose, and the tool is clearly distinguished from siblings like cohort_scores, get_cohort, and compare_providers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context: this answers cross-market comparison questions 'a written report cannot answer, because a report only ever covers one market.' It also instructs the agent to 'CHECK depth_confounded FIRST' to decide whether the delta is interpretable. However, it never names sibling alternatives explicitly or states when-not-to-use conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.1/5.0
Disambiguation3/5

Most tools are clearly separated by artifact type or resource (find_mcp vs find_openapi vs get_provider vs get_api), but the sheer volume creates some genuinely confusable clusters: apis_io_search vs find_apis vs find_artifacts, and insights_adoption vs insights_dimensions vs find_company_insights. Several readiness-related tools (what_can_i_fix, simulate_fixes, readiness_gates) also share a conceptual boundary, though their descriptions do help.

Naming Consistency3/5

The dominant patterns (find_*, get_*, cohort_*, compare_*) are consistent and predictable, but the set mixes in irregular names like apis_io_search, tag_group_tags, what_can_i_fix, whats_changed, and resolve. These deviations are readable but break the otherwise regular verb_noun convention.

Tool Count2/5

106 tools is far beyond the typical well-scoped server and will impose a heavy selection burden on agents. The server covers a genuinely broad domain (catalog search, ratings, cohorts, agent readiness, lists, exports, feedback), so the count is defensible in scope, but it is still too many to navigate efficiently.

Completeness5/5

The surface is remarkably complete: search and browse, single-entity detail, comparisons, cohort analytics, agent-readiness assessment, saved searches, list management, feedback/correction flows, and full dataset exports are all covered. There are no obvious dead ends, and even minor operations like re-running saved searches or simulating fixes are present.

Resources