Skip to main content
Glama

aggregate

Compute a grouped aggregate over one topic — the aggregation-first path. agg is one of avg/sum/min/max/count/weighted_rate; metric is a metric column (DEFAULTS to the topic key_metric; ignored for count); group_by is a list of grain columns (year, FK codes like district_code or county_fips, or categoricals — see describe_dataset). weighted_rate computes a true population-weighted SUM(numerator)/SUM(denominator) for a rate key metric (when the contract declares the components) — prefer it over avg for a rate across multiple places/years, since avg means the per-row rates and ignores population. Supports the same filters / year / year_min-year_max / detail as query_dataset, plus order_by+order for top-N (order_by 'value' for the aggregated column; NULL cells sort LAST in either direction). Returns one small row per group with <metric>_<agg> (or row_count) plus coverage diagnostics (input_rows / non-null counts) so suppression is visible; aggregation_scope flags whether rows are source-published at this grain or recomputed from a finer detail (prefer source-published — see the advisory). A sum/weighted_rate ACROSS a column the contract declares non_additive (overlapping rows, e.g. county rows that count one application in several counties) still returns, but carries a top-level non_additive block and a leading non_additive_sum advisory: that figure is NOT a total — group_by the column or use the published total row instead. Aggregates SKIP NULLs and NULL means SUPPRESSED not zero. No raw SQL: all identifiers are contract-allowlisted.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
aggNoavg
yearNo
limitNo
orderNoasc
topicYes
detailNo
metricNo
offsetNo
filtersNo
group_byNo
order_byNo
year_maxNo
year_minNo
main_topicNoeducation

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly. It discloses NULL handling ('Aggregates SKIP NULLs and NULL means SUPPRESSED not zero'), non-additive behavior with a top-level block, aggregation_scope flags, suppression diagnostics, missing-value semantics, and security constraints (allowlisted identifiers, no raw SQL). This is far beyond basic expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence adds value: it covers purpose, parameter nuances, output shape, edge cases, and security. It is front-loaded with the core definition, then parameters, then caveats. It could be restructured into bullets for scannability, but there is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter tool with no annotations and no schema descriptions, this description is exceptionally complete. It explains what the output looks like, how diagnostics work, when aggregation_scope matters, how non-additive results are flagged, and how NULLs behave. It also covers the security model. There is a clear mental model an agent needs to call the tool correctly, and the description provides it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the key parameters: agg (allowed values), metric (defaults to key_metric, ignored for count), group_by (grain types), order_by/order (top-N, NULL sort), and the meaning of weighted_rate. It omits limit/offset and main_topic, and only references 'same as query_dataset' for filters/year/detail rather than explaining them, but the coverage of the most semantically risky parameters is strong.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear, specific statement: 'Compute a grouped aggregate over one topic — the aggregation-first path.' It names the verb ('compute'), the resource ('one topic'), and the aggregation focus. It also differentiates from siblings by referencing query_dataset and describe_dataset, and by explaining how weighted_rate differs from avg.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit guidance for choosing weighted_rate over avg ('prefer it over avg for a rate across multiple places/years'), for handling non-additive columns ('group_by the column or use the published total row instead'), and for discovering valid group_by grains ('see describe_dataset'). It does not explicitly state when to use query_dataset instead of aggregate, leaving that to inference, but it does reference query_dataset for shared parameters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources