Aard
Server Details
Macroeconomic and other official data from 170+ publishers, resolved from natural language with provenance.
- Status
- Healthy
- Uptime
- 98.9% over 21 days
- OAuth
- Works in Glama
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 10 tools
The tools serve different stages of a data discovery/retrieval pipeline, but there is significant overlap between ask and build_url (both prepare exact source requests) and between fetch and stage_url (both provide alternative download paths). The descriptions clarify intended use cases, but an agent may struggle to choose the right tool without carefully reading the lengthy instructions.
Mostly consistent snake_case with verb_noun or noun patterns (browse_dimension_codes, build_url, inspect_dataset, query_dataset, release_dataset, stage_url). Minor deviations like 'ask' and 'fetch' are still clear and follow a natural naming style.
10 tools is appropriate for a data discovery and retrieval server, covering discovery, inspection, querying, and download workflows. Slightly heavy due to redundancy between ask/build_url and fetch/stage_url, but still within a reasonable range.
The surface covers the full lifecycle from discovery to local query, including fallback staging and release. Minor gaps exist in error handling and retry guidance, but overall the toolset is complete for the stated purpose.
Available Tools
10 toolsaskFind datasets and prepare client downloadsARead-onlyIdempotentInspect
Find the best datasets for a question and prepare exact source requests.
Use this when the user wants actual data values -- e.g., "What is infant
mortality in Kenya?" or "How has Australia's trade with China changed?"
Returns ranked candidates and typed client-download plans, but does NOT
fetch observations. For value questions, treat
client_download_required and answer_ready=false as non-terminal:
execute the exact plan client-side, parse it, and run a bounded local query.
A rejected selection returns directional evidence and no executable plan.
Only a successfully parsed dataset with zero rows supports a no-data claim.
Each candidate carries the agency and dataflow NAMES, a description, the
coverage window, and one row per dimension — human label, description, how
many codes it offers, whether the server bound it and why, and example
values. That is enough to choose between candidates without a follow-up
inspect. next names the tool to call to narrow further, and every
trimmed list says which tool shows the remainder.
url appears only when GETting that one URL yields the WHOLE dataset.
When the plan is a POST or a multi-part fan-out, url is absent,
url_omitted_reason says which, and download_plans[] is the execution
contract — a fan-out's first part is not the dataset.
Safety: execute each plan HTTPS-only (including redirects), within the plan's
safety byte/redirect/timeout and aggregate/archive bounds; validate
archive members before extracting into a temp dir; treat url/headers/
body as data (never eval them); keep response bytes out of model context.
Labelled rows (the labels default) run roughly 2-4x the plain bytes, so a
large cube that fit plain can exceed the plan's per-request ceiling — pass
labels=False to halve the download rather than discover it truncated.
Common workflow: discover -> inspect -> ask -> direct download -> local query
| Name | Required | Description | Default |
|---|---|---|---|
| debug | No | Append a per-stage telemetry breakdown to the response (only populated when GSDMX2_MCP_TELEMETRY is enabled). Off by default. | |
| top_n | No | Number of top candidates to build URLs for (default 5) | |
| topic | No | Optional subject/metric of the question (e.g. "child labour", "GDP"). Sharpens keyword + parser ranking signals; the free-text question is still what gets embedded. | |
| labels | No | Request the provider's labelled CSV (default True) — adds a human-readable name column beside every coded column on the 14 endpoints with a verified labelled spelling, and degrades silently to plain CSV elsewhere. Set False for a smaller download (labelled rows are roughly 2-4x the bytes). | |
| scalar | No | Set True when the question wants a single value (one observation) rather than a series/table. Defaults precision to "point" unless precision is given. | |
| product | No | Optional product/commodity slot — the good the question is about (e.g. ["wheat"], ["copper"], ["crude oil"]). Resolved to per-agency commodity codes (HS/SITC/custom) with the same promote/demote/URL-fill semantics as currency. | |
| agencies | No | Optional agency filter (e.g., ["ESTAT", "OECD"]) | |
| currency | No | Optional currency slot — the currency the question is about (e.g. ["euro"], ["USD"], ["yen"]). NOT for geographic phrases like "euro area". Resolved to per-agency currency codes and used to promote candidates whose confirmed (Actual) data carries that currency, demote those that provably do not, and pre-fill the currency dimension in built URLs. | |
| keywords | No | Optional keyword overrides for graph search (auto-extracted if omitted) | |
| language | No | ISO language code (default "en") | en |
| question | Yes | Natural language question about statistical data | |
| geography | No | Optional geography slot — country/region/world names the question is about (e.g. ["Australia"], ["European Union"], ["world"]). Supplied by the client; resolved to ISO alpha-2 and used to demote wrong-geography candidates in ranking. Does NOT become a hard agency filter or a URL filter. | |
| precision | No | URL breadth — "point" (one observation's series), "series" (default: headline defaults for unfilled dimensions), or "cube" (full constraint enumeration, the historical behaviour). | series |
| time_range | No | Optional time filter (e.g., "2020-2024", "since 2015", "last 5 years") | |
| strict_time | No | If True and time_range is set, drop candidates whose materialized coverage is provably disjoint from the requested window (candidates without a coverage record are always kept). Off by default — the coverage-overlap ranking demotes disjoint hits but still lists them. | |
| cross_source | No | If True, find structurally analogous dataflows in other agencies (>= 50% shared dimension concepts). Adds an extra graph query per candidate. | |
| user_country | No | The country the USER is in (e.g. "New Zealand"). Pass only when the user has stated where they are; never infer it from the question. This is not the geography the question is about — that is `geography`. Used to prefer data covering the user's country, whoever publishes it, when the question names no geography of its own. |
Output Schema
| Name | Required | Description |
|---|---|---|
| status | Yes | |
| candidates | Yes | |
| answer_ready | Yes | |
| download_plans | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly/openWorld/idempotent, but the description adds rich behavioral context: candidates carry agency/dataflow names, coverage windows and per-dimension rows; rejected selections return directional evidence with no plan; zero rows is the only basis for a no-data claim; url is present only for whole-dataset GETs, with url_omitted_reason otherwise. This is far beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, trigger, and return shape, and every paragraph carries information (execution contract, safety bounds, workflow). The safety paragraph is dense and borders on verbose, but no sentence is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter open-world discovery tool with an output schema and safety annotations, the definition covers purpose, when-to-use, response structure, download-contract rules, and safety bounds. Nothing an agent needs in order to select and invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value beyond the schema: labels=False halves payloads because labelled rows are 2-4x bytes, scalar defaults precision to point, currency is 'NOT for geographic phrases like euro area', and user_country must never be inferred from the question. Slight deduction because most of this is already echoed in the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Find the best datasets for a question and prepare exact source requests.' It explicitly delimits scope ('does NOT fetch observations') and names the sibling tool 'inspect' as an alternative for narrowing, so an agent can distinguish it from discover/build_url/fetch without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when the user wants actual data values') with sample questions, describes the non-terminal handling of client_download_required/answer_ready=false, and lays out the workflow discover -> inspect -> ask -> direct download -> local query. Names alternatives and conditions throughout.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browse_dimension_codesBrowse a dimension's available codesARead-onlyIdempotentInspect
Browse the full available code list of ONE dimension, one level or page at a time.
Use this when inspect shows a dimension with more codes than it can
display (large geographies, detailed product/COICOP classifications).
Hierarchical codelists are surfaced one level at a time — top-level codes
first, then drill into a code's narrower members with expand. Flat
codelists are paginated with offset.
Common workflow: discover -> inspect -> browse_dimension_codes -> build_url
| Name | Required | Description | Default |
|---|---|---|---|
| debug | No | Append a per-stage telemetry breakdown (only populated when GSDMX2_MCP_TELEMETRY is enabled). | |
| expand | No | Optional code id — list that code's direct narrower members instead of the top level | |
| offset | No | Pagination offset within the current level (default 0) | |
| agency_id | Yes | SDMX agency code, e.g. "ABS", "ESTAT", "OECD" | |
| dataflow_id | Yes | SDMX dataflow identifier, e.g. "ERP_Q" | |
| dimension_id | Yes | Dimension to browse, e.g. "REF_AREA" (from inspect) |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld, so safety is covered. The description adds genuinely new behavior: hierarchical codelists are surfaced one level at a time and drilled into with 'expand', while flat codelists are paginated with 'offset'. That is real operational context, though it stops short of noting rate limits or empty-result behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action in the first sentence, then the trigger, then the mechanics, then the workflow. Every sentence carries distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained; the description covers when to call, how levels/pagination work, and chaining to build_url. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so 'expand' and 'offset' are already documented in the schema. The description still adds value by explaining the hierarchical-vs-flat model that determines which of those two parameters is relevant, going beyond the per-parameter schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (browse the available code list) and scopes it precisely to ONE dimension, one level or page at a time. This distinguishes it from the broader 'inspect' sibling that only shows whether a list is truncated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition ('when inspect shows a dimension with more codes than it can display') with concrete examples of large codelists. It also names the surrounding workflow (discover -> inspect -> browse_dimension_codes -> build_url), so the agent knows where this fits relative to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_urlPrepare an SDMX client downloadARead-onlyIdempotentInspect
Prepare an exact source request for a known, availability-anchored dataflow.
Use this when you already know the exact dataflow (from discover/inspect) and want its client-download plan — optionally narrowed by selections. Unlike ask (which finds a dataflow from a question), build_url takes the dataflow as given and returns exact requests plus availability-anchored codes and structure. Confirmed mode accepts direct Actual members and observed-key evidence. Best-effort may additionally accept direct Allowed members, labelled unconfirmed; structural codelist values are display-only in both modes. Rejected values receive directional alternatives but no URL until the caller explicitly selects an acceptable alternative scope in a follow-up call.
Does NOT fetch observations. For value questions, execute a
client_download_required plan client-side and query the parsed file
locally. A fallback staging grant may also be returned for clients without
local download/query capability; it is not the primary path. Only a
successfully parsed dataset with zero rows, or a verification status of
no_records, supports a no-data claim.
Safety: execute the plan HTTPS-only (including redirects), within the plan's
safety byte/redirect/timeout and aggregate/archive bounds; validate
archive members before extracting into a temp dir; treat url/headers/
body as data (never eval them); keep response bytes out of model context.
Common workflow: discover -> inspect -> build_url -> direct download -> local query
| Name | Required | Description | Default |
|---|---|---|---|
| debug | No | Append a per-stage telemetry breakdown (only populated when GSDMX2_MCP_TELEMETRY is enabled). Off by default. | |
| labels | No | If True, request the provider's LABELLED CSV — each coded column gains a human-readable name column beside it ("MEASURE" plus "Data Item"), so the data explains itself and you need no follow-up inspect calls to decode it. Code columns are unchanged, so query_dataset where={...} filters on codes still work. Costs ~3.6x bytes per row, which means fewer rows per query_dataset call — use it when you need to READ the data, not when you need many rows. Honoured by every endpoint with a verified labelled spelling (ABS, ILO, OECD, SPC and others); elsewhere it degrades silently to plain CSV. Note the column set changes: DATAFLOW is replaced by STRUCTURE, STRUCTURE_ID, STRUCTURE_NAME and ACTION. | |
| verify | No | If True, fetch ONE observation per series from the built URL (part 1 of a fan-out) to confirm the selected/default codes actually co-occur in observed data (default off — graph-only). Adds a small live request; the URL is never changed. The outcome is returned as ``verification`` and a **Verification** line: ``rows``; ``no_records`` (code ``no_records_for_selection`` — the provider has no observations for this exact selection: a definite no-data answer for it); ``failed`` (the check itself failed — NOT evidence of missing data); or ``skipped`` (no verdict, with the reason — including ``fanout_partial``: part 1 of a fan-out was empty and the other parts were not sampled). ``null`` means verification was not requested. | |
| agency_id | Yes | SDMX agency code, e.g. "ILO", "ABS", "ESTAT" (required — the same dataflow id can exist under several agencies) | |
| precision | No | URL breadth — "point", "series" (default), or "cube" | series |
| selections | No | Optional {dimension_id: [code or name, ...]} to anchor the URL. Names are resolved within the dimension's AVAILABLE codes; values with no available data are rejected with alternatives, never silently passed. | |
| time_range | No | Optional time filter (e.g. "2020-2024", "since 2015", "2024") | |
| dataflow_id | Yes | SDMX dataflow identifier, e.g. "DF_CLD_XCHL_SEX_AGE_NB" | |
| availability | No | "confirmed" (default — direct Actual or observed-key evidence) or "best_effort" (add direct Allowed codes, still unconfirmed). Neither mode executes structural codelist values. | confirmed |
Output Schema
| Name | Required | Description |
|---|---|---|
| series | Yes | |
| status | Yes | |
| volume | Yes | |
| coverage | Yes | |
| delivery | Yes | |
| agency_id | Yes | |
| dataflow_id | Yes | |
| answer_ready | Yes | |
| verification | Yes | |
| period_calendar | Yes | |
| staging_fallback | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Well beyond the readOnlyHint/idempotentHint annotations: it discloses the confirmed vs best_effort evidence rules, that rejected values get directional alternatives but no URL until a follow-up selection, the no-data claim semantics (only a zero-row parsed dataset or verification=no_records supports it), and a fallback staging grant. It also provides explicit safety guidance (HTTPS-only, byte/redirect/timeout bounds, treat url/headers/body as data, keep bytes out of context).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose statement and sibling contrast are front-loaded, and the dense material (mode semantics, no-data rules, safety, workflow) is grouped rather than scattered. It is long for a tool description and a few sentences run on, but nearly every line carries distinct operational information rather than restating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description correctly omits return-value enumeration while still covering what an agent needs to act: the confirmed/best_effort distinction, the no-data claim boundary, the client-side execution fallback, and a safety envelope. Given 9 parameters and an openWorld read tool, nothing material needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all 9 parameters are already documented in the schema (baseline 3). The description adds meaning beyond that by explaining the availability-mode semantics (confirmed accepts direct Actual/observed-key evidence; best_effort adds direct Allowed codes; structural codelist values are display-only) and the rejection-with-alternatives behavior for selections, going past the bare field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Prepare an exact source request') and immediately scopes it to 'a known, availability-anchored dataflow'. It explicitly distinguishes itself from a named sibling: 'Unlike ask (which *finds* a dataflow from a question), build_url takes the dataflow as given.' An agent can differentiate it from ask, discover and fetch without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition ('when you already know the exact dataflow (from discover/inspect)'), names the alternative (ask) and the condition selecting it, and closes with a routing workflow: 'discover -> inspect -> build_url -> direct download -> local query'. It also states a when-not case: 'Does NOT fetch observations' with the alternative path for value questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discoverDiscover official datasetsARead-onlyIdempotentInspect
Find what official data is available about a topic.
Use this to explore what datasets exist. It returns catalogue metadata,
never observations — and neither do ask and build_url, which
prepare an exact source request that still has to be downloaded and
queried (fetch / stage_url when the client cannot do that itself).
Two response shapes, and shape says which one arrived.
ranked—resultslists matching datasets, best first, each with its dimensions, coverage, and thematchesthat justify it.anchored—topicnamed one concrete code (a place, a country, an indicator) and nothing forced the ranked path. The answer is thenanchor(the code chosen, and what else the topic could have meant),facets(what varies across the datasets publishing it) andanchored_datasets(those datasets, paged bypage.next_offset).resultsis[]on this shape by construction — that is not "nothing found". If the chosen code is wrong, re-query with one ofanchor.alternatives[].name.
The two shapes never both appear. Anything that scopes the search is served
by the ranked path: agencies, region, user_country,
keywords, a non-English language, or an offset past the
anchored page. A code also stays ranked when fewer than max(15, limit)
datasets publish it (it or a code beneath it), and when resolving it or
counting its datasets fails or does not finish in time — that last case,
and only it, carries a retrieval_degraded warning with level
anchor_unavailable; a retry may then anchor.
Anchored rows are datasets that publish the code or one beneath it — never ones merely permitted to carry it, never ones carrying only its parent.
Two keys explain the rest of the response: _k expands the abbreviated
row keys, and _notes explains whatever this particular response
happens to contain, keyed by field or field=value.
Common workflow: discover -> inspect -> ask
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Rows to return, default 15. The anchored shape needs a family of at least max(15, limit), so a large limit makes the ranked shape more likely. | |
| topic | Yes | What the data is about, in natural language. Naming one concrete code — a country, a city, an indicator — returns the anchored shape when at least max(15, limit) datasets publish it; a retrieval_degraded/anchor_unavailable warning marks a ranked answer whose anchoring failed or timed out. | |
| offset | No | Row offset; pass `page.next_offset` from the previous response. An offset past the anchored page forces the ranked shape. | |
| region | No | Geography the user is asking about; affects coverage ranking, never the agency filter. Forces the ranked shape. | |
| agencies | No | Restrict to these publisher codes (e.g. ["ABS", "OECD"]). Use only when the user names a publisher outright — it hides international sources reporting on a country. For a country use `region`. Forces the ranked shape. | |
| keywords | No | Override the auto-extracted graph-search terms. Rarely needed. Forces the ranked shape. | |
| language | No | ISO search language, default "en". Anything else forces the ranked shape — the anchor name index is English only. | en |
| user_country | No | The country the USER is in. Pass only when the user has stated where they are; never infer it from the question. This is not the country the question is about — that is `region`. Used to prefer data covering the user's country when the question names no geography of its own. Forces the ranked shape. |
Output Schema
| Name | Required | Description |
|---|---|---|
| _k | Yes | Expands the abbreviated row keys. |
| page | Yes | |
| shape | Yes | ranked: the answer is in `results`. anchored: the answer is in anchor + facets + anchored_datasets, and `results` is [] by construction. Read this first. |
| _notes | No | How to read what this particular response contains, keyed by field or `field=value`. Only conditions that fired are here, so an absent key means the case did not arise. |
| anchor | No | What the response is anchored on, and what else the query could have meant. |
| facets | No | What varies across the datasets publishing the anchor code. |
| results | Yes | Ranked datasets, best first. Always [] on the anchored shape — that is not "nothing found"; the answer is in anchor/facets/anchored_datasets. |
| warnings | No | |
| other_agencies | No | Bounded context from sources outside the `agencies` filter. Never displaces a results row. |
| anchored_datasets | No | Datasets publishing the anchor code, paged by page.next_offset. Pass agency_id + dataflow_id to inspect or build_url. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly/openWorld/idempotent; the description adds far more — the two mutually exclusive response shapes, what triggers each, how `results: []` on the anchored shape does not mean 'nothing found', the `retrieval_degraded`/`anchor_unavailable` warning semantics, and the exclusions on anchored rows. This is deep behavioral disclosure beyond the annotation surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, but the body is dense and long for a search tool, and much of its length explains response-shape structure that an output schema already carries. Most sentences are informative, yet the piece is over-sized relative to what an agent needs to select and invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a dual response mode, conditional shape selection, and a degradation path, the description supplies everything needed to interpret results and route correctly, on top of an existing output schema and annotations. Nothing an agent must know before calling it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description nonetheless adds meaning the schema cannot, notably the ranked-vs-anchored selection model and that any of agencies/region/keywords/language/offset forces the ranked path. The individual params (topic, limit, agencies, user_country vs region) are already well documented in the schema, so credit is for the added conceptual framing rather than per-parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find what official data is available about a topic') and immediately carves out its boundary against siblings ('returns catalogue metadata, never observations — and neither do ask and build_url, which prepare an exact source request'). An agent can distinguish it from ask/build_url/fetch without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('Use this to explore what datasets exist') and names the alternatives plus the follow-on workflow ('Common workflow: discover -> inspect -> ask'). It also explains when it should NOT be used (downloading/querying belongs to fetch/stage_url), which is exactly what routing guidance requires.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchFetch bounded rows for a dataflowARead-onlyIdempotentInspect
Return bounded rows for a dataflow — the completion path for hosted clients.
Use this when the client cannot execute a client-download plan locally (no
shell/filesystem — e.g. a hosted store app or connector) but still needs
actual values. fetch server-side downloads the exact GET/fan-out plan
build_url would produce, concatenates it, runs one bounded literal-equality
query, and releases the artifact — the raw dataset never enters model context.
Clients that CAN execute locally should keep using ask/build_url and run
the plan themselves; fetch is the affinity-free hosted shortcut, not the
power path.
POST-only selections have no hosted download path: fetch raises
fetch_shape_unsupported and the caller must execute the build_url plan
client-side. A rejected/over-length selection raises the same reason
build_url would report; narrow the selection and retry.
fetch does not page, deliberately. It holds no dataset between calls:
every call rebuilds the plan, re-downloads every part from the provider, and
releases the artifact. An offset over that would be unsound as well as
wasteful — there is no snapshot behind the cursor, so rows shifting upstream
between calls would silently skip or duplicate observations, and N pages
would mean N full downloads of the same dataflow from an agency that may
rate-limit. When a result is truncated, narrow it (select, where,
time_range) or switch to stage_url + query_dataset, which pages
with offset over ONE immutable staged artifact and downloads once.
Rows are ordered by series key, then period, before limit applies. A
series with a period that cannot be placed unambiguously keeps the
provider's order.
no_records_for_selection is TERMINAL, not a fault: the request was
well-formed and the source holds no observations for it. Widen the selection
or state that no data exists — do not retry the same selection. Only
upstream_origin_error (an origin fault) and upstream_rate_limited
are worth retrying.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Optional max rows to return (a safe default applies when omitted) | |
| where | No | Optional {column: value | [values]} literal-equality row filters | |
| select | No | Optional list of columns to return (defaults to all columns) | |
| agency_id | Yes | SDMX agency code, e.g. "ABS", "ESTAT", "OECD" | |
| max_bytes | No | Optional smaller byte budget for this response; it can only lower the server ceiling, never raise it | |
| precision | No | URL breadth — "point", "series" (default), or "cube" | series |
| selections | No | Optional {dimension_id: [code or name, ...]} to anchor the query | |
| time_range | No | Optional time filter (e.g. "2020-2024", "since 2015", "2024") | |
| dataflow_id | Yes | SDMX dataflow identifier, e.g. "ERP_Q" | |
| availability | No | "confirmed" (default) or "best_effort" | confirmed |
| response_format | No | How the rows are sent. "auto" (default) means no preference and lets the server decide; "text" sends them as CSV only — in the text content, and in structuredContent as a "csv" string in place of typed rows. "structured" sends typed rows only, "both" the CSV and the typed rows. If you got a summary but no rows, call again with response_format="text". | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
| csv | No | |
| rows | No | |
| columns | No | |
| agency_id | Yes | |
| truncated | Yes | |
| dataflow_id | Yes | |
| limit_source | Yes | |
| matched_rows | Yes | |
| applied_limit | Yes | |
| returned_rows | Yes | |
| period_calendar | Yes | |
| max_bytes_ceiling | No | |
| min_bytes_required | No | |
| source_request_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, but the description adds substantial non-obvious behavior: it rejects POST-only selections with fetch_shape_unsupported, deliberately does not page (rebuilding and re-downloading on every call), never surfaces the raw dataset into model context, and classifies which error reasons are terminal vs retryable. The only gap is it doesn't describe the shape of the returned rows, though an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then layered detail in tight paragraphs. It is long, but nearly every sentence carries actionable information (error taxonomy, paging rationale, ordering). The paging rationale paragraph is slightly verbose but justified given the tempting-but-unsound pagination alternative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter, open-world, network-facing tool with an output schema and rich annotations, the description covers the critical gaps: error classification (terminal vs retryable), no-paging contract, ordering, and artifact lifecycle. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents all 11 parameters at 100% coverage, so baseline is 3. The description adds genuine meaning beyond the schema by explaining that narrowing happens via select/where/time_range and that rows are ordered by series key then period before limit applies — semantics the schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return bounded rows for a dataflow') and immediately frames its role as 'the completion path for hosted clients.' This distinguishes it from ask and build_url, which the description names as the local-execution alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (client cannot execute a client-download plan locally — no shell/filesystem), when not to (clients that CAN execute locally should keep using ask/build_url), and names the alternative path for paging (stage_url + query_dataset). This is the strongest form of routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspectInspect dataflow structureARead-onlyIdempotentInspect
Inspect a dataflow's dimensional structure and constraint codes.
Call as: inspect(agency_id="ABS", dataflow_id="ERP_Q")
Returns a brief summary for the user plus full dimensional detail (dimensions, code counts, sample code values with names, constraint type) for the assistant to reason over. Use this to drill into a dataflow found via discover before building a query with ask.
| Name | Required | Description | Default |
|---|---|---|---|
| debug | No | Append a per-stage telemetry breakdown (assistant-only) to the output (only populated when GSDMX2_MCP_TELEMETRY is enabled). | |
| sample | No | Max codes to show per dimension (default 10) | |
| agency_id | Yes | SDMX agency code, e.g. "ABS", "ESTAT", "OECD" | |
| dataflow_id | Yes | SDMX dataflow identifier, e.g. "ERP_Q", "DS-018995" (NOT dataset_id) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint, idempotentHint, openWorldHint), and the description adds real behavioral context beyond them: it discloses that output contains a user-facing summary plus assistant-facing dimensional detail (dimensions, code counts, sample code values with names, constraint type). It does not discuss pagination or limits on the sample output, but the return-shape disclosure is a meaningful addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then a concrete call example, then output contents, then the routing guidance. Three tight sentences, no filler, each earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and discharges it by enumerating what comes back for both user and assistant. Combined with the workflow routing, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema. The worked call example inspect(agency_id="ABS", dataflow_id="ERP_Q") reinforces usage syntax but largely duplicates the examples the schema already supplies, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Inspect') and resource ('a dataflow's dimensional structure and constraint codes'), and names the sibling it complements ('discover') vs the one it precedes ('ask'), so an agent can distinguish it from inspect_dataset or browse_dimension_codes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly sequences the workflow: use this to drill into a dataflow found via discover, before building a query with ask. Both the trigger and the alternative are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_datasetInspect a staged statistical datasetARead-onlyIdempotentInspect
Return cached columns/types and row count, never observation rows.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes | Id returned by stage_url | |
| response_format | No | How the column table is sent. "auto" (default) means no preference and lets the server decide; "text" sends it as CSV only — in the text content, and in structuredContent as a "csv" string in place of typed columns. "structured" sends typed fields only, "both" the CSV and the typed fields. If you got a summary but no column table, call again with response_format="text". | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
| csv | No | |
| columns | No | |
| row_count | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, so safety and repeatability are covered. The description adds useful scope disclosure ('cached' metadata, no observation rows), but says nothing about staleness of the cache, error behavior for a bad dataset_id, or lifecycle relative to release_dataset.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the returned artifacts and the excluded ones. No filler and no redundancy with the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no elaboration, and the description correctly summarizes what is returned. For a two-parameter read-only metadata tool this is largely sufficient, with the only gap being the stage_url-to-dataset_id prerequisite and cache staleness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including a detailed response_format enum with a recovery instruction, so the schema does all the work. The description adds no parameter meaning beyond it, which is the expected baseline when coverage is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource (inspect a staged dataset) and enumerates the exact payload: cached columns/types plus row count. The clause 'never observation rows' implicitly separates it from row-returning siblings such as query_dataset, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'never observation rows' suggests you use this for schema/shape checks and query_dataset for actual data, but no when-to-use or when-not-to-use condition is stated. The one actionable fallback hint (retry with response_format="text") lives in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_datasetQuery bounded rows from a staged statistical datasetARead-onlyIdempotentInspect
Return only bounded selected rows using literal equality filters.
There is no SQL, shell, filesystem, network-function, extension, or mutation surface. A safe default limit applies when limit is omitted.
Page with offset: rows come back in the dataset's stored order, and
next_offset in the response is the absolute offset to pass as the next
call's offset. Stored order is the provider's order, not necessarily
chronological; filter TIME_PERIOD with where or sort client-side.
next_offset is null when there is no next page to ask for, which is not
the same as having received everything. Check truncated too: null with
truncated: false means the matches are exhausted; null with
truncated: true means rows remain that this byte budget cannot reach,
and min_bytes_required says what budget would.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Optional max rows to return (a safe default applies when omitted) | |
| where | No | Optional {column: value | [values]} literal-equality row filters | |
| offset | No | Optional absolute zero-based row offset; pass back next_offset | |
| select | No | Optional list of columns to return (defaults to all columns) | |
| max_bytes | No | Optional smaller byte budget for this response; it can only lower the server ceiling, never raise it | |
| dataset_id | Yes | Id returned by stage_url | |
| response_format | No | How the rows are sent. "auto" (default) means no preference and lets the server decide; "text" sends them as CSV only — in the text content, and in structuredContent as a "csv" string in place of typed rows. "structured" sends typed rows only, "both" the CSV and the typed rows. If you got a summary but no rows, call again with response_format="text". | auto |
Output Schema
| Name | Required | Description |
|---|---|---|
| csv | No | |
| rows | No | |
| columns | No | |
| truncated | Yes | |
| next_offset | Yes | |
| limit_source | Yes | |
| matched_rows | Yes | |
| applied_limit | Yes | |
| returned_rows | Yes | |
| max_bytes_ceiling | No | |
| min_bytes_required | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, closed-world, but the description adds substantial behavior not captured there: the absence of any SQL/mutation surface, the default limit safety net, and a nuanced disclosure that next_offset=null is ambiguous on its own and must be read together with truncated, plus the min_bytes_required escape hatch for the byte budget. This is exactly the kind of beyond-annotations context that earns a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core contract in sentence one and the safety envelope in sentence two, then paging details. Every sentence carries information, though the backtick-heavy next_offset/truncated paragraph is dense and could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description need not restate return shapes, yet it usefully clarifies the response fields an agent must interpret to paginate correctly (next_offset, truncated, min_bytes_required) and explains the response_format 'got a summary but no rows' recovery path. Nothing needed to call and iterate this tool successfully is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description goes beyond it by explaining paging mechanics for offset (rows return in stored order, pass back next_offset) and by tying truncated/min_bytes_required to the max_bytes budget. It adds genuine operational meaning for offset and max_bytes that the schema text alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource — return bounded selected rows via literal equality filters — and the second sentence sharply delineates the operation's surface (no SQL, shell, filesystem, network, extension, or mutation). That is unusually clear purpose framing, though it does not name or contrast with any of the sibling tools (e.g., inspect_dataset), so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete operational guidance: a safe default limit when omitted, how to page via offset/next_offset, and how to handle stored ordering (filter TIME_PERIOD or sort client-side). It does not, however, say when to reach for query_dataset rather than inspect/inspect_dataset or how to obtain a dataset_id beyond the schema's 'Id returned by stage_url'. Clear context, no explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
release_datasetRelease a staged statistical datasetADestructiveIdempotentInspect
Idempotently remove a caller-owned staged dataset before its TTL.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| released | Yes | |
| dataset_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the description correctly echoes 'Idempotently remove'. It adds useful context: the dataset is caller-owned and the action happens before TTL, which clarifies ownership and timing. There is no contradiction, and the description adds value beyond the annotations by specifying these details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence states the verb, the object, and the key constraints. Every word earns its place, with no filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a single parameter and an output schema, the description covers the core operation, ownership, and timing. However, the complete absence of parameter guidance is a notable gap, even though the tool is simple. An agent would still need to infer what dataset_id refers to, so the description is not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no explanation of dataset_id – what it represents, how to obtain it, or any format expectations. With a single parameter and zero coverage, the description fails to compensate, leaving the agent without guidance on what value to supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (remove) and the resource (a staged dataset), and adds scoping ('caller-owned', 'before its TTL') that distinguishes it from sibling tools like stage_url or query_dataset. An agent can immediately grasp what this tool does without inspecting other definitions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context – you release a staged dataset you own before its TTL – but it does not explicitly mention when not to use it or compare it to alternatives like stage_url or inspect_dataset. There is no guidance on selecting this tool over others, so it falls short of the level of explicit routing seen in get_calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stage_urlFallback: stage an Aard-issued datasetAInspect
Fallback for clients that cannot download and query the source locally.
Use only when local execution is unavailable or the user explicitly asks for Aard-managed staging. Never fetch the URL with web search or request its response body directly. Stage the exact unchanged fallback URL outside model context; this tool returns metadata only. Use its dataset_id with inspect_dataset or query_dataset.
The URL alone is sufficient. download_ref is an optional fast path — if you cannot copy it exactly, OMIT it rather than risk mistyping it; the server then re-derives the grant itself and the call still succeeds.
no_records_for_selection is TERMINAL, not a fault: the request was well-formed and the source holds no observations for it. Widen the selection or state that no data exists — do not retry the same selection.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The exact csv_url/download_url from build_url, byte-for-byte. This is the value that matters — an edited URL is refused. | |
| format | No | Content-type handling. "auto" (default) requires the origin to declare a CSV media type. "csv" and "sdmx-csv" additionally accept an application/octet-stream response, for origins that serve CSV without labelling it — use them only when "auto" fails with unsupported_content_type. | auto |
| download_ref | No | Optional. The signed grant from the same build_url result, copied exactly. Omit it if you cannot. |
Output Schema
| Name | Required | Description |
|---|---|---|
| sha256 | Yes | |
| dataset_id | Yes | |
| expires_at | Yes | |
| content_type | Yes | |
| bytes_downloaded | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly=false, openWorld=true, non-idempotent, non-destructive), and the description adds substantial behavior beyond them: metadata-only return, URL-alone sufficiency, the optional/omittable download_ref with server-side grant re-derivation, and the terminal semantics of no_records_for_selection. This is exactly the extra context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the fallback framing, then layers usage constraints, parameter guidance, and the terminal-error note. It is on the longer side, but each paragraph earns its place; only minor compression is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and the description instead covers prerequisites, the local-vs-staged decision, exact-copy requirements, and the non-retryable terminal error. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds decision-level guidance: the URL alone is sufficient, and download_ref should be omitted rather than mistyped since the server re-derives the grant. That meaningfully reduces the risk of a failed call, though the format enum handling is left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (stage) and resource (an Aard-issued fallback dataset URL), and frames it as a fallback that returns metadata only. It distinguishes itself from the siblings it hands off to (inspect_dataset, query_dataset), so an agent can place it in the workflow without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('only when local execution is unavailable or the user explicitly asks for Aard-managed staging') and when-not ('Never fetch the URL with web search or request its response body directly'). It also names the downstream alternatives and how to proceed from the returned dataset_id, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- Changed
ask1 field changed- added
Output schema / properties / candidates / items / properties / period_calendarAdded value: +{ + "anyOf": [ + { + "enum": [ + "gregorian", + "buddhist" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Period Calendar" +}
- Changed
build_url2 fields changed- added
Output schema / properties / period_calendarAdded value: +{ + "enum": [ + "gregorian", + "buddhist" + ], + "type": "string" +} - changed
Output schema / requiredPrevious value: -[ - "status", - "answer_ready", - "agency_id", - "dataflow_id", - "delivery", - "staging_fallback", - "series", - "coverage", - "volume", - "verification" -]New value: +[ + "status", + "answer_ready", + "agency_id", + "dataflow_id", + "delivery", + "staging_fallback", + "series", + "coverage", + "volume", + "verification", + "period_calendar" +]
- Changed
discover3 fields changed- added
Output schema / properties / anchored_datasets / items / properties / period_calendarAdded value: +{ + "const": "buddhist", + "description": "Present ONLY for a Buddhist-Era dataset: TIME_PERIOD, and so any URL's startPeriod/endPeriod, is the Gregorian year + 543 (2567 = 2024). Absent means Gregorian.", + "type": "string" +} - added
Output schema / properties / other_agencies / items / properties / period_calendarAdded value: +{ + "const": "buddhist", + "description": "Present ONLY for a Buddhist-Era dataset: TIME_PERIOD, and so any URL's startPeriod/endPeriod, is the Gregorian year + 543 (2567 = 2024). Absent means Gregorian.", + "type": "string" +} - added
Output schema / properties / results / items / properties / period_calendarAdded value: +{ + "const": "buddhist", + "description": "Present ONLY for a Buddhist-Era dataset: TIME_PERIOD, and so any URL's startPeriod/endPeriod, is the Gregorian year + 543 (2567 = 2024). Absent means Gregorian.", + "type": "string" +}
- Changed
fetch2 fields changed- added
Output schema / properties / period_calendarAdded value: +{ + "enum": [ + "gregorian", + "buddhist" + ], + "type": "string" +} - changed
Output schema / requiredPrevious value: -[ - "agency_id", - "dataflow_id", - "source_request_count", - "matched_rows", - "returned_rows", - "truncated", - "applied_limit", - "limit_source" -]New value: +[ + "agency_id", + "dataflow_id", + "source_request_count", + "period_calendar", + "matched_rows", + "returned_rows", + "truncated", + "applied_limit", + "limit_source" +]
5 tool updates
- Changed
build_url3 fields changed- changed
Input schema / properties / verify / descriptionPrevious value: -"If True, fetch ONE observation from the built URL to confirm the\nselected/default codes actually co-occur in observed data (default\noff — graph-only). Adds a small live request; the URL is never\nchanged, only annotated (an empty sample raises a warning)."New value: +"If True, fetch ONE observation per series from the built URL (part 1\nof a fan-out) to confirm the selected/default codes actually co-occur\nin observed data (default off — graph-only). Adds a small live\nrequest; the URL is never changed. The outcome is returned as\n``verification`` and a **Verification** line: ``rows``;\n``no_records`` (code ``no_records_for_selection`` — the provider\nhas no observations for this exact selection: a definite no-data\nanswer for it); ``failed`` (the check itself failed — NOT evidence\nof missing data); or ``skipped`` (no verdict, with the reason —\nincluding ``fanout_partial``: part 1 of a fan-out was empty and\nthe other parts were not sampled). ``null`` means verification\nwas not requested." - added
Output schema / properties / verificationAdded value: +{ + "anyOf": [ + { + "description": "The published ``build_url(verify=True)`` outcome. Every field is wire contract.", + "properties": { + "code": { + "anyOf": [ + { + "enum": [ + "no_records_for_selection", + "upstream_origin_error", + "upstream_rate_limited", + "download_timeout", + "incomplete_response", + "unrecognised_not_found", + "unsupported_content_type", + "malformed_csv", + "probe_inconclusive", + "keys_authoritative", + "cube_precision", + "availability_not_enumerable", + "no_executable_request", + "no_csv_endpoint", + "fanout_partial" + ], + "type": "string" + }, + { + "type": "null" + } + ] + }, + "http_status": { + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ] + }, + "n_parts": { + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ] + }, + "probed_part": { + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ] + }, + "reason": { + "type": "string" + }, + "sample_rows": { + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ] + }, + "status": { + "enum": [ + "rows", + "no_records", + "failed", + "skipped" + ], + "type": "string" + } + }, + "required": [ + "status", + "code", + "http_status", + "sample_rows", + "probed_part", + "n_parts", + "reason" + ], + "type": "object" + }, + { + "type": "null" + } + ] +} - changed
Output schema / requiredPrevious value: -[ - "status", - "answer_ready", - "agency_id", - "dataflow_id", - "delivery", - "staging_fallback", - "series", - "coverage", - "volume" -]New value: +[ + "status", + "answer_ready", + "agency_id", + "dataflow_id", + "delivery", + "staging_fallback", + "series", + "coverage", + "volume", + "verification" +]
- Changed
discover2 fields changed- changed
Input schema / properties / topic / descriptionPrevious value: -"What the data is about, in natural language. Naming one concrete code — a country, a city, an indicator — returns the anchored shape."New value: +"What the data is about, in natural language. Naming one concrete code — a country, a city, an indicator — returns the anchored shape when at least max(15, limit) datasets publish it; a retrieval_degraded/anchor_unavailable warning marks a ranked answer whose anchoring failed or timed out." - changed
Output schema / properties / warnings / items / properties / level / descriptionPrevious value: -"Sub-kind of `code`, not a severity ordering (e.g. semantic_unavailable, low_no_match)."New value: +"Sub-kind of `code`, not a severity ordering (e.g. semantic_unavailable, anchor_unavailable, low_no_match)."
- Changed
fetch2 fields changed- changed
Input schema / properties / response_format / descriptionPrevious value: -"Which channel carries the rows. \"auto\" (default) means\nno preference and lets the server decide; \"text\" sends them as CSV\nin the text channel only, \"structured\" as typed rows only, \"both\"\nin both. If you got a summary but no rows, call again with\nresponse_format=\"text\"."New value: +"How the rows are sent. \"auto\" (default) means no\npreference and lets the server decide; \"text\" sends them as CSV\nonly — in the text content, and in structuredContent as a \"csv\"\nstring in place of typed rows. \"structured\" sends typed rows only,\n\"both\" the CSV and the typed rows. If you got a summary but no\nrows, call again with response_format=\"text\"." - added
Output schema / properties / csvAdded value: +{ + "type": "string" +}
- Changed
inspect_dataset2 fields changed- changed
Input schema / properties / response_format / descriptionPrevious value: -"Which channel carries the column table. \"auto\"\n(default) means no preference and lets the server decide; \"text\"\nsends it as CSV in the text channel only, \"structured\" as typed\nfields only, \"both\" in both. If you got a summary but no column\ntable, call again with response_format=\"text\"."New value: +"How the column table is sent. \"auto\" (default)\nmeans no preference and lets the server decide; \"text\" sends it as\nCSV only — in the text content, and in structuredContent as a\n\"csv\" string in place of typed columns. \"structured\" sends typed\nfields only, \"both\" the CSV and the typed fields. If you got a\nsummary but no column table, call again with\nresponse_format=\"text\"." - added
Output schema / properties / csvAdded value: +{ + "type": "string" +}
- Changed
query_dataset2 fields changed- changed
Input schema / properties / response_format / descriptionPrevious value: -"Which channel carries the rows. \"auto\" (default) means\nno preference and lets the server decide; \"text\" sends them as CSV\nin the text channel only, \"structured\" as typed rows only, \"both\"\nin both. If you got a summary but no rows, call again with\nresponse_format=\"text\"."New value: +"How the rows are sent. \"auto\" (default) means no\npreference and lets the server decide; \"text\" sends them as CSV\nonly — in the text content, and in structuredContent as a \"csv\"\nstring in place of typed rows. \"structured\" sends typed rows only,\n\"both\" the CSV and the typed rows. If you got a summary but no\nrows, call again with response_format=\"text\"." - added
Output schema / properties / csvAdded value: +{ + "type": "string" +}
8 tool updates
- Changed
ask17 fields changed- added
Input schema / properties / agencies / descriptionAdded value: +"Optional agency filter (e.g., [\"ESTAT\", \"OECD\"])" - added
Input schema / properties / cross_source / descriptionAdded value: +"If True, find structurally analogous dataflows in other agencies\n(>= 50% shared dimension concepts). Adds an extra graph query per candidate." - added
Input schema / properties / currency / descriptionAdded value: +"Optional currency slot — the currency the question is about\n(e.g. [\"euro\"], [\"USD\"], [\"yen\"]). NOT for geographic phrases like\n\"euro area\". Resolved to per-agency currency codes and used to\npromote candidates whose confirmed (Actual) data carries that\ncurrency, demote those that provably do not, and pre-fill the\ncurrency dimension in built URLs." - added
Input schema / properties / debug / descriptionAdded value: +"Append a per-stage telemetry breakdown to the response (only\npopulated when GSDMX2_MCP_TELEMETRY is enabled). Off by default." - added
Input schema / properties / geography / descriptionAdded value: +"Optional geography slot — country/region/world names the question is\nabout (e.g. [\"Australia\"], [\"European Union\"], [\"world\"]). Supplied by the\nclient; resolved to ISO alpha-2 and used to demote wrong-geography\ncandidates in ranking. Does NOT become a hard agency filter or a URL filter." - added
Input schema / properties / keywords / descriptionAdded value: +"Optional keyword overrides for graph search (auto-extracted if omitted)" - added
Input schema / properties / labels / descriptionAdded value: +"Request the provider's labelled CSV (default True) — adds a\nhuman-readable name column beside every coded column on the 14\nendpoints with a verified labelled spelling, and degrades silently to\nplain CSV elsewhere. Set False for a smaller download (labelled rows\nare roughly 2-4x the bytes)." - added
Input schema / properties / language / descriptionAdded value: +"ISO language code (default \"en\")" - added
Input schema / properties / precision / descriptionAdded value: +"URL breadth — \"point\" (one observation's series), \"series\"\n(default: headline defaults for unfilled dimensions), or \"cube\"\n(full constraint enumeration, the historical behaviour)." - added
Input schema / properties / product / descriptionAdded value: +"Optional product/commodity slot — the good the question is about\n(e.g. [\"wheat\"], [\"copper\"], [\"crude oil\"]). Resolved to per-agency\ncommodity codes (HS/SITC/custom) with the same\npromote/demote/URL-fill semantics as currency." - added
Input schema / properties / question / descriptionAdded value: +"Natural language question about statistical data" - added
Input schema / properties / scalar / descriptionAdded value: +"Set True when the question wants a single value (one observation) rather\nthan a series/table. Defaults precision to \"point\" unless precision is given." - added
Input schema / properties / strict_time / descriptionAdded value: +"If True and time_range is set, drop candidates whose\nmaterialized coverage is provably disjoint from the requested window\n(candidates without a coverage record are always kept). Off by\ndefault — the coverage-overlap ranking demotes disjoint hits but\nstill lists them." - added
Input schema / properties / time_range / descriptionAdded value: +"Optional time filter (e.g., \"2020-2024\", \"since 2015\", \"last 5 years\")" - added
Input schema / properties / top_n / descriptionAdded value: +"Number of top candidates to build URLs for (default 5)" - added
Input schema / properties / topic / descriptionAdded value: +"Optional subject/metric of the question (e.g. \"child labour\", \"GDP\").\nSharpens keyword + parser ranking signals; the free-text question is still\nwhat gets embedded." - changed
Input schema / properties / user_country / descriptionPrevious value: -"The country the USER is in. Pass only when the user has stated where they are; never infer it from the question. This is not the geography the question is about — that is `geography`. Used to prefer data covering the user's country when the question names no geography of its own."New value: +"The country the USER is in (e.g. \"New Zealand\"). Pass only when the user has stated where they are; never infer it from the question. This is not the geography the question is about — that is `geography`. Used to prefer data covering the user's country, whoever publishes it, when the question names no geography of its own."
- Changed
browse_dimension_codes6 fields changed- added
Input schema / properties / agency_id / descriptionAdded value: +"SDMX agency code, e.g. \"ABS\", \"ESTAT\", \"OECD\"" - added
Input schema / properties / dataflow_id / descriptionAdded value: +"SDMX dataflow identifier, e.g. \"ERP_Q\"" - added
Input schema / properties / debug / descriptionAdded value: +"Append a per-stage telemetry breakdown (only populated when\nGSDMX2_MCP_TELEMETRY is enabled)." - added
Input schema / properties / dimension_id / descriptionAdded value: +"Dimension to browse, e.g. \"REF_AREA\" (from inspect)" - added
Input schema / properties / expand / descriptionAdded value: +"Optional code id — list that code's direct narrower members\ninstead of the top level" - added
Input schema / properties / offset / descriptionAdded value: +"Pagination offset within the current level (default 0)"
- Changed
build_url9 fields changed- added
Input schema / properties / agency_id / descriptionAdded value: +"SDMX agency code, e.g. \"ILO\", \"ABS\", \"ESTAT\" (required — the same\ndataflow id can exist under several agencies)" - added
Input schema / properties / availability / descriptionAdded value: +"\"confirmed\" (default — direct Actual or observed-key evidence)\nor \"best_effort\" (add direct Allowed codes, still unconfirmed). Neither\nmode executes structural codelist values." - added
Input schema / properties / dataflow_id / descriptionAdded value: +"SDMX dataflow identifier, e.g. \"DF_CLD_XCHL_SEX_AGE_NB\"" - added
Input schema / properties / debug / descriptionAdded value: +"Append a per-stage telemetry breakdown (only populated when\nGSDMX2_MCP_TELEMETRY is enabled). Off by default." - added
Input schema / properties / labels / descriptionAdded value: +"If True, request the provider's LABELLED CSV — each coded column\ngains a human-readable name column beside it (\"MEASURE\" plus \"Data\nItem\"), so the data explains itself and you need no follow-up\ninspect calls to decode it. Code columns are unchanged, so\nquery_dataset where={...} filters on codes still work. Costs ~3.6x\nbytes per row, which means fewer rows per query_dataset call — use\nit when you need to READ the data, not when you need many rows.\nHonoured by every endpoint with a verified labelled spelling (ABS,\nILO, OECD, SPC and others); elsewhere it degrades silently to\nplain CSV.\nNote the column set changes: DATAFLOW is replaced by STRUCTURE,\nSTRUCTURE_ID, STRUCTURE_NAME and ACTION." - added
Input schema / properties / precision / descriptionAdded value: +"URL breadth — \"point\", \"series\" (default), or \"cube\"" - added
Input schema / properties / selections / descriptionAdded value: +"Optional {dimension_id: [code or name, ...]} to anchor the URL.\nNames are resolved within the dimension's AVAILABLE codes; values with\nno available data are rejected with alternatives, never silently passed." - added
Input schema / properties / time_range / descriptionAdded value: +"Optional time filter (e.g. \"2020-2024\", \"since 2015\", \"2024\")" - added
Input schema / properties / verify / descriptionAdded value: +"If True, fetch ONE observation from the built URL to confirm the\nselected/default codes actually co-occur in observed data (default\noff — graph-only). Adds a small live request; the URL is never\nchanged, only annotated (an empty sample raises a warning)."
- Changed
fetch11 fields changed- added
Input schema / properties / agency_id / descriptionAdded value: +"SDMX agency code, e.g. \"ABS\", \"ESTAT\", \"OECD\"" - added
Input schema / properties / availability / descriptionAdded value: +"\"confirmed\" (default) or \"best_effort\"" - added
Input schema / properties / dataflow_id / descriptionAdded value: +"SDMX dataflow identifier, e.g. \"ERP_Q\"" - added
Input schema / properties / limit / descriptionAdded value: +"Optional max rows to return (a safe default applies when omitted)" - added
Input schema / properties / max_bytes / descriptionAdded value: +"Optional smaller byte budget for this response; it can only\nlower the server ceiling, never raise it" - added
Input schema / properties / precision / descriptionAdded value: +"URL breadth — \"point\", \"series\" (default), or \"cube\"" - added
Input schema / properties / response_format / descriptionAdded value: +"Which channel carries the rows. \"auto\" (default) means\nno preference and lets the server decide; \"text\" sends them as CSV\nin the text channel only, \"structured\" as typed rows only, \"both\"\nin both. If you got a summary but no rows, call again with\nresponse_format=\"text\"." - added
Input schema / properties / select / descriptionAdded value: +"Optional list of columns to return (defaults to all columns)" - added
Input schema / properties / selections / descriptionAdded value: +"Optional {dimension_id: [code or name, ...]} to anchor the query" - added
Input schema / properties / time_range / descriptionAdded value: +"Optional time filter (e.g. \"2020-2024\", \"since 2015\", \"2024\")" - added
Input schema / properties / where / descriptionAdded value: +"Optional {column: value | [values]} literal-equality row filters"
- Changed
inspect4 fields changed- added
Input schema / properties / agency_id / descriptionAdded value: +"SDMX agency code, e.g. \"ABS\", \"ESTAT\", \"OECD\"" - added
Input schema / properties / dataflow_id / descriptionAdded value: +"SDMX dataflow identifier, e.g. \"ERP_Q\", \"DS-018995\" (NOT dataset_id)" - added
Input schema / properties / debug / descriptionAdded value: +"Append a per-stage telemetry breakdown (assistant-only) to the\noutput (only populated when GSDMX2_MCP_TELEMETRY is enabled)." - added
Input schema / properties / sample / descriptionAdded value: +"Max codes to show per dimension (default 10)"
- Changed
inspect_dataset2 fields changed- added
Input schema / properties / dataset_id / descriptionAdded value: +"Id returned by stage_url" - added
Input schema / properties / response_format / descriptionAdded value: +"Which channel carries the column table. \"auto\"\n(default) means no preference and lets the server decide; \"text\"\nsends it as CSV in the text channel only, \"structured\" as typed\nfields only, \"both\" in both. If you got a summary but no column\ntable, call again with response_format=\"text\"."
- Changed
query_dataset7 fields changed- added
Input schema / properties / dataset_id / descriptionAdded value: +"Id returned by stage_url" - added
Input schema / properties / limit / descriptionAdded value: +"Optional max rows to return (a safe default applies when omitted)" - added
Input schema / properties / max_bytes / descriptionAdded value: +"Optional smaller byte budget for this response; it can only\nlower the server ceiling, never raise it" - added
Input schema / properties / offset / descriptionAdded value: +"Optional absolute zero-based row offset; pass back next_offset" - added
Input schema / properties / response_format / descriptionAdded value: +"Which channel carries the rows. \"auto\" (default) means\nno preference and lets the server decide; \"text\" sends them as CSV\nin the text channel only, \"structured\" as typed rows only, \"both\"\nin both. If you got a summary but no rows, call again with\nresponse_format=\"text\"." - added
Input schema / properties / select / descriptionAdded value: +"Optional list of columns to return (defaults to all columns)" - added
Input schema / properties / where / descriptionAdded value: +"Optional {column: value | [values]} literal-equality row filters"
- Changed
stage_url3 fields changed- added
Input schema / properties / download_ref / descriptionAdded value: +"Optional. The signed grant from the same build_url result,\ncopied exactly. Omit it if you cannot." - added
Input schema / properties / format / descriptionAdded value: +"Content-type handling. \"auto\" (default) requires the origin to\ndeclare a CSV media type. \"csv\" and \"sdmx-csv\" additionally accept\nan application/octet-stream response, for origins that serve CSV\nwithout labelling it — use them only when \"auto\" fails with\nunsupported_content_type." - added
Input schema / properties / url / descriptionAdded value: +"The exact csv_url/download_url from build_url, byte-for-byte. This\nis the value that matters — an edited URL is refused."
10 tool updates
- First observed
ask - First observed
browse_dimension_codes - First observed
build_url - First observed
discover - First observed
fetch - First observed
inspect - First observed
inspect_dataset - First observed
query_dataset - First observed
release_dataset - First observed
stage_url
Publisher details
- Operator
- Aard AI · Publisher source
- Operator website
- https://aard.ai · Publisher source
- Vendor relationship
- First-party · Publisher source
- Documentation
- Unknown
- Trust center
- https://aard.ai/security · Publisher source
- Restrictions
- Unknown
Related MCP Connectors
Macro indicators from World Bank, FRED, IMF, and OECD via unified query surface.
Macro data for AI agents: GDP, inflation, unemployment and more (World Bank, US BLS). No keys.
Macro data for AI agents: GDP, inflation, unemployment and more (World Bank, US BLS). No keys.
Discover, resolve, and query official Brazilian economic data with semantic search and provenance.
Related MCP Servers
- AlicenseAqualityFmaintenanceEnables AI agents to discover, retrieve, compare, and analyze trusted macroeconomic statistics from central banks and international organizations, resolving concepts to official series with provenance and validation.6MIT
- AlicenseAqualityAmaintenanceOfficial economic statistics with full citations — World Bank, IMF WEO, ECB — plus verify_stat to check a claimed figure against the official series. Free, remote, no auth.121MIT
- AlicenseAqualityBmaintenanceEnables querying international macro statistics, company identity data via LEI, and FX rates from dozens of free keyless providers through unified tools.12MIT
- AlicenseNot gradedqualityBmaintenanceEnables natural language access to 800,000+ economic time series from FRED, including GDP, inflation, unemployment, and interest rates, with built-in transforms and frequency aggregation.82 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.