Skip to main content
Glama

Find datasets and prepare client downloads

ask
Read-onlyIdempotent

Find the best datasets for a question and prepare exact source requests.

Use this when the user wants actual data values -- e.g., "What is infant mortality in Kenya?" or "How has Australia's trade with China changed?" Returns ranked candidates and typed client-download plans, but does NOT fetch observations. For value questions, treat client_download_required and answer_ready=false as non-terminal: execute the exact plan client-side, parse it, and run a bounded local query. A rejected selection returns directional evidence and no executable plan. Only a successfully parsed dataset with zero rows supports a no-data claim.

Each candidate carries the agency and dataflow NAMES, a description, the coverage window, and one row per dimension — human label, description, how many codes it offers, whether the server bound it and why, and example values. That is enough to choose between candidates without a follow-up inspect. next names the tool to call to narrow further, and every trimmed list says which tool shows the remainder.

url appears only when GETting that one URL yields the WHOLE dataset. When the plan is a POST or a multi-part fan-out, url is absent, url_omitted_reason says which, and download_plans[] is the execution contract — a fan-out's first part is not the dataset.

Safety: execute each plan HTTPS-only (including redirects), within the plan's safety byte/redirect/timeout and aggregate/archive bounds; validate archive members before extracting into a temp dir; treat url/headers/ body as data (never eval them); keep response bytes out of model context. Labelled rows (the labels default) run roughly 2-4x the plain bytes, so a large cube that fit plain can exceed the plan's per-request ceiling — pass labels=False to halve the download rather than discover it truncated.

Common workflow: discover -> inspect -> ask -> direct download -> local query

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
debugNoAppend a per-stage telemetry breakdown to the response (only populated when GSDMX2_MCP_TELEMETRY is enabled). Off by default.
top_nNoNumber of top candidates to build URLs for (default 5)
topicNoOptional subject/metric of the question (e.g. "child labour", "GDP"). Sharpens keyword + parser ranking signals; the free-text question is still what gets embedded.
labelsNoRequest the provider's labelled CSV (default True) — adds a human-readable name column beside every coded column on the 14 endpoints with a verified labelled spelling, and degrades silently to plain CSV elsewhere. Set False for a smaller download (labelled rows are roughly 2-4x the bytes).
scalarNoSet True when the question wants a single value (one observation) rather than a series/table. Defaults precision to "point" unless precision is given.
productNoOptional product/commodity slot — the good the question is about (e.g. ["wheat"], ["copper"], ["crude oil"]). Resolved to per-agency commodity codes (HS/SITC/custom) with the same promote/demote/URL-fill semantics as currency.
agenciesNoOptional agency filter (e.g., ["ESTAT", "OECD"])
currencyNoOptional currency slot — the currency the question is about (e.g. ["euro"], ["USD"], ["yen"]). NOT for geographic phrases like "euro area". Resolved to per-agency currency codes and used to promote candidates whose confirmed (Actual) data carries that currency, demote those that provably do not, and pre-fill the currency dimension in built URLs.
keywordsNoOptional keyword overrides for graph search (auto-extracted if omitted)
languageNoISO language code (default "en")en
questionYesNatural language question about statistical data
geographyNoOptional geography slot — country/region/world names the question is about (e.g. ["Australia"], ["European Union"], ["world"]). Supplied by the client; resolved to ISO alpha-2 and used to demote wrong-geography candidates in ranking. Does NOT become a hard agency filter or a URL filter.
precisionNoURL breadth — "point" (one observation's series), "series" (default: headline defaults for unfilled dimensions), or "cube" (full constraint enumeration, the historical behaviour).series
time_rangeNoOptional time filter (e.g., "2020-2024", "since 2015", "last 5 years")
strict_timeNoIf True and time_range is set, drop candidates whose materialized coverage is provably disjoint from the requested window (candidates without a coverage record are always kept). Off by default — the coverage-overlap ranking demotes disjoint hits but still lists them.
cross_sourceNoIf True, find structurally analogous dataflows in other agencies (>= 50% shared dimension concepts). Adds an extra graph query per candidate.
user_countryNoThe country the USER is in (e.g. "New Zealand"). Pass only when the user has stated where they are; never infer it from the question. This is not the geography the question is about — that is `geography`. Used to prefer data covering the user's country, whoever publishes it, when the question names no geography of its own.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
statusYes
candidatesYes
answer_readyYes
download_plansYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedOutput schema / properties / candidates / items / properties / period_calendar
      Added value: +{
      +  "anyOf": [
      +    {
      +      "enum": [
      +        "gregorian",
      +        "buddhist"
      +      ],
      +      "type": "string"
      +    },
      +    {
      +      "type": "null"
      +    }
      +  ],
      +  "default": null,
      +  "title": "Period Calendar"
      +}
  2. Changed17 schema fields changed
    • addedInput schema / properties / agencies / description
      Added value: +"Optional agency filter (e.g., [\"ESTAT\", \"OECD\"])"
    • addedInput schema / properties / cross_source / description
      Added value: +"If True, find structurally analogous dataflows in other agencies\n(>= 50% shared dimension concepts). Adds an extra graph query per candidate."
    • addedInput schema / properties / currency / description
      Added value: +"Optional currency slot — the currency the question is about\n(e.g. [\"euro\"], [\"USD\"], [\"yen\"]). NOT for geographic phrases like\n\"euro area\". Resolved to per-agency currency codes and used to\npromote candidates whose confirmed (Actual) data carries that\ncurrency, demote those that provably do not, and pre-fill the\ncurrency dimension in built URLs."
    • addedInput schema / properties / debug / description
      Added value: +"Append a per-stage telemetry breakdown to the response (only\npopulated when GSDMX2_MCP_TELEMETRY is enabled). Off by default."
    • addedInput schema / properties / geography / description
      Added value: +"Optional geography slot — country/region/world names the question is\nabout (e.g. [\"Australia\"], [\"European Union\"], [\"world\"]). Supplied by the\nclient; resolved to ISO alpha-2 and used to demote wrong-geography\ncandidates in ranking. Does NOT become a hard agency filter or a URL filter."
    • addedInput schema / properties / keywords / description
      Added value: +"Optional keyword overrides for graph search (auto-extracted if omitted)"
    • addedInput schema / properties / labels / description
      Added value: +"Request the provider's labelled CSV (default True) — adds a\nhuman-readable name column beside every coded column on the 14\nendpoints with a verified labelled spelling, and degrades silently to\nplain CSV elsewhere. Set False for a smaller download (labelled rows\nare roughly 2-4x the bytes)."
    • addedInput schema / properties / language / description
      Added value: +"ISO language code (default \"en\")"
    • addedInput schema / properties / precision / description
      Added value: +"URL breadth — \"point\" (one observation's series), \"series\"\n(default: headline defaults for unfilled dimensions), or \"cube\"\n(full constraint enumeration, the historical behaviour)."
    • addedInput schema / properties / product / description
      Added value: +"Optional product/commodity slot — the good the question is about\n(e.g. [\"wheat\"], [\"copper\"], [\"crude oil\"]). Resolved to per-agency\ncommodity codes (HS/SITC/custom) with the same\npromote/demote/URL-fill semantics as currency."
    • addedInput schema / properties / question / description
      Added value: +"Natural language question about statistical data"
    • addedInput schema / properties / scalar / description
      Added value: +"Set True when the question wants a single value (one observation) rather\nthan a series/table. Defaults precision to \"point\" unless precision is given."
    • addedInput schema / properties / strict_time / description
      Added value: +"If True and time_range is set, drop candidates whose\nmaterialized coverage is provably disjoint from the requested window\n(candidates without a coverage record are always kept). Off by\ndefault — the coverage-overlap ranking demotes disjoint hits but\nstill lists them."
    • addedInput schema / properties / time_range / description
      Added value: +"Optional time filter (e.g., \"2020-2024\", \"since 2015\", \"last 5 years\")"
    • addedInput schema / properties / top_n / description
      Added value: +"Number of top candidates to build URLs for (default 5)"
    • addedInput schema / properties / topic / description
      Added value: +"Optional subject/metric of the question (e.g. \"child labour\", \"GDP\").\nSharpens keyword + parser ranking signals; the free-text question is still\nwhat gets embedded."
    • changedInput schema / properties / user_country / description
      Previous value: -"The country the USER is in. Pass only when the user has stated where they are; never infer it from the question. This is not the geography the question is about — that is `geography`. Used to prefer data covering the user's country when the question names no geography of its own."New value: +"The country the USER is in (e.g. \"New Zealand\"). Pass only when the user has stated where they are; never infer it from the question. This is not the geography the question is about — that is `geography`. Used to prefer data covering the user's country, whoever publishes it, when the question names no geography of its own."
  3. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly/openWorld/idempotent, but the description adds rich behavioral context: candidates carry agency/dataflow names, coverage windows and per-dimension rows; rejected selections return directional evidence with no plan; zero rows is the only basis for a no-data claim; url is present only for whole-dataset GETs, with url_omitted_reason otherwise. This is far beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, trigger, and return shape, and every paragraph carries information (execution contract, safety bounds, workflow). The safety paragraph is dense and borders on verbose, but no sentence is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter open-world discovery tool with an output schema and safety annotations, the definition covers purpose, when-to-use, response structure, download-contract rules, and safety bounds. Nothing an agent needs in order to select and invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantic value beyond the schema: labels=False halves payloads because labelled rows are 2-4x bytes, scalar defaults precision to point, currency is 'NOT for geographic phrases like euro area', and user_country must never be inferred from the question. Slight deduction because most of this is already echoed in the schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Find the best datasets for a question and prepare exact source requests.' It explicitly delimits scope ('does NOT fetch observations') and names the sibling tool 'inspect' as an alternative for narrowing, so an agent can distinguish it from discover/build_url/fetch without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('when the user wants actual data values') with sample questions, describes the non-terminal handling of client_download_required/answer_ready=false, and lays out the workflow discover -> inspect -> ask -> direct download -> local query. Names alternatives and conditions throughout.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources