Skip to main content
Glama

faostat-mcp-server: query observations

faostat_query_observations
Read-only

Query a FAOSTAT domain's data cube by area(s), item(s), element(s), and year range, returning observations (area, item, element, year, value, unit, and the data-quality flag). Resolve codes first with faostat_resolve_codes — the cube is unqueryable without them. Aggregate regions (World, continents, economic groupings) are EXCLUDED by default so a naive SUM does not double-count a region with its member countries; set include_aggregates=true to get the regional roll-ups, or pass explicit area_codes to query exactly what you name. Small result sets return inline; large ones spill to a DataCanvas table (returned canvas_id + table_name) for GROUP BY / ranking / time-series analysis via faostat_dataframe_query. Every row carries its flag — commonly A=Official, B=time-series break, E=Estimated, I=Imputed, M=Missing (value cannot exist), T=Unofficial, X=from an international organization, plus others FAOSTAT defines per domain — so honor it, treat any unrecognized flag as informational, and never assume an estimated, imputed, or unrecognized value is official.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoMax observations returned inline — also the preview size when the result stages to a canvas table. Rows past it are never dropped silently: a match larger than limit stages in full to a DataCanvas table (canvas_id + table_name), and when no table is staged the notice reports how many matched so you can raise limit or narrow the filters. Max 1000.
domainYesFAOSTAT domain code (e.g. "QCL"). Must be indexed locally.
year_endNoInclusive end year (e.g. 2022).
canvas_idNoCanvas ID to stage onto, as returned by a prior faostat_query_observations / faostat_commodity_profile call — exactly 10 characters of letters, digits, hyphens, and underscores. Omit to stage onto this session’s canvas, created on the first spill and reused by every later call, so tables staged earlier in the session sit alongside this one.
area_codesNoArea codes from faostat_resolve_codes. When set, aggregates are NOT auto-excluded — the codes are honored verbatim.
item_codesNoItem codes from faostat_resolve_codes.
year_startNoInclusive start year (e.g. 2000).
element_codesNoElement codes from faostat_resolve_codes (e.g. 5510 Production).
include_aggregatesNoWhen false (default), exclude aggregate-region rows (codes ≥ 5000 plus a few curated sub-threshold roll-ups such as China=351) so sums are not double-counted. Set true for World/continent/grouping roll-ups. Ignored when explicit area_codes are passed.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoPresent when the call failed. Absent on success.
domainNoThe domain code echoed back.
noticeNoGuidance on empty results, aggregate exclusion, or how to reach the spilled set.
spilledNoTrue when the full result was staged on a DataCanvas table.
canvas_idNoCanvas ID holding the staged result — pass to faostat_dataframe_query / _describe.
truncatedNoTrue when the staged table hit the 50,000-row staging cap — the staged set is a PREFIX of the match, not the complete result. Partition the query by year or code ranges to capture the rest.
table_nameNoCanvas table name holding the full result set (present when spilled).
totalCountNoObservations matched. Exact when the result was returned inline or fully staged. A floor — more matched — in two cases, both named by the notice: the match exceeded the 50,000-row staging cap (truncated is then true), or staging failed and the response fell back to an inline page.
observationsNoInline observations (preview when the full set spilled to a canvas table).
staged_row_countNoRows actually staged on the canvas table (present when spilled). Equals the full match count unless truncated, in which case it is the 50,000-row cap.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed8 schema fields changed
    • changedInput schema / properties / canvas_id / description
      Previous value: -"Canvas ID from a prior call to stage onto. Omit to start a fresh canvas (a new id is returned)."New value: +"Canvas ID to stage onto, as returned by a prior faostat_query_observations / faostat_commodity_profile call — exactly 10 characters of letters, digits, hyphens, and underscores. Omit to stage onto this session’s canvas, created on the first spill and reused by every later call, so tables staged earlier in the session sit alongside this one."
    • addedInput schema / properties / canvas_id / pattern
      Added value: +"^[A-Za-z0-9_-]{10}$"
    • removedOutput schema / properties / observations / items / properties / flag / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedOutput schema / properties / observations / items / properties / flag / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • removedOutput schema / properties / observations / items / properties / unit / anyOf
      Removed value: -[
      -  {
      -    "type": "string"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedOutput schema / properties / observations / items / properties / unit / type
      Added value: +[
      +  "string",
      +  "null"
      +]
    • removedOutput schema / properties / observations / items / properties / value / anyOf
      Removed value: -[
      -  {
      -    "type": "number"
      -  },
      -  {
      -    "type": "null"
      -  }
      -]
    • addedOutput schema / properties / observations / items / properties / value / type
      Added value: +[
      +  "number",
      +  "null"
      +]
  2. Changed6 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / additionalProperties
      Added value: +false
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedOutput schema / anyOf
      Added value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "domain",
      +      "observations",
      +      "spilled",
      +      "truncated",
      +      "totalCount"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • addedOutput schema / properties / error
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "Present when the call failed. Absent on success.",
      +  "properties": {
      +    "code": {
      +      "description": "JSON-RPC error code for this failure.",
      +      "maximum": 9007199254740991,
      +      "minimum": -9007199254740991,
      +      "type": "integer"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "properties": {
      +        "reason": {
      +          "description": "Machine-readable failure mode. Declared by this tool: `domain_not_indexed`: The domain is not in the local mirror selection (FAOSTAT_DOMAINS) — whether a valid FAOSTAT code or not. `index_not_ready`: The domain mirror is cold — its initial sync has never completed. `canvas_disabled`: The result is too large to inline but DataCanvas is off, so it cannot be staged for SQL. `invalid_year_range`: year_start is greater than year_end — a self-contradictory range that can never match. Other values are possible when a failure originates below the handler.",
      +          "examples": [
      +            "domain_not_indexed",
      +            "index_not_ready",
      +            "canvas_disabled",
      +            "invalid_year_range"
      +          ],
      +          "type": "string"
      +        },
      +        "recovery": {
      +          "additionalProperties": {},
      +          "description": "Actionable next step for the caller.",
      +          "properties": {
      +            "hint": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "hint"
      +          ],
      +          "type": "object"
      +        },
      +        "retryable": {
      +          "description": "Whether retrying may succeed.",
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "message": {
      +      "description": "Human-readable description of what went wrong.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "code",
      +    "message"
      +  ],
      +  "type": "object"
      +}
    • removedOutput schema / required
      Removed value: -[
      -  "domain",
      -  "observations",
      -  "spilled",
      -  "truncated",
      -  "totalCount"
      -]
  3. Changed2 schema fields changed
    • changedInput schema / properties / limit / description
      Previous value: -"Max observations returned inline when the result does not spill. Max 1000."New value: +"Max observations returned inline — also the preview size when the result stages to a canvas table. Rows past it are never dropped silently: a match larger than limit stages in full to a DataCanvas table (canvas_id + table_name), and when no table is staged the notice reports how many matched so you can raise limit or narrow the filters. Max 1000."
    • changedOutput schema / properties / totalCount / description
      Previous value: -"Observations matched. Exact when the result was returned inline or fully staged; when the match set exceeded the 50,000-row staging cap this is that cap — a floor, not the exact count (truncated is then true)."New value: +"Observations matched. Exact when the result was returned inline or fully staged. A floor — more matched — in two cases, both named by the notice: the match exceeded the 50,000-row staging cap (truncated is then true), or staging failed and the response fell back to an inline page."
  4. Changed2 schema fields changed
    • changedInput schema / properties / include_aggregates / description
      Previous value: -"When false (default), exclude aggregate-region rows (codes ≥ 5000) so sums are not double-counted. Set true for World/continent/grouping roll-ups. Ignored when explicit area_codes are passed."New value: +"When false (default), exclude aggregate-region rows (codes ≥ 5000 plus a few curated sub-threshold roll-ups such as China=351) so sums are not double-counted. Set true for World/continent/grouping roll-ups. Ignored when explicit area_codes are passed."
    • changedOutput schema / properties / observations / items / properties / flag / description
      Previous value: -"Data-quality flag (A=Official, E=Estimated, I=Imputed, B=break, X=external); null when unflagged."New value: +"Data-quality flag — commonly A=Official, B=time-series break, E=Estimated, I=Imputed, M=Missing (value cannot exist), T=Unofficial, X=from an international organization, plus others FAOSTAT defines per domain; treat any unrecognized flag as informational, never assume official. Null when unflagged."
  5. Changed1 schema field changed
    • changedOutput schema / properties / totalCount / description
      Previous value: -"Total observations matched before any inline cap."New value: +"Observations matched. Exact when the result was returned inline or fully staged; when the match set exceeded the 50,000-row staging cap this is that cap — a floor, not the exact count (truncated is then true)."
  6. Changed3 schema fields changed
    • addedOutput schema / properties / staged_row_count
      Added value: +{
      +  "description": "Rows actually staged on the canvas table (present when spilled). Equals the full match count unless truncated, in which case it is the 50,000-row cap.",
      +  "type": "number"
      +}
    • addedOutput schema / properties / truncated
      Added value: +{
      +  "description": "True when the staged table hit the 50,000-row staging cap — the staged set is a PREFIX of the match, not the complete result. Partition the query by year or code ranges to capture the rest.",
      +  "type": "boolean"
      +}
    • changedOutput schema / required
      Previous value: -[
      -  "domain",
      -  "observations",
      -  "spilled",
      -  "totalCount"
      -]New value: +[
      +  "domain",
      +  "observations",
      +  "spilled",
      +  "truncated",
      +  "totalCount"
      +]
  7. First observed

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses that large results are staged to a DataCanvas table (returned canvas_id + table_name), that aggregate regions are excluded by default to prevent double-counting, and that quality flags must be honored (e.g., never assume an estimated or imputed value is official). These are non-obvious behavioral traits that materially affect how an agent should interpret results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core query action and return fields, then flows logically through prerequisite, aggregate handling, spill behavior, and flag warnings. It is somewhat long, especially the flag enumeration, but every sentence carries substantive guidance, so the length is justified by the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 9-parameter schema with 100% coverage and an existing output schema, the description covers all necessary operational aspects: what the tool returns, the hard dependency on faostat_resolve_codes, the aggregate behavior and its double-counting risk, the staging mechanism for large results, and the meaning and caution around data-quality flags. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaningful workflow context beyond the schema: it ties all code parameters to the prerequisite faostat_resolve_codes and explains why include_aggregates defaults to false (avoid double-counting in a naive SUM). It also clarifies the semantic relationship between explicit area_codes and aggregate exclusion, which the schema states but the description reinforces with rationale.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: "Query a FAOSTAT domain's data cube by area(s), item(s), element(s), and year range, returning observations..." This fully distinguishes it from siblings like faostat_resolve_codes (which resolves names to codes) and faostat_dataframe_query (which runs SQL on staged tables). The return fields are enumerated, making the tool's role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: "Resolve codes first with faostat_resolve_codes — the cube is unqueryable without them." It also names the alternative path for large results: "spill to a DataCanvas table ... for GROUP BY / ranking / time-series analysis via faostat_dataframe_query." The aggregate-exclusion caveat further instructs when to set include_aggregates=true versus passing explicit area_codes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.