Skip to main content
Glama

Cdc Query Dataset

cdc_query_dataset
Read-only

Execute a SoQL query against any CDC dataset. Supports filtering, aggregation, sorting, full-text search, and field selection. Use cdc_discover_datasets to find dataset IDs and cdc_get_dataset_schema to inspect columns before querying.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
groupNoSoQL GROUP BY clause. Requires aggregate functions in select.
limitNoMax rows to return (default 100, max 5000). Fewer come back when the page would carry the response past its 200,000-character budget, counted over the whole result — the rows as JSON and as the rendered table together; the response says so and gives a nextOffset to resume from.
orderNoSoQL ORDER BY clause. Field name with optional ASC/DESC: "total_deaths DESC". Set one whenever paging with offset: SODA does not order results implicitly, so consecutive offsets without a deterministic order can skip or repeat rows. When the dataset has no natural unique column, Socrata's documented minimum tie-breaker is the system field `:id`, present on every dataset — order=":id".
whereNoSoQL WHERE clause. Strings must be single-quoted: "state='California' AND year=2020". If a column name matches a SoQL keyword (group, select, where, order, limit, offset, having, search), wrap it in backticks: "`group`='By Year'".
domainNoCDC Socrata host to query. "data.cdc.gov" (default) and "chronicdata.cdc.gov" front the same catalog, so a four-by-four ID returns the same rows from either and the default works whichever host the dataset was found on.data.cdc.gov
havingNoSoQL HAVING clause. Filters aggregated results.
offsetNoRow offset for pagination (max 1,000,000). Pair with a deterministic order clause — an offset walk over unordered results can skip or repeat rows.
searchNoFull-text search across all text columns. For precise filtering use the where parameter instead.
selectNoSoQL SELECT clause — column names, aliases, or aggregates: "state, sum(deaths) as total_deaths". Omit for all columns. To enumerate distinct values of a column, set select to "{column}, count(*) as count" with group="{column}" and order="count DESC".
datasetIdYesFour-by-four dataset identifier (e.g., "bi63-dtpu"). Obtain from cdc_discover_datasets.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
capNoThe requested limit that bounded this response.
rowsNoResult rows with requested fields. Most values are strings (including numbers/dates); geo columns return GeoJSON objects.
errorNoPresent when the call failed. Absent on success.
shownNoNumber of rows returned in this response.
noticeNoGuidance when no rows matched, when offset ran past the end of the result set, when further rows remain, or when the response budget cut the page short — how to verify filters, lower offset, resume paging, or broaden the query.
rowCountNoNumber of rows returned in this response.
truncatedNoTrue when rows exist beyond the ones returned, established by fetching one row more than the limit rather than inferred from the row count. Absent means this response is the complete remainder of the result set.
nextOffsetNoOffset to pass on the next call to resume immediately after the last row returned. Present only when further rows exist and the resume point is within the offset ceiling; a deterministic order clause is what makes the walk gap-free.
effectiveQueryNoThe SoQL clauses sent to Socrata, as `$clause=value` pairs joined by "&". Values read exactly as they were supplied — not URL-encoded — so a clause can be copied back into the matching parameter of another call.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • changedInput schema / properties / limit / description
      Previous value: -"Max rows to return (default 100, max 5000). Fewer come back when the page would cross the 200,000-character response budget; the response says so and gives a nextOffset to resume from."New value: +"Max rows to return (default 100, max 5000). Fewer come back when the page would carry the response past its 200,000-character budget, counted over the whole result — the rows as JSON and as the rendered table together; the response says so and gives a nextOffset to resume from."
    • changedOutput schema / properties / notice / description
      Previous value: -"Guidance when no rows matched, when further rows remain, or when the response budget cut the page short — how to verify filters, resume paging, or broaden the query."New value: +"Guidance when no rows matched, when offset ran past the end of the result set, when further rows remain, or when the response budget cut the page short — how to verify filters, lower offset, resume paging, or broaden the query."
  2. Changed2 schema fields changed
    • changedOutput schema / properties / error / properties / data / properties / reason / description
      Previous value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset ID does not exist or has been retired. `no_such_column`: WHERE/SELECT/GROUP/ORDER references a column that does not exist on this dataset. `type_mismatch`: Filter value type does not match the column data type (e.g., quoting a number). `invalid_query`: Socrata rejected the SoQL query for other syntax or semantic reasons. `access_denied`: Socrata returned 403 — typically an ID naming a chart, map, story, file, or external link rather than a tabular dataset. `rate_limited`: Socrata API returns 429 Too Many Requests. `upstream_error`: Socrata data API returned a 5xx server error. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset ID does not exist or has been retired. `no_such_column`: WHERE/SELECT/GROUP/ORDER references a column that does not exist on this dataset. `type_mismatch`: Filter value type does not match the column data type (e.g., quoting a number). `invalid_query`: Socrata rejected the SoQL query for other syntax or semantic reasons. `not_queryable`: The ID names a non-tabular asset: the data endpoint answered with rows carrying no fields, and the asset reports no columns. `access_denied`: Socrata returned 403 — typically an ID naming a story, file, or external link rather than a tabular dataset. `rate_limited`: Socrata API returns 429 Too Many Requests. `upstream_error`: Socrata data API returned a 5xx server error. Other values are possible when a failure originates below the handler."
    • changedOutput schema / properties / error / properties / data / properties / reason / examples
      Previous value: -[
      -  "dataset_not_found",
      -  "no_such_column",
      -  "type_mismatch",
      -  "invalid_query",
      -  "access_denied",
      -  "rate_limited",
      -  "upstream_error"
      -]New value: +[
      +  "dataset_not_found",
      +  "no_such_column",
      +  "type_mismatch",
      +  "invalid_query",
      +  "not_queryable",
      +  "access_denied",
      +  "rate_limited",
      +  "upstream_error"
      +]
  3. Changed6 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / additionalProperties
      Added value: +false
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedOutput schema / anyOf
      Added value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "rows",
      +      "rowCount",
      +      "effectiveQuery"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • addedOutput schema / properties / error
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "Present when the call failed. Absent on success.",
      +  "properties": {
      +    "code": {
      +      "description": "JSON-RPC error code for this failure.",
      +      "maximum": 9007199254740991,
      +      "minimum": -9007199254740991,
      +      "type": "integer"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "properties": {
      +        "reason": {
      +          "description": "Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset ID does not exist or has been retired. `no_such_column`: WHERE/SELECT/GROUP/ORDER references a column that does not exist on this dataset. `type_mismatch`: Filter value type does not match the column data type (e.g., quoting a number). `invalid_query`: Socrata rejected the SoQL query for other syntax or semantic reasons. `access_denied`: Socrata returned 403 — typically an ID naming a chart, map, story, file, or external link rather than a tabular dataset. `rate_limited`: Socrata API returns 429 Too Many Requests. `upstream_error`: Socrata data API returned a 5xx server error. Other values are possible when a failure originates below the handler.",
      +          "examples": [
      +            "dataset_not_found",
      +            "no_such_column",
      +            "type_mismatch",
      +            "invalid_query",
      +            "access_denied",
      +            "rate_limited",
      +            "upstream_error"
      +          ],
      +          "type": "string"
      +        },
      +        "recovery": {
      +          "additionalProperties": {},
      +          "description": "Actionable next step for the caller.",
      +          "properties": {
      +            "hint": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "hint"
      +          ],
      +          "type": "object"
      +        },
      +        "retryable": {
      +          "description": "Whether retrying may succeed.",
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "message": {
      +      "description": "Human-readable description of what went wrong.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "code",
      +    "message"
      +  ],
      +  "type": "object"
      +}
    • removedOutput schema / required
      Removed value: -[
      -  "rows",
      -  "rowCount",
      -  "effectiveQuery"
      -]
  4. Changed6 schema fields changed
    • changedInput schema / properties / limit / description
      Previous value: -"Max rows to return (default 100, max 5000)."New value: +"Max rows to return (default 100, max 5000). Fewer come back when the page would cross the 200,000-character response budget; the response says so and gives a nextOffset to resume from."
    • changedInput schema / properties / offset / description
      Previous value: -"Row offset for pagination (max 1,000,000)."New value: +"Row offset for pagination (max 1,000,000). Pair with a deterministic order clause — an offset walk over unordered results can skip or repeat rows."
    • changedInput schema / properties / order / description
      Previous value: -"SoQL ORDER BY clause. Field name with optional ASC/DESC: \"total_deaths DESC\"."New value: +"SoQL ORDER BY clause. Field name with optional ASC/DESC: \"total_deaths DESC\". Set one whenever paging with offset: SODA does not order results implicitly, so consecutive offsets without a deterministic order can skip or repeat rows. When the dataset has no natural unique column, Socrata's documented minimum tie-breaker is the system field `:id`, present on every dataset — order=\":id\"."
    • addedOutput schema / properties / nextOffset
      Added value: +{
      +  "description": "Offset to pass on the next call to resume immediately after the last row returned. Present only when further rows exist and the resume point is within the offset ceiling; a deterministic order clause is what makes the walk gap-free.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / notice / description
      Previous value: -"Guidance when no rows matched or results were truncated — how to verify filters, paginate, or broaden the query."New value: +"Guidance when no rows matched, when further rows remain, or when the response budget cut the page short — how to verify filters, resume paging, or broaden the query."
    • changedOutput schema / properties / truncated / description
      Previous value: -"True when the result row count hit the requested limit and may be incomplete."New value: +"True when rows exist beyond the ones returned, established by fetching one row more than the limit rather than inferred from the row count. Absent means this response is the complete remainder of the result set."
  5. Changed2 schema fields changed
    • changedInput schema / properties / domain / description
      Previous value: -"CDC Socrata portal hosting the dataset. Must match the portal the dataset lives on: \"data.cdc.gov\" (default) or \"chronicdata.cdc.gov\" (PLACES and other chronic-disease/small-area datasets)."New value: +"CDC Socrata host to query. \"data.cdc.gov\" (default) and \"chronicdata.cdc.gov\" front the same catalog, so a four-by-four ID returns the same rows from either and the default works whichever host the dataset was found on."
    • changedOutput schema / properties / effectiveQuery / description
      Previous value: -"Assembled SoQL query string sent to the Socrata API."New value: +"The SoQL clauses sent to Socrata, as `$clause=value` pairs joined by \"&\". Values read exactly as they were supplied — not URL-encoded — so a clause can be copied back into the matching parameter of another call."
  6. Changed2 schema fields changed
    • changedInput schema / properties / offset / description
      Previous value: -"Row offset for pagination."New value: +"Row offset for pagination (max 1,000,000)."
    • changedInput schema / properties / offset / maximum
      Previous value: -9007199254740991New value: +1000000
  7. Changed1 schema field changed
    • addedInput schema / properties / domain
      Added value: +{
      +  "default": "data.cdc.gov",
      +  "description": "CDC Socrata portal hosting the dataset. Must match the portal the dataset lives on: \"data.cdc.gov\" (default) or \"chronicdata.cdc.gov\" (PLACES and other chronic-disease/small-area datasets).",
      +  "enum": [
      +    "data.cdc.gov",
      +    "chronicdata.cdc.gov"
      +  ],
      +  "type": "string"
      +}
  8. Changed1 schema field changed
    • changedInput schema / properties / where / description
      Previous value: -"SoQL WHERE clause. Strings must be single-quoted: \"state='California' AND year=2020\"."New value: +"SoQL WHERE clause. Strings must be single-quoted: \"state='California' AND year=2020\". If a column name matches a SoQL keyword (group, select, where, order, limit, offset, having, search), wrap it in backticks: \"`group`='By Year'\"."
  9. Changed4 schema fields changed
    • addedOutput schema / properties / cap
      Added value: +{
      +  "description": "The requested limit that bounded this response.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / notice / description
      Previous value: -"Guidance when no rows matched — suggests how to verify filter values or broaden the WHERE clause."New value: +"Guidance when no rows matched or results were truncated — how to verify filters, paginate, or broaden the query."
    • addedOutput schema / properties / shown
      Added value: +{
      +  "description": "Number of rows returned in this response.",
      +  "type": "number"
      +}
    • addedOutput schema / properties / truncated
      Added value: +{
      +  "description": "True when the result row count hit the requested limit and may be incomplete.",
      +  "type": "boolean"
      +}
  10. Changed1 schema field changed
    • changedInput schema / properties / datasetId / description
      Previous value: -"Four-by-four dataset identifier (e.g., \"bi63-dtpu\")."New value: +"Four-by-four dataset identifier (e.g., \"bi63-dtpu\"). Obtain from cdc_discover_datasets."
  11. Changed4 schema fields changed
    • addedOutput schema / properties / effectiveQuery
      Added value: +{
      +  "description": "Assembled SoQL query string sent to the Socrata API.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / notice
      Added value: +{
      +  "description": "Guidance when no rows matched — suggests how to verify filter values or broaden the WHERE clause.",
      +  "type": "string"
      +}
    • removedOutput schema / properties / query
      Removed value: -{
      -  "description": "Assembled SoQL query string sent to Socrata.",
      -  "type": "string"
      -}
    • changedOutput schema / required
      Previous value: -[
      -  "rows",
      -  "rowCount",
      -  "query"
      -]New value: +[
      +  "rows",
      +  "rowCount",
      +  "effectiveQuery"
      +]
  12. Changed4 schema fields changed
    • changedInput schema / properties / limit / default
      Previous value: -1000New value: +100
    • changedInput schema / properties / limit / description
      Previous value: -"Max rows to return (default 1000, max 5000)."New value: +"Max rows to return (default 100, max 5000)."
    • changedInput schema / properties / search / description
      Previous value: -"Full-text search across all text columns (maps to $q). For precise filtering use where instead."New value: +"Full-text search across all text columns. For precise filtering use the where parameter instead."
    • changedOutput schema / properties / query / description
      Previous value: -"Assembled SoQL query string (for debugging)."New value: +"Assembled SoQL query string sent to Socrata."
  13. Changed1 schema field changed
    • changedInput schema / properties / select / description
      Previous value: -"SoQL SELECT clause. Column names, aliases, aggregates: \"state, sum(deaths) as total_deaths\". Omit for all columns."New value: +"SoQL SELECT clause — column names, aliases, or aggregates: \"state, sum(deaths) as total_deaths\". Omit for all columns. To enumerate distinct values of a column, set select to \"{column}, count(*) as count\" with group=\"{column}\" and order=\"count DESC\"."
  14. First observed

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotation readOnlyHint=true already establishes that this is a safe read operation, so the bar for additional disclosure is lower. The description adds little operational behavior beyond capabilities; pagination limits, offset-ordering caveats, and default-limit behavior are documented in the parameter schemas rather than the description itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The main action and supported SoQL capabilities are front-loaded, and the second sentence gives the prerequisite tool workflow succinctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and 100% parameter schema coverage, the description need not re-explain return values or parameter formats. It names the right companion tools for dataset-ID discovery and column inspection. It is slightly incomplete because it never distinguishes this tool from cdc_query_wonder, the other query-oriented sibling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed parameter documentation for where, order, offset, limit, select, search, domain, and datasetId. The description only lightly maps capabilities to parameters (e.g., 'filtering' implies where), so it adds no meaningful parameter semantics beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-object pair, 'Execute a SoQL query against any CDC dataset,' and enumerates the supported capabilities: filtering, aggregation, sorting, full-text search, and field selection. This clearly identifies the tool's role and separates it from the discovery and schema-inspection siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit workflow guidance: use cdc_discover_datasets to find dataset IDs and cdc_get_dataset_schema to inspect columns before querying. It does not explicitly contrast itself with cdc_query_wonder, so the boundary against the other query tool is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.