Skip to main content
Glama

Cdc Discover Datasets

cdc_discover_datasets
Read-only

Search the CDC dataset catalog by keyword, category, or tag. Returns IDs, names, truncated descriptions, asset types, column counts, and update timestamps. The catalog also holds charts, maps, stories, files, and links; an entry whose columnCount is 0 is one of those and yields no data from the other tools. Use cdc_get_dataset_schema for the full column list of a chosen dataset.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
tagsNoFilter by domain tags (e.g., ["covid19", "surveillance"]). Tags widen the search instead of narrowing it — a dataset matches when it carries any one of them, so every tag added returns more results, and an unrecognized tag matches nothing and leaves the result set unchanged. Values match the catalog's own tag vocabulary, case-insensitively; call cdc_list_catalog_vocabulary for the values in use with their entry counts, or read the tags field on any result. To narrow, combine tags with query or category, which intersect with the tag set.
limitNoResults to return (default 10, max 100). offset plus limit must not exceed 10000.
orderNoResult ordering. "dataset_id" (default) sorts deterministically by each dataset's unique catalog ID — required for stable offset pagination, since consecutive pages form a gap-free, duplicate-free traversal. "relevance" returns best-match ranking for keyword search but is not stably paginable across pages, so walking offsets can skip or repeat datasets.dataset_id
queryNoFull-text search across dataset names and descriptions (e.g., "diabetes mortality", "lead exposure children").
domainNoCDC Socrata host to search. "data.cdc.gov" (default) and "chronicdata.cdc.gov" front the same catalog and return the same entries, so switching hosts neither widens nor narrows a search — chronic-disease and small-area collections such as PLACES, the Heart Disease & Stroke Atlas, and Environmental Public Health Tracking are found from either.data.cdc.gov
offsetNoPagination offset for browsing beyond first page (max 9999). offset plus limit must not exceed 10000; both CDC portals hold well under two thousand entries, so offsets near that ceiling page past the end of the catalog.
categoryNoFilter by domain category (e.g., "NNDSS", "Vaccinations", "Behavioral Risk Factors"). Values come from the catalog's own vocabulary and are matched exactly, including case — "Vaccinations" matches 89 entries while "Vaccination" matches none. Call cdc_list_catalog_vocabulary for every value with its entry count rather than guessing at one.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoPresent when the call failed. Absent on success.
noticeNoGuidance when the page came back empty — the catalog values closest to a category or tag filter that matched nothing, each with its entry count; how to broaden a search when no value is close; or the size of the result set when the offset ran past its end.
datasetsNoMatching datasets.
totalCountNoTotal matching datasets in the catalog (for pagination).
appliedFiltersNoFilters applied to this query; absent fields indicate no filter on that dimension. Query, category, and tags intersect with each other, but multiple tags union.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changed
    • changedInput schema / properties / category / description
      Previous value: -"Filter by domain category (e.g., \"NNDSS\", \"Vaccinations\", \"Behavioral Risk Factors\")."New value: +"Filter by domain category (e.g., \"NNDSS\", \"Vaccinations\", \"Behavioral Risk Factors\"). Values come from the catalog's own vocabulary and are matched exactly, including case — \"Vaccinations\" matches 89 entries while \"Vaccination\" matches none. Call cdc_list_catalog_vocabulary for every value with its entry count rather than guessing at one."
    • changedInput schema / properties / tags / description
      Previous value: -"Filter by domain tags (e.g., [\"covid19\", \"surveillance\"]). Tags widen the search instead of narrowing it — a dataset matches when it carries any one of them, so every tag added returns more results, and an unrecognized tag matches nothing and leaves the result set unchanged. Values match the catalog's own tag vocabulary, case-insensitively; the tags field on each result shows which values are in use. To narrow, combine tags with query or category, which intersect with the tag set."New value: +"Filter by domain tags (e.g., [\"covid19\", \"surveillance\"]). Tags widen the search instead of narrowing it — a dataset matches when it carries any one of them, so every tag added returns more results, and an unrecognized tag matches nothing and leaves the result set unchanged. Values match the catalog's own tag vocabulary, case-insensitively; call cdc_list_catalog_vocabulary for the values in use with their entry counts, or read the tags field on any result. To narrow, combine tags with query or category, which intersect with the tag set."
    • changedOutput schema / properties / datasets / items / properties / description / description
      Previous value: -"Dataset description when provided by the catalog, truncated to 300 characters. Fetch the full text via cdc_get_dataset_schema."New value: +"Dataset description when provided by the catalog, as plain text — markup stripped and entity references decoded — then truncated to 300 characters, so the limit bounds visible text. Fetch the full text via cdc_get_dataset_schema."
    • changedOutput schema / properties / notice / description
      Previous value: -"Guidance when the page came back empty — how to broaden a search that matched nothing, where to check tag values when a tag filter was applied, or the size of the result set when the offset ran past its end."New value: +"Guidance when the page came back empty — the catalog values closest to a category or tag filter that matched nothing, each with its entry count; how to broaden a search when no value is close; or the size of the result set when the offset ran past its end."
  2. Changed6 schema fields changed
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / additionalProperties
      Added value: +false
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedOutput schema / anyOf
      Added value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "datasets",
      +      "totalCount",
      +      "appliedFilters"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • addedOutput schema / properties / error
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "Present when the call failed. Absent on success.",
      +  "properties": {
      +    "code": {
      +      "description": "JSON-RPC error code for this failure.",
      +      "maximum": 9007199254740991,
      +      "minimum": -9007199254740991,
      +      "type": "integer"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "properties": {
      +        "reason": {
      +          "description": "Machine-readable failure mode. Declared by this tool: `rate_limited`: Socrata API returns 429 Too Many Requests. `dataset_not_found`: Socrata returned 404 for the catalog endpoint itself — the Discovery API address is wrong or the service moved. `access_denied`: Socrata returned 403 — the catalog refused this request rather than failing to serve it. `upstream_error`: Socrata catalog API returned a 5xx server error. `page_out_of_range`: offset plus limit exceeds 10000, which Socrata's catalog rejects outright. `invalid_query`: Catalog API returned 400 — typically a malformed query or invalid filter value. Other values are possible when a failure originates below the handler.",
      +          "examples": [
      +            "rate_limited",
      +            "dataset_not_found",
      +            "access_denied",
      +            "upstream_error",
      +            "page_out_of_range",
      +            "invalid_query"
      +          ],
      +          "type": "string"
      +        },
      +        "recovery": {
      +          "additionalProperties": {},
      +          "description": "Actionable next step for the caller.",
      +          "properties": {
      +            "hint": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "hint"
      +          ],
      +          "type": "object"
      +        },
      +        "retryable": {
      +          "description": "Whether retrying may succeed.",
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "message": {
      +      "description": "Human-readable description of what went wrong.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "code",
      +    "message"
      +  ],
      +  "type": "object"
      +}
    • removedOutput schema / required
      Removed value: -[
      -  "datasets",
      -  "totalCount",
      -  "appliedFilters"
      -]
  3. Changed9 schema fields changed
    • changedInput schema / properties / domain / description
      Previous value: -"CDC Socrata portal to search. \"data.cdc.gov\" (default) is the main CDC catalog; \"chronicdata.cdc.gov\" hosts chronic-disease and small-area datasets (PLACES, the Heart Disease & Stroke Atlas, Environmental Public Health Tracking)."New value: +"CDC Socrata host to search. \"data.cdc.gov\" (default) and \"chronicdata.cdc.gov\" front the same catalog and return the same entries, so switching hosts neither widens nor narrows a search — chronic-disease and small-area collections such as PLACES, the Heart Disease & Stroke Atlas, and Environmental Public Health Tracking are found from either."
    • changedInput schema / properties / limit / description
      Previous value: -"Results to return (default 10, max 100)."New value: +"Results to return (default 10, max 100). offset plus limit must not exceed 10000."
    • changedInput schema / properties / offset / description
      Previous value: -"Pagination offset for browsing beyond first page (max 9999)."New value: +"Pagination offset for browsing beyond first page (max 9999). offset plus limit must not exceed 10000; both CDC portals hold well under two thousand entries, so offsets near that ceiling page past the end of the catalog."
    • changedInput schema / properties / tags / description
      Previous value: -"Filter by domain tags (e.g., [\"covid19\", \"surveillance\"])."New value: +"Filter by domain tags (e.g., [\"covid19\", \"surveillance\"]). Tags widen the search instead of narrowing it — a dataset matches when it carries any one of them, so every tag added returns more results, and an unrecognized tag matches nothing and leaves the result set unchanged. Values match the catalog's own tag vocabulary, case-insensitively; the tags field on each result shows which values are in use. To narrow, combine tags with query or category, which intersect with the tag set."
    • changedOutput schema / properties / appliedFilters / description
      Previous value: -"Filters applied to this query; absent fields indicate no filter on that dimension."New value: +"Filters applied to this query; absent fields indicate no filter on that dimension. Query, category, and tags intersect with each other, but multiple tags union."
    • changedOutput schema / properties / appliedFilters / properties / tags / description
      Previous value: -"Tag filters used."New value: +"Tag filters used — a dataset matched when it carried any one of them."
    • addedOutput schema / properties / datasets / items / properties / assetType
      Added value: +{
      +  "description": "Catalog asset type as Socrata reports it — \"dataset\", \"filter\", \"chart\", \"map\", \"story\", \"file\", or \"href\". Descriptive only: \"filter\" entries carry real columns and query normally, while \"chart\" and \"map\" entries do not. Read columnCount, not this field, to decide whether an entry is queryable.",
      +  "type": "string"
      +}
    • changedOutput schema / properties / datasets / items / properties / columnCount / description
      Previous value: -"Number of columns in the dataset when reported by the catalog."New value: +"Number of columns in the dataset when reported by the catalog. A count of 0 means the entry is not tabular — cdc_get_dataset_schema and cdc_query_dataset return no usable data for it."
    • changedOutput schema / properties / notice / description
      Previous value: -"Guidance when no datasets matched — echoes the applied filters and suggests how to broaden the search."New value: +"Guidance when the page came back empty — how to broaden a search that matched nothing, where to check tag values when a tag filter was applied, or the size of the result set when the offset ran past its end."
  4. Changed1 schema field changed
    • addedInput schema / properties / order
      Added value: +{
      +  "default": "dataset_id",
      +  "description": "Result ordering. \"dataset_id\" (default) sorts deterministically by each dataset's unique catalog ID — required for stable offset pagination, since consecutive pages form a gap-free, duplicate-free traversal. \"relevance\" returns best-match ranking for keyword search but is not stably paginable across pages, so walking offsets can skip or repeat datasets.",
      +  "enum": [
      +    "dataset_id",
      +    "relevance"
      +  ],
      +  "type": "string"
      +}
  5. Changed6 schema fields changed
    • addedInput schema / properties / domain
      Added value: +{
      +  "default": "data.cdc.gov",
      +  "description": "CDC Socrata portal to search. \"data.cdc.gov\" (default) is the main CDC catalog; \"chronicdata.cdc.gov\" hosts chronic-disease and small-area datasets (PLACES, the Heart Disease & Stroke Atlas, Environmental Public Health Tracking).",
      +  "enum": [
      +    "data.cdc.gov",
      +    "chronicdata.cdc.gov"
      +  ],
      +  "type": "string"
      +}
    • addedOutput schema / properties / datasets / items / properties / columnCount
      Added value: +{
      +  "description": "Number of columns in the dataset when reported by the catalog.",
      +  "type": "number"
      +}
    • removedOutput schema / properties / datasets / items / properties / columnNames
      Removed value: -{
      -  "description": "Available column field names when provided.",
      -  "items": {
      -    "type": "string"
      -  },
      -  "type": "array"
      -}
    • addedOutput schema / properties / datasets / items / properties / columnSample
      Added value: +{
      +  "description": "First 8 column field names as a preview. Call cdc_get_dataset_schema for the full column list with data types.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • removedOutput schema / properties / datasets / items / properties / columnTypes
      Removed value: -{
      -  "description": "Column data types (parallel to columnNames) when provided.",
      -  "items": {
      -    "type": "string"
      -  },
      -  "type": "array"
      -}
    • changedOutput schema / properties / datasets / items / properties / description / description
      Previous value: -"Dataset description when provided by the catalog."New value: +"Dataset description when provided by the catalog, truncated to 300 characters. Fetch the full text via cdc_get_dataset_schema."
  6. Changed3 schema fields changed
    • changedOutput schema / properties / appliedFilters / description
      Previous value: -"Filters that were applied to this query; absent fields indicate no filter on that dimension."New value: +"Filters applied to this query; absent fields indicate no filter on that dimension."
    • addedOutput schema / properties / notice
      Added value: +{
      +  "description": "Guidance when no datasets matched — echoes the applied filters and suggests how to broaden the search.",
      +  "type": "string"
      +}
    • changedOutput schema / properties / totalCount / description
      Previous value: -"Total matching datasets (for pagination)."New value: +"Total matching datasets in the catalog (for pagination)."
  7. Changed2 schema fields changed
    • changedOutput schema / properties / appliedFilters / description
      Previous value: -"Filters applied to this search (echoed for diagnostics)."New value: +"Filters that were applied to this query; absent fields indicate no filter on that dimension."
    • changedOutput schema / properties / datasets / items / properties / name / description
      Previous value: -"Dataset name."New value: +"Dataset display name from the catalog (e.g., \"Provisional COVID-19 Deaths by Sex and Age\")."
  8. Changed1 schema field changed
    • addedOutput schema / properties / datasets / items / description
      Added value: +"A single dataset catalog entry."
  9. Changed11 schema fields changed
    • changedOutput schema / properties / datasets / items / properties / category / description
      Previous value: -"Domain category."New value: +"Domain category when provided."
    • changedOutput schema / properties / datasets / items / properties / columnNames / description
      Previous value: -"Available column field names."New value: +"Available column field names when provided."
    • removedOutput schema / properties / datasets / items / properties / columnNames / items / description
      Removed value: -"Column name"
    • changedOutput schema / properties / datasets / items / properties / columnTypes / description
      Previous value: -"Column data types (parallel to columnNames)."New value: +"Column data types (parallel to columnNames) when provided."
    • removedOutput schema / properties / datasets / items / properties / columnTypes / items / description
      Removed value: -"Column type"
    • changedOutput schema / properties / datasets / items / properties / description / description
      Previous value: -"Dataset description."New value: +"Dataset description when provided by the catalog."
    • changedOutput schema / properties / datasets / items / properties / pageViews / description
      Previous value: -"Total page views."New value: +"Total page views when provided."
    • changedOutput schema / properties / datasets / items / properties / tags / description
      Previous value: -"Domain tags."New value: +"Domain tags when provided."
    • removedOutput schema / properties / datasets / items / properties / tags / items / description
      Removed value: -"Tag"
    • changedOutput schema / properties / datasets / items / properties / updatedAt / description
      Previous value: -"Last data update timestamp."New value: +"Last data update timestamp when provided."
    • changedOutput schema / properties / datasets / items / required
      Previous value: -[
      -  "id",
      -  "name",
      -  "description",
      -  "category",
      -  "tags",
      -  "columnNames",
      -  "columnTypes",
      -  "updatedAt",
      -  "pageViews"
      -]New value: +[
      +  "id",
      +  "name"
      +]
  10. First observed

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses what the result contains, that descriptions returned are truncated, and that the catalog includes non-dataset assets such as charts, maps, stories, files, and links. The columnCount 0 caveat is particularly valuable behavioral context that annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The main description is only three sentences, front-loads the action and scope, and immediately states the return payload. The supporting parameter descriptions are verbose but earn their length by explaining edge cases like tag widening, offset stability, and exact-match categories.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description fully covers search scope, result contents, the non-dataset caveat, and the next-step sibling tool. Since an output schema exists and annotations plus parameter descriptions are already rich, nothing an agent needs to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already gives richly detailed semantics for every parameter, including enums, defaults, pagination constraints, and tagging behavior. The main description adds little parameter-level meaning beyond what the schema documents, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and object: 'Search the CDC dataset catalog by keyword, category, or tag,' and then enumerates the returned fields. It also names the sibling cdc_get_dataset_schema as the follow-up for full column details, making the tool's role distinct from the other catalog tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly routes the agent to cdc_get_dataset_schema when a full column list is needed and warns that columnCount 0 entries yield no data from other tools. This tells the agent not only what this tool is for, but when to move to a sibling. The parameter descriptions additionally direct to cdc_list_catalog_vocabulary for vocabulary values.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.