cdc-health-mcp-server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cdc-health-mcp-serverfind CDC datasets on diabetes prevalence"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Public Hosted Server: https://cdc.caseyjhand.com/mcp
Overview
CDC public health data — the Socrata-based CDC Open Data portal, plus CDC WONDER, a separate CDC system for national mortality statistics. Search the catalog, inspect dataset schemas, and run SoQL queries across vaccination, surveillance, and behavioral-risk data, or query WONDER for deaths, population, and death rates by year, age, sex, and race. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.
Tools
Tool | Description |
| Search the CDC dataset catalog by keyword, category, or tag |
| Fetch column schema, row count, and metadata for a dataset |
| List the catalog's category and tag values with their entry counts |
| Execute SoQL queries — filter, aggregate, sort, full-text search, select fields |
| Query CDC WONDER for national mortality, population, and death rates across five databases |
Resources
Resource | Description |
| Top 50 most-viewed catalog entries, for orientation |
| Dataset metadata and the first 100 columns for a specific dataset |
Both resources mirror data also reachable via cdc_discover_datasets and cdc_get_dataset_schema, for clients that surface resources but not tools.
Prompts
Prompt | Description |
| Guided workflow for investigating a public health question across CDC data |
Related MCP server: mcp-data-connecticut
Capability reference
cdc_discover_datasets tool
domainselectsdata.cdc.gov(default) orchronicdata.cdc.gov— both front the same catalog, so switching hosts neither widens nor narrows a searchquery,category, andtagsfilters — tags union (a dataset matches on any one tag, so each tag added widens the result set), whilequeryandcategoryintersect with the tag setUp to 100 results per page (default 10);
offsetis capped at 9999, andoffset + limitmust not exceed 10,000 — Socrata's catalog ceilingorder:dataset_id(default) sorts deterministically for stable pagination;relevanceranks by best match but is not stably paginable across pagesEach result carries
columnCount— a value of 0 marks a non-tabular asset (chart, map, story, file, or href) that yields no data from the other tools;assetTypeis descriptive onlydescriptionis converted to plain text before it is cut to 300 characters, so the budget buys visible text rather than the HTML tags catalog entries arrive wrapped inEnrichment carries
totalCountandappliedFilters; anoticedistinguishes an offset past the end of the result set from a search that matched nothing, and resolves acategory/tagsvalue that matched nothing against the catalog vocabulary —category: "Vaccination"comes back naming"Vaccinations" (89 datasets)rather than advising a broader search. The vocabulary is read only on that branch, so an ordinary search costs no extra request
cdc_get_dataset_schema tool
Accepts a four-by-four
datasetId(e.g.bi63-dtpu) and the samedomainenum as the other Socrata toolsReturns the first 100 columns by default (
column_limit, max 500) — catalog schemas run 3 to 322 columns, so ordinary datasets arrive whole; wider ones reporttotalCount,truncated, andnextOffsetto pass back ascolumn_offsetA
column_offsetat or past the column count returns an empty window rather than an errorFails with
not_queryablewhen the ID names a non-tabular catalog asset, rather than returning an empty column listrowCountprefers a livecount(*)fetched alongside the metadata;rowCountSourcesaysliveorcached, since Socrata's cached figure is built once and can understate an actively-updated dataset by a third or more. A failed count falls back to the cached figure and never fails the schema responsedescriptionis returned in full as plain text — markup stripped, entity references decoded; onlycdc_discover_datasetstruncates it
cdc_list_catalog_vocabulary tool
The controlled vocabularies
cdc_discover_datasets'categoryandtagsfilters are matched against — a value the catalog does not carry matches nothing, which is indistinguishable from a real value with no resultsAll 55 categories return whole (~3 KB); the tag vocabulary runs to 1,583 values, so tags are ranked by entry count and paged with
tag_limit(default 50, max 500) andtag_offsetThe service asks the tag endpoint for the whole vocabulary explicitly — left to its default it returns 100 values and reports
resultSetSize: 100beside them, so the under-count reads as completefilternarrows both vocabularies before the page is cut, matching on whole words in either direction ("vaccin"reachesVaccinationsandcovid-19 vaccination). Not fuzzy — a misspelling returns nothing rather than a guessEnrichment carries
vocabularySize, the matchedcategoryCount/tagCount, andtruncated/shown/cap/nextOffset;truncationCeilingbounds every omitted tag, since the list is ranked by the same countBoth hosts publish the same vocabulary — measured live,
data.cdc.govandchronicdata.cdc.govreturn the identical 55 categories and 1,583 tags
cdc_query_dataset tool
Full SoQL support —
select,where,group,having,order, plus full-textsearchacross text columnsUp to 5,000 rows per request (default 100);
offsetcapped at 1,000,000truncatedis measured by an over-fetch probe (one row past the limit), never guessed from the row count; the whole response —structuredContentandcontent[]together — is bounded by a 200,000-character budget, so a wide page can end short oflimitwith anextOffsetAn empty page at
offset > 0is diagnosed with one probe at offset 0: thenoticesays whether the offset ran past the end or the query matches nothing, and names both causes if the probe failseffectiveQueryechoes the SoQL clauses sent in their original text, not URL-encoded, so a clause can be copied back into the parameter it came fromFails with
not_queryablewhen every returned row carries no fields and the asset reports no columns — a chart or map ID, which Socrata answers 200 with a body of empty objects. When the asset does have columns, the same shape is a null-only projection and comes back as a success with a noticeAll response values are strings (SODA v2.1) — parse per the column's
dataTypefrom the schema
cdc_query_wonder tool
database selects which of five mortality databases answers the query:
Value | CDC database | Years | Race groups |
|
| D76 — Underlying Cause of Death | 1999–2020 | 4 bridged | — |
| D176 — Provisional Mortality Statistics | 2018 → current year | 6 single-race | yes |
| D158 — Underlying Cause of Death, Single Race | 2018–2024 | 6 single-race | — |
| D77 — Multiple Cause of Death | 1999–2020 | 4 bridged | yes |
| D157 — Multiple Cause of Death, Single Race | 2018–2024 | 6 single-race | yes |
group_by: 1–4 ofyear,age_group,sex,race, each at most once; national totals only — no sub-national breakdown at any settingcause_icd10andmcd_icd10take one ICD-10 code or range, or a list of up to 50 matched as one union — e.g. the drug-overdose set["X40","X41","X42","X43","X44","X60","X61","X62","X63","X64","X85","Y10","Y11","Y12","Y13","Y14"]returns one series with one set of rates. A range must be a chapter or block of WONDER's ICD-10 tree (X40-X49); any other span (X40-X44) is rejected, so list its codes insteadmcd_icd10matches a cause recorded anywhere on the death certificate rather than only the underlying cause; accepted only byprovisional,multiple_1999_2020, andmultiple_2018_2024— the others reject itage_groupsmust include"NS"(age not recorded) to match an unfiltered total; ayear_rangeoutside the selected database's span is rejected with that span namedMeasure cells CDC withholds or flags (
Suppressed,Unreliable,Not Applicable) readnullinrowsand are named per cell incellNotes; whole rows CDC hides (zero or suppressed deaths) are absent fromrowswith no gap marker — checkmessagesEach response is bounded by a 200,000-character budget over
structuredContentandcontent[]together, so a large grouping comes back a page at a time with an exacttotalCount,truncated, and anextOffsetto continue from;limit(max 5,000) takes smaller pages andoffset(max 10,000) resumesConsecutive requests are spaced 16 seconds automatically — CDC rejects anything sent less than 15 seconds after the prior response finished, measured across all five databases. Concurrent calls queue and run one at a time, each queued call adding about 16 seconds
CDC's caveats and notices keep their methodology links as Markdown links
cdc://datasets resource
Top 50 CDC catalog entries by popularity, each carrying
assetTypeandcolumnCountfor orientationcolumnCount: 0marks a non-tabular entry (chart, map, story, file, or href); usecdc_discover_datasetsfor full catalog search with filtering and pagination
cdc://datasets/{datasetId} resource
Dataset metadata plus the first 100 columns as
application/json;datasetIdis a four-by-four identifier fromcdc_discover_datasetsCarries the dataset's total
columnCountand atruncatedflag; wider schemas continue viacdc_get_dataset_schemawithcolumn_offsetTakes no query-parameter selector — an RFC 6570
{?column_limit,column_offset}template would stop the barecdc://datasets/{datasetId}form from matching at all
analyze_health_trend prompt
Arguments:
topicrequired;timeRangeandgeographyoptionalReturns one user message that routes the question to CDC WONDER (national mortality, 1999–current, ICD-10-filterable) or the Socrata catalog (everything else), then walks discover → inspect → baseline query → compare → synthesize
Routing is prose for the reader to act on — the handler does not classify the topic itself
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
CDC-specific:
Wraps the Socrata SODA API v2.1 (the CDC Open Data portal, ~1,080 datasets) — no auth required, optional app token for higher rate limits
Adds CDC WONDER mortality access (
cdc_query_wonder) — a separate XML-over-HTTP CDC system, spanning five mortality databases from 1999 through the current yearDiscovery-first workflow for a heterogeneous catalog — discover, inspect schema, then query
Two Socrata hosts via the
domaininput (data.cdc.gov,chronicdata.cdc.gov), allowlisted at the schema level — both front one tenant, so assets like PLACES and the Heart Disease & Stroke Atlas are reachable from eitherConservative request spacing for both APIs — no rate-limit headers from Socrata, and CDC WONDER requests are spaced 16 seconds apart automatically
Agent-friendly output:
Pagination and truncation disclosed on every tool —
totalCount/truncated/nextOffset(orshown/cap) rather than a bare row count, so an agent can tell a complete result from a page of oneTyped error contracts with a
recoveryhint on every declared reason (e.g.not_queryable,page_out_of_range) — actionable next steps, not just an error codeUpstream data gaps stay visible rather than silently dropped — CDC's status tokens (
Suppressed,Unreliable,Not Applicable) are named per cell incellNotes, and hidden-row notices surface inmessageseffectiveQueryechoes the exact query sent (SoQL clauses or a WONDER summary), so a result is reproducible and a clause can be copied back into its parameter
Getting started
Public Hosted Instance
A public instance is available at https://cdc.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "streamable-http",
"url": "https://cdc.caseyjhand.com/mcp"
}
}
}Self-Hosted / Local
Add the following to your MCP client configuration file.
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/cdc-health-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with npx (no Bun required):
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/cdc-health-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with Docker:
{
"mcpServers": {
"cdc-health-mcp-server": {
"type": "stdio",
"command": "docker",
"args": ["run", "-i", "--rm", "-e", "MCP_TRANSPORT_TYPE=stdio", "ghcr.io/cyanheads/cdc-health-mcp-server:latest"]
}
}
}For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcpPrerequisites
Bun v1.4.0 or higher.
Optional: Socrata app token for higher rate limits.
Installation
Clone the repository:
git clone https://github.com/cyanheads/cdc-health-mcp-server.gitNavigate into the directory:
cd cdc-health-mcp-serverInstall dependencies:
bun installConfigure environment:
cp .env.example .env
# edit .env and set optional overridesConfiguration
Variable | Description | Default |
| Transport: |
|
| HTTP server port |
|
| HTTP session posture: |
|
| Authentication: |
|
| Log level ( |
|
| Directory for log files (Node.js only) |
|
| Storage backend: |
|
| Socrata app token for higher rate limits | — |
| SODA host for requests that name no |
|
| Base URL for Socrata Discovery API |
|
| Enable OpenTelemetry instrumentation (spans, metrics, completion logs) |
|
See .env.example for the full list of optional overrides.
Running the server
Local development
Build and run the production version:
# One-time build bun run rebuild # Run the built server bun run start:http # or bun run start:stdioRun checks and tests:
bun run devcheck # Lints, formats, type-checks, and more bun run test # Runs the test suite
Docker
docker build -t cdc-health-mcp-server .
docker run --rm -p 3010:3010 cdc-health-mcp-serverThe Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/cdc-health-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
Directory | Purpose |
|
|
| Server-specific environment variable parsing and validation with Zod. |
| Tool definitions ( |
| Resource definitions. Catalog overview and dataset detail. |
| Prompt definitions. Health trend analysis workflow. |
| Socrata SODA API service layer — HTTP client, catalog search, metadata, queries. |
| CDC WONDER service layer — XML request builder and response parser. |
| Shared helpers, including |
| Unit and integration tests mirroring |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
Handlers throw, framework catches — no
try/catchin tool logicUse
ctx.logfor logging,ctx.statefor storageRegister new tools and resources in the
createApp()arraysWrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run testLicense
Apache-2.0 — see LICENSE for details.
This server cannot be deployed
Maintenance
Related MCP Connectors
CDC MCP — wraps CDC open data via Socrata API (data.cdc.gov)
Search and query government open-data portals (Socrata SODA API).
Query and explore Nova Scotia open datasets via the Socrata SODA API.
Utah state open data via Socrata SoQL — government, grants, health, demographics.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides access to 73 CDC public health datasets covering disease surveillance, vaccination tracking, behavioral risk factors, environmental health, and outbreak detection across 18 surveillance systems through the Socrata Open Data API.2MIT
- AlicenseNot gradedqualityBmaintenanceEnables querying and searching Connecticut Open Data via Socrata SoQL API, allowing AI agents to access state agency, public health, education, and transportation datasets using natural language or structured queries.249 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables querying and exploring FCC Open Data datasets via Socrata SoQL, including dataset search, metadata retrieval, and data querying.242 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables searching and querying Orlando Open Data datasets via Socrata SoQL, providing access to metadata and data rows through natural language or direct tool calls.262 npmMIT