Skip to main content
Glama
cyanheads

census-mcp-server

by cyanheads

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework

Public Hosted Server: https://census.caseyjhand.com/mcp


Overview

U.S. Census Bureau data — datasets, variables, and geography — via the Census Data API, TIGERweb, and the Census Geocoder. Discover datasets and variables, resolve place names or addresses to FIPS codes, and query or rank demographic, economic, and housing estimates across geographies from any MCP client. Runs as a stdio process, a local Streamable HTTP server, or the public hosted endpoint above.

Tools

Tool

Description

census_list_datasets

Browse available Census Bureau datasets (ACS5 and ACS1 with their profile, subject, and comparison tables, ACS Supplemental Estimates and Selected Population Profiles, Population Estimates, Decennial, County Business Patterns, Economic Census, Nonemployer Statistics) with vintage years and dataset codes.

census_list_geographies

List the geography levels supported by a dataset and year, with parent requirements and example FIPS values.

census_search_variables

Keyword search across variable labels and concept groups. On ACS, returns estimate and margin-of-error codes together.

census_get_variable

Fetch full metadata for one or more variable codes — label, concept, predicate type, universe, MOE sibling — including the annotation and flag columns the data tools accept.

census_list_predicate_values

List the codes a filter dimension accepts (EMPSZES, LFO, POPGROUP, NAICS2017…), from the dataset dictionary or a live wildcard enumeration.

census_resolve_geography

Convert place names (e.g., "King County, WA"), ZIP codes, or street addresses to Census FIPS identifiers via TIGERweb and Census Geocoder.

census_query_data

Query a Census dataset for variables at a specific geography. Returns estimates with MOE, Census sentinel values and withheld business values resolved to their published meanings, and predicate filtering for the business datasets.

census_compare_geographies

Rank and compare variables across multiple geographies — all counties in a state, all states nationally, or a named set. Sorted table output, with the same predicate filtering.


Related MCP server: Census MCP Server

Capability reference

census_list_datasets tool

  • Returns dataset codes, names, descriptions, and available vintage years

  • Covers ACS5 with its Data Profiles, Subject Tables, and Comparison Profiles; ACS1 with its Data Profiles, Subject Tables, Comparison Profiles, and Selected Population Profiles (acs/acs1/spp); ACS 1-Year Supplemental Estimates (acs/acsse); Population Estimates; the 2020 Decennial files — Redistricting (P.L. 94-171), DHC (dec/dhc), Demographic Profile (dec/dp), Supplemental DHC (dec/sdhc), and DDHC-A (dec/ddhca); County Business Patterns (cbp); Economic Census (ecnbasic); and Nonemployer Statistics (nonemp)

  • Each description names the filter predicates the dataset requires and the geography levels it publishes — both vary by dataset

  • Accepts an optional keyword filter

  • Dataset codes (e.g., acs/acs5) are the values to pass to other tools. Every tool ignores their case and takes a two-part code by its last part alone (acs5 is acs/acs5, pl is dec/pl), echoing the resolved code; three-part codes such as acs/acs5/profile must be given in full, and a bare profile fails naming the codes that end in it

  • available_years is exhaustive, not a sample: any other year fails with year_not_available before a request goes out, naming the years that do work. It is narrower than what the Census API hosts — pep/charv reaches its 2020-2022 estimates through the YEAR filter inside the 2023 vintage, the cbp/nonemp vintages left out reject the NAME column every query here sends, and the Census API answers acs/acs1/spp 2008 and 2010 with server errors


census_list_geographies tool

  • Returns one row per geography level — geography_level, whether a parent is required, required_parent_levels, and an example FIPS value

  • geography_level values are the exact inputs to geography_level in census_query_data and census_compare_geographies

  • year defaults to the dataset's latest available vintage

  • dataset_not_found when the dataset code is blank or unrecognized; year_not_available when the dataset has no geography data for the requested year


census_search_variables tool

  • Whole-word search across label and concept: every query word must match (rate never matches "separated"), and when no variable contains them all, the variables with the most words come back with a notice saying so

  • Ranked by where the query appears — the label's last !! segment or the whole concept equal to it first, then the phrase in label and concept, label only, concept only — then by fewer !! segments, shorter concept, and code, so a table total leads its breakdown rows and an estimate leads its margin of error

  • A column shared across tables, such as GEO_ID, is matched on its label only and returned without a concept

  • On ACS datasets, returns estimate (E suffix) and margin-of-error (M suffix) codes together so both can be requested in one query — the ACS comparison profiles and the other families publish no margins of error, and an E-final code there is an ordinary code. A margin's label is the one the Census publishes (Margin of Error!!Median household income…), and search matches a margin on its estimate's label, so the two rank side by side

  • Also surfaces the predicate codes a dataset filters on, such as NAICS2017 in cbp

  • limit is an integer from 1 to 100 (default 20) — out-of-range values are rejected, not clamped; totalMatches says how many matched before the limit

  • Cache-backed: variables.json is fetched once per dataset+year with a configurable TTL (default 24h)


census_get_variable tool

  • Accepts one or more variable codes (trimmed and matched regardless of case, then echoed in the dataset's own spelling) and returns metadata in the same order — label, concept, predicate type, and the table's universe when its groups.json entry publishes one

  • Resolves annotation and flag columns such as B19013_001EA and EMP_F from the Census per-variable endpoint, with attribute_of naming the column each belongs to and attribute_type its kind

  • A column shared across tables, such as GEO_ID, carries no concept — its published one joins every table's

  • On ACS datasets, returns estimate_code/moe_code sibling references, and a margin-of-error code carries its published label with attribute_of naming its estimate and attribute_type MARGIN_OF_ERROR; the comparison profiles and the other families publish no margins of error and carry none of these

  • Also resolves predicate/filter dimension codes (e.g., NAICS2017, SEX) to confirm a dimension exists in a dataset — census_list_predicate_values lists the values it accepts

  • dataset defaults to acs/acs5, year defaults to the dataset's latest available vintage

  • variable_not_found when a code isn't defined in the dataset and year


census_list_predicate_values tool

  • Two routes, picked by where the answer lives: a dimension with a published value list is read from the dataset dictionary, one without is enumerated live by wildcarding it on the data endpoint. NAICS* and POPGROUP always publish one (thousands of codes — narrow them with query); on the current vintages EMPSZES, LFO, RCPSZES, TAXSTAT, and TYPOP publish none, so the live route is the only place their codes appear

  • A dictionary value list is a classification shared across Census products, not a record of what one dataset serves — dec/ddhca declares 5,543 POPGROUP codes and publishes 2,996, cbp declares 6,694 NAICS2017 codes and publishes 2,003. The declared list is checked against the dataset's own published rows and the dead codes are dropped; source says whether that check ran and the notice says how many were withheld

  • Keyword query matches code and label; results are sorted by code and a truncated list is disclosed rather than passed off as complete (limit an integer from 1 to 500, default 50; totalCount says how many matched)

  • ecnbasic publishes TAXSTAT and TYPOP per industry, so within_naics scopes the enumeration — and the notice says the result is complete for that industry alone

  • Live enumerations are cached per dataset, year, dimension, industry scope, and probe measure


census_resolve_geography tool

  • Named places (e.g., "King County, WA") resolve via TIGERweb; street addresses resolve to tract level via Census Geocoder, with the address's block_group_fips and incorporated place_fips alongside

  • The place level covers incorporated places and census-designated places together ("Bethesda, MD" → Bethesda CDP, flagged census_designated_place). A CDP answers a name only when no incorporated place or county has it exactly, so "Paradise, CA" is Paradise town and "Arlington, VA" Arlington County; a CDP's full name ("Arlington CDP, VA") or geography_type: "place" reaches it

  • A 5-digit ZIP (or ZIP+4) resolves to its ZIP Code Tabulation Area (zip code tabulation area) — the ACS's ZIP-shaped area, not cbp's zip code level, which takes the ZIP itself with no resolution

  • The state after a comma can be an abbreviation in either case, a full name ("Chatham County, Georgia"), or a hyphenated list ("NE-IA", scoped by its first state)

  • Auto-detects geography_type for state, county, place, tract, and ZIP; metropolitan/micropolitan statistical areas, combined statistical areas, consolidated cities, and economic places are never auto-detected and need an explicit geography_type, since their names overlap city names

  • economic place returns the 8-digit code ecnbasic 2022 publishes a place under — its county, or 000 when it spans counties, then its place code (Seattle 03363000, Auburn, WA 00003180)

  • Optional county_fips scopes resolution to the county and tract levels only — required when a tract name matches more than one county; county_scope_unsupported when paired with any other level or a street address

  • Matching ignores case. A name no level matches is retried with Saint/St. respelled ("Saint Louis, MO") and with accents ignored ("Dona Ana County, NM"); a statistical area is also retried by its leading city, so a name from an earlier delineation ("Denver-Aurora-Lakewood, CO") still resolves

  • Prefers an exactly-named match over a partial one (e.g., "Kansas City, MO" does not resolve to North Kansas City), across levels too ("King, WA" is King County, not Kingston CDP)

  • A name matching more than one geography returns ambiguous_name, with every candidate's FIPS code and the state that separates them

  • Returns state_fips (→ parent_fips) and fips_summary (→ geography_fips) ready to pass to other tools; a statistical area omits state_fips since it can span several states, and a ZCTA omits it because its source layer carries no state


census_query_data tool

  • Requires FIPS codes (use census_resolve_geography for place names); geography_fips: "*" returns every geography at the level within the parent, and each row carries both geography_fips and the nationally-unique geography_geoid

  • A wildcard returns up to limit rows (default 50, max 500) in GEOID order, and offset pages through the rest; totalCount and truncated say how many rows matched, and the notice names the range returned and the next offset. Every row counts, including each pep/charv record and each category of a "*" predicate

  • Up to 49 variable codes per call, fewer on datasets where label or record columns are added: the Census API accepts 50 columns per request and every query also sends NAME. too_many_variables states the exact maximum before any request goes out. Codes are case-insensitive, and an unknown one is variable_not_found

  • Level and parent are checked against the dataset's own geography metadata before querying — parent_required and parent_not_accepted name what's missing or unaccepted rather than surfacing a raw Census 400

  • Optional tract_fips (exactly 6 digits, with a concrete county_fips) scopes a block-group or decennial block query to one tract, so the block group around an address is one call: block group 2 in 53/033/007101

  • Optional predicates map filters the business/pep/dec datasets and acs/acs1/spp (e.g., {"NAICS2017": "5112"}); a dimension left unset applies a Census-chosen default — an all-categories total on some datasets, a single category on others — echoed per row in applied_filters. Keys are case-insensitive and a blank value counts as omitted; "*" returns one row per category, each labelled in record

  • A dataset that publishes more than one record per geography (pep/charv) returns multiple rows, each carrying a record field; pin one with predicates (e.g., {"MONTH": "7"})

  • ACS sentinel values resolve to the Census's published meanings, a controlled estimate's margin of error reads as 0, and a median in an open-ended interval is flagged open_ended. On cbp, ecnbasic, and nonemp, a value the Census withheld (stored as 0 beside a flag such as D) is reported as suppressed with the flag's meaning. A null estimate means the value is either suppressed, a text cell (returned under value), or genuinely empty

  • Requires CENSUS_API_KEY


census_compare_geographies tool

  • Ranks all geographies at a level, or a named geographies list of GEOIDs/bare level codes, in one call; within/within_county scope to a state/county, omit for a national comparison

  • Ranks on one variable's value: sort_by (default the first code; it must be one of the requested codes, or the call fails with sort_by_not_requested), sort_dir (default desc), and limit (an integer from 1 to 500, default 50); totalCount reports how many geographies matched before the limit

  • A count ranks by size, not rate — rank a published percentage for a rate, e.g. S1701_C03_001E (percent below poverty, acs/acs5/subject) or DP04_0047PE (percent renter-occupied, acs/acs5/profile), both available down to tract

  • Same predicates map, variable limit, geography validation, and applied_filters default-echoing as census_query_data, applied to every geography in the ranking

  • A dataset that publishes more than one record per geography (pep/charv), or a "*" predicate, fails with ambiguous_rows unless predicates pins one (e.g., {"MONTH": "7"})

  • Suppressed values carry the same reasons as census_query_data and sort to the end in either direction, withheld business values included; a text value has no ordering, so sorting on it leaves rows tied, and the notice says the rows are not ranked

  • Requires CENSUS_API_KEY


Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Census-specific:

  • In-process variable cache with configurable TTL — variables.json fetched once per dataset+year, searched client-side

  • Three-API backend: Census Data API for data queries, TIGERweb for named-place resolution, Census Geocoder for address-to-tract

  • Automatic retry with backoff on all external API calls

  • FIPS formatting helpers — zero-padded state, county, and tract codes ready to pass between tools

Agent-friendly output:

  • Workflow-oriented tool surface — fips_summary and state_fips return values are ready to pass as geography_fips and parent_fips to the next tool

  • Suppression codes decoded — Census negative sentinel values (e.g., -666666666) and business-dataset withholding flags (e.g., D) surfaced as their published meanings instead of raw numbers or false zeros

  • Recovery hints on errors — ambiguous geography names include candidate lists; missing API key errors include registration URL


Getting started

Public Hosted Instance

A public instance is available at https://census.caseyjhand.com/mcp — no installation required. Point any MCP client at it via Streamable HTTP:

{
  "mcpServers": {
    "census-mcp-server": {
      "type": "streamable-http",
      "url": "https://census.caseyjhand.com/mcp"
    }
  }
}

Self-Hosted / Local

API key: Register a free key at api.census.gov/data/key_signup.html. Variable search and geography resolution work without a key; data queries (census_query_data, census_compare_geographies) require one.

Add the following to your MCP client configuration file:

{
  "mcpServers": {
    "census-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/census-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info",
        "CENSUS_API_KEY": "your-census-api-key"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "census-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/census-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info",
        "CENSUS_API_KEY": "your-census-api-key"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "census-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "-e", "CENSUS_API_KEY=your-census-api-key",
        "ghcr.io/cyanheads/census-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 CENSUS_API_KEY=... bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

Installation

  1. Clone the repository:

git clone https://github.com/cyanheads/census-mcp-server.git
  1. Navigate into the directory:

cd census-mcp-server
  1. Install dependencies:

bun install
  1. Configure environment:

cp .env.example .env
# edit .env and set CENSUS_API_KEY

Configuration

Variable

Description

Default

CENSUS_API_KEY

Required for data queries. Register free at api.census.gov/data/key_signup.html.

—

CENSUS_DEFAULT_YEAR

Default vintage year when no year is specified.

2024

CENSUS_VARIABLE_CACHE_TTL_HOURS

Hours to cache variables.json per dataset+year in memory.

24

MCP_TRANSPORT_TYPE

Transport: stdio or http.

stdio

MCP_SESSION_MODE

HTTP session mode: stateful, stateless, or auto. The server declares stateless in src/index.ts; set this only to override it.

stateless

MCP_HTTP_PORT

Port for HTTP server.

3010

MCP_AUTH_MODE

Auth mode: none, jwt, or oauth.

none

MCP_LOG_LEVEL

Log level (debug, info, notice, warning, error).

info

OTEL_ENABLED

Enable OpenTelemetry instrumentation.

false

See .env.example for the full list of optional overrides.


Running the server

Local development

# One-time build
bun run rebuild

# Run the built server
bun run start:stdio
# or
bun run start:http

Run checks and tests:

bun run devcheck   # Lint, format, typecheck, security audit
bun run test       # Vitest test suite
bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t census-mcp-server .
docker run --rm -e CENSUS_API_KEY=your-key -p 3010:3010 census-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/census-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.


Project structure

Path

Purpose

src/index.ts

createApp() entry point — registers tools and initializes services.

src/config/server-config.ts

Census-specific env var parsing and validation with Zod.

src/mcp-server/tools/definitions/

Tool definitions (*.tool.ts).

src/services/census-api/

Census Data API client — data queries, suppression code mapping, retry logic.

src/services/geography/

Geography resolution — TIGERweb named-place lookup and Census Geocoder address-to-tract.

src/services/variable-cache/

In-process variables.json cache with TTL and keyword search.

tests/

Vitest tests mirroring src/ structure.


Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic

  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage

  • Register new tools via the barrel in src/mcp-server/tools/definitions/index.ts

  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields


Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables access to U.S. Census Bureau data including demographics, population, income, and housing statistics. Users can query specific variables, search datasets, and retrieve geographic FIPS codes across various surveys like the American Community Survey and Decennial Census.
    1
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    A production-grade MCP server for querying U.S. Census Bureau data (ACS 5-Year and Decennial) with tools for geographic fuzzy matching, variable search, and batched data retrieval, backed by a PostgreSQL cache for performance.
    -
  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that exposes the IPUMS API as LLM tools for browsing metadata, creating and downloading extracts, and generating reproducible R/Python code.
    23
    MIT