census-mcp-server
Server Details
Query U.S. Census Bureau data, variables, and geography via MCP.
- Status
- Healthy
- Uptime
- 99.8% over 22 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- cyanheads/census-mcp-server
- GitHub Stars
- 2
- Server Listing
- census-mcp-server
TDQS
Scored across 8 tools
Each tool targets a distinct operation: discovery (list_datasets, list_geographies, list_predicate_values), variable lookup (search_variables, get_variable), querying (query_data, compare_geographies), and geocoding (resolve_geography). No overlap in purpose.
All tools follow a consistent census_ + verb_noun pattern (e.g., list_datasets, resolve_geography, search_variables), with clear action-first naming and no stylistic deviations.
8 tools is well-scoped for a Census data server, covering the full workflow from dataset discovery to data retrieval and comparison without redundancy or bloat.
The surface covers the complete lifecycle: discover datasets, list geographies, search and inspect variables, resolve place names, query data, compare geographies, and enumerate filter values. No obvious gaps for the stated purpose.
Available Tools
8 toolscensus_compare_geographiesCompare Census GeographiesARead-onlyInspect
Compare one or more variables across multiple geographies at the same level — all counties in a state, all states nationally, or a named set of specific geographies — ranked on the value of one of them. Covers queries like "compare median income across WA counties" or "which states have the most people below the poverty line." A count ranks geographies by size, not by rate, so to rank a rate, rank a published percentage: S1701_C03_001E (percent below the poverty level, dataset acs/acs5/subject), DP03_0128PE (the same percentage, acs/acs5/profile), or DP04_0047PE (percent of occupied housing units that are renter-occupied, acs/acs5/profile). Profile and subject tables reach tracts but not block groups. Omit within to compare all geographies nationally at the level. Suppressed values are decoded to human-readable labels rather than passed through as raw negative sentinels. On the business datasets (cbp, ecnbasic, nonemp), pep/charv, dec/ddhca, and acs/acs1/spp, use predicates to rank within one industry, size class, or population group — a comparison that omits one ranks on a default the Census API picks, which is an all-categories total on some dimensions and a single category on others. Each row names the defaults that were applied in applied_filters, and census_list_predicate_values enumerates the codes a dimension accepts. A dataset that publishes several records per geography cannot be ranked until one is pinned: pep/charv publishes an April estimates base and a July estimate, so a comparison that pins neither fails with ambiguous_rows rather than giving every geography two ranks — pass predicates {"MONTH": "7"} for the July estimate.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Vintage year (default: latest available for the dataset). | |
| limit | No | Maximum geographies to return (default: 50, max: 500). When results are truncated, totalCount says how many matched. | |
| within | No | State FIPS to constrain results (e.g., "53" to compare counties or tracts within WA only). Omit to compare all geographies at the level nationally. Use census_resolve_geography to get state_fips. Pass "*" to span every state. Blank is treated as omitted. | |
| dataset | No | Dataset to query (default: "acs/acs5"). Use census_list_datasets for valid values. Case is ignored, and a two-part code can be given by its last part alone — "acs5" is acs/acs5, "pl" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code. | |
| sort_by | No | Variable code to rank on (default: the first code in variables), uppercased like the variables. Must be one of the requested codes, or the call fails with sort_by_not_requested. Geographies rank on that code's own value, so a count ranks by size and only a published percentage such as S1701_C03_001E or DP03_0128PE ranks by rate. | |
| sort_dir | No | Sort direction (default: "desc" — highest value first). | |
| variables | Yes | Variable codes to compare (e.g., ["B19013_001E", "B19013_001M"]); the ranking is on one of them, set by sort_by. Codes are uppercased before the request, and each row is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca, acs/acs1/spp), the maximum is lower, and too_many_variables states the exact number for the comparison. On ACS datasets, add the margin-of-error counterpart of a code (same code, E suffix swapped for M) for reliability context. The ACS comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile) and the other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error. | |
| predicates | No | Filter values keyed by variable code, applied to every geography in the comparison — e.g. {"NAICS2017": "5112"} to rank counties by their software-publisher establishment count in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, dec/ddhca, and acs/acs1/spp declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so a ranking can read like an overall one without being it. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of "*" returns every geography once per category of that dimension, which a ranking cannot hold, so it fails with ambiguous_rows naming the dimension to pin — use census_query_data for a per-category breakdown. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers). | |
| geographies | No | Optional list of specific geographies to include; only these are returned. Prefer full GEOIDs — the level concatenated with its parents, e.g. "53033" for King County WA and "06037" for Los Angeles County CA — which are nationally unique and so work across states. Bare level codes ("033") are also accepted but match that code in every state unless within scopes them to one. A GEOID is easiest taken from the geography_geoid field of a census_query_data or census_compare_geographies row; from census_resolve_geography, concatenate state_fips, then county_fips when it is present, then fips_summary. Entries that match nothing, and bare codes that match more than one state, are named in the response notice. | |
| within_county | No | County FIPS to constrain tract or block-group comparisons to a single county within the state specified by within (e.g., "033" for King County). Required when geography_level is "tract" or "block group" and you want county-scoped results. census_resolve_geography returns this as county_fips. Pass "*" to span every county in the state, which is the only way a block-group comparison reaches a whole state. Blank is treated as omitted. | |
| geography_level | Yes | The level to compare across (e.g., "state", "county", "tract"). Use census_list_geographies to see valid values for the dataset. |
Output Schema
| Name | Required | Description |
|---|---|---|
| rows | No | Geographies sorted by the requested variable. Suppressed values are labeled. |
| year | No | Vintage year queried. |
| error | No | Present when the call failed. Absent on success. |
| notice | No | Guidance when results were truncated, when the sort column holds no number on any row (so the rows are in the order the Census returned them rather than ranked), when geographies entries matched no row, when a bare level code matched more than one state, or when the dataset declares filter dimensions the comparison left unset — how to narrow scope, raise the limit, correct the FIPS codes, or add the predicates that pin what the ranking covers. For an unset dimension it also quotes the label of the default the Census API applied, which is what says whether the ranking is on a total or on one category. Also names any variable codes whose flags could not be checked because the request had no room left under the Census 50-column limit — a withheld value there reads as 0 and ranks as one. |
| dataset | No | Dataset queried. |
| truncated | No | True when totalCount exceeds the limit and results were cut off. |
| totalCount | No | Total number of geographies matched before the limit was applied. |
| sortVariable | No | Variable code the rows are ranked on, uppercased as it appears in variables. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses many non-obvious behaviors: suppressed values are decoded to labels, unpinned predicate dimensions get Census-API defaults that are echoed in applied_filters, and some datasets publish multiple records per geography requiring a pinned predicate. These go well beyond the readOnlyHint and openWorldHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense and almost every sentence carries operational guidance. It is front-loaded with the core purpose before diving into dataset-specific caveats. Slight length is justified by the high complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers error cases, default behaviors, dataset-specific ranking hazards, and links to sibling tools for supporting tasks. With an output schema present and readOnlyHint set, nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description meaningfully extends parameter understanding: sort_by must be one of the requested variables, at most 49 variables are allowed because NAME consumes a column, GEOID construction is explained for geographies, and predicate defaults vary by dataset. This is substantial value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Compare one or more variables across multiple geographies at the same level,' and immediately gives concrete example queries. It clearly distinguishes this from siblings like census_query_data by focusing on ranked cross-geography comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names when to use alternatives: census_resolve_geography for obtaining FIPS codes, census_list_predicate_values for filter codes, and census_query_data for per-category breakdowns. It also explains when a comparison is invalid and why, such as the ambiguous_rows failure when predicates are unpinned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_get_variableGet Census Variable MetadataARead-onlyInspect
Fetch full metadata for one or more Census variable codes — label, concept group, predicate type, the table's universe, and margin-of-error sibling references. Use to confirm a variable code before building a query, or to look up what a known code means. On ACS datasets it returns estimate_code and moe_code sibling references so you can request both without a separate search, and a margin-of-error code carries attribute_of and attribute_type MARGIN_OF_ERROR as the Census publishes them; the ACS comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile) and the other dataset families publish no margins of error and carry none of these fields. It also resolves the annotation and flag columns the data tools accept, such as B19013_001EA or EMP_F, naming the column each one belongs to, and predicate codes such as NAICS2017 or SEX, confirming a filter dimension exists in a dataset before a query uses it — for the values a dimension accepts rather than the dimension itself, call census_list_predicate_values.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Vintage year (default: latest available for the dataset). | |
| dataset | No | Dataset the variables belong to (default: "acs/acs5"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — "acs5" is acs/acs5, "pl" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code. | |
| variables | Yes | One or more variable codes to look up (e.g., ["B19013_001E", "B19013_001M"]). Codes are trimmed and matched regardless of case, and the response echoes the dataset's own spelling — uppercase everywhere except the comparison profiles' significance columns (e.g., CP03_2024to2019_062SS). |
Output Schema
| Name | Required | Description |
|---|---|---|
| year | No | Vintage year queried. |
| error | No | Present when the call failed. Absent on success. |
| dataset | No | Dataset queried. |
| variables | No | Variable metadata in the same order as the input array. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the read-only nature is covered. The description adds substantial behavioral context beyond annotations: it discloses that ACS datasets return estimate_code and moe_code sibling references, that MOE codes carry attribute_of/attribute_type MARGIN_OF_ERROR, that comparison profiles and other dataset families omit MOE fields entirely, and that annotation/flag columns are resolved to their owning columns. This is rich, detailed behavior an agent cannot infer from the schema or annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense, multi-clause sentence with several parenthetical asides and em-dash digressions. While every detail is arguably relevant, the structure is not front-loaded or easily skimmable; the key usage guidance appears only after a long list of return fields. It is thorough but not concise, and a more bulleted or scannable layout would improve clarity without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the tool definition does not need to explain return values in detail. The description covers all essential non-obvious contexts: dataset-family differences (ACS vs. comparison profiles vs. others), MOE behavior, annotation/flag column resolution, and predicate dimension validation. With the read-only annotation and full parameter schema, nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter (year, dataset, variables) already has a clear description in the input schema. The main description adds no further parameter-level meaning – it repeats the variable-code matching behavior that the schema already documents. Baseline 3 is appropriate because the schema does the heavy lifting and the description does not need to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Fetch') and resource ('full metadata for one or more Census variable codes') followed by a concrete list of returned metadata (label, concept group, predicate type, universe, MOE references). It also explicitly distinguishes itself from siblings: 'look up what a known code means' versus census_search_variables, and 'call census_list_predicate_values' for values rather than dimensions. An agent can instantly tell which tool does what.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'Use to confirm a variable code before building a query, or to look up what a known code means.' It also names a direct alternative and the condition that selects it: 'for the values a dimension accepts rather than the dimension itself, call census_list_predicate_values.' This leaves no ambiguity about when to pick this over a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_list_datasetsList Census DatasetsARead-onlyInspect
Browse available Census Bureau datasets with their supported vintage years. Use as the starting point when the right dataset is unknown — ACS5, ACS1, and their profile, subject, and comparison tables, population estimates, the decennial census files, and the business datasets (County Business Patterns, Economic Census, Nonemployer Statistics) serve different use cases. Pass the dataset_id value to the dataset parameter in other census tools. Each description names the predicates a dataset requires and the geography levels it publishes, both of which vary by dataset.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | Keyword to filter datasets by name or description. Omit to list all datasets. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | Present when the call failed. Absent on success. |
| notice | No | Guidance when no datasets matched the filter keyword. |
| datasets | No | Matching Census datasets. |
| totalCount | No | Total number of matching datasets. |
| filterApplied | No | Filter keyword applied to the dataset list, when provided. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful behavioral context beyond that: dataset descriptions name required predicates and geography levels, both of which vary by dataset. It also notes that each dataset supports vintage years, which informs the agent about what the listing contains. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and every sentence earns its place: scope, usage context, downstream integration, and dataset-specific behavioral notes. The key purpose is front-loaded, and no redundant or filler content appears.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional parameter, strong annotations, and an output schema present, the description fully covers what an agent needs to select and invoke this tool. It explains when to use it, what the results contain, how results vary, and how to use the returned dataset_id elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage for its single optional filter parameter, describing it as a keyword to filter by name or description. The description does not add parameter-specific detail, but that is acceptable because the schema fully documents the parameter. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Browse available Census Bureau datasets with their supported vintage years.' It clearly differentiates this from sibling tools by framing it as the starting point when the right dataset is unknown, and by explaining how the results feed into other census tools via dataset_id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it: 'Use as the starting point when the right dataset is unknown.' It also explains that the dataset_id returned should be passed to the dataset parameter in other census tools, giving clear downstream context. It does not name sibling tools as alternatives or state explicit exclusions, but the guidance is strong enough for an agent to select this tool appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_list_geographiesList Census Geography LevelsARead-onlyInspect
List the geography levels available for a given Census dataset and year, along with the parent geographies each level requires. Use before querying to confirm that the target geography level exists — ACS1 omits many sub-state levels, and not all datasets support tracts or block groups. The geography_level values returned here are the valid inputs to the geography_level parameter in census_query_data and census_compare_geographies.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Vintage year. Defaults to the latest available year for the dataset. | |
| dataset | Yes | Dataset code (e.g., "acs/acs5", "acs/acs1"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — "acs5" is acs/acs5, "pl" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code. |
Output Schema
| Name | Required | Description |
|---|---|---|
| year | No | Vintage year queried. |
| error | No | Present when the call failed. Absent on success. |
| dataset | No | Dataset queried. |
| totalLevels | No | Total number of geography levels available for this dataset and year. |
| geography_levels | No | Geography levels supported by this dataset and year. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description doesn't need to restate that. It adds value by explaining the relationship to other tools (the returned geography_level values are valid inputs elsewhere) and the caveat about ACS1 omissions, which helps the agent set expectations. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The first sentence states purpose and output; the second gives usage guidance and cross-tool relevance. Front-loaded and efficient, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (2 params, 1 required), has an output schema, and the description covers purpose and usage. The note about ACS1 omissions and the cross-reference to other tools makes it complete for an agent to call correctly without needing additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description repeats the dataset schema text exactly and adds no new parameter semantics beyond the schema. Baseline 3 applies because the schema carries the weight.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('geography levels'), and clearly distinguishes itself from siblings by noting it is used to confirm existence before querying and that its output feeds into census_query_data and census_compare_geographies. This makes it easy for an agent to understand what the tool does and how it differs from other census tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use before querying to confirm that the target geography level exists' and notes that ACS1 omits many sub-state levels, providing clear context for when to invoke it. However, it does not explicitly state when not to use it or name an alternative tool for other scenarios, which keeps it just below a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_list_predicate_valuesList Census Predicate ValuesARead-onlyInspect
List the codes a Census filter dimension accepts, so a predicates map can be written without guessing. Answers the question left open when census_query_data or census_compare_geographies reports that a dimension was left unset. Which route a dimension takes depends on the vintage: NAICS and POPGROUP always publish a value list in the dataset dictionary, and on the current vintages EMPSZES, LFO, RCPSZES, TAXSTAT, and TYPOP publish none and are enumerated here against the live data endpoint instead. A dictionary value list is a classification shared across Census products rather than a list of what one dataset serves, and roughly half of its codes typically return no rows anywhere — those are checked against the dataset's own published rows and dropped, and the response source field says whether that check ran. The dictionary lists run to thousands of codes and are best narrowed with query. Pass the returned code as the dimension's value in predicates.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Vintage year (default: latest available for the dataset). | |
| limit | No | Maximum codes to return (default: 50, max: 500). totalCount says how many matched. | |
| query | No | Keyword to narrow the list, matched case-insensitively against each code and label (e.g., "software" against NAICS2017, "exempt" against TAXSTAT). Omit to list from the start. NAICS and POPGROUP run to thousands of codes, so a keyword is the practical way to use them. | |
| dataset | Yes | Dataset the dimension belongs to (e.g., "cbp", "nonemp", "ecnbasic", "dec/ddhca", "pep/charv", "acs/acs1/spp"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — "ddhca" is dec/ddhca, "charv" is pep/charv. Three-part codes such as acs/acs1/spp must be given in full. The response echoes the resolved code. Dimension codes are vintage-specific, so the dataset and year must match the query the values are for. | |
| predicate | Yes | Filter dimension code to enumerate (e.g., "EMPSZES", "LFO", "POPGROUP", "NAICS2017"). Trimmed and uppercased, and the response echoes that spelling. The response notice of census_query_data names the dimensions a dataset declares, and census_search_variables finds them by keyword. | |
| within_naics | No | Industry code to scope the enumeration by, for dimensions the Census publishes per industry. On ecnbasic, TAXSTAT and TYPOP return only the all-establishments row until a NAICS sector is named — pass a sector code such as "62" (Health Care) or "42" (Wholesale Trade) and the result is complete for that industry alone. Ignored for dimensions with a published value list. Get sector codes by calling this tool on the dataset's own NAICS dimension. Blank is treated as omitted. |
Output Schema
| Name | Required | Description |
|---|---|---|
| year | No | Vintage year queried. |
| error | No | Present when the call failed. Absent on success. |
| notice | No | Guidance when the list was truncated, when query matched no published code, when codes from the dataset dictionary could not be checked against its published rows, or when the codes returned are complete only for a named industry rather than for the dimension as a whole. |
| source | No | Where the codes came from. "live_query" is a wildcard group-by against the data endpoint, which returns only codes the dataset publishes. "dataset_dictionary_verified" is the dataset's published value map with the codes it serves no rows for removed. Plain "dataset_dictionary" is that map unchecked — every code in it is declared by the dataset, but some of them return nothing at any geography, and the notice says why the check did not run. |
| values | No | Codes the dimension accepts, sorted by code. |
| dataset | No | Dataset queried. |
| predicate | No | Filter dimension enumerated. |
| truncated | No | True when totalCount exceeds the limit and the list was cut. |
| totalCount | No | Codes matched before the limit was applied. |
| predicate_label | No | Label of the dimension itself (e.g., "Employment size of establishments code"). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true and openWorldHint=false already in annotations, the description goes well beyond them: it discloses that NAICS and POPGROUP use dictionary value lists while other dimensions are enumerated live, that roughly half of dictionary codes return no rows and are dropped, that the response source field indicates whether that check ran, and that dictionary lists can run to thousands of codes. This is substantial behavioral context no structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, then adds necessary behavioral nuance in a dense but efficient way. Every sentence earns its place: routing logic, dictionary filtering behavior, query guidance, and output usage. There is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity — six parameters, vintage-dependent behavior, and schema with full coverage and output schema — the description covers the key decision points: which dimensions follow which enumeration route, why codes are dropped, how to handle large result sets, and how to feed results back into predicates. The read-only annotation covers safety, and the output schema removes the need to describe return shape. Nothing an agent needs to invoke it correctly seems missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3, but the description adds cross-parameter meaning: it explains how the predicate parameter relates to the tool's output ('Pass the returned code as the dimension's value in predicates'), and it contextualizes query as the practical narrowing mechanism for large dictionaries. It does not restate each schema field, but it clarifies the purpose and output usage of the parameters beyond the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List the codes a Census filter dimension accepts.' It immediately ties the purpose to the unresolved question left by census_query_data or census_compare_geographies, and names the sibling tools (census_list_datasets, census_search_variables) where relevant. An agent can clearly distinguish this from searching variables or querying data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool: when a predicates map needs to be written without guessing or when another tool reports a dimension was left unset. It gives concrete routing guidance based on vintage and dimension, explains when a keyword is the practical way to use it, and tells the agent to pass the returned code into predicates. It also indirectly tells when not to rely on this tool's dictionary path by noting dictionary lists are shared classifications, not dataset-specific result sets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_query_dataQuery Census DataARead-onlyInspect
Query a Census dataset for one or more variables at a specific geography. Accepts FIPS codes for the target geography — use census_resolve_geography to convert place names to FIPS when needed. On ACS datasets, labeled estimates and margin-of-error values are returned together (the comparison profiles publish no margins), and the negative sentinel values the Census writes for an estimate or margin of error it cannot publish are decoded into the meanings the Census gives them rather than passed through as raw numbers. A value cbp, ecnbasic, or nonemp withheld is stored as 0 beside a flag, and is reported as withheld, with the meaning of its flag, rather than as a zero. Pass geography_fips as "*" for every geography at the level within the parent: rows come back in GEOID order, up to limit per call (default 50, max 500), with totalCount giving how many matched and offset reaching the rest — the order is not a ranking, so use census_compare_geographies to rank. On the business datasets (cbp, ecnbasic, nonemp), pep/charv, dec/ddhca, and acs/acs1/spp, use predicates to filter by industry, size class, or population group — a query that omits one is answered with a default the Census API picks, which is an all-categories total on some dimensions and a single category on others. Each row names the defaults that were applied in applied_filters, and census_list_predicate_values enumerates the codes a dimension accepts. One geography can also come back on more than one row: pep/charv publishes an April estimates base alongside its July estimate, and each row carries a record field saying which it is.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Vintage year (default: latest available for the dataset). | |
| limit | No | Most rows to return (default: 50, max: 500). Rows come in GEOID order, a geography's records or categories in code order, and each one counts, so a geography returned as several records (pep/charv April and July) or as one row per category of a "*" predicate takes one row each. totalCount says how many rows matched. | |
| offset | No | Rows to skip before returning up to limit (default: 0). Pages run in GEOID order, so offset 50 with limit 50 returns rows 51–100, and the notice names the offset of the next page. An offset at or past totalCount returns no rows. | |
| dataset | No | Dataset to query (default: "acs/acs5"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — "acs5" is acs/acs5, "pl" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code. | |
| variables | Yes | Variable codes to retrieve (e.g., ["B19013_001E", "B19013_001M"]). Codes are uppercased before the request, so "b19013_001e" reads as B19013_001E and the response is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca, acs/acs1/spp), the maximum is lower, and too_many_variables states the exact number for the query. Use census_search_variables to find codes. On ACS datasets only, apart from the comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile), which publish none, each estimate has a margin-of-error counterpart at the same code with the E suffix swapped for M — request both to get the margin alongside the estimate. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error, and an E-final code there is an ordinary code with no M sibling. A code can also name a text column rather than a measure — GEO_ID, on every dataset, is the nationally unique geography identifier and comes back under value with estimate null, which is the code to request when a stable join key is what is wanted. | |
| predicates | No | Filter values keyed by variable code, sent as extra query parameters — e.g. {"NAICS2017": "5112"} to count only software publishers in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, dec/ddhca, and acs/acs1/spp declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so an unfiltered value can read like a total without being one. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of "*" returns one row per category of that dimension for each geography, each row labelled with its category in record (e.g. {"NAICS2017": "*"} gives King County one row per industry) — a breakdown that can run to over a thousand rows. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers). | |
| tract_fips | No | Census tract code scoping the query to one tract (e.g., "007101" for Census Tract 71.01), for the levels that sit within a tract — block group on acs/acs5, block group and block on dec/pl. census_resolve_geography returns it as tract_fips, and for a street address also returns the block_group_fips to pass as geography_fips. A tract code is unique only within its county, so it needs parent_fips and a concrete county_fips (not "*"). It is exactly 6 digits and is not padded, since "7101" and "71" do not name one tract. A level that does not sit within a tract rejects it. Blank is treated as omitted. | |
| county_fips | No | County FIPS code when querying tracts or block groups within a specific county (e.g., "033" for King County within WA). Required for tract and block-group queries scoped to a county — use alongside parent_fips (state). census_resolve_geography returns this as county_fips. Pass "*" to span every county in the state, which is the only way a block-group query reaches a whole state. Blank is treated as omitted. | |
| parent_fips | No | State FIPS code when querying sub-state levels (e.g., "53" for Washington). Required for county, tract, and block-group queries. census_resolve_geography returns this as state_fips. Pass "*" to span every state. Blank is treated as omitted. | |
| geography_fips | Yes | FIPS code for the target geography (e.g., "033" for a county, "*" for every geography at the level within the parent, returned up to limit rows per call and paged with offset). Use census_resolve_geography to obtain this value — it is returned as fips_summary. The Census API matches this literally and its width follows geography_level, so it is passed through unpadded: a county is 3 digits ("051", not "51") and a tract is 6. parent_fips and county_fips are zero-padded for you; this one is not. | |
| geography_level | Yes | Level of the target geography (e.g., "county", "tract", "state", "zip code tabulation area"). Use census_list_geographies to see valid values for the dataset. |
Output Schema
| Name | Required | Description |
|---|---|---|
| rows | No | One row per geography, or per record or category where a geography has several. When geography_fips is "*", the rows from offset up to limit of every geography at the level within the parent, in GEOID order. |
| year | No | Vintage year queried. |
| error | No | Present when the call failed. Absent on success. |
| notice | No | Warning that the dataset declares filter dimensions the query left unset, naming each one alongside the label of the default the Census API applied to it. That default is an all-categories total on some dimensions and one ordinary category on others, so the label is what says which. Also carries the warning that a geography came back on more than one row, naming the column that separates the records and the values it took; the range of rows returned when offset or limit left some out, with the offset of the next page; and any variable codes whose flags could not be checked because the request had no room left under the Census 50-column limit — a withheld value there reads as 0. |
| dataset | No | Dataset queried. |
| totalRows | No | Number of rows returned. |
| truncated | No | True when rows were left out by offset or limit — totalCount exceeds the rows returned. |
| totalCount | No | Number of rows the query matched, before offset and limit were applied. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses sentinel-value decoding, withheld-value handling, default predicate behavior, GEOID ordering, pagination via totalCount/offset, and the possibility of multiple rows per geography. These are non-obvious behaviors that materially affect interpretation of results and are not available from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries distinct information: purpose, FIPS resolution, sentinel decoding, wildcard semantics, predicate defaults, and row multiplicity. It is front-loaded with the core purpose and progressively adds edge-case detail without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description covers ordering, pagination, defaults, sentinel values, withheld values, multi-row results, and sibling routing. An output schema exists, so return-value structure need not be repeated, and the description still explains the non-obvious return semantics that the schema cannot convey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds some cross-parameter context, such as the '*' wildcard behavior and predicate defaults, but most parameter-level details (limit, offset, FIPS formatting, variable limits) are already fully documented in the schema. The description's extra value is more behavioral than parameter-specific.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Query a Census dataset for one or more variables at a specific geography.' It also distinguishes itself from siblings by naming census_resolve_geography for FIPS conversion and census_compare_geographies for ranking, making the tool's role clear relative to alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly routes to alternatives: use census_resolve_geography for place-name-to-FIPS conversion, census_compare_geographies when ranking is needed, census_list_predicate_values for dimension codes, census_search_variables for variable codes, and census_list_datasets/list_geographies for valid values. This gives an agent concrete when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_resolve_geographyResolve Census GeographyARead-onlyInspect
Resolve a place name, ZIP code, or street address to Census FIPS identifiers. Converts names like "King County, WA", "Seattle, WA", or "Seattle-Tacoma-Bellevue, WA", and ZIPs like "98109", to the codes required by census_query_data and census_compare_geographies. Use before querying when you have a place name rather than raw FIPS codes — state_fips maps to parent_fips and fips_summary maps to geography_fips in downstream tools, and geography_type is itself the geography_level to query at.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Place name (e.g., "King County, WA", "Seattle, WA", "California"), 5-digit ZIP code (e.g., "98109", resolved to its ZIP Code Tabulation Area), or street address (e.g., "1600 Pennsylvania Ave NW, Washington, DC 20500"). Include the state after a comma — its abbreviation or full name, as in "Chatham County, Georgia" — to disambiguate places with common names. It narrows a statistical area as well, matching any state the area spans, so "Kansas City, MO" and "Kansas City, KS" both reach the MO-KS metro area. Matching ignores case, reads "Saint" and "St." as the same word, and accepts unaccented spellings ("Dona Ana County, NM"). For a statistical area a leading city is enough ("Denver, CO" for the Denver-Aurora-Centennial metro area), and an older full name resolves through its leading city ("Denver-Aurora-Lakewood, CO"). | |
| county_fips | No | County FIPS code to resolve within — 1 to 3 digits, zero-padded here to the 3 the Census stores. A tract name is unique only inside its county, so a bare tract name matching two counties comes back as ambiguous_name until this is set: take the countyFips of the candidate you want from that error and re-call. Only county and tract sit within a county, so this restricts resolution to those two levels — pairing it with any other geography_type, or with a street address, is a county_scope_unsupported error rather than a scope quietly dropped. census_query_data takes the same code as its own county_fips but pads nothing, so hand it the 3-digit county_fips returned here, not the shorter value. | |
| geography_type | No | Geography level to resolve to, named exactly as census_query_data's geography_level and census_list_geographies name it. Auto-detection covers only state, county, place, tract, and zip code tabulation area: zip code tabulation area for a 5-digit ZIP or ZIP+4 — the ACS's ZIP-shaped area, not cbp's "zip code" level, which takes the ZIP itself with no resolution — state for a two-letter abbreviation or a spelled-out state name, county when the name contains the word "County" or "Parish", county then place for the word "Borough" (an Alaska borough is a county, a PA or NJ borough a place), tract for the word "Tract", otherwise place (incorporated places and census-designated places together) with a fallback to county, where a census-designated place answers only when no incorporated place or county has the exact name — "Arlington, VA" is Arlington County, and set "place" to reach the Arlington CDP. The other four are never auto-detected and must be set explicitly, because their names overlap city names — "metropolitan statistical area/micropolitan statistical area" covers both metro and micro areas and yields a 5-digit code, "combined statistical area" yields a 3-digit code, "consolidated city" covers the eight merged city-county governments (Nashville-Davidson, Louisville/Jefferson County, Indianapolis, Athens-Clarke County, Augusta-Richmond County, Butte-Silver Bow, Milford CT, Greeley County KS), and "economic place" yields the 8-digit code ecnbasic 2022 publishes a place under: the 3-digit county it lies in (000 when it spans counties) followed by its 5-digit place code. Economic places are the incorporated places, census-designated places, and county subdivisions the 2022 Economic Census tabulates, plus each county's remainder ("Balance of Adams County, WA"). Setting it explicitly also overrides auto-detection — "New York" auto-detects as the state, so New York City needs "place". |
Output Schema
| Name | Required | Description |
|---|---|---|
| name | No | Canonical name of the resolved geography. |
| error | No | Present when the call failed. Absent on success. |
| place_fips | No | 5-digit place FIPS code when the resolved geography is a place — incorporated or census-designated. For an economic place, the place part of its 8-digit code, which is also the place code the 2017 and 2012 ecnbasic vintages take. For a street address, the incorporated place the address sits in (query it at the place level with state_fips as parent_fips); absent when the address is outside every incorporated place, and a census-designated place is never reported for an address. |
| state_fips | No | 2-digit state FIPS code. Use as parent_fips in census_query_data for sub-state queries. Absent for a metropolitan/micropolitan or combined statistical area, which can span several states and needs no parent_fips, and for a zip code tabulation area, whose source layer carries no state. |
| tract_fips | No | 6-digit census tract FIPS code when the resolved geography is a tract — from a street address, or from a tract name. |
| county_fips | No | 3-digit county FIPS code when the resolved geography is a county or sub-county level. |
| fips_summary | No | Pre-formatted FIPS value ready to use as geography_fips in census_query_data (e.g., "033" for King County with state_fips "53" as parent_fips, "42660" for the Seattle-Tacoma-Bellevue metro area with no parent at all, "03363000" for Seattle as an ecnbasic 2022 economic place). |
| geography_type | No | Resolved geography level. Pass this straight through as geography_level in census_query_data and census_compare_geographies. |
| block_group_fips | No | 1-digit block group within tract_fips, for a street address only. Query it as geography_level "block group" with this as geography_fips, alongside parent_fips, county_fips, and tract_fips. geography_type and fips_summary stay at the tract. |
| census_designated_place | No | Present (true) when the resolved place or economic place is a census-designated place (CDP): an unincorporated community the Census delineates for statistics, with no municipal government. Its code is queried like an incorporated place's. Absent for an incorporated place and every other level. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses detailed matching behavior: case-insensitivity, 'Saint'/'St.' equivalence, unaccented spellings, auto-detection coverage for specific geography types, explicit override behavior, and error conditions like ambiguous_name and county_scope_unsupported. It also mentions county_fips zero-padding and how to pass the value to census_query_data, adding substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: the first states the core purpose with examples, the second gives the usage context and mapping, and the third reinforces the downstream connection. No filler or redundant restating of the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (auto-detection rules, multiple geography types, error states, downstream mapping), the description covers all essential selection and invocation information. The existence of an output schema means the return structure does not need to be spelled out here, and the annotations plus description suffice for safe, correct calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The tool description adds value by explaining how the resolved outputs map to downstream parameters (state_fips→parent_fips, fips_summary→geography_fips, geography_type→geography_level), which is not present in the parameter schemas. This enriches the medication of the parameters' purpose in the broader workflow.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Resolve a place name, ZIP code, or street address to Census FIPS identifiers.' It gives concrete examples and explicitly ties the output to downstream tools (census_query_data, census_compare_geographies), distinguishing this from siblings by stating it converts names to codes rather than querying data directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states exactly when to use the tool: 'Use before querying when you have a place name rather than raw FIPS codes.' It also explains how the output maps to parameters of downstream tools (state_fips→parent_fips, fips_summary→geography_fips, geography_type→geography_level), giving clear routing guidance without needing to inspect sibling schemas.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
census_search_variablesSearch Census VariablesARead-onlyInspect
Search Census variables by keyword across variable labels and concept groups. Returns variable codes with human-readable labels — use this to go from a concept like "median household income" to the variable code B19013_001E needed for data queries. On ACS datasets it returns both estimate (E suffix) and margin-of-error (M suffix) codes so you can request both; the ACS comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile) and the other dataset families publish no margins of error. Also use it to find the predicate codes a dataset filters on, such as NAICS2017 in cbp. Adding a word narrows the results, since every word must match; when totalMatches exceeds the limit, a more specific query reaches the rest.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Vintage year to search (default: latest available for the dataset). | |
| limit | No | Maximum results to return (default: 20, max: 100). Increase if totalMatches greatly exceeds the limit. | |
| query | Yes | Keywords to search (e.g., "median household income", "poverty", "bachelor's degree"). Each word must match a whole word of the label or of the concept, ignoring case and punctuation, so "rate" does not match "separated". A column shared across tables, such as GEO_ID, is matched on its label only, and a margin of error on its estimate's label. When no variable contains every word, the results are the variables containing the most words, and the notice says how many that was. | |
| dataset | No | Dataset to search within (default: "acs/acs5"). Use census_list_datasets to discover options. Case is ignored, and a two-part code can be given by its last part alone — "acs5" is acs/acs5, "pl" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code, and the default year is that dataset's latest. |
Output Schema
| Name | Required | Description |
|---|---|---|
| cap | No | The limit that was applied. |
| year | No | Vintage year that was searched. |
| error | No | Present when the call failed. Absent on success. |
| shown | No | Number of variables returned after the limit. |
| notice | No | Guidance when no variables matched, when no variable contained every query word and the results hold only some of them, or when results were truncated — suggests other keywords, a narrower query, or a higher limit. |
| dataset | No | Dataset that was searched. |
| truncated | No | True when totalMatches exceeded the limit and results were cut off. |
| variables | No | Matching variables, best first: a variable whose label's last !!-separated segment or whose whole concept equals the query, then the query as a phrase in both label and concept, in the label only, in the concept only, then every word present but not as a phrase; ties go to fewer !! segments in the label, then a shorter concept, then the code, which puts an E estimate before its M margin of error. On ACS datasets, codes ending in E are estimates and M are their margins of error, except on the comparison profiles, which publish no M codes; on other datasets the suffix carries no such meaning. |
| totalMatches | No | Variables that contain every query word, before the limit was applied — or, when none does, the variables that contain the most words. |
| effectiveQuery | No | Query as the server parsed it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only, and the description adds substantial behavioral detail beyond what annotations provide: ACS returns both estimate and margin-of-error codes, comparison profiles publish no MOEs, matching is whole-word with a most-words fallback, and totalMatches influences query refinement. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, and each sentence adds distinct high-value information: output codes, profile exceptions, predicate-code use, and matching behavior. It is dense but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and a readOnlyHint annotation, the description still covers key edge cases: MOE suffixes, ACS comparison profiles, predicate codes, whole-word matching and fallback behavior, and totalMatches guidance. An agent has enough context to invoke the tool correctly in its primary use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description lightly reinforces query and limit behavior ('Adding a word narrows the results', 'when totalMatches exceeds the limit'), but it does not add meaningfully new parameter semantics beyond the already detailed schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action: searches Census variables by keyword across labels and concept groups and returns variable codes with labels. The description frames its purpose as bridging a concept like 'median household income' to a code like B19013_001E for data queries, and also as finding predicate codes, clearly distinguishing it from related tools like census_get_variable and census_query_data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear use contexts: converting a concept to a variable code for data queries, and finding predicate codes for datasets such as cbp. It does not explicitly state when not to use this tool or name sibling alternatives like census_get_variable or census_list_predicate_values, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
- Changed
census_compare_geographies4 fields changed- changed
Input schema / properties / dataset / descriptionPrevious value: -"Dataset to query (default: \"acs/acs5\"). Use census_list_datasets for valid values."New value: +"Dataset to query (default: \"acs/acs5\"). Use census_list_datasets for valid values. Case is ignored, and a two-part code can be given by its last part alone — \"acs5\" is acs/acs5, \"pl\" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code." - changed
Input schema / properties / predicates / descriptionPrevious value: -"Filter values keyed by variable code, applied to every geography in the comparison — e.g. {\"NAICS2017\": \"5112\"} to rank counties by their software-publisher establishment count in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, and dec/ddhca declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so a ranking can read like an overall one without being it. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of \"*\" returns every geography once per category of that dimension, which a ranking cannot hold, so it fails with ambiguous_rows naming the dimension to pin — use census_query_data for a per-category breakdown. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)."New value: +"Filter values keyed by variable code, applied to every geography in the comparison — e.g. {\"NAICS2017\": \"5112\"} to rank counties by their software-publisher establishment count in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, dec/ddhca, and acs/acs1/spp declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so a ranking can read like an overall one without being it. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of \"*\" returns every geography once per category of that dimension, which a ranking cannot hold, so it fails with ambiguous_rows naming the dimension to pin — use census_query_data for a per-category breakdown. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)." - changed
Input schema / properties / variables / descriptionPrevious value: -"Variable codes to compare (e.g., [\"B19013_001E\", \"B19013_001M\"]); the ranking is on one of them, set by sort_by. Codes are uppercased before the request, and each row is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca), the maximum is lower, and too_many_variables states the exact number for the comparison. On ACS datasets, add the margin-of-error counterpart of a code (same code, E suffix swapped for M) for reliability context. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error."New value: +"Variable codes to compare (e.g., [\"B19013_001E\", \"B19013_001M\"]); the ranking is on one of them, set by sort_by. Codes are uppercased before the request, and each row is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca, acs/acs1/spp), the maximum is lower, and too_many_variables states the exact number for the comparison. On ACS datasets, add the margin-of-error counterpart of a code (same code, E suffix swapped for M) for reliability context. The ACS comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile) and the other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS but within, or within_county, was not provided. `parent_not_accepted`: within or within_county names a parent the geography level does not sit within. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for this dataset and year. It names only the first unknown code in a request. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `sort_by_not_requested`: sort_by names a code that is not among the requested variables, so no column exists to rank on. `ambiguous_rows`: The dataset publishes several records per geography and the comparison pinned none of them, so every geography would occupy several ranks with different values. `no_data`: No geographies were returned for the query, or no row matched any entry in the geographies list. `upstream_error`: Census API was unreachable or returned an error. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized, even after case and shorthand resolution. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS but within, or within_county, was not provided. `parent_not_accepted`: within or within_county names a parent the geography level does not sit within. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for this dataset and year. It names only the first unknown code in a request. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `sort_by_not_requested`: sort_by names a code that is not among the requested variables, so no column exists to rank on. `ambiguous_rows`: The dataset publishes several records per geography and the comparison pinned none of them, so every geography would occupy several ranks with different values. `no_data`: No geographies were returned for the query, or no row matched any entry in the geographies list. `upstream_error`: Census API was unreachable or returned an error. Other values are possible when a failure originates below the handler."
- Changed
census_get_variable10 fields changed- changed
Input schema / properties / dataset / descriptionPrevious value: -"Dataset the variables belong to (default: \"acs/acs5\"). Use census_list_datasets to discover valid values."New value: +"Dataset the variables belong to (default: \"acs/acs5\"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — \"acs5\" is acs/acs5, \"pl\" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code." - changed
Input schema / properties / variables / descriptionPrevious value: -"One or more variable codes to look up (e.g., [\"B19013_001E\", \"B19013_001M\"]). Variable codes are case-sensitive."New value: +"One or more variable codes to look up (e.g., [\"B19013_001E\", \"B19013_001M\"]). Codes are trimmed and matched regardless of case, and the response echoes the dataset's own spelling — uppercase everywhere except the comparison profiles' significance columns (e.g., CP03_2024to2019_062SS)." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `variable_not_found`: One or more variable codes were not found in the dataset and year. `dataset_not_found`: Dataset code is not recognized. `year_not_available`: The dataset does not serve the requested vintage year. `variables_unavailable`: Variable metadata endpoint is unreachable or returned an unparseable response. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `variable_not_found`: One or more variable codes were not found in the dataset and year. `dataset_not_found`: Dataset code is not recognized, even after case and shorthand resolution. `year_not_available`: The dataset does not serve the requested vintage year. `variables_unavailable`: Variable metadata — the dataset's variables.json, groups.json, or an attribute column's own entry — could not be fetched or parsed. Other values are possible when a failure originates below the handler." - added
Output schema / properties / variables / items / properties / attribute_ofAdded value: +{ + "description": "For an annotation, flag, or margin-of-error column, the column it belongs to, as the Census publishes it (e.g., \"B19013_001E\" for B19013_001EA or B19013_001M). Absent on ordinary variables.", + "type": "string" +} - added
Output schema / properties / variables / items / properties / attribute_typeAdded value: +{ + "description": "For an annotation, flag, or margin-of-error column, its kind as the Census publishes it (e.g., \"ANNOTATION\", \"FLAG\", \"MARGIN_OF_ERROR\"). Absent on ordinary variables.", + "type": "string" +} - changed
Output schema / properties / variables / items / properties / concept / descriptionPrevious value: -"Concept group the variable belongs to."New value: +"Concept of the table the variable belongs to. Absent for a column shared across tables, such as GEO_ID, whose concept joins every table it appears in, and for a column the dataset publishes no concept for, such as STATE." - changed
Output schema / properties / variables / items / properties / estimate_code / descriptionPrevious value: -"Estimate sibling variable code when this is a margin-of-error variable. ACS datasets only — no other family publishes margins of error."New value: +"Estimate sibling variable code when this is a margin-of-error variable. ACS datasets only, apart from the comparison profiles — no other dataset publishes margins of error." - changed
Output schema / properties / variables / items / properties / moe_code / descriptionPrevious value: -"Margin-of-error sibling code when this is an estimate variable. Include both in census_query_data for complete data. ACS datasets only — on other families an E-final code is an ordinary code with no margin-of-error sibling, so the field is absent."New value: +"Margin-of-error sibling code when this is an estimate variable. Include both in census_query_data for complete data. ACS datasets only, apart from the comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile) — there and on other families an E-final code has no margin-of-error sibling, so the field is absent." - changed
Output schema / properties / variables / items / properties / universe / descriptionPrevious value: -"Universe the variable applies to (e.g., \"Households\", \"People 25 years and over\")."New value: +"Universe of the variable's table (e.g., \"Households\", \"Population 25 years and over\"). Absent when the table publishes none — the ACS subject, profile, and selected population profile tables, dec/dp, the business datasets, pep/charv, and every ACS and dec/pl vintage before 2020 publish none — and for a column that belongs to no single table." - changed
Output schema / properties / variables / items / requiredPrevious value: -[ - "variable_code", - "label", - "concept", - "predicate_type" -]New value: +[ + "variable_code", + "label", + "predicate_type" +]
- Changed
census_list_datasets1 field changed- changed
Output schema / properties / datasets / items / properties / available_years / descriptionPrevious value: -"Vintage years this dataset can be queried for. Passing any other year to census_query_data, census_compare_geographies, or census_search_variables fails with year_not_available rather than returning data — the list is exhaustive, not a sample. It is narrower than what the Census API hosts: pep/charv publishes its 2020-2022 estimates inside the 2023 vintage under the YEAR filter, and the cbp and nonemp vintages left out reject the NAME column every query here sends."New value: +"Vintage years this dataset can be queried for. Passing any other year to census_query_data, census_compare_geographies, or census_search_variables fails with year_not_available rather than returning data — the list is exhaustive, not a sample. It is narrower than what the Census API hosts: pep/charv publishes its 2020-2022 estimates inside the 2023 vintage under the YEAR filter, the cbp and nonemp vintages left out reject the NAME column every query here sends, and the Census API answers the acs/acs1/spp 2008 and 2010 vintages with server errors."
- Changed
census_list_geographies2 fields changed- changed
Input schema / properties / dataset / descriptionPrevious value: -"Dataset code (e.g., \"acs/acs5\", \"acs/acs1\"). Use census_list_datasets to discover valid values."New value: +"Dataset code (e.g., \"acs/acs5\", \"acs/acs1\"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — \"acs5\" is acs/acs5, \"pl\" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `year_not_available`: Dataset exists but the requested year has no data. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is missing or not recognized, even after case and shorthand resolution. `year_not_available`: Dataset exists but the requested year has no data. Other values are possible when a failure originates below the handler."
- Changed
census_list_predicate_values3 fields changed- changed
Input schema / properties / dataset / descriptionPrevious value: -"Dataset the dimension belongs to (e.g., \"cbp\", \"nonemp\", \"ecnbasic\", \"dec/ddhca\", \"pep/charv\"). Use census_list_datasets to discover valid values. Dimension codes are vintage-specific, so the dataset and year must match the query the values are for."New value: +"Dataset the dimension belongs to (e.g., \"cbp\", \"nonemp\", \"ecnbasic\", \"dec/ddhca\", \"pep/charv\", \"acs/acs1/spp\"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — \"ddhca\" is dec/ddhca, \"charv\" is pep/charv. Three-part codes such as acs/acs1/spp must be given in full. The response echoes the resolved code. Dimension codes are vintage-specific, so the dataset and year must match the query the values are for." - changed
Input schema / properties / predicate / descriptionPrevious value: -"Filter dimension code to enumerate (e.g., \"EMPSZES\", \"LFO\", \"POPGROUP\", \"NAICS2017\"). Case-sensitive. The response notice of census_query_data names the dimensions a dataset declares, and census_search_variables finds them by keyword."New value: +"Filter dimension code to enumerate (e.g., \"EMPSZES\", \"LFO\", \"POPGROUP\", \"NAICS2017\"). Trimmed and uppercased, and the response echoes that spelling. The response notice of census_query_data names the dimensions a dataset declares, and census_search_variables finds them by keyword." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `year_not_available`: The dataset does not serve the requested vintage year. `predicate_not_supported`: The predicate code is not a variable in this dataset and year. `not_a_filter_dimension`: The code is not one of the dimensions the dataset filters on, so it takes no value list. `no_values`: The dimension returned no codes for the scope requested. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is missing or not recognized, even after case and shorthand resolution. `year_not_available`: The dataset does not serve the requested vintage year. `predicate_not_supported`: The predicate code is not a variable in this dataset and year. `not_a_filter_dimension`: The code is not one of the dimensions the dataset filters on, so it takes no value list. `no_values`: The dimension returned no codes for the scope requested. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."
- Changed
census_query_data4 fields changed- changed
Input schema / properties / dataset / descriptionPrevious value: -"Dataset to query (default: \"acs/acs5\"). Use census_list_datasets to discover valid values."New value: +"Dataset to query (default: \"acs/acs5\"). Use census_list_datasets to discover valid values. Case is ignored, and a two-part code can be given by its last part alone — \"acs5\" is acs/acs5, \"pl\" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code." - changed
Input schema / properties / predicates / descriptionPrevious value: -"Filter values keyed by variable code, sent as extra query parameters — e.g. {\"NAICS2017\": \"5112\"} to count only software publishers in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, and dec/ddhca declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so an unfiltered value can read like a total without being one. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of \"*\" returns one row per category of that dimension for each geography, each row labelled with its category in record (e.g. {\"NAICS2017\": \"*\"} gives King County one row per industry) — a breakdown that can run to over a thousand rows. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)."New value: +"Filter values keyed by variable code, sent as extra query parameters — e.g. {\"NAICS2017\": \"5112\"} to count only software publishers in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, dec/ddhca, and acs/acs1/spp declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so an unfiltered value can read like a total without being one. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of \"*\" returns one row per category of that dimension for each geography, each row labelled with its category in record (e.g. {\"NAICS2017\": \"*\"} gives King County one row per industry) — a breakdown that can run to over a thousand rows. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)." - changed
Input schema / properties / variables / descriptionPrevious value: -"Variable codes to retrieve (e.g., [\"B19013_001E\", \"B19013_001M\"]). Codes are uppercased before the request, so \"b19013_001e\" reads as B19013_001E and the response is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca), the maximum is lower, and too_many_variables states the exact number for the query. Use census_search_variables to find codes. On ACS datasets only, each estimate has a margin-of-error counterpart at the same code with the E suffix swapped for M — request both to get the margin alongside the estimate. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error, and an E-final code there is an ordinary code with no M sibling. A code can also name a text column rather than a measure — GEO_ID, on every dataset, is the nationally unique geography identifier and comes back under value with estimate null, which is the code to request when a stable join key is what is wanted."New value: +"Variable codes to retrieve (e.g., [\"B19013_001E\", \"B19013_001M\"]). Codes are uppercased before the request, so \"b19013_001e\" reads as B19013_001E and the response is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca, acs/acs1/spp), the maximum is lower, and too_many_variables states the exact number for the query. Use census_search_variables to find codes. On ACS datasets only, apart from the comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile), which publish none, each estimate has a margin-of-error counterpart at the same code with the E suffix swapped for M — request both to get the margin alongside the estimate. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error, and an E-final code there is an ordinary code with no M sibling. A code can also name a text column rather than a measure — GEO_ID, on every dataset, is the nationally unique geography identifier and comes back under value with estimate null, which is the code to request when a stable join key is what is wanted." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for the requested dataset and year. It names only the first unknown code in a request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS code but parent_fips was not provided, a tract/block-group level requires county_fips but it was omitted, a single block group requires tract_fips, or tract_fips was set without a concrete county_fips. `parent_not_accepted`: parent_fips, county_fips, or tract_fips names a parent the geography level does not sit within. `no_data`: The query returned no rows. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized, even after case and shorthand resolution. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for the requested dataset and year. It names only the first unknown code in a request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS code but parent_fips was not provided, a tract/block-group level requires county_fips but it was omitted, a single block group requires tract_fips, or tract_fips was set without a concrete county_fips. `parent_not_accepted`: parent_fips, county_fips, or tract_fips names a parent the geography level does not sit within. `no_data`: The query returned no rows. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."
- Changed
census_search_variables10 fields changed- changed
Input schema / properties / dataset / descriptionPrevious value: -"Dataset to search within (default: \"acs/acs5\"). Use census_list_datasets to discover options."New value: +"Dataset to search within (default: \"acs/acs5\"). Use census_list_datasets to discover options. Case is ignored, and a two-part code can be given by its last part alone — \"acs5\" is acs/acs5, \"pl\" is dec/pl. Three-part codes such as acs/acs5/profile must be given in full. The response echoes the resolved code, and the default year is that dataset's latest." - changed
Input schema / properties / query / descriptionPrevious value: -"Keyword to search (e.g., \"median household income\", \"poverty\", \"bachelor's degree\"). Multi-word queries search for all terms."New value: +"Keywords to search (e.g., \"median household income\", \"poverty\", \"bachelor's degree\"). Each word must match a whole word of the label or of the concept, ignoring case and punctuation, so \"rate\" does not match \"separated\". A column shared across tables, such as GEO_ID, is matched on its label only, and a margin of error on its estimate's label. When no variable contains every word, the results are the variables containing the most words, and the notice says how many that was." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `year_not_available`: The dataset does not serve the requested vintage year. `variables_unavailable`: Variable metadata could not be fetched or parsed from the Census API. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized, even after case and shorthand resolution. `year_not_available`: The dataset does not serve the requested vintage year. `variables_unavailable`: Variable metadata could not be fetched or parsed from the Census API. Other values are possible when a failure originates below the handler." - changed
Output schema / properties / notice / descriptionPrevious value: -"Guidance when no variables matched, or when results were truncated — suggests broader keywords, a narrower query, or a higher limit."New value: +"Guidance when no variables matched, when no variable contained every query word and the results hold only some of them, or when results were truncated — suggests other keywords, a narrower query, or a higher limit." - changed
Output schema / properties / totalMatches / descriptionPrevious value: -"Total variables matching the query before the limit was applied."New value: +"Variables that contain every query word, before the limit was applied — or, when none does, the variables that contain the most words." - changed
Output schema / properties / variables / descriptionPrevious value: -"Matching variables sorted by relevance. On ACS datasets, codes ending in E are estimates and M are their margins of error; on other datasets the suffix carries no such meaning."New value: +"Matching variables, best first: a variable whose label's last !!-separated segment or whose whole concept equals the query, then the query as a phrase in both label and concept, in the label only, in the concept only, then every word present but not as a phrase; ties go to fewer !! segments in the label, then a shorter concept, then the code, which puts an E estimate before its M margin of error. On ACS datasets, codes ending in E are estimates and M are their margins of error, except on the comparison profiles, which publish no M codes; on other datasets the suffix carries no such meaning." - changed
Output schema / properties / variables / items / properties / concept / descriptionPrevious value: -"Concept group the variable belongs to (e.g., \"MEDIAN HOUSEHOLD INCOME\")."New value: +"Concept of the table the variable belongs to (e.g., \"Median Household Income in the Past 12 Months\"). Absent for a column shared across tables, such as GEO_ID, whose concept joins every table it appears in, and for a column the dataset publishes no concept for, such as STATE." - changed
Output schema / properties / variables / items / properties / estimate_code / descriptionPrevious value: -"Corresponding estimate variable code when this is a margin-of-error variable. ACS datasets only — no other family publishes margins of error."New value: +"Corresponding estimate variable code when this is a margin-of-error variable. ACS datasets only, apart from the comparison profiles — no other dataset publishes margins of error." - changed
Output schema / properties / variables / items / properties / moe_code / descriptionPrevious value: -"Corresponding margin-of-error variable code when this is an estimate variable. Request both estimate and MOE in census_query_data for complete data. ACS datasets only — on other families an E-final code is an ordinary code with no margin-of-error sibling, so the field is absent."New value: +"Corresponding margin-of-error variable code when this is an estimate variable. Request both estimate and MOE in census_query_data for complete data. ACS datasets only, apart from the comparison profiles (acs/acs5/cprofile, acs/acs1/cprofile) — there and on other families an E-final code has no margin-of-error sibling, so the field is absent." - changed
Output schema / properties / variables / items / requiredPrevious value: -[ - "variable_code", - "label", - "concept", - "predicate_type" -]New value: +[ + "variable_code", + "label", + "predicate_type" +]
2 tool updates
- Changed
census_query_data2 fields changed- added
Input schema / properties / tract_fipsAdded value: +{ + "anyOf": [ + { + "const": "", + "type": "string" + }, + { + "description": "Exactly 6 digits — never padded here, and never \"*\".", + "pattern": "^\\d{6}$", + "type": "string" + } + ], + "description": "Census tract code scoping the query to one tract (e.g., \"007101\" for Census Tract 71.01), for the levels that sit within a tract — block group on acs/acs5, block group and block on dec/pl. census_resolve_geography returns it as tract_fips, and for a street address also returns the block_group_fips to pass as geography_fips. A tract code is unique only within its county, so it needs parent_fips and a concrete county_fips (not \"*\"). It is exactly 6 digits and is not padded, since \"7101\" and \"71\" do not name one tract. A level that does not sit within a tract rejects it. Blank is treated as omitted." +} - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for the requested dataset and year. It names only the first unknown code in a request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS code but parent_fips was not provided, or a tract/block-group level requires county_fips but it was omitted. `parent_not_accepted`: parent_fips or county_fips names a parent the geography level does not sit within. `no_data`: The query returned no rows. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for the requested dataset and year. It names only the first unknown code in a request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS code but parent_fips was not provided, a tract/block-group level requires county_fips but it was omitted, a single block group requires tract_fips, or tract_fips was set without a concrete county_fips. `parent_not_accepted`: parent_fips, county_fips, or tract_fips names a parent the geography level does not sit within. `no_data`: The query returned no rows. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."
- Changed
census_resolve_geography8 fields changed- changed
Input schema / properties / geography_type / descriptionPrevious value: -"Geography level to resolve to, named exactly as census_query_data's geography_level and census_list_geographies name it. Auto-detection covers only state, county, place, and tract: state for a two-letter abbreviation or a spelled-out state name, county when the name contains \"County\"/\"Borough\"/\"Parish\", tract when it contains \"Tract\", otherwise place with a fallback to county. The other three are never auto-detected and must be set explicitly, because their names overlap city names — \"metropolitan statistical area/micropolitan statistical area\" covers both metro and micro areas and yields a 5-digit code, \"combined statistical area\" yields a 3-digit code, and \"consolidated city\" covers the eight merged city-county governments (Nashville-Davidson, Louisville/Jefferson County, Indianapolis, Athens-Clarke County, Augusta-Richmond County, Butte-Silver Bow, Milford CT, Greeley County KS). Setting it explicitly also overrides auto-detection — \"New York\" auto-detects as the state, so New York City needs \"place\"."New value: +"Geography level to resolve to, named exactly as census_query_data's geography_level and census_list_geographies name it. Auto-detection covers only state, county, place, tract, and zip code tabulation area: zip code tabulation area for a 5-digit ZIP or ZIP+4 — the ACS's ZIP-shaped area, not cbp's \"zip code\" level, which takes the ZIP itself with no resolution — state for a two-letter abbreviation or a spelled-out state name, county when the name contains the word \"County\" or \"Parish\", county then place for the word \"Borough\" (an Alaska borough is a county, a PA or NJ borough a place), tract for the word \"Tract\", otherwise place (incorporated places and census-designated places together) with a fallback to county, where a census-designated place answers only when no incorporated place or county has the exact name — \"Arlington, VA\" is Arlington County, and set \"place\" to reach the Arlington CDP. The other four are never auto-detected and must be set explicitly, because their names overlap city names — \"metropolitan statistical area/micropolitan statistical area\" covers both metro and micro areas and yields a 5-digit code, \"combined statistical area\" yields a 3-digit code, \"consolidated city\" covers the eight merged city-county governments (Nashville-Davidson, Louisville/Jefferson County, Indianapolis, Athens-Clarke County, Augusta-Richmond County, Butte-Silver Bow, Milford CT, Greeley County KS), and \"economic place\" yields the 8-digit code ecnbasic 2022 publishes a place under: the 3-digit county it lies in (000 when it spans counties) followed by its 5-digit place code. Economic places are the incorporated places, census-designated places, and county subdivisions the 2022 Economic Census tabulates, plus each county's remainder (\"Balance of Adams County, WA\"). Setting it explicitly also overrides auto-detection — \"New York\" auto-detects as the state, so New York City needs \"place\"." - changed
Input schema / properties / geography_type / enumPrevious value: -[ - "state", - "county", - "place", - "tract", - "metropolitan statistical area/micropolitan statistical area", - "combined statistical area", - "consolidated city" -]New value: +[ + "state", + "county", + "place", + "tract", + "metropolitan statistical area/micropolitan statistical area", + "combined statistical area", + "consolidated city", + "zip code tabulation area", + "economic place" +] - changed
Input schema / properties / name / descriptionPrevious value: -"Place name (e.g., \"King County, WA\", \"Seattle, WA\", \"California\") or street address (e.g., \"1600 Pennsylvania Ave NW, Washington, DC 20500\"). Include the state abbreviation to disambiguate places with common names — it narrows a statistical area as well, matching any state the area spans, so \"Kansas City, MO\" and \"Kansas City, KS\" both reach the MO-KS metro area. For a statistical area, the name is the full hyphenated one the Census publishes (\"Seattle-Tacoma-Bellevue, WA\" for the metro area, \"Seattle-Tacoma, WA\" for the combined one) — a single city name matches it too when only one area contains that city."New value: +"Place name (e.g., \"King County, WA\", \"Seattle, WA\", \"California\"), 5-digit ZIP code (e.g., \"98109\", resolved to its ZIP Code Tabulation Area), or street address (e.g., \"1600 Pennsylvania Ave NW, Washington, DC 20500\"). Include the state after a comma — its abbreviation or full name, as in \"Chatham County, Georgia\" — to disambiguate places with common names. It narrows a statistical area as well, matching any state the area spans, so \"Kansas City, MO\" and \"Kansas City, KS\" both reach the MO-KS metro area. Matching ignores case, reads \"Saint\" and \"St.\" as the same word, and accepts unaccented spellings (\"Dona Ana County, NM\"). For a statistical area a leading city is enough (\"Denver, CO\" for the Denver-Aurora-Centennial metro area), and an older full name resolves through its leading city (\"Denver-Aurora-Lakewood, CO\")." - added
Output schema / properties / block_group_fipsAdded value: +{ + "description": "1-digit block group within tract_fips, for a street address only. Query it as geography_level \"block group\" with this as geography_fips, alongside parent_fips, county_fips, and tract_fips. geography_type and fips_summary stay at the tract.", + "type": "string" +} - added
Output schema / properties / census_designated_placeAdded value: +{ + "const": true, + "description": "Present (true) when the resolved place or economic place is a census-designated place (CDP): an unincorporated community the Census delineates for statistics, with no municipal government. Its code is queried like an incorporated place's. Absent for an incorporated place and every other level.", + "type": "boolean" +} - changed
Output schema / properties / fips_summary / descriptionPrevious value: -"Pre-formatted FIPS value ready to use as geography_fips in census_query_data (e.g., \"033\" for King County with state_fips \"53\" as parent_fips, \"42660\" for the Seattle-Tacoma-Bellevue metro area with no parent at all)."New value: +"Pre-formatted FIPS value ready to use as geography_fips in census_query_data (e.g., \"033\" for King County with state_fips \"53\" as parent_fips, \"42660\" for the Seattle-Tacoma-Bellevue metro area with no parent at all, \"03363000\" for Seattle as an ecnbasic 2022 economic place)." - changed
Output schema / properties / place_fips / descriptionPrevious value: -"Place FIPS code when the resolved geography is an incorporated place."New value: +"5-digit place FIPS code when the resolved geography is a place — incorporated or census-designated. For an economic place, the place part of its 8-digit code, which is also the place code the 2017 and 2012 ecnbasic vintages take. For a street address, the incorporated place the address sits in (query it at the place level with state_fips as parent_fips); absent when the address is outside every incorporated place, and a census-designated place is never reported for an address." - changed
Output schema / properties / state_fips / descriptionPrevious value: -"2-digit state FIPS code. Use as parent_fips in census_query_data for sub-state queries. Absent for a metropolitan/micropolitan or combined statistical area, which can span several states and needs no parent_fips."New value: +"2-digit state FIPS code. Use as parent_fips in census_query_data for sub-state queries. Absent for a metropolitan/micropolitan or combined statistical area, which can span several states and needs no parent_fips, and for a zip code tabulation area, whose source layer carries no state."
4 tool updates
- Changed
census_compare_geographies12 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum geographies to return (default: 50, max: 500). When results are truncated, total_count indicates how many matched."New value: +"Maximum geographies to return (default: 50, max: 500). When results are truncated, totalCount says how many matched." - added
Input schema / properties / limit / maximumAdded value: +500 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - changed
Input schema / properties / predicates / descriptionPrevious value: -"Filter values keyed by variable code, applied to every geography in the comparison — e.g. {\"NAICS2017\": \"5112\"} to rank counties by their software-publisher establishment count in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, and dec/ddhca declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so a ranking can read like an overall one without being it. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)."New value: +"Filter values keyed by variable code, applied to every geography in the comparison — e.g. {\"NAICS2017\": \"5112\"} to rank counties by their software-publisher establishment count in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, and dec/ddhca declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so a ranking can read like an overall one without being it. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of \"*\" returns every geography once per category of that dimension, which a ranking cannot hold, so it fails with ambiguous_rows naming the dimension to pin — use census_query_data for a per-category breakdown. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)." - changed
Input schema / properties / sort_by / descriptionPrevious value: -"Variable code to sort by (default: first variable in the list). Must be one of the requested variable codes."New value: +"Variable code to rank on (default: the first code in variables), uppercased like the variables. Must be one of the requested codes, or the call fails with sort_by_not_requested. Geographies rank on that code's own value, so a count ranks by size and only a published percentage such as S1701_C03_001E or DP03_0128PE ranks by rate." - changed
Input schema / properties / variables / descriptionPrevious value: -"Variable codes to compare (e.g., [\"B17001_002E\", \"B17001_001E\"]). On ACS datasets, add the margin-of-error counterpart of a code (same code, E suffix swapped for M) for reliability context. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error."New value: +"Variable codes to compare (e.g., [\"B19013_001E\", \"B19013_001M\"]); the ranking is on one of them, set by sort_by. Codes are uppercased before the request, and each row is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca), the maximum is lower, and too_many_variables states the exact number for the comparison. On ACS datasets, add the margin-of-error counterpart of a code (same code, E suffix swapped for M) for reliability context. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error." - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS but within, or within_county, was not provided. `parent_not_accepted`: within or within_county names a parent the geography level does not sit within. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: One or more variable codes are not found in this dataset and year. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `ambiguous_rows`: The dataset publishes several records per geography and the comparison pinned none of them, so every geography would occupy several ranks with different values. `no_data`: No geographies were returned for the query, or no row matched any entry in the geographies list. `upstream_error`: Census API was unreachable or returned an error. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS but within, or within_county, was not provided. `parent_not_accepted`: within or within_county names a parent the geography level does not sit within. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for this dataset and year. It names only the first unknown code in a request. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `sort_by_not_requested`: sort_by names a code that is not among the requested variables, so no column exists to rank on. `ambiguous_rows`: The dataset publishes several records per geography and the comparison pinned none of them, so every geography would occupy several ranks with different values. `no_data`: No geographies were returned for the query, or no row matched any entry in the geographies list. `upstream_error`: Census API was unreachable or returned an error. Other values are possible when a failure originates below the handler." - changed
Output schema / properties / error / properties / data / properties / reason / examplesPrevious value: -[ - "dataset_not_found", - "missing_api_key", - "geography_not_supported", - "parent_required", - "parent_not_accepted", - "year_not_available", - "variable_not_found", - "variables_unavailable", - "predicate_not_supported", - "ambiguous_rows", - "no_data", - "upstream_error" -]New value: +[ + "dataset_not_found", + "missing_api_key", + "geography_not_supported", + "parent_required", + "parent_not_accepted", + "year_not_available", + "variable_not_found", + "too_many_variables", + "variables_unavailable", + "predicate_not_supported", + "sort_by_not_requested", + "ambiguous_rows", + "no_data", + "upstream_error" +] - changed
Output schema / properties / notice / descriptionPrevious value: -"Guidance when results were truncated, when geographies entries matched no row, when a bare level code matched more than one state, or when the dataset declares filter dimensions the comparison left unset — how to narrow scope, raise the limit, correct the FIPS codes, or add the predicates that pin what the ranking covers. For an unset dimension it also quotes the label of the default the Census API applied, which is what says whether the ranking is on a total or on one category."New value: +"Guidance when results were truncated, when the sort column holds no number on any row (so the rows are in the order the Census returned them rather than ranked), when geographies entries matched no row, when a bare level code matched more than one state, or when the dataset declares filter dimensions the comparison left unset — how to narrow scope, raise the limit, correct the FIPS codes, or add the predicates that pin what the ranking covers. For an unset dimension it also quotes the label of the default the Census API applied, which is what says whether the ranking is on a total or on one category. Also names any variable codes whose flags could not be checked because the request had no room left under the Census 50-column limit — a withheld value there reads as 0 and ranks as one." - changed
Output schema / properties / rows / items / properties / variables / descriptionPrevious value: -"Map of variable code to value entry. Each key is a variable code from the variables input; each value has: estimate (number|null), moe (number|null, optional), label (string), suppressed (boolean), value (string, optional). An estimate of null means one of three things and the other fields say which: suppressed true is a number the Census withheld, a value field is a cell holding text rather than a number (GEO_ID returns \"0500000US53033\"; the older ACS profile vintages write not-applicable as \"(X)\" in a column that is a number elsewhere), and neither is a cell with nothing in it. Text has no ordering, so sorting on a column of it leaves every row tied and ranked in the order the Census returned them."New value: +"Map of variable code to value entry. Each key is a variable code from the variables input, uppercased; each value has: estimate (number|null), moe (number|null, optional), label (string), suppressed (boolean), suppression_reason (string, optional), open_ended (true, optional), flag ({code, meaning}, optional), value (string, optional). An estimate of null means one of three things and the other fields say which: suppressed true is a number the Census withheld, a value field is a cell holding text rather than a number (GEO_ID returns \"0500000US53033\"; the older ACS profile vintages write not-applicable as \"(X)\" in a column that is a number elsewhere), and neither is a cell with nothing in it. suppression_reason carries the meaning the Census publishes for the sentinel or flag, and a suppressed value ranks after every number in either sort direction. On ACS, a margin of error the Census treats as zero (a controlled estimate) is moe 0. open_ended true marks an ACS median that falls in the lowest or highest interval of an open-ended distribution, so the estimate is that interval's boundary (250001 for \"250,000+\") — it ranks by that figure, so geographies sharing it are tied, and it appears only when the matching M code was requested. flag is the symbol a business dataset (cbp, ecnbasic, nonemp) published beside the value: a withholding flag comes with suppressed true, and so does a noise or data-quality band beside a 0, which is the range a range column such as EMP_N or RCPTOT_IMP publishes in place of a number; a quality note keeps the estimate. Text has no ordering, so sorting on a column of it leaves every row tied and ranked in the order the Census returned them, and the notice says so." - changed
Output schema / properties / sortVariable / descriptionPrevious value: -"Variable code used for sorting."New value: +"Variable code the rows are ranked on, uppercased as it appears in variables."
- Changed
census_list_predicate_values4 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum codes to return (default: 50, max: 500)."New value: +"Maximum codes to return (default: 50, max: 500). totalCount says how many matched." - added
Input schema / properties / limit / maximumAdded value: +500 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer"
- Changed
census_query_data14 fields changed- changed
Input schema / properties / geography_fips / descriptionPrevious value: -"FIPS code for the target geography (e.g., \"033\" for a county, \"*\" for all geographies at the level within the parent). Use census_resolve_geography to obtain this value — it is returned as fips_summary. The Census API matches this literally and its width follows geography_level, so it is passed through unpadded: a county is 3 digits (\"051\", not \"51\") and a tract is 6. parent_fips and county_fips are zero-padded for you; this one is not."New value: +"FIPS code for the target geography (e.g., \"033\" for a county, \"*\" for every geography at the level within the parent, returned up to limit rows per call and paged with offset). Use census_resolve_geography to obtain this value — it is returned as fips_summary. The Census API matches this literally and its width follows geography_level, so it is passed through unpadded: a county is 3 digits (\"051\", not \"51\") and a tract is 6. parent_fips and county_fips are zero-padded for you; this one is not." - added
Input schema / properties / limitAdded value: +{ + "description": "Most rows to return (default: 50, max: 500). Rows come in GEOID order, a geography's records or categories in code order, and each one counts, so a geography returned as several records (pep/charv April and July) or as one row per category of a \"*\" predicate takes one row each. totalCount says how many rows matched.", + "maximum": 500, + "minimum": 1, + "type": "integer" +} - added
Input schema / properties / offsetAdded value: +{ + "description": "Rows to skip before returning up to limit (default: 0). Pages run in GEOID order, so offset 50 with limit 50 returns rows 51–100, and the notice names the offset of the next page. An offset at or past totalCount returns no rows.", + "maximum": 9007199254740991, + "minimum": 0, + "type": "integer" +} - changed
Input schema / properties / predicates / descriptionPrevious value: -"Filter values keyed by variable code, sent as extra query parameters — e.g. {\"NAICS2017\": \"5112\"} to count only software publishers in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, and dec/ddhca declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so an unfiltered value can read like a total without being one. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)."New value: +"Filter values keyed by variable code, sent as extra query parameters — e.g. {\"NAICS2017\": \"5112\"} to count only software publishers in cbp. The business datasets (cbp, ecnbasic, nonemp), pep/charv, and dec/ddhca declare filter dimensions such as industry (NAICS2017/NAICS2022), legal form (LFO), size class (EMPSZES/RCPSZES), tax status (TAXSTAT), operation type (TYPOP), sex (SEX), age (AGE), and population group (POPGROUP). Leaving one unset is not an error: the Census API substitutes its own default, which is the all-categories total on cbp NAICS2017 but a single population group on dec/ddhca POPGROUP and a single sector on ecnbasic NAICS2022 — so an unfiltered value can read like a total without being one. Every unset dimension is named in the response notice and its applied default is echoed per row in applied_filters. Keys are matched case-insensitively, and a blank value is treated as omitted. A value of \"*\" returns one row per category of that dimension for each geography, each row labelled with its category in record (e.g. {\"NAICS2017\": \"*\"} gives King County one row per industry) — a breakdown that can run to over a thousand rows. Code names vary by dataset and vintage — cbp 2023 uses NAICS2017 while nonemp 2023 uses NAICS2022 — so read them from the notice or from census_search_variables. Call census_list_predicate_values for the codes a dimension accepts; NAICS values are standard North American Industry Classification System codes at any depth (51 information, 5112 software publishers)." - changed
Input schema / properties / variables / descriptionPrevious value: -"Variable codes to retrieve (e.g., [\"B19013_001E\", \"B19013_001M\"]). Max 50 per request. Use census_search_variables to find codes. On ACS datasets only, each estimate has a margin-of-error counterpart at the same code with the E suffix swapped for M — request both to get the margin alongside the estimate. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error, and an E-final code there is an ordinary code with no M sibling. A code can also name a text column rather than a measure — GEO_ID, on every dataset, is the nationally unique geography identifier and comes back under value with estimate null, which is the code to request when a stable join key is what is wanted."New value: +"Variable codes to retrieve (e.g., [\"B19013_001E\", \"B19013_001M\"]). Codes are uppercased before the request, so \"b19013_001e\" reads as B19013_001E and the response is keyed by the uppercase code. At most 49 per call: the Census API accepts 50 columns per request and every query also sends NAME. On datasets where a label column is added for each filter dimension left unset, or record columns are added (cbp, ecnbasic, nonemp, pep/charv, dec/ddhca), the maximum is lower, and too_many_variables states the exact number for the query. Use census_search_variables to find codes. On ACS datasets only, each estimate has a margin-of-error counterpart at the same code with the E suffix swapped for M — request both to get the margin alongside the estimate. Other dataset families (pep, dec, cbp, ecnbasic, nonemp) publish no margins of error, and an E-final code there is an ordinary code with no M sibling. A code can also name a text column rather than a measure — GEO_ID, on every dataset, is the nationally unique geography identifier and comes back under value with estimate null, which is the code to request when a stable join key is what is wanted." - changed
Output schema / anyOfPrevious value: -[ - { - "not": { - "required": [ - "error" - ] - }, - "required": [ - "rows", - "totalRows", - "dataset", - "year" - ] - }, - { - "required": [ - "error" - ] - } -]New value: +[ + { + "not": { + "required": [ + "error" + ] + }, + "required": [ + "rows", + "totalRows", + "totalCount", + "truncated", + "dataset", + "year" + ] + }, + { + "required": [ + "error" + ] + } +] - changed
Output schema / properties / error / properties / data / properties / reason / descriptionPrevious value: -"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: One or more variable codes do not exist in the requested dataset and year. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS code but parent_fips was not provided, or a tract/block-group level requires county_fips but it was omitted. `parent_not_accepted`: parent_fips or county_fips names a parent the geography level does not sit within. `no_data`: The query returned no rows. `too_many_variables`: More than 50 variable codes were requested. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `dataset_not_found`: Dataset code is not recognized. `missing_api_key`: CENSUS_API_KEY is not configured or the key is invalid. `year_not_available`: The dataset does not serve the requested vintage year. `variable_not_found`: The Census API rejected a variable code as unknown for the requested dataset and year. It names only the first unknown code in a request. `variables_unavailable`: The variable metadata endpoint returned an unparseable response for this dataset and year. `geography_not_supported`: The requested geography level does not exist in this dataset and year. `parent_required`: The geography level requires a parent FIPS code but parent_fips was not provided, or a tract/block-group level requires county_fips but it was omitted. `parent_not_accepted`: parent_fips or county_fips names a parent the geography level does not sit within. `no_data`: The query returned no rows. `too_many_variables`: The variable codes plus NAME and the label and record columns added for the dataset exceed the 50 columns the Census API accepts per request. `predicate_not_supported`: A key in predicates is not a variable in this dataset and year. `upstream_error`: Census API returned an error or was unreachable. Other values are possible when a failure originates below the handler." - changed
Output schema / properties / notice / descriptionPrevious value: -"Warning that the dataset declares filter dimensions the query left unset, naming each one alongside the label of the default the Census API applied to it. That default is an all-categories total on some dimensions and one ordinary category on others, so the label is what says which. Also carries the warning that a geography came back on more than one row, naming the column that separates the records and the values it took."New value: +"Warning that the dataset declares filter dimensions the query left unset, naming each one alongside the label of the default the Census API applied to it. That default is an all-categories total on some dimensions and one ordinary category on others, so the label is what says which. Also carries the warning that a geography came back on more than one row, naming the column that separates the records and the values it took; the range of rows returned when offset or limit left some out, with the offset of the next page; and any variable codes whose flags could not be checked because the request had no room left under the Census 50-column limit — a withheld value there reads as 0." - changed
Output schema / properties / rows / descriptionPrevious value: -"One row per geography. When geography_fips is \"*\", includes all geographies at the level within the parent."New value: +"One row per geography, or per record or category where a geography has several. When geography_fips is \"*\", the rows from offset up to limit of every geography at the level within the parent, in GEOID order." - changed
Output schema / properties / rows / items / properties / record / descriptionPrevious value: -"Which record this row is, for a dataset that publishes more than one per geography — keyed by the column that separates them, each value carrying a code and a label (e.g. {\"MONTH\": {\"code\": \"7\", \"label\": \"July\"}}). pep/charv publishes an April estimates base and a July estimate, so one geography comes back on two rows whose numbers differ; this field is what says which is which. Pass the code back in predicates (e.g. {\"MONTH\": \"7\"}) to return that record alone. Absent on the datasets that return one row per geography."New value: +"Which record this row is, when one geography comes back on more than one row — keyed by the column that separates them, each value carrying a code and a label (e.g. {\"MONTH\": {\"code\": \"7\", \"label\": \"July\"}}). pep/charv publishes an April estimates base and a July estimate, so one geography comes back on two rows whose numbers differ; this field is what says which is which. A dimension set to \"*\" in predicates lands here too, one row per category (e.g. {\"NAICS2017\": {\"code\": \"11\", \"label\": \"Agriculture, forestry, fishing and hunting\"}}), with the code as its label when the dimension publishes no label column. Pass the code back in predicates (e.g. {\"MONTH\": \"7\"}) to return that record alone. Absent on the datasets that return one row per geography." - changed
Output schema / properties / rows / items / properties / variables / descriptionPrevious value: -"Map of variable code to value entry. Each key is a variable code from the variables input; each value has: estimate (number|null), moe (number|null, optional), label (string), suppressed (boolean), suppression_reason (string, optional), value (string, optional). An estimate of null means one of three things and the other fields say which: suppressed true is a number the Census withheld, a value field is a cell holding text rather than a number (GEO_ID returns \"0500000US53033\"; the older ACS profile vintages write not-applicable as \"(X)\" in a column that is a number elsewhere), and neither is a cell with nothing in it."New value: +"Map of variable code to value entry. Each key is a variable code from the variables input, uppercased; each value has: estimate (number|null), moe (number|null, optional), label (string), suppressed (boolean), suppression_reason (string, optional), open_ended (true, optional), flag ({code, meaning}, optional), value (string, optional). An estimate of null means one of three things and the other fields say which: suppressed true is a number the Census withheld, a value field is a cell holding text rather than a number (GEO_ID returns \"0500000US53033\"; the older ACS profile vintages write not-applicable as \"(X)\" in a column that is a number elsewhere), and neither is a cell with nothing in it. suppression_reason carries the meaning the Census publishes for the sentinel or flag. On ACS, a margin of error the Census treats as zero (a controlled estimate) is moe 0, not a suppression. open_ended true marks an ACS median that falls in the lowest or highest interval of an open-ended distribution, so the estimate is that interval's boundary (250001 for \"250,000+\", 9999 for \"10,000-\") rather than the median itself — it appears only when the matching M code was requested, since that margin of error is the only signal, and it does not say which end. flag is the symbol a business dataset (cbp, ecnbasic, nonemp) published beside the value: a withholding flag (D, S, an employment or sales range letter) comes with suppressed true, and so does a noise or data-quality band (G/H/J, 0-9) beside a 0, which is the range a range column such as EMP_N or RCPTOT_IMP publishes in place of a number; a quality note (r revised, s high relative standard error) keeps the estimate." - added
Output schema / properties / totalCountAdded value: +{ + "description": "Number of rows the query matched, before offset and limit were applied.", + "type": "number" +} - changed
Output schema / properties / totalRows / descriptionPrevious value: -"Number of geography rows returned."New value: +"Number of rows returned." - added
Output schema / properties / truncatedAdded value: +{ + "description": "True when rows were left out by offset or limit — totalCount exceeds the rows returned.", + "type": "boolean" +}
- Changed
census_search_variables5 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum results to return (default: 20, max: 100). Increase if total_matches greatly exceeds the limit."New value: +"Maximum results to return (default: 20, max: 100). Increase if totalMatches greatly exceeds the limit." - added
Input schema / properties / limit / maximumAdded value: +100 - added
Input schema / properties / limit / minimumAdded value: +1 - changed
Input schema / properties / limit / typePrevious value: -"number"New value: +"integer" - changed
Output schema / properties / truncated / descriptionPrevious value: -"True when total_matches exceeded the limit and results were cut off."New value: +"True when totalMatches exceeded the limit and results were cut off."
8 tool updates
- First observed
census_compare_geographies - First observed
census_get_variable - First observed
census_list_datasets - First observed
census_list_geographies - First observed
census_list_predicate_values - First observed
census_query_data - First observed
census_resolve_geography - First observed
census_search_variables
Related MCP Connectors
Query US Census Bureau data: demographics, economics, and housing statistics.
Query US Treasury national debt, interest rates, exchange rates, and fiscal datasets via MCP.
Census MCP — U.S. Census Bureau housing-relevant APIs.
Access U.S. congressional data - bills, votes, members, committees - via MCP.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceQuery SEC EDGAR filings, XBRL financials, and company data through MCP.681 npm10Apache 2.0
- AlicenseNot gradedqualityAmaintenanceSearch and query government open-data portals (Socrata SODA API) via MCP.326 npm3Apache 2.0
- FlicenseNot gradedqualityCmaintenanceA production-grade MCP server for querying U.S. Census Bureau data (ACS 5-Year and Decennial) with tools for geographic fuzzy matching, variable search, and batched data retrieval, backed by a PostgreSQL cache for performance.-
- AlicenseNot gradedqualityBmaintenanceAccess FCC broadband availability, coverage analysis, and digital divide data for US geographies and census blocks via MCP.83 npm1Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.