nyc-open-data-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@nyc-open-data-mcpWhich ramen spots in 10003 have an A grade?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
nyc-open-data-mcp
An MCP server that gives Claude, Cursor, or any MCP client read-only access to NYC Open Data: 2,400+ city datasets served through the Socrata SODA API. It includes two ready-made tools for the questions people ask most (restaurant health grades and 311 complaints) and two general tools that let the model find and query any other dataset. Inputs are validated, every value placed in a query is escaped, results are paginated and capped in size, upstream errors come back as plain-language hints the model can act on. No API key is required.
It runs two ways: over stdio on your own machine (npx -y github:frogr/nyc-open-data-mcp), or as a remote server over Streamable HTTP (the current MCP transport for servers on the web) with a web playground where you can try every tool in a browser. Screenshots of the playground are in docs/screenshots/ and what was verified is in PROOF.md.

What you can ask
Which ramen spots in 10003 have an A grade?
What were the top 311 complaints in the East Village last month?
Are there any restaurants on St. Marks Place with critical violations in their latest inspection?
Find the dataset for NYC street tree census and tell me the most common species in Brooklyn.
How did noise complaints in 11211 change between June and August?
Related MCP server: DataSF MCP
Install
Requires Node.js 20 or newer.
The package is not on npm yet. The commands below install it straight from GitHub with npx -y github:frogr/nyc-open-data-mcp: npm clones the repo, installs dependencies and builds it (a prepare script runs npm run build). The first start takes about 20 seconds while that happens, so run it once in a terminal before adding it to a client. If you'd rather not run a build through npx, use From source.
After the package is published to npm, npx -y nyc-open-data-mcp will do the same thing. Until then, don't run that name: nothing has been published under it by this project.
Claude Desktop
Add this to claude_desktop_config.json (macOS: ~/Library/Application Support/Claude/, Windows: %APPDATA%\Claude\), then restart Claude Desktop:
{
"mcpServers": {
"nyc-open-data": {
"command": "npx",
"args": ["-y", "github:frogr/nyc-open-data-mcp"],
"env": {
"SOCRATA_APP_TOKEN": "optional-but-recommended"
}
}
}
}Claude Code
claude mcp add --transport stdio nyc-open-data -- npx -y github:frogr/nyc-open-data-mcp
# with an app token, available in every project:
claude mcp add --env SOCRATA_APP_TOKEN=your-token --transport stdio --scope user nyc-open-data -- npx -y github:frogr/nyc-open-data-mcpCursor
Add to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (this project):
{
"mcpServers": {
"nyc-open-data": {
"command": "npx",
"args": ["-y", "github:frogr/nyc-open-data-mcp"]
}
}
}From source
git clone https://github.com/frogr/nyc-open-data-mcp && cd nyc-open-data-mcp
npm ci # also builds dist/ through the prepare script
# then use "command": "node", "args": ["/absolute/path/to/nyc-open-data-mcp/dist/index.js"]Configuration
Env var | Default | Purpose |
| none | Free Socrata app token. Unauthenticated requests share a small IP-based rate limit; a token raises it a lot. |
|
| Per-request timeout. |
The HTTP server also reads PORT (3000), HOST (0.0.0.0), RATE_LIMIT_PER_MINUTE (30), DAILY_REQUEST_LIMIT (5000), MAX_BODY_BYTES (65536), REQUEST_TIMEOUT_MS (30000), CORS_ORIGINS (*) and TRUST_PROXY (0, set to 1 behind one reverse proxy such as Render's). All are listed in .env.example.
Use it remotely
npm start runs the HTTP server. It serves:
Route | What it is |
| The MCP endpoint (Streamable HTTP, stateless, JSON responses). |
| The playground: example questions, a form for each tool built from its input schema, results as a table or raw JSON, and copy-paste client config. |
| Status, version, whether an app token is set, and the current limits. Never calls Socrata. |
Once it is deployed (see Deploy), point a client at https://<your-host>/mcp:
# Claude Code
claude mcp add --transport http nyc-open-data https://<your-host>/mcp// Cursor: ~/.cursor/mcp.json
{ "mcpServers": { "nyc-open-data": { "url": "https://<your-host>/mcp" } } }// Claude Desktop without a custom connector: claude_desktop_config.json, via the mcp-remote bridge
{ "mcpServers": { "nyc-open-data": { "command": "npx", "args": ["-y", "mcp-remote", "https://<your-host>/mcp"] } } }In Claude Desktop or claude.ai you can also add it under Settings > Connectors > Add custom connector, if your plan has custom connectors. To poke at it by hand, run npx @modelcontextprotocol/inspector, choose Streamable HTTP and paste the URL.
Limits on the public endpoint. These protect the shared Socrata quota and keep one visitor from using it all up.
Per-IP token bucket on
POST /mcp, 30 requests per minute by default. A blocked call gets HTTP 429 withRetry-After.A global cap per UTC day (5000 by default), also 429.
Request bodies over 64 KB get 413, checked while streaming, so a missing
Content-Lengthdoesn't get around it.Each
/mcprequest has a 30 second wall-clock limit (504), and each Socrata call a 15 second timeout. Slow clients are cut off by Node's header and request timeouts.CORS is open (
*) by default so browser-based clients work. SetCORS_ORIGINSto a list to restrict it.Internal errors are logged and visitors get a generic message, never a stack trace.
The limiter lives in memory, so counts reset on restart and are per instance. That is fine for one free-tier instance.
Tools
Tool | What it does | Key inputs | Returns |
| DOHMH restaurant grades (43nn-pn8j) |
| Per restaurant: latest grade and what it means, latest inspection date, type, score, and deduplicated violations with critical flags |
| 311 complaint summary (erm2-nwe9) |
| Total requests, top complaint types with counts and share, remainder count, most recent example requests |
| Search the NYC catalog |
| Dataset id, name, short description, category, last-updated date, URL, |
| Read-only SoQL against any dataset |
| Rows, |
Every tool is marked readOnlyHint: true and declares an outputSchema. Results come back as JSON text, which works in any client, and as structuredContent for clients that use it.
Design notes
Why these four tools. The two specific tools cover the questions people actually ask, and they do the awkward parts on the server. The restaurant dataset stores one row per violation, so the tool first groups by restaurant to page through restaurants, then fetches the history for only that page and works out the latest grade. That grade isn't always from the latest inspection: a re-inspection can leave a grade pending. The 311 dataset has about 22.7 million rows, so the tool runs three small aggregate queries (total, group-by, recent samples) on Socrata's side instead of downloading rows. The two general tools are the fallback for everything else, and search_datasets returns column names so the model can write a valid where clause on its first try.
Pagination. Every list result includes next_offset, which is null on the last page. query_dataset fetches limit + 1 rows so has_more is exact rather than guessed.
Limits. query_dataset caps limit at 500 rows. Each response is also capped at about 60 KB of JSON; when that cap removes rows, the response says so and gives the offset to continue from. 311 date ranges max out at 366 days so queries don't time out upstream. Long descriptions are shortened.
Safety.
Every user value that goes into SoQL (names, ZIPs, complaint types, dates, ids) is passed through
soqlString(), which doubles single quotes, sox' OR '1'='1stays an ordinary string. In "contains" searches, the LIKE wildcards%and_are removed from user input.Dataset ids must match
xxxx-xxxx. ZIP codes must be 5 digits. Dates must be realYYYY-MM-DDdates. Boroughs come from a fixed list.Parameters are encoded with
URLSearchParams, so a&inside a clause can't add extra query parameters.query_datasetpasses SoQL clauses through as written, which is the point of that tool. That's safe because the Socrata endpoint is read-only public data and the tool can only send GET requests to/resource/{id}.json.
Rate limits and reliability.
Every request has a timeout.
429 and 5xx responses are retried up to 2 times with exponential backoff, following
Retry-Afterup to 5 seconds.Identical requests are cached in memory for 60 seconds, so an agent that asks the same thing again doesn't send another request.
An optional app token raises the rate limit.
Errors. Upstream failures are mapped to tool errors (isError: true) that say what went wrong and what to try next, for example:
Dataset 'zzzz-zzzz' was not found on data.cityofnewyork.us. (HTTP 404, dataset.missing)
Hint: Use search_datasets to find a valid dataset id (format: xxxx-xxxx).Deploy
The HTTP server needs no database and no secrets. A free tier is enough.
Render (free web service; Render is a hosting platform that reads render.yaml from the repo):
Push this repo to GitHub.
In Render: New > Blueprint, pick the repo. It reads
render.yaml(planfree, buildnpm ci && npm run build, startnpm start, health check/health,TRUST_PROXY=1).Optional: set
SOCRATA_APP_TOKENwhen Render asks for it (it is markedsync: false, so it is never stored in the repo).Open
https://<service>.onrender.com/for the playground. The MCP URL is the same host plus/mcp.
Free Render services sleep after a while without traffic, so the first request after a quiet spell takes longer.
Docker (any host that runs containers):
docker build -t nyc-open-data-mcp .
docker run -p 3000:3000 -e TRUST_PROXY=1 nyc-open-data-mcpAnywhere with Node 20+:
npm ci && npm run build
PORT=3000 npm start
curl localhost:3000/healthDevelopment
npm install
npm test # vitest, recorded fixtures only, no network
npm run build # compiles to dist/
npm run smoke # spawns the server over stdio, runs initialize + tools/list
node scripts/smoke.mjs --live # also makes one real call to Socrata
npm run smoke:http # starts the HTTP server, checks /health, CORS, initialize, tools/list, /
node scripts/smoke-http.mjs --live # also makes real tools/calls over HTTP
npm start # HTTP server + playground on PORT (default 3000)
npm run screenshots # playground screenshots with Chromium, needs networkTests use hand-built fixtures in test/fixtures/ that match Socrata's response shapes, plus a mocked fetch that throws on any request it doesn't recognize. test/server.test.ts runs the whole MCP protocol in memory: tool listing, schema validation, output-schema checks, and error mapping. test/http.test.ts starts the HTTP server on a real local socket and talks to it with the official SDK client, then checks CORS, limits, timeouts and error responses. test/rateLimit.test.ts covers the limiters with a fake clock.
src/
index.ts stdio entrypoint (the package bin)
http.ts HTTP entrypoint (npm start): Node adapter, body limit, client IP
app.ts routes: /mcp, /health, /, CORS, rate limits, timeouts
rateLimit.ts per-IP token bucket and daily cap
server.ts tool registration
socrata.ts HTTP client: timeouts, retries, cache, error mapping
soql.ts literal escaping and predicate builders
tools/ one file per tool: zod input/output schemas + handler
public/index.html the playground (one file, no build step)
test/ vitest suites + fixtures/
scripts/smoke.mjs raw JSON-RPC stdio smoke test
scripts/smoke-http.mjs same, over HTTP
scripts/screenshots.mjs Playwright screenshots of the playgroundLicense
MIT © Austin French
Need an MCP server for your own API? austn.net
Available Tools
4 toolsquery_datasetQuery a dataset (SoQL)ARead-onlyIdempotent
Run a read-only SoQL query against any NYC Open Data dataset by id. Supports select / where / order / group / full-text q, with limit (max 500) and offset paging; the response says when more rows exist. Get the dataset id and column names from search_datasets first. For counts, prefer select="count(*)" or a group-by over pulling raw rows.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | Full-text search across all text columns. | |
| group | No | SoQL $group, required when $select mixes aggregates and plain columns. | |
| limit | No | Rows to return (1-500, default 100). | |
| order | No | SoQL $order, e.g. "inspection_date DESC". | |
| where | No | SoQL $where, e.g. "zipcode = '10003' AND grade = 'A'". Single-quote string literals; double any quote inside ('O''Brien'). | |
| offset | No | Rows to skip; pass next_offset from a previous call to page. Use a stable order when paging. | |
| select | No | SoQL $select, e.g. "borough, count(*) as n". Default: all columns. | |
| dataset_id | Yes | Socrata dataset id, e.g. 43nn-pn8j (restaurant inspections) or erm2-nwe9 (311). |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| rows | Yes | |
| offset | Yes | |
| has_more | Yes | |
| returned | Yes | |
| dataset_id | Yes | |
| next_offset | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context on top: the 500-row limit ceiling, offset-based paging, and that the response signals when more rows exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and scope, followed by capabilities and then the routing tip. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value explanation is unnecessary, and the description still notes the paging signal in responses. The one omission is that it never routes users to the specialized restaurant/311 siblings, which an agent working in that domain would benefit from knowing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each SoQL clause already carries its own documented example and quoting rules, so the schema does the heavy lifting. The description largely restates the clause list (select/where/order/group/q/limit/offset) and the max-500 limit, adding little syntax or format detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (run) and resource (read-only SoQL query against any NYC Open Data dataset by id), and immediately scopes it as read-only. It is clearly distinguishable from search_datasets, which is named as the discovery step rather than the query execution tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit prerequisite guidance (get dataset id and column names from search_datasets first) and a concrete recommendation for count queries (select="count(*)" or group-by rather than raw rows). It stops short of explaining when to prefer the specialized restaurant_inspections and service_requests_311 siblings over a generic SoQL query, which is a real routing gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restaurant_inspectionsNYC restaurant inspection gradesARead-onlyIdempotent
Look up NYC restaurant health inspections (DOHMH). Filter by name fragment, ZIP code, borough and/or cuisine (at least one). Returns each restaurant's latest letter grade, latest inspection date, score and a violations summary. Paginated, most recently inspected first.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Restaurant name or part of it, case-insensitive, e.g. "ramen" or "Joe's Pizza". | |
| limit | No | Restaurants per page (1-50, default 10). | |
| offset | No | Pagination offset; pass next_offset from a previous call. | |
| borough | No | Borough name. | |
| cuisine | No | Cuisine contains, e.g. "Japanese", "Pizza", "Thai". | |
| zip_code | No | 5-digit NYC ZIP code, e.g. 10003. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| has_more | Yes | |
| returned | Yes | |
| next_offset | Yes | |
| restaurants | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world, non-destructive, so the safety profile is covered. The description adds genuinely new behavioral context beyond the annotations: results are 'most recently inspected first' and results are paginated, which the agent needs to page correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: identity, filtering constraint, return shape and ordering. The critical 'at least one' constraint is front-loaded rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and full schema coverage, the description only needs to add calling context, and it does: it names the data source, the required filter condition, the sort order and the pagination model. Nothing blocking a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with examples and constraints for name, cuisine, zip_code, borough enum, limit and offset already documented. The description restates the filter set but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Look up NYC restaurant health inspections (DOHMH)') and enumerates the filter dimensions, so the agent knows exactly what it returns. It does not explicitly distinguish itself from generic dataset siblings like search_datasets or query_dataset, so sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(at least one)' establishes a real constraint on how to invoke the tool, which is useful guidance. However, it never says when to prefer this over search_datasets/query_dataset, so usage is only implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_datasetsSearch NYC Open Data catalogARead-onlyIdempotent
Search the NYC Open Data catalog (data.cityofnewyork.us) by keyword. Returns dataset ids, names, short descriptions, last-updated dates and column names. Use this first to find a dataset id and its columns before calling query_dataset.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Results per page (1-25, default 10). | |
| query | Yes | Keywords, e.g. "restaurant inspections", "bike lanes", "rat sightings". | |
| offset | No | Pagination offset; pass next_offset from a previous call. | |
| category | No | Optional NYC Open Data category, e.g. "Health", "Transportation", "Housing & Development". | |
| include_columns | No | Include each dataset's column names and types (needed to write query_dataset filters). |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| total | Yes | Total matching datasets. |
| offset | Yes | |
| datasets | Yes | |
| returned | Yes | |
| next_offset | Yes | Pass as offset for the next page; null when there are no more. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds valuable context by listing the exact fields returned and the downstream workflow purpose. It does not mention rate limits or pagination semantics, which would be a further improvement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler: scope, return payload, then the workflow directive. The most actionable guidance (use before query_dataset) is placed last as a clear call to action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and the annotations carry the safety profile. Combined with 100% schema coverage, everything an agent needs to invoke this tool and route correctly afterward is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (query, limit, offset, category, include_columns) is already fully documented in the schema with examples and ranges. The description only restates the keyword-search nature of the query parameter, adding no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (NYC Open Data catalog), names the host domain, and enumerates what is returned (ids, names, descriptions, dates, column names). It is clearly distinguishable from the query sibling, which executes queries rather than discovering datasets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs the agent to 'use this first to find a dataset id and its columns before calling query_dataset', establishing a clear ordering relationship with the named alternative. The condition that selects this tool over query_dataset is stated outright.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_requests_311NYC 311 complaint summaryARead-onlyIdempotent
Summarize NYC 311 service requests for an area and time window. Filter by ZIP codes, borough and/or complaint type (substring); dates default to the last 30 days (max range 366 days). Returns the total, top complaint types with counts and share, and a few recent example requests.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | How many complaint types to rank (1-50, default 10). | |
| borough | No | Borough: Manhattan, Brooklyn, Queens, Bronx or Staten Island. | |
| end_date | No | Inclusive end date YYYY-MM-DD. Default: today (New York time). | |
| zip_codes | No | One or more 5-digit ZIP codes. Neighborhoods span several, e.g. East Village = ["10003","10009"]. | |
| start_date | No | Inclusive start date YYYY-MM-DD. Default: 30 days before end_date. | |
| sample_size | No | Most recent example requests to include (0-25, default 5). | |
| complaint_type | No | Complaint type contains (case-insensitive), e.g. "noise", "heat", "rodent", "illegal parking". |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| filters | Yes | |
| samples | Yes | |
| date_range | Yes | |
| total_requests | Yes | |
| other_types_count | Yes | Requests not in the top_n types. |
| top_complaint_types | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and openWorld, so safety is covered. The description adds real behavioral context beyond them: the 30-day default window, the 366-day hard range cap, and that complaint_type is a substring match rather than exact -- useful constraints an agent cannot get from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: what it summarizes, how to scope it, and what comes back. Filtering and the default date window are front-loaded, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a 100%-documented schema, full annotation coverage and an output schema, the description need not explain return fields, and it correctly gestures at them briefly. The remaining gap is that it does not clarify how multiple filters combine or whether ZIP and borough are intersected, which matters for a faceted summary tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description contributes value the schema does not: the 366-day maximum range and the implied 'and/or' combination of ZIP, borough and complaint_type filters. It still leaves the AND/OR interaction between filters ambiguous, keeping it below 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: summarize NYC 311 service requests, with scope (area + time window) and output shape. It is clearly more specific than the generic sibling query_dataset, but it never names or contrasts those siblings, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description (aggregate/summarize view of 311 complaints), and the filtering options give a sense of when it applies. However there is no explicit when-to-use vs when-not guidance and no routing to query_dataset or search_datasets for raw-record needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.2.0- First observed
query_dataset - First observed
restaurant_inspections - First observed
search_datasets - First observed
service_requests_311
TDQS
Scored across 4 tools
search_datasets (catalog discovery) and query_dataset (SoQL execution) are clearly separated by a documented workflow. However, restaurant_inspections and service_requests_311 overlap with what query_dataset could do against those same datasets, so an agent may wonder when to use the specialized tools versus the generic query tool.
All names are snake_case, which is good, but conventions are mixed: search_datasets and query_dataset use verb_noun, while restaurant_inspections and service_requests_311 are noun phrases with no verb. Readable but not a predictable pattern.
Four tools is on the lean side but each earns its place: one discovery, one general query, and two high-value pre-built domain queries. A reasonable, well-scoped set for the server's purpose.
search_datasets plus query_dataset give broad coverage of the whole NYC Open Data catalog, and the two specialized tools add convenience aggregations. Minor gaps like dataset metadata/column-listing or write operations exist but read-only SoQL over any dataset covers most needs.
Maintenance
Related MCP Connectors
Search and query government open-data portals (Socrata SODA API).
Agent-ready NYC public records. Hosted, source-backed civic data organized around durable anchors.
Query and explore Nova Scotia open datasets via the Socrata SODA API.
Provides access to Civic Plus - See Click Fix, allowing you to interact with your data via an LLM.…
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables AI assistants to search, explore, and query San Francisco's open data portal through a standardized interface for public datasets. It supports SQL-like querying via the Socrata platform and includes features like fuzzy column matching and schema caching.4-
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to query and retrieve San Francisco open data from data.sfgov.org via the Socrata SODA API.246 npmMIT
- AlicenseNot gradedqualityBmaintenanceProvides access to Cincinnati open data via the Socrata SODA API, allowing users to query datasets using natural language or direct tool calls.371 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables querying NYC Open Data datasets using natural language, with tools to search datasets, run SoQL queries, and retrieve metadata, all without an API key.254 npmMIT