Skip to main content
Glama
cliwant

mcp-sam-gov

ckan_discover_datasets

Read-only

Search CKAN government data portals by keyword to find queryable datastore resource IDs, returning dataset titles, formats, and datastore status for direct follow-up queries.

Instructions

Find CKAN datastore resource ids by keyword via package_search (keyless). Input host (allowlisted enum), q (e.g. 'procurement', 'checkbook'), limit (≤100, def 20). Returns per-resource rows [{ resourceId, name, datasetTitle, format, datastoreActive }] + totalAvailable = the matching DATASET count. Feed a datastoreActive:true result's resourceId to ckan_query (a datastoreActive:false resource is a raw file blob NOT in the datastore, not queryable).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
qYesKeyword(s) to find datasets, e.g. 'procurement', 'checkbook', 'vendor'.
hostYesWhich allowlisted CKAN portal to query (curated .gov hosts — the SSRF host allowlist, no free host): data.ca.gov (CA), data.virginia.gov (VA — eVA), data.boston.gov (City of Boston Checkbook).
limitNoMax datasets (packages) to return, 1..100, default 20.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv1.16.0
    • changedInput schema / properties / host / enum
      Previous value: -[
      -  "data.ca.gov",
      -  "data.virginia.gov",
      -  "data.boston.gov"
      -]New value: +[
      +  "data.ca.gov",
      +  "data.ok.gov",
      +  "data.virginia.gov",
      +  "data.boston.gov"
      +]
  2. Addedv1.12.0

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint. The description adds substantial behavioral context: keyless operation, the meaning of totalAvailable as a dataset count, the per-resource row shape, and the critical distinction that datastoreActive:false resources are raw file blobs not in the datastore. This goes well beyond what the annotations or schema already provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, dense but purposeful. The key purpose is front-loaded, and the inline result shape and downstream workflow earn their place. Slightly long due to the embedded JSON snippet, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of explaining the return value, and it does so precisely: per-resource rows with fields and totalAvailable semantics. It also covers the key edge case (datastoreActive:false) and connects to ckan_query. For this tool's complexity, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description repeats the host allowlist, q examples, and limit constraints without adding meaningful new parameter semantics beyond the schema. This matches the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Find CKAN datastore resource ids by keyword') and the underlying mechanism (package_search, keyless). It is immediately distinguishable from siblings like socrata_discover_datasets by naming CKAN and the allowlisted .gov hosts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly explains when this tool fits: it searches CKAN datastore resources and produces resourceIds for downstream use. It names ckan_query as the follow-up tool and warns that datastoreActive:false results are not queryable. It does not explicitly say when NOT to use it versus socrata_discover_datasets, but the workflow guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools