Skip to main content
Glama

Find Malaysian Public Data

search_datasets
Read-onlyIdempotent

Use for discovery only: find DataPulse's 418 Malaysian public datasets by topic, source, or licence—for example, 'Malaysian public data inflation', licence and attribution, or a government dataset source. Returns ranked matches with id, title, source, licence, published status, and score. This is not trust verification: a status is published pipeline context, not proof that a dataset is current or reliable. For pre-trust use search_datasets → verify_dataset → get_provenance. Use it to discover candidates; do not use it for a trust decision—use verify_dataset instead. It reads published catalogue data, so no match means the published catalogue has no matching entry; DataPulse is read-only, requires no API key, and the edge limits clients to roughly one request per second with a small burst, so pace or retry.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum highest-ranked discovery matches to return, e.g. 10; omitted defaults to 10 and does not change ranking.
queryYesTopic or task phrasing used to rank discovery candidates, e.g. 'Malaysian public data inflation'; use a returned id with get_dataset.
sourceNoOptional publisher/source substring filter applied before ranking, e.g. 'OpenDOSM'; omit it to search every source.
licenceNoOptional exact licence name or supported alias applied before ranking, e.g. 'CC BY 4.0'; omit it to search every licence.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changed
    • changedInput schema / properties / licence / description
      Previous value: -"Optional exact licence name or supported alias for reuse discovery, e.g. 'CC BY 4.0'; this does not verify attribution compliance."New value: +"Optional exact licence name or supported alias applied before ranking, e.g. 'CC BY 4.0'; omit it to search every licence."
    • changedInput schema / properties / limit / description
      Previous value: -"Maximum discovery matches to return; integer from 1 to 50, e.g. 10."New value: +"Maximum highest-ranked discovery matches to return, e.g. 10; omitted defaults to 10 and does not change ranking."
    • changedInput schema / properties / query / description
      Previous value: -"Topic or task phrasing for Malaysian public-data discovery only, e.g. 'Malaysian public data inflation'; verify a result separately."New value: +"Topic or task phrasing used to rank discovery candidates, e.g. 'Malaysian public data inflation'; use a returned id with get_dataset."
    • changedInput schema / properties / source / description
      Previous value: -"Optional case-insensitive publisher/source filter, e.g. 'OpenDOSM'."New value: +"Optional publisher/source substring filter applied before ranking, e.g. 'OpenDOSM'; omit it to search every source."
  2. Changed5 schema fields changed
    • changedInput schema / properties / licence / description
      Previous value: -"Optional exact licence name or supported alias, e.g. 'CC BY 4.0'."New value: +"Optional exact licence name or supported alias for reuse discovery, e.g. 'CC BY 4.0'; this does not verify attribution compliance."
    • changedInput schema / properties / limit / description
      Previous value: -"Maximum ranked matches to return; integer from 1 to 50, e.g. 10."New value: +"Maximum discovery matches to return; integer from 1 to 50, e.g. 10."
    • changedInput schema / properties / query / description
      Previous value: -"Free-text search terms; natural language is allowed, e.g. 'inflation cpi'."New value: +"Topic or task phrasing for Malaysian public-data discovery only, e.g. 'Malaysian public data inflation'; verify a result separately."
    • changedInput schema / properties / query / examples
      Previous value: -[
      -  "inflation cpi"
      -]New value: +[
      +  "Malaysian public data inflation"
      +]
    • changedInput schema / properties / source / description
      Previous value: -"Optional case-insensitive source-name substring, e.g. 'OpenDOSM'."New value: +"Optional case-insensitive publisher/source filter, e.g. 'OpenDOSM'."
  3. Changed3 schema fields changed
    • addedInput schema / properties / licence / examples
      Added value: +[
      +  "CC BY 4.0",
      +  "Open Government Licence (Malaysia)"
      +]
    • addedInput schema / properties / query / examples
      Added value: +[
      +  "inflation cpi"
      +]
    • addedInput schema / properties / source / examples
      Added value: +[
      +  "OpenDOSM",
      +  "data.gov.my",
      +  "MET Malaysia"
      +]
  4. Changed4 schema fields changed
    • addedInput schema / properties / licence / description
      Added value: +"Optional exact licence name or supported alias, e.g. 'CC BY 4.0'."
    • addedInput schema / properties / limit / description
      Added value: +"Maximum ranked matches to return; integer from 1 to 50, e.g. 10."
    • addedInput schema / properties / query / description
      Added value: +"Free-text search terms; natural language is allowed, e.g. 'inflation cpi'."
    • addedInput schema / properties / source / description
      Added value: +"Optional case-insensitive source-name substring, e.g. 'OpenDOSM'."
  5. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond that: it is discovery-only, not trust verification, and it discloses the rate limit ('roughly one request per second with a small burst, so pace or retry'). It also explains that 'published status' is pipeline context, not a reliability proof. Minor gap: it doesn't describe pagination or what happens when limit is exceeded, but the rate-limit disclosure is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it front-loads the core purpose, then the trust caveat, then the pipeline guidance, then the rate-limit note. Every sentence earns its place, though the final sentence is long and packs several distinct facts (read-only, no API key, rate limit) into one clause, which slightly reduces scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a discovery tool with a rich output schema, full parameter documentation, and safety annotations, the description covers the essential context: what it returns, how to use it in the larger workflow, what it does not do, and operational constraints. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters with examples and defaults. The description adds a little context by explaining that source and licence filters are applied before ranking, and that query is topic/task phrasing, but it mostly restates what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('find'), a resource ('DataPulse's 418 Malaysian public datasets'), and the dimensions by which to search (topic, source, licence). It also explicitly contrasts with trust verification and names the sibling verify_dataset, so an agent can distinguish it from the many sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance ('Use it to discover candidates; do not use it for a trust decision—use verify_dataset instead') and even names the intended pre-trust pipeline: search_datasets → verify_dataset → get_provenance. It also clarifies the meaning of no matches, which is a usage-relevant edge case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.