Skip to main content
Glama
VladyslavMykhailyshyn

opendata-ua-mcp

opendata-mcp — data.gov.ua MCP server

Українською

An open-source MCP server that lets any MCP-compatible AI agent search and analyze Ukraine's national open-data portal data.gov.ua in natural language.

It speaks the Model Context Protocol, so it works with any client that supports MCP — including Claude (Desktop / Code), ChatGPT (Developer Mode / custom connectors), Google Gemini, Cursor, Cline / Continue, VS Code Copilot, local models via Ollama / LM Studio, and your own agents built on the Python/TypeScript MCP SDKs. Not tied to any single vendor.

Phase 1: read-only, public endpoints — no API token, no auth required.

Design

Tools are jobs-to-be-done, not thin API wrappers. Each tool does a complete user task, composes several CKAN calls internally, hides portal quirks (UUID slugs, dirty formats, sparse DataStore, Unicode homoglyphs), and returns a token-efficient slim result (a search hit is ~150 B vs ~17 KB raw).

Tool

What it does

find_datasets

Find datasets by topic; ranked slim candidates + refine hints

explore_catalog

Aggregate view (who publishes / how much) — counts only

inspect_dataset

Full dataset card: license, freshness, resources

get_dataset_data

Get the actual data — auto-picks best resource; DataStore or download+parse CSV/JSON/XLSX/XML, including files inside ZIP archives (ЄДР, debtors registries)

filter_data

Filter rows / read-only SQL over a structured (DataStore) resource

track_updates

Recently updated datasets, filterable by topic/org

Related MCP server: CKAN MCP Server

Install

The server runs over stdio, the transport every MCP client understands. The config below is the same everywhere — only the file/menu where you paste it differs per client.

Any MCP client (npm) — universal

{
  "mcpServers": {
    "opendata-ua": {
      "command": "npx",
      "args": ["-y", "@opendata-ua/mcp-server"]
    }
  }
}

Where to put it:

Client

Location

Claude Desktop

Settings → Developer → Edit Config (claude_desktop_config.json)

Claude Code

claude mcp add opendata-ua -- npx -y @opendata-ua/mcp-server

ChatGPT

Settings → Connectors → add MCP server (Developer Mode)

Google Gemini

Gemini CLI / SDK mcpServers config

Cursor

Settings → MCP → Add Server (~/.cursor/mcp.json)

Cline / Continue / VS Code

the extension's MCP settings

Custom agent

point your MCP SDK at npx -y @opendata-ua/mcp-server

Claude Desktop — one-click DXT

DXT is Claude Desktop's drag-and-drop bundle format:

  1. Download opendata-ua-mcp.dxt from Releases.

  2. Drag it into Claude Desktop → Settings → Extensions. Done. No config.

From source

npm install
npm run build
node dist/stdio.js   # any MCP client can spawn this

Try it

"Find ecology datasets from Lviv published in 2024" "Which organizations publish the most procurement data?" "Show me the first rows of the stolen-vehicles register"

Configuration (optional)

Env var

Default

Purpose

DATA_GOV_UA_BASE_URL

https://data.gov.ua/api/3/action

API base (mirror/dev portal)

CACHE_TTL_SECONDS

300

LRU cache TTL for catalogs

HTTP_TIMEOUT_MS

30000

Request timeout

MAX_RESPONSE_CHARS

60000

Per-response context-budget ceiling

MAX_DOWNLOAD_BYTES

10000000

Cap for files downloaded + parsed locally

LOG_LEVEL

info

debug/info/warn/error

Develop

npm test          # vitest
npm run typecheck
npm run lint
npm run smoke      # live test against data.gov.ua
npm run build:dxt  # build the .dxt package

The portal runs CKAN 2.7.2. DataStore (queryable rows) covers only ~0.3 % of resources, so get_dataset_data downloads + parses files locally when needed; filter_data/SQL apply to the DataStore-active minority.

License

MIT © Open Data UA Community

Available Tools

6 tools
explore_catalogA

Огляд каталогу data.gov.ua агрегатами (без видачі самих датасетів): скільки всього, хто публікує найбільше, розподіл за категоріями/форматами. Дешево за токенами. Використовуй для «хто публікує найбільше даних про X», «скільки датасетів про Y».

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoЗвузити до теми перед агрегуванням
group_byNoЗа чим рахувати: розпорядник / категорія / форматorganization
categoryNoSlug категорії-фільтра
organizationNoSlug розпорядника-фільтра
limitNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions 'cheap on tokens' which is helpful, but does not explicitly state that the operation is read-only or idempotent. The description covers the behavioral scope adequately.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus a usage hint, with no redundant information. Every sentence adds value, and the structure is front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description explains the return type (aggregates, not datasets). It covers purpose, usage, and parameters. Could detail the response format (e.g., fields returned) but is sufficient for selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 80% with descriptions for 4 of 5 parameters. The description adds context (e.g., 'query' narrows topic, group_by options), but does not significantly enhance understanding beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it provides aggregate catalog overview (count, publishers, categories/formats) and explicitly distinguishes itself from siblings by noting it does not return datasets themselves. The Ukrainian text gives concrete usage examples.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit use cases (e.g., 'who publishes most data about X') and implies when not to use (when you need actual datasets). It does not explicitly name sibling alternatives but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

filter_dataA

Фільтрувати/агрегувати рядки структурованого ресурсу: точні фільтри (колонка=значення) або read-only SQL (SELECT). Працює ЛИШЕ для DataStore-активних ресурсів (≈0.3% на порталі — перевір has_datastore через inspect_dataset). Для решти спершу візьми дані через get_dataset_data.

ParametersJSON Schema
NameRequiredDescriptionDefault
resource_idYesID ресурсу (має бути DataStore-активним)
filtersNoТочні фільтри колонка=значення, напр. {"region":"Київ","year":2024}
qNoПовнотекстовий пошук по рядках
sqlNoRead-only SQL замість filters (тільки SELECT). Має пріоритет.
fieldsNoКолонки на вихід
limitNo
offsetNo

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description mentions read-only SQL, indicating safe operation. Though no annotations are provided, it implies non-destructive behavior. Missing explicit statement about side effects or idempotency, but sufficient given the context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with core purpose, then crucial usage restriction. No unnecessary words. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, output schema, and 7 parameters, the description covers all needed context: what it does, constraints, and integration with sibling tools. Complete and actionable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Adds meaning beyond the schema by clarifying that filters are exact column=value mappings, and that SQL overrides filters. Provides an example. Schema coverage is high, but the description enhances understanding without redundancy.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: to filter/aggregate rows of a structured resource using exact filters or read-only SQL. It specifies the resource type (DataStore-active) and differentiates from siblings by noting the prerequisite check via inspect_dataset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use (DataStore-active resources) and when not to (other resources, use get_dataset_data instead). Provides actionable guidance to verify has_datastore via inspect_dataset, and names specific alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_datasetsA

Знайти датасети на data.gov.ua за темою. Повертає компактний ранжований список кандидатів (назва, розпорядник, формати, свіжість, посилання) + підказки для звуження пошуку (топ розпорядники/категорії/формати). Використовуй це для будь-якого «знайди дані про …».

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNoТема або ключові слова (напр. «закупівлі громади 2024»)
categoryNoSlug тематичної категорії (напр. 'ekolohiia')
organizationNoSlug розпорядника
formatNoФормат файлу: CSV, JSON, XLSX, …
tagsNo
limitNo
offsetNo
sortNorelevance

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool returns a compact ranked list with title, manager, formats, freshness, link, and search-narrowing hints, which is transparent for a search tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with one main sentence and additional detail, front-loading the primary purpose. It is well-structured but could be slightly more structured for clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter search tool with no output schema or annotations, the description explains the return format and hints. It lacks details on pagination or sorting but is otherwise complete for typical search usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (87.5%), so the schema already documents most parameters. The description adds context about search narrowing hints but does not significantly enhance parameter understanding beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool finds datasets on data.gov.ua by topic, specifies the return format (ranked list with details and hints), and distinguishes it from siblings by instructing to use it for any 'find data about' query.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use this for any 'find data about' query, providing clear context. It does not explicitly mention when not to use it, but the purpose is well-defined enough to infer appropriate use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_dataset_dataA

Отримати самі дані (перші рядки + схема колонок) з датасету чи ресурсу. Автоматично обирає найкращий машиночитний ресурс. Якщо є DataStore — читає звідти; інакше завантажує файл і парсить локально (CSV/JSON/XLSX). Повертає прев'ю (обмежене), оцінку кількості рядків і посилання на повний файл. Це твій основний інструмент для «покажи дані».

ParametersJSON Schema
NameRequiredDescriptionDefault
datasetNoID/slug/назва датасету (автовибір найкращого ресурсу)
resource_idNoID конкретного ресурсу (має пріоритет над dataset)
columnsNoОбмежити колонки
limitNoСкільки рядків-прев'ю

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses that it reads from DataStore or parses locally (CSV/JSON/XLSX), returns a preview with row estimate and full file link. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the verb, no wasted words. Each sentence provides essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 4 parameters, no output schema, and no annotations, the description covers core functionality, automatic behavior, output components, and usage context. Could mention error handling or permissions, but adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions. Description adds context: automatic resource selection for 'dataset', priority for 'resource_id', and default for 'limit'. Adds value beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it retrieves data (first rows + column schema) from a dataset or resource. It distinguishes itself from siblings like explore_catalog or filter_data by focusing on data preview and automatic best-resource selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'This is your main tool for "show data"' and explains automatic resource selection and fallback behavior. It could compare to siblings more, but the guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

inspect_datasetA

Детальна картка одного датасету перед використанням: опис, розпорядник, ліцензія (+URL), свіжість, частота оновлення, і список ресурсів з форматом, розміром та ознакою machine_readable. Приймає ID/slug або назву (автопошук). Далі бери дані через get_dataset_data.

ParametersJSON Schema
NameRequiredDescriptionDefault
datasetYesID/slug датасету АБО його назва (якщо назва — буде автопошук найкращого збігу)

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided; description discloses input flexibility (ID/slug or name with auto-search) and output fields. However, it lacks details on read-only nature, side effects, or error cases. Adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is 3-4 sentences in Ukrainian, front-loading the purpose and then listing included fields. No fluff, but slightly longer than necessary due to detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers input, output contents, and suggests next steps (use get_dataset_data). It is complete for a metadata inspection tool, though error handling is omitted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a detailed description for the single parameter. The description adds value by explaining that a name triggers auto-search for best match, going beyond the schema's basic string definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool provides a detailed card of a dataset including description, manager, license, freshness, and resources. It distinguishes from siblings like find_datasets and get_dataset_data by focusing on inspection before data retrieval.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Description specifies this tool is for inspection before using get_dataset_data, implying a workflow. It does not explicitly mention when not to use or list alternatives, but the context is clear enough for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

track_updatesA

Стрічка нещодавно оновлених датасетів на data.gov.ua (моніторинг). Можна звузити за темою або розпорядником. Повертає компактний список: назва, розпорядник, коли змінено, тип зміни, посилання.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoФільтр за словом у назві датасету
organizationNoФільтр за назвою/slug розпорядника
limitNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the output format (compact list with name, administrator, change time, change type, link). However, it does not mention any behavioral traits like rate limits, authentication needs, or side effects, which is acceptable for a read-like monitoring tool but still a gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with two sentences. The first sentence states the purpose, and the second describes the output. No extraneous information, front-loaded effectively.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given three parameters and no output schema, the description covers purpose, filter options, and return fields. It omits details like ordering or default limit behavior, but is largely complete for a simple monitoring tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67% (topic and organization have descriptions; limit has default/min/max but no description). The description adds context about the output but does not elaborate on parameter semantics beyond what the schema provides. It partially compensates for the missing limit description by implying its role in the list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it tracks recent dataset updates on data.gov.ua (monitoring). It uses specific verb-resource (track updates) and distinguishes from siblings like 'explore_catalog' and 'find_datasets' by focusing on change monitoring.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description mentions narrowing by topic or organization, implying when to filter. However, it does not provide explicit when-to-use or when-not-to-use guidance, nor does it name alternatives among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedexplore_catalog
    • First observedfilter_data
    • First observedfind_datasets
    • First observedget_dataset_data
    • First observedinspect_dataset
    • First observedtrack_updates

TDQS

A4.2/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: catalog overview, data filtering, dataset search, data retrieval, dataset inspection, and update tracking. No overlapping functionality.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern (e.g., explore_catalog, find_datasets), making them predictable and easy to distinguish.

Tool Count5/5

With 6 tools, the server is well-scoped for interacting with an open data portal. Each tool covers a essential operation without redundancy or excessive complexity.

Completeness4/5

The tool set covers major read operations: search, inspect, retrieve, filter, and monitor updates. Missing a direct 'list all datasets' function, but find_datasets with a broad query can approximate it.

Maintenance

ActivityStale
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    B
    maintenance
    Enables AI assistants to search, discover, and analyze thousands of datasets from Israel's national open data portal. It provides tools for querying government ministries, municipalities, and public bodies using the CKAN API.
    9
    39 npm
    109
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI assistants to search, explore, and query any CKAN open data portal through natural language, making public datasets accessible without requiring knowledge of the portal's API.
    20
    1,372 npm
    57
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to search and retrieve metadata and data files from Peru's National Open Data Platform, and generate Jupyter notebooks for data analysis.
    3
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to search, retrieve metadata, and query tabular resources from Latvia's Open Data portal (data.gov.lv) via CKAN.
    5 npm
    MIT