opendata-ua-mcp
Enables Google Gemini to search and analyze Ukraine's open data portal using natural language queries.
Enables local AI models via Ollama to search and analyze Ukraine's open data portal using natural language queries.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@opendata-ua-mcpFind ecology datasets from Lviv 2024"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
opendata-mcp — data.gov.ua MCP server
An open-source MCP server that lets any MCP-compatible AI agent search and analyze Ukraine's national open-data portal data.gov.ua in natural language.
It speaks the Model Context Protocol, so it works with any client that supports MCP — including Claude (Desktop / Code), ChatGPT (Developer Mode / custom connectors), Google Gemini, Cursor, Cline / Continue, VS Code Copilot, local models via Ollama / LM Studio, and your own agents built on the Python/TypeScript MCP SDKs. Not tied to any single vendor.
Phase 1: read-only, public endpoints — no API token, no auth required.
Design
Tools are jobs-to-be-done, not thin API wrappers. Each tool does a complete user task, composes several CKAN calls internally, hides portal quirks (UUID slugs, dirty formats, sparse DataStore, Unicode homoglyphs), and returns a token-efficient slim result (a search hit is ~150 B vs ~17 KB raw).
Tool | What it does |
| Find datasets by topic; ranked slim candidates + refine hints |
| Aggregate view (who publishes / how much) — counts only |
| Full dataset card: license, freshness, resources |
| Get the actual data — auto-picks best resource; DataStore or download+parse CSV/JSON/XLSX/XML, including files inside ZIP archives (ЄДР, debtors registries) |
| Filter rows / read-only SQL over a structured (DataStore) resource |
| Recently updated datasets, filterable by topic/org |
Related MCP server: CKAN MCP Server
Install
The server runs over stdio, the transport every MCP client understands. The config below is the same everywhere — only the file/menu where you paste it differs per client.
Any MCP client (npm) — universal
{
"mcpServers": {
"opendata-ua": {
"command": "npx",
"args": ["-y", "@opendata-ua/mcp-server"]
}
}
}Where to put it:
Client | Location |
Claude Desktop | Settings → Developer → Edit Config ( |
Claude Code |
|
ChatGPT | Settings → Connectors → add MCP server (Developer Mode) |
Google Gemini | Gemini CLI / SDK |
Cursor | Settings → MCP → Add Server ( |
Cline / Continue / VS Code | the extension's MCP settings |
Custom agent | point your MCP SDK at |
Claude Desktop — one-click DXT
DXT is Claude Desktop's drag-and-drop bundle format:
Download
opendata-ua-mcp.dxtfrom Releases.Drag it into Claude Desktop → Settings → Extensions. Done. No config.
From source
npm install
npm run build
node dist/stdio.js # any MCP client can spawn thisTry it
"Find ecology datasets from Lviv published in 2024" "Which organizations publish the most procurement data?" "Show me the first rows of the stolen-vehicles register"
Configuration (optional)
Env var | Default | Purpose |
|
| API base (mirror/dev portal) |
|
| LRU cache TTL for catalogs |
|
| Request timeout |
|
| Per-response context-budget ceiling |
|
| Cap for files downloaded + parsed locally |
|
|
|
Develop
npm test # vitest
npm run typecheck
npm run lint
npm run smoke # live test against data.gov.ua
npm run build:dxt # build the .dxt packageThe portal runs CKAN 2.7.2. DataStore (queryable rows) covers only ~0.3 % of resources, so get_dataset_data downloads + parses files locally when needed; filter_data/SQL apply to the DataStore-active minority.
License
MIT © Open Data UA Community
Available Tools
6 toolsexplore_catalogA
Огляд каталогу data.gov.ua агрегатами (без видачі самих датасетів): скільки всього, хто публікує найбільше, розподіл за категоріями/форматами. Дешево за токенами. Використовуй для «хто публікує найбільше даних про X», «скільки датасетів про Y».
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Звузити до теми перед агрегуванням | |
| group_by | No | За чим рахувати: розпорядник / категорія / формат | organization |
| category | No | Slug категорії-фільтра | |
| organization | No | Slug розпорядника-фільтра | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions 'cheap on tokens' which is helpful, but does not explicitly state that the operation is read-only or idempotent. The description covers the behavioral scope adequately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences plus a usage hint, with no redundant information. Every sentence adds value, and the structure is front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description explains the return type (aggregates, not datasets). It covers purpose, usage, and parameters. Could detail the response format (e.g., fields returned) but is sufficient for selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80% with descriptions for 4 of 5 parameters. The description adds context (e.g., 'query' narrows topic, group_by options), but does not significantly enhance understanding beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it provides aggregate catalog overview (count, publishers, categories/formats) and explicitly distinguishes itself from siblings by noting it does not return datasets themselves. The Ukrainian text gives concrete usage examples.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit use cases (e.g., 'who publishes most data about X') and implies when not to use (when you need actual datasets). It does not explicitly name sibling alternatives but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filter_dataA
Фільтрувати/агрегувати рядки структурованого ресурсу: точні фільтри (колонка=значення) або read-only SQL (SELECT). Працює ЛИШЕ для DataStore-активних ресурсів (≈0.3% на порталі — перевір has_datastore через inspect_dataset). Для решти спершу візьми дані через get_dataset_data.
| Name | Required | Description | Default |
|---|---|---|---|
| resource_id | Yes | ID ресурсу (має бути DataStore-активним) | |
| filters | No | Точні фільтри колонка=значення, напр. {"region":"Київ","year":2024} | |
| q | No | Повнотекстовий пошук по рядках | |
| sql | No | Read-only SQL замість filters (тільки SELECT). Має пріоритет. | |
| fields | No | Колонки на вихід | |
| limit | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions read-only SQL, indicating safe operation. Though no annotations are provided, it implies non-destructive behavior. Missing explicit statement about side effects or idempotency, but sufficient given the context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with core purpose, then crucial usage restriction. No unnecessary words. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, output schema, and 7 parameters, the description covers all needed context: what it does, constraints, and integration with sibling tools. Complete and actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond the schema by clarifying that filters are exact column=value mappings, and that SQL overrides filters. Provides an example. Schema coverage is high, but the description enhances understanding without redundancy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to filter/aggregate rows of a structured resource using exact filters or read-only SQL. It specifies the resource type (DataStore-active) and differentiates from siblings by noting the prerequisite check via inspect_dataset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use (DataStore-active resources) and when not to (other resources, use get_dataset_data instead). Provides actionable guidance to verify has_datastore via inspect_dataset, and names specific alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_datasetsA
Знайти датасети на data.gov.ua за темою. Повертає компактний ранжований список кандидатів (назва, розпорядник, формати, свіжість, посилання) + підказки для звуження пошуку (топ розпорядники/категорії/формати). Використовуй це для будь-якого «знайди дані про …».
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Тема або ключові слова (напр. «закупівлі громади 2024») | |
| category | No | Slug тематичної категорії (напр. 'ekolohiia') | |
| organization | No | Slug розпорядника | |
| format | No | Формат файлу: CSV, JSON, XLSX, … | |
| tags | No | ||
| limit | No | ||
| offset | No | ||
| sort | No | relevance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool returns a compact ranked list with title, manager, formats, freshness, link, and search-narrowing hints, which is transparent for a search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with one main sentence and additional detail, front-loading the primary purpose. It is well-structured but could be slightly more structured for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter search tool with no output schema or annotations, the description explains the return format and hints. It lacks details on pagination or sorting but is otherwise complete for typical search usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (87.5%), so the schema already documents most parameters. The description adds context about search narrowing hints but does not significantly enhance parameter understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool finds datasets on data.gov.ua by topic, specifies the return format (ranked list with details and hints), and distinguishes it from siblings by instructing to use it for any 'find data about' query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this for any 'find data about' query, providing clear context. It does not explicitly mention when not to use it, but the purpose is well-defined enough to infer appropriate use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dataset_dataA
Отримати самі дані (перші рядки + схема колонок) з датасету чи ресурсу. Автоматично обирає найкращий машиночитний ресурс. Якщо є DataStore — читає звідти; інакше завантажує файл і парсить локально (CSV/JSON/XLSX). Повертає прев'ю (обмежене), оцінку кількості рядків і посилання на повний файл. Це твій основний інструмент для «покажи дані».
| Name | Required | Description | Default |
|---|---|---|---|
| dataset | No | ID/slug/назва датасету (автовибір найкращого ресурсу) | |
| resource_id | No | ID конкретного ресурсу (має пріоритет над dataset) | |
| columns | No | Обмежити колонки | |
| limit | No | Скільки рядків-прев'ю |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses that it reads from DataStore or parses locally (CSV/JSON/XLSX), returns a preview with row estimate and full file link. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the verb, no wasted words. Each sentence provides essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 4 parameters, no output schema, and no annotations, the description covers core functionality, automatic behavior, output components, and usage context. Could mention error handling or permissions, but adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions. Description adds context: automatic resource selection for 'dataset', priority for 'resource_id', and default for 'limit'. Adds value beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves data (first rows + column schema) from a dataset or resource. It distinguishes itself from siblings like explore_catalog or filter_data by focusing on data preview and automatic best-resource selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'This is your main tool for "show data"' and explains automatic resource selection and fallback behavior. It could compare to siblings more, but the guidance is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_datasetA
Детальна картка одного датасету перед використанням: опис, розпорядник, ліцензія (+URL), свіжість, частота оновлення, і список ресурсів з форматом, розміром та ознакою machine_readable. Приймає ID/slug або назву (автопошук). Далі бери дані через get_dataset_data.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset | Yes | ID/slug датасету АБО його назва (якщо назва — буде автопошук найкращого збігу) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description discloses input flexibility (ID/slug or name with auto-search) and output fields. However, it lacks details on read-only nature, side effects, or error cases. Adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is 3-4 sentences in Ukrainian, front-loading the purpose and then listing included fields. No fluff, but slightly longer than necessary due to detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers input, output contents, and suggests next steps (use get_dataset_data). It is complete for a metadata inspection tool, though error handling is omitted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a detailed description for the single parameter. The description adds value by explaining that a name triggers auto-search for best match, going beyond the schema's basic string definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool provides a detailed card of a dataset including description, manager, license, freshness, and resources. It distinguishes from siblings like find_datasets and get_dataset_data by focusing on inspection before data retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description specifies this tool is for inspection before using get_dataset_data, implying a workflow. It does not explicitly mention when not to use or list alternatives, but the context is clear enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
track_updatesA
Стрічка нещодавно оновлених датасетів на data.gov.ua (моніторинг). Можна звузити за темою або розпорядником. Повертає компактний список: назва, розпорядник, коли змінено, тип зміни, посилання.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | Фільтр за словом у назві датасету | |
| organization | No | Фільтр за назвою/slug розпорядника | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the output format (compact list with name, administrator, change time, change type, link). However, it does not mention any behavioral traits like rate limits, authentication needs, or side effects, which is acceptable for a read-like monitoring tool but still a gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise with two sentences. The first sentence states the purpose, and the second describes the output. No extraneous information, front-loaded effectively.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters and no output schema, the description covers purpose, filter options, and return fields. It omits details like ordering or default limit behavior, but is largely complete for a simple monitoring tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (topic and organization have descriptions; limit has default/min/max but no description). The description adds context about the output but does not elaborate on parameter semantics beyond what the schema provides. It partially compensates for the missing limit description by implying its role in the list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it tracks recent dataset updates on data.gov.ua (monitoring). It uses specific verb-resource (track updates) and distinguishes from siblings like 'explore_catalog' and 'find_datasets' by focusing on change monitoring.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions narrowing by topic or organization, implying when to filter. However, it does not provide explicit when-to-use or when-not-to-use guidance, nor does it name alternatives among the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
explore_catalog - First observed
filter_data - First observed
find_datasets - First observed
get_dataset_data - First observed
inspect_dataset - First observed
track_updates
TDQS
Scored across 6 tools
Each tool has a clearly distinct purpose: catalog overview, data filtering, dataset search, data retrieval, dataset inspection, and update tracking. No overlapping functionality.
All tool names follow a consistent snake_case verb_noun pattern (e.g., explore_catalog, find_datasets), making them predictable and easy to distinguish.
With 6 tools, the server is well-scoped for interacting with an open data portal. Each tool covers a essential operation without redundancy or excessive complexity.
The tool set covers major read operations: search, inspect, retrieve, filter, and monitor updates. Missing a direct 'list all datasets' function, but find_datasets with a broad query can approximate it.
Maintenance
Related MCP Connectors
Public Data Ukraine Mcp connects AI agents to real public APIs via MCP. Tools include
Search a Ukrainian catalog of 21,000+ AI tools — search tools, get details, list categories.
Verified Polish open data for AI agents: debt, budget, 460 MPs, votings, judiciary search, RAG.
Related MCP Servers
- AlicenseCqualityBmaintenanceEnables AI assistants to search, discover, and analyze thousands of datasets from Israel's national open data portal. It provides tools for querying government ministries, municipalities, and public bodies using the CKAN API.939 npm109MIT
- AlicenseAqualityAmaintenanceEnables AI assistants to search, explore, and query any CKAN open data portal through natural language, making public datasets accessible without requiring knowledge of the portal's API.201,372 npm57MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to search and retrieve metadata and data files from Peru's National Open Data Platform, and generate Jupyter notebooks for data analysis.3Apache 2.0

mcp-data-lvofficial
AlicenseNot gradedqualityCmaintenanceEnables AI agents to search, retrieve metadata, and query tabular resources from Latvia's Open Data portal (data.gov.lv) via CKAN.5 npmMIT