Skip to main content
Glama
hyeonseo2

dataset-search-mcp

by hyeonseo2

dataset-search-mcp

Unified Model Context Protocol (MCP) server for open-dataset discovery. Search across Hugging Face, Zenodo, and optionally Kaggle, then generate ready-to-run Colab starter code for any result.


Live demo

You can try the search & ranking logic in a simple UI here: Open Dataset Finder (Hugging Face Spaces)


Related MCP server: HF Dataset MCP

Features

  • Multi-source search: Hugging Face / Zenodo / Kaggle (when credentials are available)

  • Sensible ranking (BM25 + fuzzy + light recency weighting)

  • Kaggle API with automatic CLI fallback

  • Safe by default: the server returns metadata only (no server-side downloads)

  • One-click starter snippets for quick experimentation


Repository layout

dataset-search-mcp/
├─ src/
│  └─ dataset_search_mcp/
│     ├─ __init__.py
│     └─ server.py          # MCP server + tools
├─ examples/
│  ├─ claude-desktop.settings.json
│  └─ cursor.settings.json
├─ .github/workflows/
│  ├─ ci.yml
│  └─ release.yml
├─ Dockerfile
├─ pyproject.toml
├─ .dockerignore
├─ .gitignore
├─ LICENSE
└─ README.md

Install (local)

Requires Python 3.9+

pip install -e .
dataset-search-mcp

This starts the MCP server over stdio (awaiting an MCP client).


Docker

Build

docker build -t dataset-search-mcp:local .

Quick smoke test (import only)

docker run --rm --entrypoint python dataset-search-mcp:local -c \
"import importlib; m=importlib.import_module('dataset_search_mcp.server'); print('OK', hasattr(m,'main'))"
# Expected: OK True

Manual run (server waits for a client)

docker run -it --rm dataset-search-mcp:local

Using with Claude Desktop

Add the server to Settings → MCP Servers.

Simplest (Docker):

{
  "mcpServers": {
    "dataset-search-mcp": {
      "command": "docker",
      "args": ["run","-i","--rm","dataset-search-mcp:local"]
    }
  }
}

With Kaggle credentials:

{
  "mcpServers": {
    "dataset-search-mcp": {
      "command": "docker",
      "args": [
        "run","-i","--rm",
        "-e","KAGGLE_USERNAME=your_username",
        "-e","KAGGLE_KEY=your_api_key",
        "dataset-search-mcp:local"
      ]
    }
  }
}

Restart Claude Desktop and open a new chat.


Tools (overview)

search_datasets

Search public datasets across the supported sources.

Args (common):

  • query (string, required)

  • sources (optional): e.g. ["huggingface","zenodo"] Note: the server is tolerant—string forms like "huggingface, zenodo" also work.

  • limit (optional, default 40): per-source cap before ranking

  • format_filter (optional): e.g. "csv" or "json"

Example call (as JSON):

{"query":"korean weather","sources":["huggingface","zenodo"],"limit":10}

Returns: a ranked array of items, each with: source, id, title, description, updated, url, download_url, formats, score.


starter_code

Generate a small Python snippet to quickly try the selected dataset in Colab.

Args (typical):

  • source (e.g., "huggingface", "zenodo", "kaggle")

  • id

  • url (optional)

  • download_url (optional; if present and CSV, the snippet loads it directly)

  • formats (optional; used to choose the best snippet)


Kaggle credentials (brief)

Provide either environment variables:

export KAGGLE_USERNAME=your_username
export KAGGLE_KEY=your_api_key

or a file:

~/.kaggle/kaggle.json
{"username":"your_username","key":"your_api_key"}

Some Kaggle datasets require accepting terms on the website first.


How it works (short)

  • Hugging Face: list_datasets() plus optional dataset_info() for card details

  • Zenodo: REST search via GET /api/records

  • Kaggle: API first; fallback to CLI datasets list --csv

  • Ranking: BM25 + fuzzy matching + light recency factor; duplicates merged on (source,id)


Quick examples (in chat)

  • Search HF + Zenodo:

{"query":"korean weather","sources":["huggingface","zenodo"],"limit":10}
  • CSV-only on Zenodo:

{"query":"traffic accident Korea","sources":["zenodo"],"limit":20,"format_filter":"csv"}
  • Then request starter code using one of the returned items’ fields (source, id, url, download_url, formats).

A
license - permissive license
Not graded
quality - not tested
D
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An unofficial MCP server that provides semantic search capabilities for Hugging Face models and datasets, enabling Claude and other MCP-compatible clients to search, discover, and explore the Hugging Face ecosystem using natural language queries.
    20
    MIT
  • F
    license
    A
    quality
    D
    maintenance
    An MCP server for the Hugging Face Dataset Viewer API that enables searching, fetching, and filtering datasets on the Hugging Face Hub. It allows users to explore schemas, perform full-text searches, and analyze dataset statistics through natural language.
    10
  • A
    license
    Not graded
    quality
    B
    maintenance
    Search, preview, and analyze datasets from 20+ platforms and millions of datasets via a single MCP connector. Works with Claude instantly
    2
    27
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Connects open data to LLMs via MCP, enabling easy access to public datasets and publishing new datasets with community help.
    MIT

View all related MCP servers

Related MCP Connectors

  • All HasData scraping tools in one MCP server: Google, TikTok, Instagram, maps, e-commerce and more.

  • MCP Hub: AI service discovery, per-user OAuth, and multi-service workflow orchestration

  • Reddit & X data for AI agents over MCP. Semantic search, hosted, no Reddit API.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/hyeonseo2/dataset-search-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server