Skip to main content
Glama
cfahlgren1

HF Dataset MCP

by cfahlgren1

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
HF_TOKENNoHugging Face API token (required for private/gated datasets)
HF_DATASETS_SERVERNoCustom Dataset Viewer API URLhttps://datasets-server.huggingface.co

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}

Tools

Functions exposed to the LLM to take actions

NameDescription
search_datasetsA

Find datasets on the Hugging Face Hub by name, tag, or author

validate_datasetC

Check if a dataset is accessible and which viewer features are available

list_splitsC

Get all available configurations and splits for a dataset

get_dataset_infoB

Get the schema, metadata, and row counts for a dataset configuration

get_rowsC

Fetch a slice of rows from a dataset split

search_datasetB

Full-text search within a dataset split using BM25 ranking

filter_rowsC

Filter dataset rows using SQL-like WHERE conditions

get_dataset_sizeB

Get row counts and byte sizes for all configs and splits

list_parquet_filesA

Get URLs for the dataset's Parquet files for direct download or processing

get_statisticsB

Get descriptive statistics for each column in a dataset split

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.7/5.0

Scored across 10 tools

Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: filtering rows, getting metadata, fetching rows, searching datasets, etc. The descriptions clearly differentiate operations like get_rows vs. filter_rows vs. search_dataset, preventing misselection.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern with snake_case (e.g., filter_rows, get_dataset_info, list_splits). The naming is predictable and readable throughout the set, with no deviations in style.

Tool Count5/5

With 10 tools, this is well-scoped for a dataset management server. Each tool serves a specific function in exploring, querying, and validating datasets, with no redundant or missing tools that would make the set feel too thin or bloated.

Completeness5/5

The toolset provides complete coverage for dataset operations: discovery (search_datasets, list_splits), inspection (get_dataset_info, get_statistics), access (get_rows, list_parquet_files), querying (filter_rows, search_dataset), and validation (validate_dataset). No obvious gaps exist for the domain.

Maintenance

ActivityInactive
ResponsivenessNo issues