HF Dataset MCP
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| HF_TOKEN | No | Hugging Face API token (required for private/gated datasets) | |
| HF_DATASETS_SERVER | No | Custom Dataset Viewer API URL | https://datasets-server.huggingface.co |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| search_datasetsA | Find datasets on the Hugging Face Hub by name, tag, or author |
| validate_datasetC | Check if a dataset is accessible and which viewer features are available |
| list_splitsC | Get all available configurations and splits for a dataset |
| get_dataset_infoB | Get the schema, metadata, and row counts for a dataset configuration |
| get_rowsC | Fetch a slice of rows from a dataset split |
| search_datasetB | Full-text search within a dataset split using BM25 ranking |
| filter_rowsC | Filter dataset rows using SQL-like WHERE conditions |
| get_dataset_sizeB | Get row counts and byte sizes for all configs and splits |
| list_parquet_filesA | Get URLs for the dataset's Parquet files for direct download or processing |
| get_statisticsB | Get descriptive statistics for each column in a dataset split |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 10 tools
Each tool has a clearly distinct purpose with no overlap: filtering rows, getting metadata, fetching rows, searching datasets, etc. The descriptions clearly differentiate operations like get_rows vs. filter_rows vs. search_dataset, preventing misselection.
All tools follow a consistent verb_noun pattern with snake_case (e.g., filter_rows, get_dataset_info, list_splits). The naming is predictable and readable throughout the set, with no deviations in style.
With 10 tools, this is well-scoped for a dataset management server. Each tool serves a specific function in exploring, querying, and validating datasets, with no redundant or missing tools that would make the set feel too thin or bloated.
The toolset provides complete coverage for dataset operations: discovery (search_datasets, list_splits), inspection (get_dataset_info, get_statistics), access (get_rows, list_parquet_files), querying (filter_rows, search_dataset), and validation (validate_dataset). No obvious gaps exist for the domain.