kaggle-mcp
Kaggle MCP(モデルコンテキストプロトコル)サーバー
このリポジトリには、 fastmcpライブラリを使用して構築されたMCP(モデルコンテキストプロトコル)サーバー( server.py )が含まれています。Kaggle APIと連携して、データセットの検索とダウンロードのためのツール、およびEDAノートブック生成のためのプロンプトを提供します。
プロジェクト構造
server.py: FastMCP サーバーアプリケーション。Kaggle とやり取りするためのリソース、ツール、プロンプトを定義します。.env.example: 環境変数(Kaggle API 認証情報)のサンプルファイル。ファイル名を.envに変更し、必要な情報を入力してください。requirements.txt: 必要な Python パッケージをリストします。pyproject.toml&uv.lock:uvパッケージ マネージャーのプロジェクト メタデータとロックされた依存関係。datasets/: ダウンロードした Kaggle データセットが保存されるデフォルトのディレクトリ。
Related MCP server: Kaggle-MCP
設定
リポジトリをクローンします。
git clone <repository-url> cd <repository-directory>仮想環境を作成します (推奨):
python -m venv venv source venv/bin/activate # On Windows use `venv\Scripts\activate` # Or use uv: uv venv依存関係をインストールします: pip を使用します:
pip install -r requirements.txtまたはuvを使用します:
uv syncKaggle API 資格情報を設定します。
方法1(推奨): 環境変数
.envファイルを作成する.envファイルを開き、Kaggle ユーザー名と API キーを追加します。KAGGLE_USERNAME=your_kaggle_username KAGGLE_KEY=your_kaggle_api_keyAPIキーはKaggleアカウントページ(
Account>API>Create New API Token)から取得できます。ユーザー名とキーを含むkaggle.jsonファイルがダウンロードされます。
方法2:
kaggle.jsonファイルKaggle アカウントから
kaggle.jsonファイルをダウンロードします。kaggle.jsonファイルを適切な場所(Linux/macOS の場合は通常~/.kaggle/kaggle.json、Windows の場合はC:\Users\<Your User Name>\.kaggle\kaggle.json)に配置します。環境変数が設定されていない場合、kaggleライブラリはこのファイルを自動的に検出します。
サーバーの実行
仮想環境がアクティブであることを確認します。
MCP サーバーを実行します。
uv run kaggle-mcpサーバーが起動し、リソース、ツール、プロンプトが登録されます。MCPクライアントまたは互換ツールを使用してサーバーと対話できます。
Dockerコンテナの実行
1. Kaggle API 認証情報を設定する
このプロジェクトでは、Kaggle データセットにアクセスするために Kaggle API 認証情報が必要です。
https://www.kaggle.com/settingsにアクセスし、「Create New API Token」をクリックして
kaggle.jsonファイルをダウンロードします。kaggle.jsonファイルを開き、ユーザー名とキーをプロジェクト ルートの新しい.envファイルにコピーします。
KAGGLE_USERNAME=your_username
KAGGLE_KEY=your_key2. Dockerイメージをビルドする
docker build -t kaggle-mcp-test .3. .envファイルを使用してDockerコンテナを実行する
docker run --rm -it --env-file .env kaggle-mcp-testこれにより、Kaggle の認証情報がコンテナ内の環境変数として自動的に読み込まれます。
サーバー機能
サーバーは、モデル コンテキスト プロトコルを通じて次の機能を公開します。
ツール
search_kaggle_datasets(query: str):指定されたクエリ文字列に一致する Kaggle 上のデータセットを検索します。
参照、タイトル、ダウンロード数、最終更新日などの詳細を含む、一致する上位 10 個のデータセットの JSON リストを返します。
download_kaggle_dataset(dataset_ref: str, download_path: str | None = None):特定の Kaggle データセットのファイルをダウンロードして解凍します。
dataset_ref:username/dataset-slugの形式のデータセット識別子 (例:kaggle/titanic)。download_path(オプション): データセットのダウンロード先を指定します。省略した場合は、サーバースクリプトの場所を基準とした相対パス./datasets/<dataset_slug>/がデフォルトとなります。
プロンプト
generate_eda_notebook(dataset_ref: str):指定された Kaggle データセット参照の基本的な探索的データ分析 (EDA) ノートブックを作成するための AI モデル (Gemini など) に適したプロンプト メッセージを生成し、
プロンプトでは、データの読み込み、欠損値のチェック、視覚化、基本的な統計をカバーする Python コードが求められます。
Claudeデスクトップに接続しています
Claude > 設定 > 開発者 > 構成の編集 > claude_desktop_config.json に移動して、以下を追加します。
{
"mcpServers": {
"kaggle-mcp": {
"command": "kaggle-mcp",
"cwd": "<path-to-their-cloned-repo>/kaggle-mcp"
}
}
}使用例
AI エージェントまたは MCP クライアントは、次のようにしてこのサーバーと対話できます。
エージェント: 「Kaggle で「心臓病」に関するデータセットを検索してください」
サーバーは
search_kaggle_datasets(query='heart disease')を実行します。
エージェント: 「データセット 'user/heart-disease-dataset' をダウンロードしてください」
サーバーは
download_kaggle_dataset(dataset_ref='user/heart-disease-dataset')を実行します。
エージェント: 「'user/heart-disease-dataset' の EDA ノートブックプロンプトを生成します」
サーバーは
generate_eda_notebook(dataset_ref='user/heart-disease-dataset')を実行します。サーバーは構造化されたプロンプト メッセージを返します。
エージェント: (コード生成モデルにプロンプトを送信します) -> EDA Python コードを受信します。
Available Tools
2 toolsdownload_kaggle_datasetC
Downloads files for a specific Kaggle dataset. Args: dataset_ref: The reference of the dataset (e.g., 'username/dataset-slug'). download_path: Optional. The path to download the files to. Defaults to '/datasets/'.
| Name | Required | Description | Default |
|---|---|---|---|
| dataset_ref | Yes | ||
| download_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states the action but lacks critical details: whether authentication is required (Kaggle typically needs API credentials), what happens if files already exist at the path, error handling, or any rate limits. The description is minimal beyond the basic operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with a clear purpose statement followed by parameter explanations. It avoids unnecessary fluff, though the formatting with 'Args:' could be more integrated. Every sentence adds value, making it appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of downloading datasets (which often involves authentication, file management, and error cases), no annotations, and no output schema, the description is insufficient. It misses key contextual details like authentication requirements, response format, or handling of large downloads, leaving significant gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful context for both parameters: it explains the format of 'dataset_ref' with an example and clarifies the default behavior and path structure for 'download_path'. With 0% schema description coverage, this compensates somewhat, but it doesn't fully detail constraints (e.g., path validity, dataset accessibility).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Downloads files') and resource ('for a specific Kaggle dataset'), making the purpose immediately understandable. It distinguishes from the sibling tool 'search_kaggle_datasets' by focusing on downloading rather than searching, though it doesn't explicitly contrast them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. While it's implied this is for downloading after a dataset is identified (versus searching with the sibling tool), there's no explicit mention of prerequisites, dependencies, or when-not-to-use scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_kaggle_datasetsC
Searches for datasets on Kaggle matching the query using the Kaggle API.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions using the Kaggle API but doesn't disclose behavioral traits such as authentication requirements, rate limits, pagination, or what the search returns (e.g., format, fields). This leaves significant gaps for an agent to understand how to use it effectively.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero waste. It is appropriately sized and front-loaded, directly stating the tool's purpose without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't cover key aspects like authentication, rate limits, return format, or error handling. For a search tool with no structured support, more context is needed to guide an agent effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implies the 'query' parameter is used for searching datasets, but doesn't add meaning beyond what the schema's title ('Query') and type suggest. No details on query syntax, examples, or constraints are provided, resulting in minimal added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Searches for datasets') and target resource ('on Kaggle'), specifying it uses the Kaggle API. It distinguishes from the sibling tool 'download_kaggle_dataset' by focusing on search rather than download, though it doesn't explicitly mention this distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description doesn't mention the sibling tool 'download_kaggle_dataset' or any other search methods, nor does it specify prerequisites like authentication or rate limits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools have clearly distinct purposes: one downloads a specific dataset, while the other searches for datasets. There is no overlap in functionality, making it easy for an agent to choose the correct tool for each task without confusion.
Both tools follow a consistent verb_noun pattern (download_kaggle_dataset and search_kaggle_datasets), using snake_case and clear action verbs. This consistency makes the tool set predictable and easy to understand at a glance.
With only two tools, the server feels thin for a Kaggle integration, lacking essential operations like listing datasets, uploading data, or managing competitions. While the tools are functional, the scope is incomplete for typical Kaggle workflows, making the count too low for the domain.
The tool set is severely incomplete for a Kaggle MCP server. It covers downloading and searching datasets but misses critical operations such as uploading datasets, accessing competition data, or interacting with notebooks. This creates significant gaps that will hinder agents from performing common Kaggle tasks.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Machine-readable utilities and datasets for AI agents.
List datasets, schemas, run APL queries, and use prompts for exploration, anomalies, and monitoring.
The all-in-one data stack for agents. Upload files, run SQL, evolve tables, and render charts.
Gateway between LLM agents and world data through eight tools and a bundled endpoint catalog.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables interaction with the Hugging Face Dataset Viewer API, allowing users to browse, search, filter, and analyze datasets hosted on the Hugging Face Hub.831MIT
- AlicenseNot gradedqualityDmaintenanceConnects Claude AI to the Kaggle API through the Model Context Protocol, enabling users to browse competitions, search and download datasets, analyze kernels, and access pre-trained models through natural language interactions.MIT
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to interact with Kaggle competitions, including listing competitions, downloading files, submitting predictions, and viewing submission history.10
- AlicenseNot gradedqualityDmaintenanceA full-featured MCP server with 96 tools for the Kaggle API, enabling users to manage competitions, datasets, notebooks, models, discussions, and workflows via natural language.MIT
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/arrismo/kaggle-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server