Skip to main content
Glama

значок кузнеца

Сервер Kaggle MCP (протокол контекста модели)

Этот репозиторий содержит сервер MCP (Model Context Protocol) ( server.py ), созданный с использованием библиотеки fastmcp . Он взаимодействует с API Kaggle для предоставления инструментов для поиска и загрузки наборов данных, а также подсказки для генерации блокнотов EDA.

Структура проекта

  • server.py : Приложение сервера FastMCP. Оно определяет ресурсы, инструменты и подсказки для взаимодействия с Kaggle.

  • .env.example : Пример файла для переменных среды (учетные данные API Kaggle). Переименуйте в .env и заполните свои данные.

  • requirements.txt : список необходимых пакетов Python.

  • pyproject.toml и uv.lock : метаданные проекта и заблокированные зависимости для менеджера пакетов uv .

  • datasets/ : каталог по умолчанию, в котором будут храниться загруженные наборы данных Kaggle.

Related MCP server: Kaggle-MCP

Настраивать

  1. Клонируйте репозиторий:

    git clone <repository-url>
    cd <repository-directory>
  2. Создать виртуальную среду (рекомендуется):

    python -m venv venv
    source venv/bin/activate  # On Windows use `venv\Scripts\activate`
    # Or use uv: uv venv
  3. Установка зависимостей: Используя pip:

    pip install -r requirements.txt

    Или с помощью УФ:

    uv sync
  4. Настройте учетные данные API Kaggle:

    • Метод 1 (рекомендуемый): переменные среды

      • Создать файл .env

      • Откройте файл .env и добавьте свое имя пользователя Kaggle и ключ API:

        KAGGLE_USERNAME=your_kaggle_username
        KAGGLE_KEY=your_kaggle_api_key
      • Вы можете получить свой ключ API на странице вашего аккаунта Kaggle ( Account > API > Create New API Token ). Это загрузит файл kaggle.json , содержащий ваше имя пользователя и ключ.

    • Метод 2: файл kaggle.json

      • Загрузите файл kaggle.json из своей учетной записи Kaggle.

      • Поместите файл kaggle.json в ожидаемое место (обычно ~/.kaggle/kaggle.json в Linux/macOS или C:\Users\<Your User Name>\.kaggle\kaggle.json в Windows). Библиотека kaggle автоматически обнаружит этот файл, если переменные среды не установлены.

Запуск сервера

  1. Убедитесь, что ваша виртуальная среда активна.

  2. Запустите MCP-сервер:

    uv run kaggle-mcp

    Сервер запустится и зарегистрирует свои ресурсы, инструменты и подсказки. Вы можете взаимодействовать с ним с помощью клиента MCP или совместимых инструментов.

Запуск Docker-контейнера

1. Настройте учетные данные API Kaggle

Для доступа к наборам данных Kaggle этому проекту требуются учетные данные API Kaggle.

  • Перейдите по ссылке https://www.kaggle.com/settings и нажмите «Создать новый токен API», чтобы загрузить файл kaggle.json .

  • Откройте файл kaggle.json и скопируйте свое имя пользователя и ключ в новый файл .env в корне проекта:

KAGGLE_USERNAME=your_username
KAGGLE_KEY=your_key

2. Создайте образ Docker

docker build -t kaggle-mcp-test .

3. Запустите контейнер Docker, используя файл .env.

docker run --rm -it --env-file .env kaggle-mcp-test

Это автоматически загрузит ваши учетные данные Kaggle в качестве переменных среды внутри контейнера.


Возможности сервера

Сервер предоставляет следующие возможности через протокол контекста модели:

Инструменты

  • search_kaggle_datasets(query: str) :

    • Выполняет поиск наборов данных на Kaggle, соответствующих предоставленной строке запроса.

    • Возвращает список JSON из 10 лучших соответствующих наборов данных с такими подробностями, как ссылка, название, количество загрузок и дата последнего обновления.

  • download_kaggle_dataset(dataset_ref: str, download_path: str | None = None) :

    • Загружает и распаковывает файлы для определенного набора данных Kaggle.

    • dataset_ref : идентификатор набора данных в формате username/dataset-slug (например, kaggle/titanic ).

    • download_path (Необязательно): Указывает, куда загрузить набор данных. Если не указано, по умолчанию используется ./datasets/<dataset_slug>/ относительно расположения скрипта сервера.

Подсказки

  • generate_eda_notebook(dataset_ref: str) :

    • Генерирует подсказку, подходящую для модели ИИ (например, Gemini), для создания базовой записной книжки разведочного анализа данных (EDA) для указанного эталонного набора данных Kaggle.

    • В приглашении запрашивается код Python, охватывающий загрузку данных, проверку пропущенных значений, визуализацию и базовую статистику.

Подключение к Claude Desktop

Перейдите в Claude > Настройки > Разработчик > Изменить конфигурацию > claude_desktop_config.json, чтобы включить следующее:

{
  "mcpServers": {
    "kaggle-mcp": {
      "command": "kaggle-mcp",
      "cwd": "<path-to-their-cloned-repo>/kaggle-mcp"
    }
  }
}

Пример использования

Агент ИИ или клиент MCP может взаимодействовать с этим сервером следующим образом:

  1. Агент: «Поиск на Kaggle наборов данных о «болезнях сердца»»

    • Сервер выполняет search_kaggle_datasets(query='heart disease')

  2. Агент: «Загрузить набор данных 'user/heart-disease-dataset'»

    • Сервер выполняет download_kaggle_dataset(dataset_ref='user/heart-disease-dataset')

  3. Агент: «Создать запрос блокнота EDA для 'user/heart-disease-dataset'»

    • Сервер выполняет generate_eda_notebook(dataset_ref='user/heart-disease-dataset')

    • Сервер возвращает структурированное сообщение-подсказку.

  4. Агент: (Отправляет запрос в модель генерации кода) -> Получает код EDA Python.

Available Tools

2 tools
download_kaggle_datasetC

Downloads files for a specific Kaggle dataset. Args: dataset_ref: The reference of the dataset (e.g., 'username/dataset-slug'). download_path: Optional. The path to download the files to. Defaults to '/datasets/'.

ParametersJSON Schema
NameRequiredDescriptionDefault
dataset_refYes
download_pathNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states the action but lacks critical details: whether authentication is required (Kaggle typically needs API credentials), what happens if files already exist at the path, error handling, or any rate limits. The description is minimal beyond the basic operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently structured with a clear purpose statement followed by parameter explanations. It avoids unnecessary fluff, though the formatting with 'Args:' could be more integrated. Every sentence adds value, making it appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of downloading datasets (which often involves authentication, file management, and error cases), no annotations, and no output schema, the description is insufficient. It misses key contextual details like authentication requirements, response format, or handling of large downloads, leaving significant gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaningful context for both parameters: it explains the format of 'dataset_ref' with an example and clarifies the default behavior and path structure for 'download_path'. With 0% schema description coverage, this compensates somewhat, but it doesn't fully detail constraints (e.g., path validity, dataset accessibility).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Downloads files') and resource ('for a specific Kaggle dataset'), making the purpose immediately understandable. It distinguishes from the sibling tool 'search_kaggle_datasets' by focusing on downloading rather than searching, though it doesn't explicitly contrast them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. While it's implied this is for downloading after a dataset is identified (versus searching with the sibling tool), there's no explicit mention of prerequisites, dependencies, or when-not-to-use scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_kaggle_datasetsC

Searches for datasets on Kaggle matching the query using the Kaggle API.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions using the Kaggle API but doesn't disclose behavioral traits such as authentication requirements, rate limits, pagination, or what the search returns (e.g., format, fields). This leaves significant gaps for an agent to understand how to use it effectively.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It is appropriately sized and front-loaded, directly stating the tool's purpose without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't cover key aspects like authentication, rate limits, return format, or error handling. For a search tool with no structured support, more context is needed to guide an agent effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It implies the 'query' parameter is used for searching datasets, but doesn't add meaning beyond what the schema's title ('Query') and type suggest. No details on query syntax, examples, or constraints are provided, resulting in minimal added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Searches for datasets') and target resource ('on Kaggle'), specifying it uses the Kaggle API. It distinguishes from the sibling tool 'download_kaggle_dataset' by focusing on search rather than download, though it doesn't explicitly mention this distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description doesn't mention the sibling tool 'download_kaggle_dataset' or any other search methods, nor does it specify prerequisites like authentication or rate limits.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.1/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: one downloads a specific dataset, while the other searches for datasets. There is no overlap in functionality, making it easy for an agent to choose the correct tool for each task without confusion.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern (download_kaggle_dataset and search_kaggle_datasets), using snake_case and clear action verbs. This consistency makes the tool set predictable and easy to understand at a glance.

Tool Count2/5

With only two tools, the server feels thin for a Kaggle integration, lacking essential operations like listing datasets, uploading data, or managing competitions. While the tools are functional, the scope is incomplete for typical Kaggle workflows, making the count too low for the domain.

Completeness2/5

The tool set is severely incomplete for a Kaggle MCP server. It covers downloading and searching datasets but misses critical operations such as uploading datasets, accessing competition data, or interacting with notebooks. This creates significant gaps that will hinder agents from performing common Kaggle tasks.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Connects Claude AI to the Kaggle API through the Model Context Protocol, enabling users to browse competitions, search and download datasets, analyze kernels, and access pre-trained models through natural language interactions.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to interact with Kaggle competitions, including listing competitions, downloading files, submitting predictions, and viewing submission history.
    10

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/arrismo/kaggle-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server