tslab-mcp
tslab-mcp
MCP-сервер, который предоставляет детерминированное прогнозирование временных рядов в виде инструментов, так что ваш агент является механизмом рассуждения, а каждое число поступает из обычного, воспроизводимого Python.
Никакая LLM не вызывается нигде в этом пакете. Никакой API-ключ не требуется (если только вы не запросите TimeGPT, который вызывает Nixtla API).
Зачем
Некоторые библиотеки прогнозирования поставляют агента, который читает признаки, выбирает модель и объясняет результат с помощью LLM в цикле. Вызов такого агента из вашего собственного агента вкладывает агента внутрь агента — два запроса, два счета, два источника недетерминизма и непрозрачный промежуточный слой, который делает обоснование выбора модели неаудитируемым.
Поэтому здесь управление инвертировано: библиотека прогнозирования — это инструмент, а ваш агент — тот, кто рассуждает. Он читает признаки, аргументирует выбор семейства моделей, перекрестно проверяет кандидатов и записывает обоснование в манифест. Каждое число на этом пути создается вызовом библиотеки, который можно перезапустить без LLM в цепочке.
Это разделение распространяется и на то, как построен сам пакет. Базовая установка запускает одиннадцать статистических моделей — AutoARIMA, AutoETS, Theta, CrostonClassic и другие — через statsforecast: примерно 340 МБ, без PyTorch, и запускается за секунды. Дополнительный пакет foundation добавляет предобученные модели TimeCopilot — Chronos, Moirai, TimesFM, TiRex, Toto и другие — плюс Prophet, для случаев, когда статистического базиса недостаточно. Запрос, который называет только статистические модели, никогда не импортирует TimeCopilot или torch; запрос, который называет хотя бы одну фундаментальную модель, полностью выполняется через TimeCopilot, который также несет в себе статистические модели. В любом случае tsf_list_models сообщает, что на самом деле установлено, прежде чем вы примете решение о модели.
Related MCP server: timeseries-mcp
Установка
Требуется Python 3.10+ (рекомендуется 3.13, см. Версия Python).
uvx tslab-mcp # run without installing
uv tool install tslab-mcp # or install the CLIБазовая установка запускает одиннадцать статистических моделей через statsforecast: примерно 340 МБ, без PyTorch, и запускается мгновенно. Для предобученных фундаментальных моделей — Chronos, Moirai, TimesFM, Toto, TiRex — и Prophet добавьте дополнительный пакет:
uvx --from 'tslab-mcp[foundation]' tslab-mcpДополнительный пакет
foundationподтягивает TimeCopilot, который приносит torch, transformers и lightning: примерно 2 ГБ при первой установке, и первый вызов инструмента, который его касается, тратит ~30 секунд на импорт. Оба действия одноразовые, и ни одно из них не является платным, если только вы не запросите модель, которая их требует.
Из GitHub
uv и uvx оба принимают git-URL вместо имени пакета, что устанавливает текущий main без ожидания релиза:
uvx --from git+https://github.com/pedrobtz/tslab-mcp tslab-mcp
uv tool install git+https://github.com/pedrobtz/tslab-mcp # or install the CLI
# with the foundation extra
uvx --from 'tslab-mcp[foundation] @ git+https://github.com/pedrobtz/tslab-mcp' tslab-mcpЗафиксируйте ревизию для чего-либо, кроме случайного тестирования — ветка может перемещаться под вами. Коммит работает сегодня; тег версии тоже будет работать, как только он будет создан:
uv tool install "git+https://github.com/pedrobtz/tslab-mcp@136824c1cc2a"Из локальной копии
git clone https://github.com/pedrobtz/tslab-mcp
cd tslab-mcp
uv sync # base
uv sync --extra foundation # with the pretrained models
uv run tslab-mcpНастройка
Добавьте сервер в конфигурацию вашего MCP-клиента. Файл различается в зависимости от клиента — часто это .mcp.json в корне проекта — но сама запись имеет одинаковую форму:
{
"mcpServers": {
"tslab": {
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "~/.tslab-mcp"
}
}
}
}TSLAB_MCP_HOME задает, куда записываются артефакты; по умолчанию это ~/.tslab-mcp, а результаты выполнения попадают в <home>/runs.
Транспорт только stdio, по замыслу: ваши данные считаются чувствительными и никогда не покидают машину. Сервер не делает исходящих запросов, кроме загрузки весов моделей, которые TimeCopilot сам выполняет для фундаментальных моделей, и вызовов Nixtla API, которые делает TimeGPT, если вы специально его запросите.
GitHub Copilot
Copilot обнаруживает MCP-серверы из файла mcp.json и предоставляет их инструменты в режиме агента — инструменты не появляются в режиме ask или edit.
VS Code. Поместите сервер в .vscode/mcp.json, чтобы поделиться им с репозиторием, или выполните MCP: Open User Configuration из палитры команд, чтобы сохранить его в своем профиле для всех рабочих областей. Обратите внимание, что ключ — servers, а не mcpServers:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "${userHome}/.tslab-mcp"
}
}
}
}Из локальной копии укажите путь к рабочему дереву:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uv",
"args": ["run", "--directory", "${workspaceFolder}", "tslab-mcp"]
}
}
}Затем: откройте Chat, переключите селектор режима на Agent и с помощью кнопки Tools убедитесь, что восемь инструментов tsf_* перечислены и включены. MCP: List Servers показывает статус сервера и его логи, где объясняется неудачный запуск. Copilot ограничивает количество одновременно активных инструментов, поэтому, если вы запускаете несколько MCP-серверов, возможно, придется отключить некоторые, чтобы уместить все восемь.
Visual Studio. Та же форма JSON, в .mcp.json в корне решения (или %USERPROFILE%\.mcp.json для всех решений), затем включите инструменты из панели выбора инструментов в режиме агента Copilot Chat.
JetBrains, Eclipse и Xcode. Откройте панель выбора инструментов в режиме агента Copilot Chat, выберите Edit MCP configuration и добавьте ту же запись servers в открывшийся mcp.json.
Copilot coding agent (облачный агент на github.com) плохо подходит для этого сервера: он запускает ваши MCP-серверы в эфемерном окружении GitHub Actions, что означает оплату ~2 ГБ установки TimeCopilot при каждом запуске, и у него нет доступа к локальным файлам данных. Используйте его из своего редактора.
Инструменты
Инструмент | Назначение | Возвращает |
| Чтение CSV/Parquet, проверка контракта | JSON-сводка + SHA-256 |
| Признаки для каждого ряда для выбора семейства моделей | Таблица Markdown или JSON, ограниченная по строкам |
| Проверка, какие модели на самом деле импортируются здесь |
|
| Сравнение с скользящим началом по нескольким моделям | Таблица метрик, рейтинг, путь к parquet |
| Подгонка и прогнозирование с интервалами прогноза | Путь к parquet + ограниченный предпросмотр |
| Флагирование на основе перекрестно проверенных интервалов | Количество, ограниченный список флагов, путь к parquet |
| Фиксация сессии в перезапускаемом манифесте | Путь к манифесту |
| Преобразование каждого шага в читаемый отчет | Путь к HTML или Markdown |
Все, кроме двух инструментов tsf_export_*, помечены как read-only; здесь ничего не удаляется, так что очистка ~/.tslab-mcp/runs — ваша забота, а не агента.
Начало сессии
Инструменты не навязывают порядок, поэтому начальный запрос превращает восемь вызываемых функций в анализ. Что-то вроде этого работает хорошо:
Используй инструменты tslab для прогнозирования рядов в
/Users/me/data/deposits.csv, на 12 месяцев вперед.Работай в таком порядке и показывай свои рассуждения на каждом шаге:
Загрузи файл и расскажи, что ты нашел — сколько рядов, какая частота, есть ли пропуски или отсутствующие значения.
Опиши признаки и скажи, какие семейства моделей они обосновывают и почему.
Проверь, какие модели на самом деле установлены, прежде чем предлагать.
Проведи перекрестную проверку твоего шорт-листа против базовой SeasonalNaive на 4 окнах. Пока только статистические модели.
Спрогнозируй с победителем, с интервалами 80% и 95%.
Экспортируй манифест запуска и HTML-отчет, и помести обоснование выбора модели в заметку: что ты выбрал, что показала таблица метрик и что ты отклонил.
Обобщи результаты и дай мне пути к parquet — не вставляй целые таблицы в чат.
Четыре вещи в этом запросе действительно важны:
Абсолютный путь. Относительные пути разрешаются относительно рабочего каталога сервера, который выбирает ваш MCP-клиент, и вы обычно не можете его предсказать.
Горизонт, соответствующий решению.
hуправляет как прогнозом, так и тем, сколько истории потребляет каждое окно CV; 12 месячных шагов — это год планирования, а не произвольное значение по умолчанию.«Пока только статистические модели.» Без этого агент может потянуться к фундаментальной модели и потратить несколько минут на загрузку весов, чтобы ответить на вопрос, который
AutoETSрешил бы за секунды. Снимите ограничение, когда дешевые модели установят базовый уровень.Запрос обоснования в заметке манифеста. Стенограмма чата одноразова; манифест — это та часть, которую кто-то может перезапустить и проверить. Если рассуждения существуют только в разговоре, они фактически потеряны.
Более короткие начальные запросы, когда вы знаете, что хотите:
Загрузи
/Users/me/data/sales.parquetи опиши признаки. Пока не прогнозируй — я хочу сначала увидеть, с чем мы имеем дело.
Сравни SeasonalNaive, AutoETS и AutoARIMA на загруженном дескрипторе
depositsна 6 окнах при h=12, затем скажи, превосходит ли что-то базовую модель настолько, чтобы стоить дополнительной сложности.
Вызовы только статистических моделей отвечают за секунды. Первый вызов, который называет фундаментальную модель, тратит ~30 секунд на импорт TimeCopilot, прежде чем сделать что-либо еще — эта пауза ожидаема, это не зависание, и она происходит только в том случае, если установлен дополнительный пакет foundation и запрос действительно обращается к такой модели.
Пример сессии
Начнем с CSV в длинном формате Nixtla:
unique_id,ds,y
branch_01,2018-01-01,1043.2
branch_01,2018-02-01,1102.7
...1. Загрузите его. Панель остается в процессе сервера; дескриптор — это все, что несет сессия.
{"handle": "deposits", "n_series": 12, "n_obs": 864, "freq": "MS",
"start": "2018-01-01T00:00:00", "end": "2023-12-01T00:00:00",
"obs_per_series": {"min": 72, "median": 72, "max": 72},
"n_missing_y": 0, "sha256": "9f2c…"}2. Опишите его. Это числа, над которыми вы рассуждаете.
| id | n | mean | cv | %zero | trend | seasonal | acf1(diff) |
|-----------|----|--------|-------|-------|-------|----------|------------|
| branch_01 | 72 | 1180.4 | 0.112 | 0.0 | 0.83 | 0.62 | -0.31 |Высокая сезонная сила и четкий тренд говорят в пользу AutoETS и AutoARIMA по сравнению с наивным базовым уровнем; высокий %zero говорил бы в пользу ADIDA или CrostonClassic.
seasonal — это сила STL — сезонная составляющая, измеренная относительно того, что остается после удаления тренда, — поэтому растущий ряд все равно честно сообщает о своей сезонности. Он имеет шумовой порог примерно 0.3–0.5: оценки в этом диапазоне означают «нет доказательств», а не «умеренно сезонный».
3. Проверьте, что установлено с помощью tsf_list_models, чтобы никогда не предлагать модель, которую эта машина не может запустить.
4. Проведите перекрестную проверку кандидатов — всегда включая SeasonalNaive, так как модель, которая не может его превзойти, не стоит развертывания:
{"kind": "cross_validation", "models": ["SeasonalNaive", "AutoETS", "AutoARIMA"],
"h": 12, "n_windows": 4, "seasonality_used_for_mase": 12,
"metrics": {"mase": {"SeasonalNaive": 1.0, "AutoETS": 0.71, "AutoARIMA": 0.68}},
"ranking": {"mase": ["AutoARIMA", "AutoETS", "SeasonalNaive"]},
"artifact": "~/.tslab-mcp/runs/cv_deposits_3f1a9c02.parquet"}5. Спрогнозируйте с победителем. Полный фрейм идет в parquet; ответ содержит путь, столбцы и краткий предпросмотр.
6. Экспортируйте запуск и отчет. Запишите почему в заметку — это единственная часть ваших рассуждений, которая переживет разговор:
{"manifest": "~/.tslab-mcp/runs/manifest_deposits_77b0e415.json", "n_runs": 3,
"kinds": ["cross_validation", "forecast"]}Манифест содержит исходный путь и хэш, частоту, каждый вызов с его аргументами и путями артефактов, зафиксированные версии всего, что на самом деле установлено — statsforecast, pandas и Python всегда; TimeCopilot и torch тоже, если установлен дополнительный пакет foundation — и вашу заметку. Этого достаточно, чтобы воспроизвести числа при остановленном сервере.
tsf_export_report превращает тот же манифест в то, что читает человек — признаки, таблицы метрик, упорядоченные от лучших к худшим, прогнозы, аномалии и окружение, в порядке их возникновения:
{"report": "~/.tslab-mcp/runs/report_deposits_5c31d0a7.html",
"format": "html", "n_steps": 3,
"steps": ["features", "cross_validation", "forecast"]}Отчет является чистой функцией манифеста: он не читает parquet и не вызывает модель, поэтому tsf_export_report с manifest_path перерисовывает запуск месячной давности без загруженных данных. HTML встраивает свой собственный CSS и не ссылается на внешние скрипты, таблицы стилей или шрифты, поэтому он все еще корректно открывается в офлайн-режиме.
Дизайн
Четыре инварианта и причины их существования:
Дескрипторы, а не датафреймы. Один фрейм кросс-валидации содержит n_series × h × n_windows × n_models строк. Сериализация его в результат инструмента исчерпывает контекст сессии при первом же вызове и ухудшает каждый последующий оборот. Инструменты принимают дескриптор и возвращают сводки, агрегаты и пути к файлам; каждый массовый путь ограничен и сообщает, что было пропущено, чтобы сессия знала, что нужно прочитать parquet, а не запрашивать снова.
Блокирующая работа никогда не касается цикла событий. Кросс-валидация нескольких моделей на большой панели занимает минуты процессорного времени. Каждое тело инструмента представляет собой синхронное замыкание, отправляемое через anyio.to_thread.run_sync, поэтому транспорт stdio продолжает отвечать, и клиент не обрывает сервер на середине выполнения.
Окружение определяется, а не предполагается. Модели импортируются лениво и проверяются, никогда не предполагается, что они присутствуют. tsf_list_models сообщает, что на самом деле разрешилось, поэтому запрос Chronos без дополнительного пакета возвращает сообщение, указывающее на этот пакет, а не трейсбек через десять минут выполнения.
Бэкенд выбирается в зависимости от того, что вы запрашиваете: запрос, все модели в котором статистические, выполняется через statsforecast, и только запрос, которому нужна предварительно обученная модель, обращается к TimeCopilot. Таким образом, статистические запуски никогда не импортируют torch, и сервер запускается мгновенно в любом случае.
statsforecast намеренно оставлен с настройкой по умолчанию n_jobs=1. Его параллельный режим порождает рабочие процессы, которые повторно импортируют входной модуль, что внутри MCP-сервера приводит к конкуренции и риску для stdout, а не к ускорению.
Манифест — это артефакт для записи. Проза в диалоге — это комментарий. Манифест — это то, что кто-то перезапустит через шесть месяцев, и то, что читает рецензент, чтобы увидеть, какие модели сравнивались и на каком основании.
Python version
TimeCopilot gates several models on the interpreter version, and on Python < 3.13
it pins tabpfn-time-series, which caps pandas below 2.2.
Python | Models | pandas |
3.13 | всё, кроме | ≥ 2.2 |
3.10–3.12 | добавляется | < 2.2 |
3.13 — рекомендуемая версия. В любом случае tsf_list_models сообщает, что на самом деле разрешилось, с указанием причины для всего, что не разрешилось.
Development
uv sync --all-groups
uv run pytest # fast suite
uv run pytest -m slow # exercises TimeCopilot; slower, no weight downloads
uv run ruff check src tests
uv run mypyПроверьте поверхность инструментов с помощью MCP Inspector:
npx @modelcontextprotocol/inspector uv run tslab-mcpLicense
MIT
Available Tools
8 toolstsf_cross_validateARead-onlyIdempotent
Compare models by rolling-origin cross-validation.
This is the tool that replaces guesswork about model choice: it produces the evidence, you read the table and decide. Always include SeasonalNaive as the baseline -- a model that cannot beat it is not worth deploying.
Returns a per-model metric table aggregated over series and windows, a best-first ranking per metric, and the parquet path holding every per-window prediction. LONG-RUNNING: seconds for statistical models, many minutes for foundation models on a large panel. Start with statistical models on the real horizon before reaching for anything pretrained.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, idempotent, non-destructive behavior. The description adds crucial behavioral context: runtime warning ('LONG-RUNNING'), output description (aggregated table, ranking, parquet path), and implicitly that it is safe but compute-intensive. This goes well beyond the annotations and aids agent expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, purposeful paragraphs. The first sentence immediately states the tool's function. Each subsequent section (usage advice, output details, runtime warning) earns its place with no redundancy. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (cross-validation, multiple models, windows, output schema exists), the description covers purpose, usage, baseline recommendation, output contents (aggregated metrics, rankings, prediction parquet), and runtime behavior. It is sufficiently complete for an agent to understand when and how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The nested input schema (CrossValidateInput) already provides parameter descriptions (e.g., models, metrics). The description adds high-level advice (like horizon matching the decision, baseline recommendation) but no new parameter-level semantics beyond what the schema offers. Schema coverage is effectively high despite the 0% top-level stat, so the description's incremental value here is moderate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource combination ('Compare models by rolling-origin cross-validation') and clearly distinguishes this tool from siblings like tsf_forecast (single model forecast) and tsf_list_models (model names). It states it replaces guesswork about model choice, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Always include SeasonalNaive as the baseline' and 'Start with statistical models on the real horizon before reaching for anything pretrained.' It also explains that the tool produces evidence for model selection. However, it does not explicitly state when not to use this tool (e.g., for final forecasts) or name alternatives, slightly reducing completeness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_describe_seriesARead-onlyIdempotent
Compute the per-series features that decide which model family to try.
Returns length, mean, sd, coefficient of variation, share of zeros, trend strength (R-squared against time), seasonal strength (variance explained by the period means), and lag-1 autocorrelation of the differenced series.
Read it as evidence, not as an answer: high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models (ADIDA, IMAPA, CrostonClassic); high cv with low structure argues for keeping expectations modest. Cheap -- returns in under a second.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and nondestructive nature. The description adds 'Cheap -- returns in under a second' and 'Read it as evidence, not as an answer,' providing behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with five sentences, front-loaded with purpose, and every sentence adds value. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema (not shown but present), the description covers the semantic meaning of the features and how to interpret them. It also provides cost and time estimates, making it self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has detailed descriptions for all three parameters (handle, max_series, response_format). The tool description focuses on output features and usage advice, not parameter details. Since schema coverage is high, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a clear verb-resource pair: 'Compute the per-series features that decide which model family to try.' It lists the specific features computed, distinguishing this diagnostic tool from siblings like tsf_forecast or tsf_load_series.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit decision rules: 'high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models...' and positions the tool as 'evidence, not an answer.' This clearly guides when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_detect_anomaliesARead-onlyIdempotent
Flag historical points that fall outside a cross-validated prediction interval.
The detector model defines what "expected" means, so pick one that fits the series: a weak detector flags its own errors rather than real anomalies. Run tsf_describe_series or tsf_cross_validate first.
Returns flagged counts per series, a capped list of flagged rows, and the parquet path with the full result.
LONG-RUNNING, and the default is the expensive one: leaving n_windows unset refits the model once per observation across the whole history, which takes minutes even for a statistical model. Pass n_windows (e.g. 12) unless you genuinely need every point tested.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses long-running nature and expensive default (n_windows unset refits per observation, taking minutes). Annotations (readOnlyHint, idempotentHint) are consistent; description adds critical performance context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact four paragraphs with clear structure: purpose, prerequisites, output summary, performance warning. No redundant sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given tool complexity (cross-validation, long-running, return includes parquet path and capped list), the description covers prerequisites, output, and performance. Output schema exists, so return values are adequately summarized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond schema by explaining the default slowness of n_windows and the risk of weak models. Schema already has clear descriptions, but description provides crucial usage context for these parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it flags historical points outside a cross-validated prediction interval, distinguishes from sibling tools (tsf_describe_series, tsf_cross_validate, etc.), and warns about weak detectors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises running tsf_describe_series or tsf_cross_validate first, warns against weak detectors, and gives concrete guidance on setting n_windows to avoid slow default behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_reportA
Render every step of the analysis as a report someone can read.
Covers the input and its hash, the features, each cross-validation with its metric table and ranking, the forecasts, any anomaly runs, and the pinned environment -- in the order they happened. HTML is self-contained, with no external stylesheet or script, so it opens correctly years later.
Call it after tsf_export_run at the end of an analysis. Pass a note: the report headlines it as the rationale, and a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.
Report from manifest_path instead of handle to re-render an older run --
it needs nothing but the manifest file.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (all hints false), so the description carries the burden. It discloses that HTML output is 'self-contained, with no external stylesheet or script' and that the report covers steps 'in the order they happened.' However, it does not clarify whether the tool writes a file to disk, returns the report content, or has other side effects. The output schema exists but is not described in the tool description, leaving behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a lead sentence defining the tool, then a list of contents, then usage order, then note advice, then alternative invocation. It is informative without being verbose. Minor inefficiency: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again' is slightly colorful but still earns its place. Could be tighter, but overall effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 optional parameters and an output schema (which handles return value documentation), the description covers its purpose, contents, usage order, and parameter trade-offs. It does not explain what happens if both handle and manifest_path are provided (mutual exclusion handled by schema? not specified). It also assumes the agent knows tsf_export_run was called, which is implied by 'at the end of an analysis.' Overall adequately complete for an AI agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has detailed descriptions for all parameters (note, format, handle, manifest_path), so baseline is 3. The description goes beyond by explaining the semantic purpose of note ('headlines it as the rationale') and the trade-off between handle and manifest_path ('re-renders an older run -- it needs nothing but the manifest file'). This adds actionable context for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description starts with a specific verb ('Render every step of the analysis as a report') and enumerates the exact contents (input hash, features, cross-validation, forecasts, anomalies, pinned environment). This immediately distinguishes it from sibling tools like tsf_export_run (which exports run data) and tsf_forecast (which only forecasts). The purpose is unambiguous and comprehensive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states ordering: 'Call it after tsf_export_run at the end of an analysis.' Provides clear alternatives: 'Report from manifest_path instead of handle to re-render an older run.' Also advises on best practice for the note parameter: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.' This gives the agent concrete when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_runA
Write a JSON manifest of everything done to this handle.
Records the source path and SHA-256, the frequency, every call with its arguments and artifact paths, the pinned package versions, and your note. This is the artifact of record: your prose in the conversation is lost, this file is not. Write the note -- say which model you picked, what the metric table showed, and what you rejected.
Call it at the end of any analysis someone might have to defend or rerun. Writes a file, so it is not read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds valuable context: 'Writes a file, so it is not read-only' and lists everything included in the manifest. It also emphasizes that the note is the only place a reasoning trace survives, which is important behavioral insight. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: starts with the core purpose, then enumerates contents, gives usage advice, and closes with a note about file writing. It is not overly long; every sentence contributes. A slight trim could improve conciseness, but it remains clear and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description correctly omits return value details. It covers when to call, what the manifest contains, and the critical role of the note. The only minor gap is no mention of potential side effects (e.g., overwriting existing files), but overall it is sufficiently complete for a tool with good annotations and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already contains descriptions for both parameters: handle ('Handle whose run log should be pinned to a manifest') and note (detailed explanation of what to write). The tool description adds further guidance for the note, specifically: 'Write the note -- say which model you picked, what the metric table showed, and what you rejected.' This enhances the schema's descriptions, making it clear how to use the note parameter effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Write a JSON manifest of everything done to this handle' clearly states the verb (write a manifest) and the resource (handle). The description further details what the manifest includes (source path, SHA-256, calls, arguments, artifact paths, pinned package versions, note), differentiating it from sibling tools like 'tsf_export_report' or 'tsf_describe_series'. This makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to call it: 'Call it at the end of any analysis someone might have to defend or rerun.' This provides clear context for use. It does not explicitly state when not to use it or name alternatives, but for a specialized export tool the guidance is sufficient and well-placed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_forecastARead-onlyIdempotent
Fit on the full history and forecast h periods ahead with intervals.
Use after tsf_cross_validate has justified the model choice. The full forecast goes to parquet; the response carries the path, the column list, the row count and a small preview. Read the parquet for anything more -- raising max_preview_rows to dump the frame into the conversation is the one thing that reliably ruins a long session.
LONG-RUNNING for foundation models.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the bar is lowered. The description adds valuable context: the full forecast goes to parquet, the response carries path/column list/row count/preview, warns against raising max_preview_rows, and flags 'LONG-RUNNING for foundation models' — all beyond what annotations provide. No contradictions with annotations (readOnlyHint=true is consistent with generating forecasts without mutating data).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: a terse three-sentence explanation that front-loads the core action, then adds usage guidance and behavioral warnings. Every sentence adds essential information without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (forecasting with multiple models and intervals), the output schema exists, so return values don't need elaboration. The description covers the critical workflow (use after cross-validation), output format (parquet with preview), and a key gotcha (don't dump full frame). It lacks explicit error conditions or prerequisite checks, but for the scope, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema provides no parameter descriptions — so the description must compensate. Although the main description does not detail parameters, the parameter `max_preview_rows` receives meaningful context: 'Rows of the forecast to inline... read that instead of raising this.' Other parameters (handle, models, h, level) have descriptions in the schema via the JSON Schema, but since coverage is 0% (likely meaning no separate param list in the description), the main text does not clarify their meaning beyond what the schema already provides. Baseline 3 is appropriate as the description adds some context for max_preview_rows but not for others.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool fits on full history and forecasts h periods ahead with intervals. It uses specific verbs like 'fit' and 'forecast' and explicitly identifies the resource as the time series forecast. However, it does not directly differentiate from siblings like tsf_cross_validate, though the usage guideline addresses that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after tsf_cross_validate has justified the model choice,' providing clear sequencing context and an alternative (cross-validation). It does not mention when not to use it or list specific alternatives for other tasks like anomaly detection, but the primary usage guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_list_modelsARead-onlyIdempotent
Probe which models actually import in this environment.
Call this before cross-validating so you never propose a model that cannot run here. Returns {available, statistical, foundation, unavailable}, where each unavailable entry carries the real reason -- some models are gated on the Python version, not merely absent.
The first call imports TimeCopilot and can take ~30 seconds; later calls are instant.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds important behavioral details beyond that: the first call may take ~30 seconds to import TimeCopilot, later calls are instant; it returns structured output with real reasons for unavailability (including Python version gating). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and structured: first sentence states purpose, second gives usage guidance, third explains return structure, and fourth notes startup latency. Every sentence adds essential information, with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (the agent can rely on structured return type details), the description covers all necessary context: what the tool does, when to use it, its runtime behavior (latency), and the high-level shape of results. No gaps remain for a list/probe tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema contains descriptions for both parameters (family and include_unavailable) that are self-explanatory. The tool description does not add new parameter information—it only repeats that statistical models are cheap and always installed, which already appears in the schema. Schema coverage via inline descriptions is present, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Probe which models actually import in this environment' and specifies the tool's role in preventing proposal of non-running models. It clearly distinguishes itself from siblings like cross_validate by giving a precise pre-check use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contains an explicit directive: 'Call this before cross-validating so you never propose a model that cannot run here.' It also notes that statistical models are always installed, which helps with decision-making. No alternatives are listed, but the use case is crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_load_seriesARead-onlyIdempotent
Read a CSV or Parquet panel from disk and register it under a handle.
Call this first; every other tool takes the handle it returns. The file must be in Nixtla long format (unique_id, ds, y). Returns a compact JSON summary -- series count, inferred frequency, date range, missing values, and the SHA-256 of the source -- and nothing else: the data stays in the server so it never consumes your context.
Read the summary before choosing a horizon. If obs_per_series.min is small, a long horizon or many CV windows will not fit.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description reveals that data stays server-side ('never consumes your context') and that the return is a compact summary with specific fields. It also hints at the side effect of reusing a handle (replacing the panel). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five sentences, each serving a distinct purpose: purpose, ordering, format, return details, caution. Front-loaded with the core action. No superfluous words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's role as the entry point for a time series workflow, the description covers purpose, required file format, return value (with summary contents), and a concrete usage caution. With an output schema present, the lack of detailed return structure is acceptable. The description is complete enough for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides detailed descriptions for all three parameters. The tool description adds the critical constraint that the file must be in 'Nixtla long format (unique_id, ds, y)', which is not in the schema. This adds meaningful value beyond the schema, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action: 'Read a CSV or Parquet panel from disk and register it under a handle.' It distinguishes from sibling tools by explicitly saying 'Call this first; every other tool takes the handle it returns.' This makes the purpose unambiguous and contextually positioned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit ordering ('Call this first'), explains the handle's role in subsequent tools, and gives a practical caution about horizon choices based on the summary output. This equips the agent with clear when-to-use and how-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
tsf_cross_validate - First observed
tsf_describe_series - First observed
tsf_detect_anomalies - First observed
tsf_export_report - First observed
tsf_export_run - First observed
tsf_forecast - First observed
tsf_list_models - First observed
tsf_load_series
TDQS
Scored across 8 tools
Each tool has a distinct and well-defined purpose within the time series forecasting workflow: loading, describing, listing models, cross-validating, forecasting, detecting anomalies, and exporting results. There is no overlap or ambiguity between tools.
All tools follow a consistent pattern: the prefix 'tsf_' followed by a verb (and optional noun), all in snake_case. Examples include tsf_load_series, tsf_describe_series, tsf_cross_validate, and tsf_export_report. The naming is predictable and uniform.
With 8 tools, the server covers a complete analysis pipeline without excess. Each tool is necessary and corresponds to a clear step in the workflow, from data loading to report generation. The count is well-scoped for the domain.
The tool surface covers the full lifecycle of a typical time series analysis: load data, explore features, check available models, cross-validate, forecast, detect anomalies, and export manifests/reports. There are no obvious gaps; the workflow feels self-contained and actionable.
Maintenance
Related MCP Connectors
Probabilistic time-series forecasts from zero-shot foundation models: routed, single or ensembled.
1PredictOracle - 12 forecasting tools: time-series, scenario analysis, risk projections.
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Deterministic time tools for AI agents: timezone conversion, business-day math, cron interpretation.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnable any AI agent to forecast time-series data (e.g., sales, traffic) using Google's TimesFM or a zero-dependency statistical baseline.3Apache 2.0
- AlicenseAqualityDmaintenanceDeterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.17MIT
- AlicenseNot gradedqualityCmaintenanceEnables time-series analysis and forecasting through a structured tool catalogue, including data loading, quality repair, diagnostics, and forecasting with ARIMA, exponential smoothing, Chronos-2, Toto 2.0, and AutoML.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to run TimesFM-3 forecasting workflows locally, including joint multivariate forecasting with known future drivers, backtesting against a baseline, what-if scenario comparison, and historical anomaly detection. It exposes the studio's tools and bundled public and synthetic demo datasets over MCP.Apache 2.0