tslab-mcp
tslab-mcp
時系列予測をツールとして公開するMCPサーバーです。予測は決定論的で再現可能なPythonによって行われ、エージェントは推論エンジンとして機能します。
このパッケージ内でLLMが呼び出されることは一切ありません。APIキーも不要です(TimeGPTを使用する場合を除きます。その場合はNixtla APIを呼び出します)。
なぜ
一部の予測ライブラリは、特徴量を読み取り、モデルを選択し、LLMを介して結果を説明するエージェントを搭載しています。そのようなライブラリを自身のエージェントから呼び出すと、エージェントの中にエージェントが入れ子になります。つまり、2つのプロンプト、2つの課金、2つの非決定論的要因、そしてモデル選択の根拠を監査不可能にする不透明な中間層が生まれます。
そこで、ここでは制御を逆転させています。予測ライブラリはツールであり、推論を行うのはあなたのエージェントです。エージェントは特徴量を読み取り、モデルファミリーについて議論し、候補を交差検証し、その根拠をマニフェストに書き込みます。途中のすべての数値は、LLMを介さずに再実行可能なライブラリ呼び出しによって生成されます。
この分割は、パッケージ自体の構築方法にも反映されています。基本インストールでは、statsforecastを介してAutoARIMA、AutoETS、Theta、CrostonClassicなどの11の統計モデルが実行されます。サイズは約340MB、PyTorchは不要で、起動は数秒です。オプションのfoundationエクストラを追加すると、TimeCopilotの事前学習済みモデル(Chronos、Moirai、TimesFM、TiRex、Totoなど)とProphetが追加され、統計的ベースラインでは不十分な場合に対応します。統計モデルのみを指定したリクエストでは、TimeCopilotやtorchはインポートされません。一方、1つでもファンデーションモデルを指定したリクエストは、TimeCopilotを介して完全に実行され、統計モデルも含まれます。どちらの場合でも、tsf_list_modelsは、モデルを確定する前に実際にインストールされているものを報告します。
Related MCP server: timeseries-mcp
インストール
Python 3.10以上が必要です(3.13推奨、Pythonバージョンを参照)。
uvx tslab-mcp # run without installing
uv tool install tslab-mcp # or install the CLI基本インストールでは、statsforecastを介して11の統計モデルが実行されます。サイズは約340MB、PyTorchは不要で、即座に起動します。事前学習済みファンデーションモデル(Chronos、Moirai、TimesFM、Toto、TiRex)とProphetを追加するには、エクストラを追加します:
uvx --from 'tslab-mcp[foundation]' tslab-mcp
foundationエクストラはTimeCopilotをプルし、torch、transformers、lightningをもたらします。初回インストール時は約2GB、それに触れる最初のツール呼び出しではインポートに約30秒かかります。どちらも1回限りであり、それらを必要とするモデルを要求しない限り、費用は発生しません。
GitHubから
uvとuvxはどちらもパッケージ名の代わりにgit URLを受け入れ、リリースを待たずに現在のmainをインストールします:
uvx --from git+https://github.com/pedrobtz/tslab-mcp tslab-mcp
uv tool install git+https://github.com/pedrobtz/tslab-mcp # or install the CLI
# with the foundation extra
uvx --from 'tslab-mcp[foundation] @ git+https://github.com/pedrobtz/tslab-mcp' tslab-mcpカジュアルなテスト以外では、refを固定してください。そうしないと、ブランチの先頭が移動する可能性があります。コミットは現在機能します。バージョンタグも、一度作成されれば機能します:
uv tool install "git+https://github.com/pedrobtz/tslab-mcp@136824c1cc2a"チェックアウトから
git clone https://github.com/pedrobtz/tslab-mcp
cd tslab-mcp
uv sync # base
uv sync --extra foundation # with the pretrained models
uv run tslab-mcp設定
MCPクライアントの設定にサーバーを追加します。ファイルはクライアントによって異なります(多くの場合、プロジェクトルートの.mcp.json)が、エントリ自体は同じ形式です:
{
"mcpServers": {
"tslab": {
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "~/.tslab-mcp"
}
}
}
}TSLAB_MCP_HOMEはアーティファクトの書き込み先を設定します。デフォルトは~/.tslab-mcpで、実行出力は<home>/runsに配置されます。
トランスポートはstdioのみです。これは設計上の意図で、データは機密性が高く、マシンから出ることはありません。サーバーが行うアウトバウンドリクエストは、TimeCopilotがファンデーションモデルに対して実行するモデル重みのダウンロードと、TimeGPTを明示的に要求した場合にNixtla APIを呼び出すことだけです。
GitHub Copilot
Copilotはmcp.jsonファイルからMCPサーバーを検出し、エージェントモードでツールを公開します。ツールはaskモードやeditモードでは表示されません。
VS Code。 サーバーを.vscode/mcp.jsonに配置してリポジトリと共有するか、コマンドパレットからMCP: Open User Configurationを実行して、すべてのワークスペースで自分のプロファイルに保持します。キーはmcpServersではなくserversであることに注意してください:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "${userHome}/.tslab-mcp"
}
}
}
}チェックアウトから使用する場合は、代わりにワーキングツリーを指定します:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uv",
"args": ["run", "--directory", "${workspaceFolder}", "tslab-mcp"]
}
}
}その後、Chatを開き、モードセレクターをAgentに切り替え、Toolsボタンを使用して8つのtsf_*ツールがリストされ、有効になっていることを確認します。MCP: List Serversはサーバーのステータスとログを表示し、起動に失敗した場合の説明がここに表示されます。Copilotは同時にアクティブにできるツールの数に上限があるため、複数のMCPサーバーを実行している場合は、8つすべてを収めるためにいくつかを選択解除する必要がある場合があります。
Visual Studio。 同じJSON形式で、ソリューションルートの.mcp.json(またはすべてのソリューションで%USERPROFILE%\.mcp.json)に配置し、Copilot Chatのエージェントモードツールピッカーからツールを有効にします。
JetBrains、Eclipse、Xcode。 Copilot Chatのエージェントモードツールピッカーを開き、Edit MCP configurationを選択し、開いたmcp.jsonに同じserversエントリを追加します。
Copilot coding agent(github.com上のクラウドエージェント)は、このサーバーには適していません。MCPサーバーを一時的なGitHub Actions環境内で実行するため、実行のたびに約2GBのTimeCopilotインストールを支払うことになり、ローカルデータファイルにアクセスできません。代わりにエディターから使用してください。
ツール
ツール | 目的 | 戻り値 |
| CSV/Parquetを読み取り、 | JSONサマリー + SHA-256 |
| モデルファミリー選択のための系列ごとの特徴量 | MarkdownテーブルまたはJSON、行数制限 |
| ここで実際にインポートされるモデルを調査 |
|
| モデル間のローリング起点比較 | メトリックテーブル、ランキング、parquetパス |
| 予測区間付きでフィットと予測 | Parquetパス + 制限付きプレビュー |
| 交差検証された区間フラグ付け | カウント、制限付きフラグリスト、parquetパス |
| セッションを再実行可能なマニフェストに固定 | マニフェストパス |
| すべてのステップを読み取り可能なレポートとしてレンダリング | HTMLまたはMarkdownパス |
2つのtsf_export_*ツールを除くすべては読み取り専用としてマークされています。ここでは何も削除しないため、~/.tslab-mcp/runsのクリーンアップはあなたの責任であり、エージェントの責任ではありません。
セッションの開始
ツールは順序を強制しないため、開始プロンプトが8つの呼び出し可能な関数を分析に変えるものです。次のようなものがうまく機能します:
tslabツールを使用して、
/Users/me/data/deposits.csvの系列を12ヶ月先まで予測してください。この順序で作業し、各ステップで推論を示してください:
ファイルをロードし、見つけたものを教えてください — 系列数、頻度、ギャップや欠損値の有無。
特徴量を説明し、それらがどのモデルファミリーを支持するか、そしてその理由を述べてください。
提案する前に、実際にインストールされているモデルを確認してください。
ショートリストをSeasonalNaiveベースラインに対して4つのウィンドウで交差検証してください。今のところ統計モデルのみ。
勝者で予測し、80%と95%の区間を付けてください。
実行マニフェストとHTMLレポートをエクスポートし、モデル選択の根拠をノートに記入してください:何を選んだか、メトリックテーブルが何を示したか、何を却下したか。
結果を要約し、parquetパスを教えてください — フレーム全体をチャットに貼り付けないでください。
そのプロンプトで実際に機能している4つの要素:
絶対パス。 相対パスはサーバーのワーキングディレクトリに対して解決されます。これはMCPクライアントが選択し、通常は予測できません。
決定に一致する地平線。
hは予測と、各CVウィンドウが消費する履歴量の両方を決定します。12ヶ月ステップは1年の計画であり、任意のデフォルトではありません。「今のところ統計モデルのみ。」 これがないと、エージェントはファンデーションモデルに手を伸ばし、
AutoETSが数秒で解決するはずの質問に答えるために重みのダウンロードに数分を費やす可能性があります。安価なモデルが基準を設定したら、制限を解除してください。マニフェストノートに根拠を求めること。 チャットのトランスクリプトは使い捨てです。マニフェストは、誰かが再実行して監査できる部分です。推論が会話にしか存在しない場合、それは事実上失われます。
何をしたいかがわかっている場合の、より短い開始プロンプト:
/Users/me/data/sales.parquetをロードし、特徴量を説明してください。まだ予測しないでください — まず何を扱っているのか見たいのです。
ロードされた
depositsハンドルに対して、SeasonalNaive、AutoETS、AutoARIMAをh=12で6つのウィンドウにわたって比較し、ベースラインを十分に上回って追加の複雑さに見合うものがあるかどうかを教えてください。
統計のみの呼び出しは数秒で応答します。ファンデーションモデルを指定した最初の呼び出しは、他のことをする前にTimeCopilotのインポートに約30秒かかります。その一時停止は予想されるものであり、ハングではありません。また、foundationエクストラがインストールされ、リクエストが実際にファンデーションモデルに手を伸ばした場合にのみ発生します。
実践セッション
Nixtlaロング形式のCSVから開始:
unique_id,ds,y
branch_01,2018-01-01,1043.2
branch_01,2018-02-01,1102.7
...1. ロードします。 パネルはサーバープロセスに残ります。ハンドルがセッションが保持するすべてです。
{"handle": "deposits", "n_series": 12, "n_obs": 864, "freq": "MS",
"start": "2018-01-01T00:00:00", "end": "2023-12-01T00:00:00",
"obs_per_series": {"min": 72, "median": 72, "max": 72},
"n_missing_y": 0, "sha256": "9f2c…"}2. 説明します。 これらはあなたが推論する数値です。
| id | n | mean | cv | %zero | trend | seasonal | acf1(diff) |
|-----------|----|--------|-------|-------|-------|----------|------------|
| branch_01 | 72 | 1180.4 | 0.112 | 0.0 | 0.83 | 0.62 | -0.31 |高い季節性の強さと明確なトレンドは、ナイーブベースラインよりもAutoETSとAutoARIMAを支持します。高い%zeroは、代わりにADIDAやCrostonClassicを支持したでしょう。
seasonalはSTLの強さです。トレンドが除去された後に残るものに対して測定された季節成分です。そのため、成長する系列でも季節性を正直に報告します。約0.3〜0.5のノイズフロアを持ちます。この帯域のスコアは「証拠なし」を意味し、「弱い季節性」ではありません。
3. インストールされているものを確認します tsf_list_modelsを使用して、このマシンで実行できないモデルを提案しないようにします。
4. 候補を交差検証します — 常にSeasonalNaiveを含めます。それを打ち負かせないモデルはデプロイする価値がないためです:
{"kind": "cross_validation", "models": ["SeasonalNaive", "AutoETS", "AutoARIMA"],
"h": 12, "n_windows": 4, "seasonality_used_for_mase": 12,
"metrics": {"mase": {"SeasonalNaive": 1.0, "AutoETS": 0.71, "AutoARIMA": 0.68}},
"ranking": {"mase": ["AutoARIMA", "AutoETS", "SeasonalNaive"]},
"artifact": "~/.tslab-mcp/runs/cv_deposits_3f1a9c02.parquet"}5. 勝者で予測します。 完全なフレームはparquetに出力されます。レスポンスにはパス、列、および短いプレビューが含まれます。
6. 実行とレポートをエクスポートします。 なぜをノートに書き留めてください。それは会話よりも長く生き残る推論の唯一の部分です:
{"manifest": "~/.tslab-mcp/runs/manifest_deposits_77b0e415.json", "n_runs": 3,
"kinds": ["cross_validation", "forecast"]}マニフェストには、ソースパスとハッシュ、頻度、すべての呼び出しとその引数およびアーティファクトパス、実際にインストールされているものの固定バージョン(statsforecast、pandas、Pythonは常に、foundationエクストラが含まれている場合はTimeCopilotとtorchも)、およびあなたのノートが含まれます。サーバーが停止していても、数値を再現するのに十分です。
tsf_export_reportは、同じマニフェストを人が読めるものに変換します。特徴量、最良順に並べられたメトリックテーブル、予測、異常、環境を、発生順に表示します:
{"report": "~/.tslab-mcp/runs/report_deposits_5c31d0a7.html",
"format": "html", "n_steps": 3,
"steps": ["features", "cross_validation", "forecast"]}レポートはマニフェストの純粋関数です。parquetを読み取らず、モデルも呼び出しません。そのため、manifest_pathを指定したtsf_export_reportは、何もロードせずに数ヶ月前の実行を再レンダリングします。HTMLは独自のCSSを埋め込み、外部スクリプト、スタイルシート、フォントを参照しないため、オフラインでも正しく開きます。
設計
4つの不変条件と、それらが存在する理由:
ハンドルであり、データフレームではない。 1つのクロスバリデーションフレームは n_series × h × n_windows × n_models 行になる。これをツール結果にシリアライズすると、最初の呼び出しでセッションのコンテキストを使い果たし、それ以降のすべてのターンを悪化させる。ツールはハンドルを受け取り、サマリー、集計、ファイルパスを返す。すべての一括パスには上限が設定され、省略した内容が報告されるため、セッションは再度問い合わせるのではなく、parquet を読むべきだと判断できる。
ブロッキング処理はイベントループに一切触れない。 大規模なパネルに対する複数モデルのクロスバリデーションは数分の CPU 時間を要する。すべてのツール本体は anyio.to_thread.run_sync 経由でディスパッチされる同期クロージャであるため、stdio トランスポートは応答し続け、クライアントが実行中にサーバーを切断することはない。
環境は推測されるのではなく、検出される。 モデルは遅延インポートされ、プローブされる。存在が前提とされることはない。tsf_list_models はここで実際に解決されたものを報告する。したがって、追加パッケージなしで Chronos を要求すると、実行開始から10分後にトレースバックが返る代わりに、不足している追加パッケージの名前を挙げたメッセージが返る。
バックエンドはリクエスト内容に応じて選択される。モデルがすべて統計モデルであるリクエストは statsforecast で実行され、事前学習済みモデルを必要とするリクエストだけが TimeCopilot を使用する。したがって、統計モデルの実行は torch をインポートすることがなく、どちらの場合でもサーバーは即座に起動する。
statsforecast は意図的にデフォルトの n_jobs=1 のままにされている。その並列モードは、エントリモジュールを再インポートするワーカープロセスを生成するが、MCP サーバー内では速度ではなく、競合と stdout の危険をもたらす。
マニフェストは記録の成果物である。 会話中の散文はコメントに過ぎない。マニフェストは、6か月後に誰かが再実行するものであり、レビュアーがどのモデルがどのような基準で比較されたかを確認するために読むものである。
Python バージョン
TimeCopilot はいくつかのモデルをインタープリタのバージョンに依存させており、Python < 3.13 では tabpfn-time-series をピン留めする。これにより、pandas は 2.2 未満に制限される。
Python | モデル | pandas |
3.13 |
| ≥ 2.2 |
3.10–3.12 |
| < 2.2 |
3.13 が推奨ターゲットである。いずれにせよ、tsf_list_models は実際に解決されたものを報告し、解決されなかったものについてはその理由も報告する。
開発
uv sync --all-groups
uv run pytest # fast suite
uv run pytest -m slow # exercises TimeCopilot; slower, no weight downloads
uv run ruff check src tests
uv run mypyMCP Inspector でツールサーフェスを検査する:
npx @modelcontextprotocol/inspector uv run tslab-mcpライセンス
MIT
Available Tools
8 toolstsf_cross_validateARead-onlyIdempotent
Compare models by rolling-origin cross-validation.
This is the tool that replaces guesswork about model choice: it produces the evidence, you read the table and decide. Always include SeasonalNaive as the baseline -- a model that cannot beat it is not worth deploying.
Returns a per-model metric table aggregated over series and windows, a best-first ranking per metric, and the parquet path holding every per-window prediction. LONG-RUNNING: seconds for statistical models, many minutes for foundation models on a large panel. Start with statistical models on the real horizon before reaching for anything pretrained.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, idempotent, non-destructive behavior. The description adds crucial behavioral context: runtime warning ('LONG-RUNNING'), output description (aggregated table, ranking, parquet path), and implicitly that it is safe but compute-intensive. This goes well beyond the annotations and aids agent expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, purposeful paragraphs. The first sentence immediately states the tool's function. Each subsequent section (usage advice, output details, runtime warning) earns its place with no redundancy. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (cross-validation, multiple models, windows, output schema exists), the description covers purpose, usage, baseline recommendation, output contents (aggregated metrics, rankings, prediction parquet), and runtime behavior. It is sufficiently complete for an agent to understand when and how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The nested input schema (CrossValidateInput) already provides parameter descriptions (e.g., models, metrics). The description adds high-level advice (like horizon matching the decision, baseline recommendation) but no new parameter-level semantics beyond what the schema offers. Schema coverage is effectively high despite the 0% top-level stat, so the description's incremental value here is moderate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource combination ('Compare models by rolling-origin cross-validation') and clearly distinguishes this tool from siblings like tsf_forecast (single model forecast) and tsf_list_models (model names). It states it replaces guesswork about model choice, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Always include SeasonalNaive as the baseline' and 'Start with statistical models on the real horizon before reaching for anything pretrained.' It also explains that the tool produces evidence for model selection. However, it does not explicitly state when not to use this tool (e.g., for final forecasts) or name alternatives, slightly reducing completeness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_describe_seriesARead-onlyIdempotent
Compute the per-series features that decide which model family to try.
Returns length, mean, sd, coefficient of variation, share of zeros, trend strength (R-squared against time), seasonal strength (variance explained by the period means), and lag-1 autocorrelation of the differenced series.
Read it as evidence, not as an answer: high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models (ADIDA, IMAPA, CrostonClassic); high cv with low structure argues for keeping expectations modest. Cheap -- returns in under a second.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and nondestructive nature. The description adds 'Cheap -- returns in under a second' and 'Read it as evidence, not as an answer,' providing behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with five sentences, front-loaded with purpose, and every sentence adds value. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema (not shown but present), the description covers the semantic meaning of the features and how to interpret them. It also provides cost and time estimates, making it self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has detailed descriptions for all three parameters (handle, max_series, response_format). The tool description focuses on output features and usage advice, not parameter details. Since schema coverage is high, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a clear verb-resource pair: 'Compute the per-series features that decide which model family to try.' It lists the specific features computed, distinguishing this diagnostic tool from siblings like tsf_forecast or tsf_load_series.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit decision rules: 'high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models...' and positions the tool as 'evidence, not an answer.' This clearly guides when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_detect_anomaliesARead-onlyIdempotent
Flag historical points that fall outside a cross-validated prediction interval.
The detector model defines what "expected" means, so pick one that fits the series: a weak detector flags its own errors rather than real anomalies. Run tsf_describe_series or tsf_cross_validate first.
Returns flagged counts per series, a capped list of flagged rows, and the parquet path with the full result.
LONG-RUNNING, and the default is the expensive one: leaving n_windows unset refits the model once per observation across the whole history, which takes minutes even for a statistical model. Pass n_windows (e.g. 12) unless you genuinely need every point tested.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses long-running nature and expensive default (n_windows unset refits per observation, taking minutes). Annotations (readOnlyHint, idempotentHint) are consistent; description adds critical performance context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact four paragraphs with clear structure: purpose, prerequisites, output summary, performance warning. No redundant sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given tool complexity (cross-validation, long-running, return includes parquet path and capped list), the description covers prerequisites, output, and performance. Output schema exists, so return values are adequately summarized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond schema by explaining the default slowness of n_windows and the risk of weak models. Schema already has clear descriptions, but description provides crucial usage context for these parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it flags historical points outside a cross-validated prediction interval, distinguishes from sibling tools (tsf_describe_series, tsf_cross_validate, etc.), and warns about weak detectors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises running tsf_describe_series or tsf_cross_validate first, warns against weak detectors, and gives concrete guidance on setting n_windows to avoid slow default behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_reportA
Render every step of the analysis as a report someone can read.
Covers the input and its hash, the features, each cross-validation with its metric table and ranking, the forecasts, any anomaly runs, and the pinned environment -- in the order they happened. HTML is self-contained, with no external stylesheet or script, so it opens correctly years later.
Call it after tsf_export_run at the end of an analysis. Pass a note: the report headlines it as the rationale, and a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.
Report from manifest_path instead of handle to re-render an older run --
it needs nothing but the manifest file.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (all hints false), so the description carries the burden. It discloses that HTML output is 'self-contained, with no external stylesheet or script' and that the report covers steps 'in the order they happened.' However, it does not clarify whether the tool writes a file to disk, returns the report content, or has other side effects. The output schema exists but is not described in the tool description, leaving behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a lead sentence defining the tool, then a list of contents, then usage order, then note advice, then alternative invocation. It is informative without being verbose. Minor inefficiency: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again' is slightly colorful but still earns its place. Could be tighter, but overall effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 optional parameters and an output schema (which handles return value documentation), the description covers its purpose, contents, usage order, and parameter trade-offs. It does not explain what happens if both handle and manifest_path are provided (mutual exclusion handled by schema? not specified). It also assumes the agent knows tsf_export_run was called, which is implied by 'at the end of an analysis.' Overall adequately complete for an AI agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has detailed descriptions for all parameters (note, format, handle, manifest_path), so baseline is 3. The description goes beyond by explaining the semantic purpose of note ('headlines it as the rationale') and the trade-off between handle and manifest_path ('re-renders an older run -- it needs nothing but the manifest file'). This adds actionable context for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description starts with a specific verb ('Render every step of the analysis as a report') and enumerates the exact contents (input hash, features, cross-validation, forecasts, anomalies, pinned environment). This immediately distinguishes it from sibling tools like tsf_export_run (which exports run data) and tsf_forecast (which only forecasts). The purpose is unambiguous and comprehensive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states ordering: 'Call it after tsf_export_run at the end of an analysis.' Provides clear alternatives: 'Report from manifest_path instead of handle to re-render an older run.' Also advises on best practice for the note parameter: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.' This gives the agent concrete when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_runA
Write a JSON manifest of everything done to this handle.
Records the source path and SHA-256, the frequency, every call with its arguments and artifact paths, the pinned package versions, and your note. This is the artifact of record: your prose in the conversation is lost, this file is not. Write the note -- say which model you picked, what the metric table showed, and what you rejected.
Call it at the end of any analysis someone might have to defend or rerun. Writes a file, so it is not read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds valuable context: 'Writes a file, so it is not read-only' and lists everything included in the manifest. It also emphasizes that the note is the only place a reasoning trace survives, which is important behavioral insight. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: starts with the core purpose, then enumerates contents, gives usage advice, and closes with a note about file writing. It is not overly long; every sentence contributes. A slight trim could improve conciseness, but it remains clear and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description correctly omits return value details. It covers when to call, what the manifest contains, and the critical role of the note. The only minor gap is no mention of potential side effects (e.g., overwriting existing files), but overall it is sufficiently complete for a tool with good annotations and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already contains descriptions for both parameters: handle ('Handle whose run log should be pinned to a manifest') and note (detailed explanation of what to write). The tool description adds further guidance for the note, specifically: 'Write the note -- say which model you picked, what the metric table showed, and what you rejected.' This enhances the schema's descriptions, making it clear how to use the note parameter effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Write a JSON manifest of everything done to this handle' clearly states the verb (write a manifest) and the resource (handle). The description further details what the manifest includes (source path, SHA-256, calls, arguments, artifact paths, pinned package versions, note), differentiating it from sibling tools like 'tsf_export_report' or 'tsf_describe_series'. This makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to call it: 'Call it at the end of any analysis someone might have to defend or rerun.' This provides clear context for use. It does not explicitly state when not to use it or name alternatives, but for a specialized export tool the guidance is sufficient and well-placed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_forecastARead-onlyIdempotent
Fit on the full history and forecast h periods ahead with intervals.
Use after tsf_cross_validate has justified the model choice. The full forecast goes to parquet; the response carries the path, the column list, the row count and a small preview. Read the parquet for anything more -- raising max_preview_rows to dump the frame into the conversation is the one thing that reliably ruins a long session.
LONG-RUNNING for foundation models.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the bar is lowered. The description adds valuable context: the full forecast goes to parquet, the response carries path/column list/row count/preview, warns against raising max_preview_rows, and flags 'LONG-RUNNING for foundation models' — all beyond what annotations provide. No contradictions with annotations (readOnlyHint=true is consistent with generating forecasts without mutating data).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: a terse three-sentence explanation that front-loads the core action, then adds usage guidance and behavioral warnings. Every sentence adds essential information without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (forecasting with multiple models and intervals), the output schema exists, so return values don't need elaboration. The description covers the critical workflow (use after cross-validation), output format (parquet with preview), and a key gotcha (don't dump full frame). It lacks explicit error conditions or prerequisite checks, but for the scope, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema provides no parameter descriptions — so the description must compensate. Although the main description does not detail parameters, the parameter `max_preview_rows` receives meaningful context: 'Rows of the forecast to inline... read that instead of raising this.' Other parameters (handle, models, h, level) have descriptions in the schema via the JSON Schema, but since coverage is 0% (likely meaning no separate param list in the description), the main text does not clarify their meaning beyond what the schema already provides. Baseline 3 is appropriate as the description adds some context for max_preview_rows but not for others.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool fits on full history and forecasts h periods ahead with intervals. It uses specific verbs like 'fit' and 'forecast' and explicitly identifies the resource as the time series forecast. However, it does not directly differentiate from siblings like tsf_cross_validate, though the usage guideline addresses that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after tsf_cross_validate has justified the model choice,' providing clear sequencing context and an alternative (cross-validation). It does not mention when not to use it or list specific alternatives for other tasks like anomaly detection, but the primary usage guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_list_modelsARead-onlyIdempotent
Probe which models actually import in this environment.
Call this before cross-validating so you never propose a model that cannot run here. Returns {available, statistical, foundation, unavailable}, where each unavailable entry carries the real reason -- some models are gated on the Python version, not merely absent.
The first call imports TimeCopilot and can take ~30 seconds; later calls are instant.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds important behavioral details beyond that: the first call may take ~30 seconds to import TimeCopilot, later calls are instant; it returns structured output with real reasons for unavailability (including Python version gating). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and structured: first sentence states purpose, second gives usage guidance, third explains return structure, and fourth notes startup latency. Every sentence adds essential information, with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (the agent can rely on structured return type details), the description covers all necessary context: what the tool does, when to use it, its runtime behavior (latency), and the high-level shape of results. No gaps remain for a list/probe tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema contains descriptions for both parameters (family and include_unavailable) that are self-explanatory. The tool description does not add new parameter information—it only repeats that statistical models are cheap and always installed, which already appears in the schema. Schema coverage via inline descriptions is present, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Probe which models actually import in this environment' and specifies the tool's role in preventing proposal of non-running models. It clearly distinguishes itself from siblings like cross_validate by giving a precise pre-check use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contains an explicit directive: 'Call this before cross-validating so you never propose a model that cannot run here.' It also notes that statistical models are always installed, which helps with decision-making. No alternatives are listed, but the use case is crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_load_seriesARead-onlyIdempotent
Read a CSV or Parquet panel from disk and register it under a handle.
Call this first; every other tool takes the handle it returns. The file must be in Nixtla long format (unique_id, ds, y). Returns a compact JSON summary -- series count, inferred frequency, date range, missing values, and the SHA-256 of the source -- and nothing else: the data stays in the server so it never consumes your context.
Read the summary before choosing a horizon. If obs_per_series.min is small, a long horizon or many CV windows will not fit.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description reveals that data stays server-side ('never consumes your context') and that the return is a compact summary with specific fields. It also hints at the side effect of reusing a handle (replacing the panel). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five sentences, each serving a distinct purpose: purpose, ordering, format, return details, caution. Front-loaded with the core action. No superfluous words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's role as the entry point for a time series workflow, the description covers purpose, required file format, return value (with summary contents), and a concrete usage caution. With an output schema present, the lack of detailed return structure is acceptable. The description is complete enough for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides detailed descriptions for all three parameters. The tool description adds the critical constraint that the file must be in 'Nixtla long format (unique_id, ds, y)', which is not in the schema. This adds meaningful value beyond the schema, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action: 'Read a CSV or Parquet panel from disk and register it under a handle.' It distinguishes from sibling tools by explicitly saying 'Call this first; every other tool takes the handle it returns.' This makes the purpose unambiguous and contextually positioned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit ordering ('Call this first'), explains the handle's role in subsequent tools, and gives a practical caution about horizon choices based on the summary output. This equips the agent with clear when-to-use and how-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
tsf_cross_validate - First observed
tsf_describe_series - First observed
tsf_detect_anomalies - First observed
tsf_export_report - First observed
tsf_export_run - First observed
tsf_forecast - First observed
tsf_list_models - First observed
tsf_load_series
TDQS
Scored across 8 tools
Each tool has a distinct and well-defined purpose within the time series forecasting workflow: loading, describing, listing models, cross-validating, forecasting, detecting anomalies, and exporting results. There is no overlap or ambiguity between tools.
All tools follow a consistent pattern: the prefix 'tsf_' followed by a verb (and optional noun), all in snake_case. Examples include tsf_load_series, tsf_describe_series, tsf_cross_validate, and tsf_export_report. The naming is predictable and uniform.
With 8 tools, the server covers a complete analysis pipeline without excess. Each tool is necessary and corresponds to a clear step in the workflow, from data loading to report generation. The count is well-scoped for the domain.
The tool surface covers the full lifecycle of a typical time series analysis: load data, explore features, check available models, cross-validate, forecast, detect anomalies, and export manifests/reports. There are no obvious gaps; the workflow feels self-contained and actionable.
Maintenance
Related MCP Connectors
Probabilistic time-series forecasts from zero-shot foundation models: routed, single or ensembled.
1PredictOracle - 12 forecasting tools: time-series, scenario analysis, risk projections.
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Deterministic time tools for AI agents: timezone conversion, business-day math, cron interpretation.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnable any AI agent to forecast time-series data (e.g., sales, traffic) using Google's TimesFM or a zero-dependency statistical baseline.3Apache 2.0
- AlicenseAqualityDmaintenanceDeterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.17MIT
- AlicenseNot gradedqualityCmaintenanceEnables time-series analysis and forecasting through a structured tool catalogue, including data loading, quality repair, diagnostics, and forecasting with ARIMA, exponential smoothing, Chronos-2, Toto 2.0, and AutoML.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to run TimesFM-3 forecasting workflows locally, including joint multivariate forecasting with known future drivers, backtesting against a baseline, what-if scenario comparison, and historical anomaly detection. It exposes the studio's tools and bundled public and synthetic demo datasets over MCP.Apache 2.0