Skip to main content
Glama

tslab-mcp

一个MCP服务器,将确定性时间序列预测暴露为工具,让你的智能体成为推理引擎,而每个数字都来自普通、可复现的Python。

此包中不调用任何LLM。无需API密钥(除非你请求TimeGPT,它会调用Nixtla API)。

为什么

一些预测库自带一个智能体,它读取特征、选择模型,并通过LLM循环解释结果。从你自己的智能体中调用这样的库,相当于在智能体内部嵌套了另一个智能体——两次提示、两次计费、两个非确定性来源,以及一个不透明的中间层,使得模型选择理由无法审计。

因此,这里的控制权被反转了:预测库是工具,你的智能体才是进行推理的一方。它读取特征、论证模型家族、交叉验证候选模型,并将推理过程写入清单。过程中的每个数字都由库调用产生,你可以在没有LLM参与的情况下重新运行。

这种分离也体现在包本身的构建方式上。基础安装通过statsforecast运行十一个统计模型——AutoARIMA、AutoETS、Theta、CrostonClassic等:约340 MB,无需PyTorch,启动只需几秒。可选的foundation扩展增加了TimeCopilot的预训练模型——Chronos、Moirai、TimesFM、TiRex、Toto等——以及Prophet,用于统计基线不够用的情况。只指定统计模型的请求永远不会导入TimeCopilot或torch;即使指定一个基础模型的请求也会完全通过TimeCopilot运行,而TimeCopilot也包含了统计模型。无论哪种情况,tsf_list_models都会在你提交模型之前报告实际安装的内容。

Related MCP server: timeseries-mcp

安装

需要Python 3.10+(推荐3.13,参见Python版本)。

uvx tslab-mcp                     # run without installing
uv tool install tslab-mcp         # or install the CLI

基础安装通过statsforecast运行十一个统计模型:约340 MB,无需PyTorch,启动即完成。对于预训练的基础模型——Chronos、Moirai、TimesFM、Toto、TiRex——以及Prophet,添加扩展:

uvx --from 'tslab-mcp[foundation]' tslab-mcp

foundation扩展会拉取TimeCopilot,它带来了torch、transformers和lightning:首次安装约2 GB,第一次调用相关工具时大约需要30秒导入。两者都是一次性的,除非你请求需要它们的模型,否则无需付费。

从GitHub安装

uv和uvx都接受git URL代替包名,这会安装当前的main分支,无需等待发布:

uvx --from git+https://github.com/pedrobtz/tslab-mcp tslab-mcp
uv tool install git+https://github.com/pedrobtz/tslab-mcp        # or install the CLI

# with the foundation extra
uvx --from 'tslab-mcp[foundation] @ git+https://github.com/pedrobtz/tslab-mcp' tslab-mcp

对于非临时测试,请固定一个引用——否则分支头可能会在你不知情的情况下移动。使用commit可以正常工作;一旦有版本标签,使用版本标签也可以:

uv tool install "git+https://github.com/pedrobtz/tslab-mcp@136824c1cc2a"

从本地检出安装

git clone https://github.com/pedrobtz/tslab-mcp
cd tslab-mcp
uv sync                              # base
uv sync --extra foundation           # with the pretrained models
uv run tslab-mcp

配置

将服务器添加到你的MCP客户端配置中。不同客户端的配置文件不同——通常是项目根目录下的.mcp.json——但条目本身的格式相同:

{
  "mcpServers": {
    "tslab": {
      "command": "uvx",
      "args": ["tslab-mcp"],
      "env": {
        "TSLAB_MCP_HOME": "~/.tslab-mcp"
      }
    }
  }
}

TSLAB_MCP_HOME设置工件写入位置;默认为~/.tslab-mcp,运行输出存放在<home>/runs下。

传输方式仅限stdio,这是有意为之:假设你的数据是敏感的,永远不会离开机器。服务器不会发出出站请求,除了TimeCopilot为基础模型执行的模型权重下载,以及如果你特别请求TimeGPT时它调用的Nixtla API。

GitHub Copilot

Copilot从mcp.json文件中发现MCP服务器,并在agent模式下暴露其工具——这些工具不会出现在ask或edit模式中。

VS Code。 将服务器放在.vscode/mcp.json中以与仓库共享,或从命令面板运行MCP: Open User Configuration以将其保留在你自己的配置文件中,跨所有工作区使用。注意键是servers,而不是mcpServers:

{
  "servers": {
    "tslab": {
      "type": "stdio",
      "command": "uvx",
      "args": ["tslab-mcp"],
      "env": {
        "TSLAB_MCP_HOME": "${userHome}/.tslab-mcp"
      }
    }
  }
}

从本地检出安装时,将其指向工作树:

{
  "servers": {
    "tslab": {
      "type": "stdio",
      "command": "uv",
      "args": ["run", "--directory", "${workspaceFolder}", "tslab-mcp"]
    }
  }
}

然后:打开Chat,将模式选择器切换到Agent,使用Tools按钮确认八个tsf_*工具已列出并启用。MCP: List Servers显示服务器的状态及其日志,这是查看启动失败原因的地方。Copilot限制了同时激活的工具数量,因此如果你运行多个MCP服务器,可能需要取消选择一些以容纳全部八个工具。

Visual Studio。 相同的JSON格式,放在解决方案根目录下的.mcp.json中(或所有解决方案的%USERPROFILE%\.mcp.json),然后从Copilot Chat agent模式的工具选择器中启用工具。

JetBrains、Eclipse和Xcode。 打开Copilot Chat agent模式的工具选择器,选择Edit MCP configuration,将相同的servers条目添加到打开的mcp.json中。

Copilot coding agent(github.com上的云智能体)不适合此服务器:它在临时的GitHub Actions环境中运行你的MCP服务器,这意味着每次运行都要支付约2 GB的TimeCopilot安装费用,而且它无法访问本地数据文件。请从你的编辑器中使用它。

工具

工具

目的

返回

tsf_load_series

读取CSV/Parquet,验证unique_id/ds/y契约,推断频率,注册句柄

JSON摘要 + SHA-256

tsf_describe_series

每个序列的特征,用于选择模型家族

Markdown表格或JSON,行数限制

tsf_list_models

探测哪些模型实际在此处导入

{available, statistical, foundation, unavailable}

tsf_cross_validate

跨模型的滚动起点比较

指标表、排名、parquet路径

tsf_forecast

拟合和预测,带预测区间

Parquet路径 + 有限预览

tsf_detect_anomalies

交叉验证的区间标记

计数、限制标记列表、parquet路径

tsf_export_run

将会话固定为可重新运行的清单

清单路径

tsf_export_report

将每一步渲染为可读报告

HTML或Markdown路径

除了两个tsf_export_*工具外,所有工具都标记为只读;这里不会删除任何内容,因此清理~/.tslab-mcp/runs是你的责任,而不是智能体的。

开始一个会话

这些工具不强制顺序,因此开场提示是将八个可调用函数转化为分析的关键。类似这样的提示效果很好:

使用tslab工具预测/Users/me/data/deposits.csv中的序列,预测12个月。

按此顺序工作,并在每一步展示你的推理:

  1. 加载文件并告诉我你发现了什么——有多少序列、什么频率、是否有任何间隙或缺失值。

  2. 描述特征,并说明它们支持哪些模型家族,以及原因。

  3. 在提出任何模型之前,检查哪些模型实际已安装。

  4. 将你的候选列表与SeasonalNaive基线在4个窗口上进行交叉验证。目前仅使用统计模型。

  5. 使用胜出模型进行预测,带80%和95%的区间。

  6. 导出运行清单和HTML报告,并将模型选择理由放在注释中:你选择了什么、指标表显示了什么、以及你拒绝了什么。

总结结果并给我parquet路径——不要将整个数据框粘贴到聊天中。

该提示中有四个要点在发挥实际作用:

  • 绝对路径。 相对路径相对于服务器的工作目录解析,而工作目录由你的MCP客户端选择,你通常无法预测。

  • 与决策匹配的预测步长。 h驱动预测以及每个CV窗口消耗的历史数据量;12个月步长是一年的规划,而不是任意默认值。

  • "目前仅使用统计模型。" 没有这个限制,智能体可能会选择基础模型,花费几分钟下载权重来回答AutoETS几秒钟就能解决的问题。在廉价模型设定了基线后再解除限制。

  • 要求将推理过程放在清单注释中。 聊天记录是可丢弃的;清单才是可以重新运行和审计的部分。如果推理只存在于对话中,它实际上就丢失了。

当你明确需求时,更简短的开场提示:

加载/Users/me/data/sales.parquet并描述特征。先不要预测——我想先看看我们面对的是什么。

比较SeasonalNaive、AutoETS和AutoARIMA在已加载的deposits句柄上,在h=12时跨6个窗口,然后告诉我是否有任何模型比基线好到值得增加额外复杂度。

仅统计模型的调用在几秒内完成。第一个指定基础模型的调用在开始任何其他操作之前需要大约30秒导入TimeCopilot——这个暂停是预期的,不是挂起,并且只有在安装了foundation扩展且请求实际使用了基础模型时才会发生。

一个完整的会话示例

从Nixtla长格式的CSV开始:

unique_id,ds,y
branch_01,2018-01-01,1043.2
branch_01,2018-02-01,1102.7
...

1. 加载它。 数据面板保留在服务器进程中;句柄是会话携带的所有内容。

{"handle": "deposits", "n_series": 12, "n_obs": 864, "freq": "MS",
 "start": "2018-01-01T00:00:00", "end": "2023-12-01T00:00:00",
 "obs_per_series": {"min": 72, "median": 72, "max": 72},
 "n_missing_y": 0, "sha256": "9f2c…"}

2. 描述它。 这些是你进行推理的数字。

| id        | n  | mean   | cv    | %zero | trend | seasonal | acf1(diff) |
|-----------|----|--------|-------|-------|-------|----------|------------|
| branch_01 | 72 | 1180.4 | 0.112 | 0.0   | 0.83  | 0.62     | -0.31      |

高季节性强度和明显的趋势支持使用AutoETS和AutoARIMA而不是朴素基线;高%zero则会支持使用ADIDA或CrostonClassic。

seasonal是STL强度——在去除趋势后剩余的季节性成分——因此增长中的序列仍然诚实地报告其季节性。它带有大约0.3–0.5的噪声基底:该范围内的分数意味着"无证据",而不是"轻微季节性"。

3. 检查已安装的内容使用tsf_list_models,这样你永远不会提出这台机器无法运行的模型。

4. 交叉验证候选模型——始终包括SeasonalNaive,因为无法击败它的模型不值得部署:

{"kind": "cross_validation", "models": ["SeasonalNaive", "AutoETS", "AutoARIMA"],
 "h": 12, "n_windows": 4, "seasonality_used_for_mase": 12,
 "metrics": {"mase": {"SeasonalNaive": 1.0, "AutoETS": 0.71, "AutoARIMA": 0.68}},
 "ranking": {"mase": ["AutoARIMA", "AutoETS", "SeasonalNaive"]},
 "artifact": "~/.tslab-mcp/runs/cv_deposits_3f1a9c02.parquet"}

5. 使用胜出模型进行预测。 完整数据框写入parquet;响应包含路径、列和简短预览。

6. 导出运行和报告。 在注释中写下原因——这是你的推理中唯一比对话存活时间更长的部分:

{"manifest": "~/.tslab-mcp/runs/manifest_deposits_77b0e415.json", "n_runs": 3,
 "kinds": ["cross_validation", "forecast"]}

清单包含源路径和哈希、频率、每次调用及其参数和工件路径、实际安装的任何内容的固定版本——始终包括statsforecast、pandas和Python;如果安装了foundation扩展,还包括TimeCopilot和torch——以及你的注释。它足以在服务器停止的情况下重现数字。

tsf_export_report将相同的清单转化为人类可读的内容——特征、按最佳优先排序的指标表、预测、异常和环境,按发生顺序排列:

{"report": "~/.tslab-mcp/runs/report_deposits_5c31d0a7.html",
 "format": "html", "n_steps": 3,
 "steps": ["features", "cross_validation", "forecast"]}

报告是清单的纯函数:它不读取parquet也不调用模型,因此使用manifest_path的tsf_export_report可以在没有任何加载内容的情况下重新渲染几个月前的运行。HTML嵌入自己的CSS,不引用任何外部脚本、样式表或字体,因此离线时也能正常打开。

设计

四个不变性,以及它们存在的原因:

处理句柄,而非数据框。 单次交叉验证框架会生成 n_series × h × n_windows × n_models 行数据。若将其序列化为工具结果,首次调用就会耗尽会话上下文,导致后续每轮交互效果变差。工具应返回句柄、摘要、聚合结果和文件路径;所有批量路径都有上限并会报告遗漏内容,以便会话主动读取 Parquet 文件而非重复请求。

阻塞操作绝不触碰事件循环。 对大型面板数据集进行多模型交叉验证可能需要数分钟 CPU 耗时。每个工具体均为通过 anyio.to_thread.run_sync 派发的同步闭包,确保 stdio 传输持续响应,客户端不会在服务器运行中途断开连接。

环境信息通过探查而非假定获取。 模型采用惰性加载和探测机制,从不假定其存在。tsf_list_models 会报告当前环境实际解析到的内容;若请求 Chronos 但缺少额外依赖,则会返回提示安装依赖包的消息,而非在运行十分钟后抛出回溯错误。

后端根据请求内容自动选择:仅含统计模型的请求使用 statsforecast 运行,只有需要预训练模型的请求才会调用 TimeCopilot。因此统计模型运行永不导入 torch,服务器在任何情况下都能瞬间启动。

statsforecast 特意保持默认 n_jobs=1 设置。其并行模式会衍生工作进程并重新导入入口模块,这在 MCP 服务器内部不仅会引发资源竞争和标准输出风险,且无法带来速度提升。

清单是记录性产物。 对话中的说明性文字仅为补充说明。清单才是六个月后他人重新运行时的依据,也是评审者查看哪些模型在何种基准下被比较的依据。

Python 版本要求

TimeCopilot 根据解释器版本限制部分模型。在 Python < 3.13 时,它会固定使用 tabpfn-time-series,这将 pandas 限制在 2.2 以下。

Python

模型范围

pandas

3.13

除 TabPFN 和 Sundial 外的所有模型

≥ 2.2

3.10–3.12

增加 TabPFN、Sundial

< 2.2

推荐使用 3.13。无论如何,tsf_list_models 都会报告实际解析到的内容,并说明未解析项的原因。

开发

uv sync --all-groups
uv run pytest                  # fast suite
uv run pytest -m slow          # exercises TimeCopilot; slower, no weight downloads
uv run ruff check src tests
uv run mypy

使用 MCP Inspector 检查工具接口:

npx @modelcontextprotocol/inspector uv run tslab-mcp

许可证

MIT

Available Tools

8 tools
tsf_cross_validateA
Read-onlyIdempotent

Compare models by rolling-origin cross-validation.

This is the tool that replaces guesswork about model choice: it produces the evidence, you read the table and decide. Always include SeasonalNaive as the baseline -- a model that cannot beat it is not worth deploying.

Returns a per-model metric table aggregated over series and windows, a best-first ranking per metric, and the parquet path holding every per-window prediction. LONG-RUNNING: seconds for statistical models, many minutes for foundation models on a large panel. Start with statistical models on the real horizon before reaching for anything pretrained.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly, idempotent, non-destructive behavior. The description adds crucial behavioral context: runtime warning ('LONG-RUNNING'), output description (aggregated table, ranking, parquet path), and implicitly that it is safe but compute-intensive. This goes well beyond the annotations and aids agent expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short, purposeful paragraphs. The first sentence immediately states the tool's function. Each subsequent section (usage advice, output details, runtime warning) earns its place with no redundancy. It is front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (cross-validation, multiple models, windows, output schema exists), the description covers purpose, usage, baseline recommendation, output contents (aggregated metrics, rankings, prediction parquet), and runtime behavior. It is sufficiently complete for an agent to understand when and how to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The nested input schema (CrossValidateInput) already provides parameter descriptions (e.g., models, metrics). The description adds high-level advice (like horizon matching the decision, baseline recommendation) but no new parameter-level semantics beyond what the schema offers. Schema coverage is effectively high despite the 0% top-level stat, so the description's incremental value here is moderate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource combination ('Compare models by rolling-origin cross-validation') and clearly distinguishes this tool from siblings like tsf_forecast (single model forecast) and tsf_list_models (model names). It states it replaces guesswork about model choice, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: 'Always include SeasonalNaive as the baseline' and 'Start with statistical models on the real horizon before reaching for anything pretrained.' It also explains that the tool produces evidence for model selection. However, it does not explicitly state when not to use this tool (e.g., for final forecasts) or name alternatives, slightly reducing completeness.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_describe_seriesA
Read-onlyIdempotent

Compute the per-series features that decide which model family to try.

Returns length, mean, sd, coefficient of variation, share of zeros, trend strength (R-squared against time), seasonal strength (variance explained by the period means), and lag-1 autocorrelation of the differenced series.

Read it as evidence, not as an answer: high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models (ADIDA, IMAPA, CrostonClassic); high cv with low structure argues for keeping expectations modest. Cheap -- returns in under a second.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and nondestructive nature. The description adds 'Cheap -- returns in under a second' and 'Read it as evidence, not as an answer,' providing behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with five sentences, front-loaded with purpose, and every sentence adds value. No redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema (not shown but present), the description covers the semantic meaning of the features and how to interpret them. It also provides cost and time estimates, making it self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already has detailed descriptions for all three parameters (handle, max_series, response_format). The tool description focuses on output features and usage advice, not parameter details. Since schema coverage is high, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a clear verb-resource pair: 'Compute the per-series features that decide which model family to try.' It lists the specific features computed, distinguishing this diagnostic tool from siblings like tsf_forecast or tsf_load_series.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit decision rules: 'high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models...' and positions the tool as 'evidence, not an answer.' This clearly guides when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_detect_anomaliesA
Read-onlyIdempotent

Flag historical points that fall outside a cross-validated prediction interval.

The detector model defines what "expected" means, so pick one that fits the series: a weak detector flags its own errors rather than real anomalies. Run tsf_describe_series or tsf_cross_validate first.

Returns flagged counts per series, a capped list of flagged rows, and the parquet path with the full result.

LONG-RUNNING, and the default is the expensive one: leaving n_windows unset refits the model once per observation across the whole history, which takes minutes even for a statistical model. Pass n_windows (e.g. 12) unless you genuinely need every point tested.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses long-running nature and expensive default (n_windows unset refits per observation, taking minutes). Annotations (readOnlyHint, idempotentHint) are consistent; description adds critical performance context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact four paragraphs with clear structure: purpose, prerequisites, output summary, performance warning. No redundant sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given tool complexity (cross-validation, long-running, return includes parquet path and capped list), the description covers prerequisites, output, and performance. Output schema exists, so return values are adequately summarized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Adds meaning beyond schema by explaining the default slowness of n_windows and the risk of weak models. Schema already has clear descriptions, but description provides crucial usage context for these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it flags historical points outside a cross-validated prediction interval, distinguishes from sibling tools (tsf_describe_series, tsf_cross_validate, etc.), and warns about weak detectors.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises running tsf_describe_series or tsf_cross_validate first, warns against weak detectors, and gives concrete guidance on setting n_windows to avoid slow default behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_export_reportA

Render every step of the analysis as a report someone can read.

Covers the input and its hash, the features, each cross-validation with its metric table and ranking, the forecasts, any anomaly runs, and the pinned environment -- in the order they happened. HTML is self-contained, with no external stylesheet or script, so it opens correctly years later.

Call it after tsf_export_run at the end of an analysis. Pass a note: the report headlines it as the rationale, and a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.

Report from manifest_path instead of handle to re-render an older run -- it needs nothing but the manifest file.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (all hints false), so the description carries the burden. It discloses that HTML output is 'self-contained, with no external stylesheet or script' and that the report covers steps 'in the order they happened.' However, it does not clarify whether the tool writes a file to disk, returns the report content, or has other side effects. The output schema exists but is not described in the tool description, leaving behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a lead sentence defining the tool, then a list of contents, then usage order, then note advice, then alternative invocation. It is informative without being verbose. Minor inefficiency: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again' is slightly colorful but still earns its place. Could be tighter, but overall effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 optional parameters and an output schema (which handles return value documentation), the description covers its purpose, contents, usage order, and parameter trade-offs. It does not explain what happens if both handle and manifest_path are provided (mutual exclusion handled by schema? not specified). It also assumes the agent knows tsf_export_run was called, which is implied by 'at the end of an analysis.' Overall adequately complete for an AI agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has detailed descriptions for all parameters (note, format, handle, manifest_path), so baseline is 3. The description goes beyond by explaining the semantic purpose of note ('headlines it as the rationale') and the trade-off between handle and manifest_path ('re-renders an older run -- it needs nothing but the manifest file'). This adds actionable context for parameter selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description starts with a specific verb ('Render every step of the analysis as a report') and enumerates the exact contents (input hash, features, cross-validation, forecasts, anomalies, pinned environment). This immediately distinguishes it from sibling tools like tsf_export_run (which exports run data) and tsf_forecast (which only forecasts). The purpose is unambiguous and comprehensive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states ordering: 'Call it after tsf_export_run at the end of an analysis.' Provides clear alternatives: 'Report from manifest_path instead of handle to re-render an older run.' Also advises on best practice for the note parameter: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.' This gives the agent concrete when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_export_runA

Write a JSON manifest of everything done to this handle.

Records the source path and SHA-256, the frequency, every call with its arguments and artifact paths, the pinned package versions, and your note. This is the artifact of record: your prose in the conversation is lost, this file is not. Write the note -- say which model you picked, what the metric table showed, and what you rejected.

Call it at the end of any analysis someone might have to defend or rerun. Writes a file, so it is not read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds valuable context: 'Writes a file, so it is not read-only' and lists everything included in the manifest. It also emphasizes that the note is the only place a reasoning trace survives, which is important behavioral insight. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: starts with the core purpose, then enumerates contents, gives usage advice, and closes with a note about file writing. It is not overly long; every sentence contributes. A slight trim could improve conciseness, but it remains clear and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, the description correctly omits return value details. It covers when to call, what the manifest contains, and the critical role of the note. The only minor gap is no mention of potential side effects (e.g., overwriting existing files), but overall it is sufficiently complete for a tool with good annotations and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already contains descriptions for both parameters: handle ('Handle whose run log should be pinned to a manifest') and note (detailed explanation of what to write). The tool description adds further guidance for the note, specifically: 'Write the note -- say which model you picked, what the metric table showed, and what you rejected.' This enhances the schema's descriptions, making it clear how to use the note parameter effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence 'Write a JSON manifest of everything done to this handle' clearly states the verb (write a manifest) and the resource (handle). The description further details what the manifest includes (source path, SHA-256, calls, arguments, artifact paths, pinned package versions, note), differentiating it from sibling tools like 'tsf_export_report' or 'tsf_describe_series'. This makes the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells when to call it: 'Call it at the end of any analysis someone might have to defend or rerun.' This provides clear context for use. It does not explicitly state when not to use it or name alternatives, but for a specialized export tool the guidance is sufficient and well-placed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_forecastA
Read-onlyIdempotent

Fit on the full history and forecast h periods ahead with intervals.

Use after tsf_cross_validate has justified the model choice. The full forecast goes to parquet; the response carries the path, the column list, the row count and a small preview. Read the parquet for anything more -- raising max_preview_rows to dump the frame into the conversation is the one thing that reliably ruins a long session.

LONG-RUNNING for foundation models.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the bar is lowered. The description adds valuable context: the full forecast goes to parquet, the response carries path/column list/row count/preview, warns against raising max_preview_rows, and flags 'LONG-RUNNING for foundation models' — all beyond what annotations provide. No contradictions with annotations (readOnlyHint=true is consistent with generating forecasts without mutating data).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise: a terse three-sentence explanation that front-loads the core action, then adds usage guidance and behavioral warnings. Every sentence adds essential information without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (forecasting with multiple models and intervals), the output schema exists, so return values don't need elaboration. The description covers the critical workflow (use after cross-validation), output format (parquet with preview), and a key gotcha (don't dump full frame). It lacks explicit error conditions or prerequisite checks, but for the scope, it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the schema provides no parameter descriptions — so the description must compensate. Although the main description does not detail parameters, the parameter `max_preview_rows` receives meaningful context: 'Rows of the forecast to inline... read that instead of raising this.' Other parameters (handle, models, h, level) have descriptions in the schema via the JSON Schema, but since coverage is 0% (likely meaning no separate param list in the description), the main text does not clarify their meaning beyond what the schema already provides. Baseline 3 is appropriate as the description adds some context for max_preview_rows but not for others.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool fits on full history and forecasts h periods ahead with intervals. It uses specific verbs like 'fit' and 'forecast' and explicitly identifies the resource as the time series forecast. However, it does not directly differentiate from siblings like tsf_cross_validate, though the usage guideline addresses that.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use after tsf_cross_validate has justified the model choice,' providing clear sequencing context and an alternative (cross-validation). It does not mention when not to use it or list specific alternatives for other tasks like anomaly detection, but the primary usage guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_list_modelsA
Read-onlyIdempotent

Probe which models actually import in this environment.

Call this before cross-validating so you never propose a model that cannot run here. Returns {available, statistical, foundation, unavailable}, where each unavailable entry carries the real reason -- some models are gated on the Python version, not merely absent.

The first call imports TimeCopilot and can take ~30 seconds; later calls are instant.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds important behavioral details beyond that: the first call may take ~30 seconds to import TimeCopilot, later calls are instant; it returns structured output with real reasons for unavailability (including Python version gating). No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and structured: first sentence states purpose, second gives usage guidance, third explains return structure, and fourth notes startup latency. Every sentence adds essential information, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (the agent can rely on structured return type details), the description covers all necessary context: what the tool does, when to use it, its runtime behavior (latency), and the high-level shape of results. No gaps remain for a list/probe tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema contains descriptions for both parameters (family and include_unavailable) that are self-explanatory. The tool description does not add new parameter information—it only repeats that statistical models are cheap and always installed, which already appears in the schema. Schema coverage via inline descriptions is present, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Probe which models actually import in this environment' and specifies the tool's role in preventing proposal of non-running models. It clearly distinguishes itself from siblings like cross_validate by giving a precise pre-check use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description contains an explicit directive: 'Call this before cross-validating so you never propose a model that cannot run here.' It also notes that statistical models are always installed, which helps with decision-making. No alternatives are listed, but the use case is crystal clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_load_seriesA
Read-onlyIdempotent

Read a CSV or Parquet panel from disk and register it under a handle.

Call this first; every other tool takes the handle it returns. The file must be in Nixtla long format (unique_id, ds, y). Returns a compact JSON summary -- series count, inferred frequency, date range, missing values, and the SHA-256 of the source -- and nothing else: the data stays in the server so it never consumes your context.

Read the summary before choosing a horizon. If obs_per_series.min is small, a long horizon or many CV windows will not fit.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, idempotent, non-destructive), the description reveals that data stays server-side ('never consumes your context') and that the return is a compact summary with specific fields. It also hints at the side effect of reusing a handle (replacing the panel). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five sentences, each serving a distinct purpose: purpose, ordering, format, return details, caution. Front-loaded with the core action. No superfluous words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's role as the entry point for a time series workflow, the description covers purpose, required file format, return value (with summary contents), and a concrete usage caution. With an output schema present, the lack of detailed return structure is acceptable. The description is complete enough for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides detailed descriptions for all three parameters. The tool description adds the critical constraint that the file must be in 'Nixtla long format (unique_id, ds, y)', which is not in the schema. This adds meaningful value beyond the schema, justifying a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action: 'Read a CSV or Parquet panel from disk and register it under a handle.' It distinguishes from sibling tools by explicitly saying 'Call this first; every other tool takes the handle it returns.' This makes the purpose unambiguous and contextually positioned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit ordering ('Call this first'), explains the handle's role in subsequent tools, and gives a practical caution about horizon choices based on the summary output. This equips the agent with clear when-to-use and how-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedtsf_cross_validate
    • First observedtsf_describe_series
    • First observedtsf_detect_anomalies
    • First observedtsf_export_report
    • First observedtsf_export_run
    • First observedtsf_forecast
    • First observedtsf_list_models
    • First observedtsf_load_series

TDQS

A4.6/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a distinct and well-defined purpose within the time series forecasting workflow: loading, describing, listing models, cross-validating, forecasting, detecting anomalies, and exporting results. There is no overlap or ambiguity between tools.

Naming Consistency5/5

All tools follow a consistent pattern: the prefix 'tsf_' followed by a verb (and optional noun), all in snake_case. Examples include tsf_load_series, tsf_describe_series, tsf_cross_validate, and tsf_export_report. The naming is predictable and uniform.

Tool Count5/5

With 8 tools, the server covers a complete analysis pipeline without excess. Each tool is necessary and corresponds to a clear step in the workflow, from data loading to report generation. The count is well-scoped for the domain.

Completeness5/5

The tool surface covers the full lifecycle of a typical time series analysis: load data, explore features, check available models, cross-validate, forecast, detect anomalies, and export manifests/reports. There are no obvious gaps; the workflow feels self-contained and actionable.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Deterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.
    17
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables time-series analysis and forecasting through a structured tool catalogue, including data loading, quality repair, diagnostics, and forecasting with ARIMA, exponential smoothing, Chronos-2, Toto 2.0, and AutoML.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to run TimesFM-3 forecasting workflows locally, including joint multivariate forecasting with known future drivers, backtesting against a baseline, what-if scenario comparison, and historical anomaly detection. It exposes the studio's tools and bundled public and synthetic demo datasets over MCP.
    Apache 2.0