tslab-mcp
tslab-mcp
시계열 예측을 도구로 노출하는 MCP 서버로, 모든 예측값은 일반적이고 재현 가능한 Python에서 비롯됩니다. 에이전트는 추론 엔진 역할을 합니다.
이 패키지 어디에서도 LLM을 호출하지 않습니다. API 키가 필요하지 않습니다(단, Nixtla API를 호출하는 TimeGPT를 요청하는 경우는 제외).
이유
일부 예측 라이브러리에는 LLM이 내장된 에이전트가 포함되어 있어, 특징을 읽고 모델을 선택한 후 결과를 설명합니다. 이러한 라이브러리를 자체 에이전트에서 호출하면 에이전트 안에 에이전트가 중첩되어, 두 번의 프롬프트, 두 번의 비용, 두 가지 비결정적 요소, 그리고 모델 선택 근거를 감사할 수 없게 만드는 불투명한 중간 계층이 생깁니다.
따라서 여기서는 제어가 역전됩니다. 예측 라이브러리는 도구가 되고, 에이전트는 추론합니다. 특징을 읽고, 모델군에 대해 논쟁하며, 후보들을 교차 검증하고, 그 근거를 매니페스트에 기록합니다. 경로에 LLM 없이도 재실행할 수 있는 라이브러리 호출로 모든 숫자가 생성됩니다.
이러한 분리는 패키지 자체의 구성 방식에도 반영됩니다. 기본 설치는 statsforecast를 통해 AutoARIMA, AutoETS, Theta, CrostonClassic 등 11개의 통계 모델을 실행합니다. 약 340MB, PyTorch 없이도 몇 초 안에 시작됩니다. 선택적 foundation 확장 기능은 TimeCopilot의 사전 훈련된 모델(Chronos, Moirai, TimesFM, TiRex, Toto 등)과 Prophet을 추가하여, 통계적 기준선만으로 충분하지 않을 때 사용합니다. 통계 모델만 지정한 요청은 TimeCopilot이나 torch를 가져오지 않으며, 하나의 기반 모델이라도 포함된 요청은 TimeCopilot을 통해 완전히 실행되며, TimeCopilot은 통계 모델도 함께 제공합니다. 어느 쪽이든 tsf_list_models는 모델을 선택하기 전에 실제로 설치된 모델을 보고합니다.
Related MCP server: timeseries-mcp
설치
Python 3.10+ 필요 (3.13 권장, Python 버전 참조).
uvx tslab-mcp # run without installing
uv tool install tslab-mcp # or install the CLI기본 설치는 statsforecast를 통해 11개의 통계 모델을 실행합니다. 약 340MB, PyTorch 없이 즉시 시작됩니다. 사전 훈련된 기반 모델(Chronos, Moirai, TimesFM, Toto, TiRex) 및 Prophet을 추가하려면 확장 기능을 설치하세요:
uvx --from 'tslab-mcp[foundation]' tslab-mcp
foundation확장 기능은 TimeCopilot을 가져오며, 여기에는 torch, transformers 및 lightning이 포함됩니다. 첫 설치 시 약 2GB, 처음 도구 호출 시 약 30초간 임포트가 발생합니다. 둘 다 일회성이며, 해당 모델이 필요하지 않으면 비용이 발생하지 않습니다.
GitHub에서 설치
uv와 uvx 모두 패키지 이름 대신 git URL을 허용하며, 릴리스를 기다리지 않고 현재 main을 설치합니다:
uvx --from git+https://github.com/pedrobtz/tslab-mcp tslab-mcp
uv tool install git+https://github.com/pedrobtz/tslab-mcp # or install the CLI
# with the foundation extra
uvx --from 'tslab-mcp[foundation] @ git+https://github.com/pedrobtz/tslab-mcp' tslab-mcp테스트가 아닌 경우에는 ref를 고정하세요. 그렇지 않으면 브랜치 헤드가 이동할 수 있습니다. 현재는 커밋이 작동하며, 버전 태그가 생성되면 태그도 사용 가능합니다:
uv tool install "git+https://github.com/pedrobtz/tslab-mcp@136824c1cc2a"체크아웃에서 설치
git clone https://github.com/pedrobtz/tslab-mcp
cd tslab-mcp
uv sync # base
uv sync --extra foundation # with the pretrained models
uv run tslab-mcp구성
MCP 클라이언트의 설정에 서버를 추가하세요. 파일은 클라이언트마다 다르지만(주로 프로젝트 루트의 .mcp.json), 항목 자체는 동일한 형식입니다:
{
"mcpServers": {
"tslab": {
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "~/.tslab-mcp"
}
}
}
}TSLAB_MCP_HOME은 아티팩트가 기록되는 위치를 설정합니다. 기본값은 ~/.tslab-mcp이며, 실행 결과는 <home>/runs에 저장됩니다.
전송 방식은 stdio만 사용합니다. 데이터가 민감하다고 가정하고 절대 머신을 벗어나지 않습니다. 서버는 TimeCopilot이 기반 모델을 위한 가중치 다운로드, 그리고 TimeGPT를 요청할 경우 Nixtla API 호출을 제외하고는 외부 요청을 보내지 않습니다.
GitHub Copilot
Copilot은 mcp.json 파일에서 MCP 서버를 발견하고 에이전트 모드에서 도구를 노출합니다. 도구는 ask 또는 edit 모드에서는 나타나지 않습니다.
VS Code. 서버를 .vscode/mcp.json에 넣어 리포지토리와 공유하거나, 명령 팔레트에서 MCP: Open User Configuration을 실행하여 모든 워크스페이스에서 개인 프로필에 유지하세요. 키는 mcpServers가 아닌 servers입니다:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "${userHome}/.tslab-mcp"
}
}
}
}체크아웃에서 사용하려면 작업 트리를 가리키도록 설정하세요:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uv",
"args": ["run", "--directory", "${workspaceFolder}", "tslab-mcp"]
}
}
}그런 다음: 채팅을 열고 모드 선택기를 Agent로 전환한 후 Tools 버튼을 사용하여 8개의 tsf_* 도구가 나열되고 활성화되었는지 확인하세요. MCP: List Servers는 서버의 상태와 로그를 보여주며, 시작 실패 시 원인을 확인할 수 있습니다. Copilot은 한 번에 활성화할 수 있는 도구 수에 제한이 있으므로, 여러 MCP 서버를 실행하는 경우 일부를 선택 해제해야 할 수 있습니다.
Visual Studio. 동일한 JSON 형식으로, 솔루션 루트(또는 모든 솔루션의 경우 %USERPROFILE%\.mcp.json)의 .mcp.json에 추가한 후, Copilot Chat 에이전트 모드 도구 선택기에서 도구를 활성화하세요.
JetBrains, Eclipse, Xcode. Copilot Chat 에이전트 모드 도구 선택기를 열고 Edit MCP configuration을 선택한 후 열리는 mcp.json에 동일한 servers 항목을 추가하세요.
Copilot coding agent(github.com의 클라우드 에이전트)는 이 서버에 적합하지 않습니다. MCP 서버를 임시 GitHub Actions 환경에서 실행하므로, 실행마다 약 2GB의 TimeCopilot 설치 비용이 발생하며 로컬 데이터 파일에 접근할 수 없습니다. 대신 에디터에서 사용하세요.
도구
도구 | 목적 | 반환값 |
| CSV/Parquet 읽기, | JSON 요약 + SHA-256 |
| 모델군 선택을 위한 계열별 특징 | Markdown 표 또는 JSON, 행 제한 |
| 실제로 임포트 가능한 모델 조사 |
|
| 모델 간 롤링 오리진 비교 | 메트릭 표, 순위, parquet 경로 |
| 예측 구간과 함께 적합 및 예측 | Parquet 경로 + 제한된 미리보기 |
| 교차 검증된 구간 플래그 지정 | 개수, 제한된 플래그 목록, parquet 경로 |
| 세션을 재실행 가능한 매니페스트로 고정 | 매니페스트 경로 |
| 모든 단계를 읽기 쉬운 보고서로 렌더링 | HTML 또는 Markdown 경로 |
tsf_export_* 두 도구를 제외한 모든 도구는 읽기 전용으로 표시됩니다. 여기서 삭제 작업은 수행되지 않으므로, ~/.tslab-mcp/runs 정리는 사용자의 책임입니다.
세션 시작하기
도구는 순서를 강제하지 않으므로, 여는 프롬프트가 8개의 호출 가능한 함수를 분석으로 전환합니다. 다음과 같은 방식이 효과적입니다:
tslab 도구를 사용하여
/Users/me/data/deposits.csv의 계열을 12개월 앞으로 예측하세요.다음 순서로 작업하고 각 단계에서 추론을 보여주세요:
파일을 로드하고 발견한 내용을 알려주세요 — 몇 개의 계열, 주기, 간격이나 누락값이 있는지.
특징을 설명하고, 어떤 모델군이 적합한지, 그 이유를 말하세요.
모델을 제안하기 전에 실제로 설치된 모델을 확인하세요.
후보 목록을 4개 윈도우에 대해 SeasonalNaive 기준선과 교차 검증하세요. 지금은 통계 모델만 사용.
승자로 80% 및 95% 예측 구간을 포함한 예측.
실행 매니페스트와 HTML 보고서를 내보내고, 모델 선택 근거를 노트에 기록하세요: 무엇을 선택했는지, 메트릭 표가 무엇을 보여주었는지, 무엇을 기각했는지.
결과를 요약하고 parquet 경로를 알려주세요 — 전체 프레임을 채팅에 붙여넣지 마세요.
해당 프롬프트에서 실제로 중요한 네 가지 사항:
절대 경로. 상대 경로는 서버의 작업 디렉터리를 기준으로 해석되며, 이는 MCP 클라이언트가 선택하고 일반적으로 예측할 수 없습니다.
결정에 맞는 예측 기간.
h는 예측 자체와 각 CV 윈도우가 소비하는 과거 길이를 결정합니다. 12개월 단계는 1년 계획으로, 임의의 기본값이 아닙니다."지금은 통계 모델만 사용." 이 제한이 없으면 에이전트가 기반 모델을 사용하려고 시도하여,
AutoETS가 몇 초 만에 해결할 질문에 가중치 다운로드에 몇 분을 소비할 수 있습니다. 저렴한 모델이 기준선을 설정한 후에 제한을 해제하세요.매니페스트 노트에 근거를 요청하세요. 채팅 기록은 일회용입니다. 매니페스트는 누군가가 재실행하고 감사할 수 있는 부분입니다. 추론이 대화에만 존재하면 사실상 손실됩니다.
원하는 것이 명확할 때 더 짧은 시작 문구:
/Users/me/data/sales.parquet를 로드하고 특징을 설명하세요. 아직 예측하지 마세요 — 먼저 데이터를 보고 싶습니다.
로드된
deposits핸들에 대해 SeasonalNaive, AutoETS, AutoARIMA를 비교하세요. h=12에서 6개 윈도우로 비교하고, 어떤 모델이 기준선을 충분히 능가하여 추가 복잡성을 정당화하는지 알려주세요.
통계 모델 전용 호출은 몇 초 안에 응답합니다. 기반 모델을 처음 호출하면 TimeCopilot 임포트에 약 30초가 소요됩니다. 이 지연은 예상된 것이며 중단이 아닙니다. foundation 확장 기능이 설치되고 요청이 실제로 기반 모델을 사용할 때만 발생합니다.
작업 예시
Nixtla 긴 형식의 CSV로 시작:
unique_id,ds,y
branch_01,2018-01-01,1043.2
branch_01,2018-02-01,1102.7
...1. 로드합니다. 패널은 서버 프로세스에 유지되며, 핸들은 세션에서만 사용됩니다.
{"handle": "deposits", "n_series": 12, "n_obs": 864, "freq": "MS",
"start": "2018-01-01T00:00:00", "end": "2023-12-01T00:00:00",
"obs_per_series": {"min": 72, "median": 72, "max": 72},
"n_missing_y": 0, "sha256": "9f2c…"}2. 설명합니다. 이 숫자가 추론의 대상입니다.
| id | n | mean | cv | %zero | trend | seasonal | acf1(diff) |
|-----------|----|--------|-------|-------|-------|----------|------------|
| branch_01 | 72 | 1180.4 | 0.112 | 0.0 | 0.83 | 0.62 | -0.31 |높은 계절성 강도와 명확한 추세는 naive 기준선보다 AutoETS와 AutoARIMA를 지지합니다. 높은 %zero는 대신 ADIDA나 CrostonClassic을 지지했을 것입니다.
seasonal은 STL 강도입니다. 추세가 제거된 후 남은 부분에 대한 계절적 구성 요소의 측정값이므로, 증가하는 계열도 계절성을 정직하게 보고합니다. 대략 0.3–0.5의 노이즈 플로어를 가지며, 이 범위의 점수는 "증거 없음"을 의미하며 "약간 계절적"이 아닙니다.
3. 설치된 모델 확인 tsf_list_models로 확인하여, 이 머신에서 실행할 수 없는 모델을 제안하지 않도록 합니다.
4. 후보 교차 검증 — 항상 SeasonalNaive를 포함합니다. 이를 능가하지 못하는 모델은 배포할 가치가 없습니다:
{"kind": "cross_validation", "models": ["SeasonalNaive", "AutoETS", "AutoARIMA"],
"h": 12, "n_windows": 4, "seasonality_used_for_mase": 12,
"metrics": {"mase": {"SeasonalNaive": 1.0, "AutoETS": 0.71, "AutoARIMA": 0.68}},
"ranking": {"mase": ["AutoARIMA", "AutoETS", "SeasonalNaive"]},
"artifact": "~/.tslab-mcp/runs/cv_deposits_3f1a9c02.parquet"}5. 승자로 예측합니다. 전체 프레임은 parquet으로 저장되며, 응답에는 경로, 열, 그리고 짧은 미리보기가 포함됩니다.
6. 실행 및 보고서 내보내기 이유를 노트에 기록하세요. 이는 대화보다 오래 지속되는 유일한 추론 부분입니다:
{"manifest": "~/.tslab-mcp/runs/manifest_deposits_77b0e415.json", "n_runs": 3,
"kinds": ["cross_validation", "forecast"]}매니페스트에는 소스 경로와 해시, 빈도, 모든 호출과 인수 및 아티팩트 경로, 실제로 설치된 항목의 고정 버전(statsforecast, pandas, Python은 항상, TimeCopilot과 torch는 foundation 확장 기능이 있으면 포함), 그리고 노트가 포함됩니다. 서버가 중지된 상태에서도 숫자를 재현하기에 충분합니다.
tsf_export_report는 동일한 매니페스트를 사람이 읽을 수 있는 형태로 변환합니다. 특징, 메트릭 표(최고 우선 순서), 예측, 이상치, 환경을 발생 순서대로 보여줍니다:
{"report": "~/.tslab-mcp/runs/report_deposits_5c31d0a7.html",
"format": "html", "n_steps": 3,
"steps": ["features", "cross_validation", "forecast"]}보고서는 매니페스트의 순수 함수입니다. parquet을 읽지 않고 모델을 호출하지 않으므로, manifest_path와 함께 tsf_export_report를 호출하면 아무것도 로드하지 않고 몇 달 전의 실행을 다시 렌더링합니다. HTML은 자체 CSS를 포함하며 외부 스크립트, 스타일시트, 글꼴을 참조하지 않으므로 오프라인에서도 정상적으로 열립니다.
설계
네 가지 불변 조건과 그 이유:
핸들이지 데이터프레임이 아닙니다. 하나의 교차 검증 프레임은
n_시계열 × h × n_윈도우 × n_모델 행으로 구성됩니다. 이를 도구 결과로 직렬화하면
첫 번째 호출에서 세션의 컨텍스트를 소진하고 이후의 모든 턴을 악화시킵니다.
도구는 핸들을 받아 요약, 집계, 파일 경로를 반환합니다. 모든 대량 경로는 상한이 있고
생략된 내용을 보고하므로, 세션은 다시 요청하는 대신 parquet 파일을 읽어야 함을 알게 됩니다.
차단 작업은 이벤트 루프를 건드리지 않습니다. 대규모 패널에 대해 여러 모델을 교차 검증하는 것은
수 분의 CPU 시간이 소요됩니다. 모든 도구 본체는 anyio.to_thread.run_sync를 통해
디스패치되는 동기 클로저이므로, stdio 전송이 계속 응답하고 클라이언트가 실행 중에 서버를
끊지 않습니다.
환경은 추정되지 않고 발견됩니다. 모델은 지연 로딩되고 탐색되며, 존재한다고 가정되지 않습니다.
tsf_list_models는 여기서 실제로 확인된 것을 보고하므로, 추가 패키지 없이 Chronos를 요청하면
실행 10분 후의 트레이스백 대신 해당 추가 패키지를 명명하는 메시지를 반환합니다.
백엔드는 요청 내용에 따라 선택됩니다. 모든 모델이 통계적 모델인 요청은 statsforecast를 통해 실행되고, 사전 학습된 모델이 필요한 요청만 TimeCopilot에 도달합니다. 따라서 통계적 실행은 절대 torch를 임포트하지 않으며, 어느 경우든 서버는 즉시 시작됩니다.
statsforecast는 의도적으로 기본값 n_jobs=1로 유지됩니다. 병렬 모드는 작업자 프로세스를 생성하여
진입 모듈을 다시 임포트하는데, MCP 서버 내부에서는 속도 향상보다는 경합과 stdout 위험을 초래합니다.
매니페스트가 기록의 대상입니다. 대화의 서술은 주석입니다. 매니페스트는 누군가가 6개월 후에 재실행할 대상이며, 리뷰어가 어떤 모델이 어떤 기준으로 비교되었는지 확인하기 위해 읽는 대상입니다.
파이썬 버전
TimeCopilot은 여러 모델을 인터프리터 버전에 따라 제한하며, Python < 3.13에서는
tabpfn-time-series를 고정하여 pandas를 2.2 미만으로 제한합니다.
Python | 모델 | pandas |
3.13 |
| ≥ 2.2 |
3.10–3.12 |
| < 2.2 |
3.13이 권장 대상입니다. 어느 쪽이든 tsf_list_models는 실제로 확인된 내용을
확인되지 않은 항목에 대한 이유와 함께 보고합니다.
개발
uv sync --all-groups
uv run pytest # fast suite
uv run pytest -m slow # exercises TimeCopilot; slower, no weight downloads
uv run ruff check src tests
uv run mypyMCP 인스펙터로 도구 표면을 검사합니다:
npx @modelcontextprotocol/inspector uv run tslab-mcp라이선스
MIT
Available Tools
8 toolstsf_cross_validateARead-onlyIdempotent
Compare models by rolling-origin cross-validation.
This is the tool that replaces guesswork about model choice: it produces the evidence, you read the table and decide. Always include SeasonalNaive as the baseline -- a model that cannot beat it is not worth deploying.
Returns a per-model metric table aggregated over series and windows, a best-first ranking per metric, and the parquet path holding every per-window prediction. LONG-RUNNING: seconds for statistical models, many minutes for foundation models on a large panel. Start with statistical models on the real horizon before reaching for anything pretrained.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, idempotent, non-destructive behavior. The description adds crucial behavioral context: runtime warning ('LONG-RUNNING'), output description (aggregated table, ranking, parquet path), and implicitly that it is safe but compute-intensive. This goes well beyond the annotations and aids agent expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, purposeful paragraphs. The first sentence immediately states the tool's function. Each subsequent section (usage advice, output details, runtime warning) earns its place with no redundancy. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (cross-validation, multiple models, windows, output schema exists), the description covers purpose, usage, baseline recommendation, output contents (aggregated metrics, rankings, prediction parquet), and runtime behavior. It is sufficiently complete for an agent to understand when and how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The nested input schema (CrossValidateInput) already provides parameter descriptions (e.g., models, metrics). The description adds high-level advice (like horizon matching the decision, baseline recommendation) but no new parameter-level semantics beyond what the schema offers. Schema coverage is effectively high despite the 0% top-level stat, so the description's incremental value here is moderate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource combination ('Compare models by rolling-origin cross-validation') and clearly distinguishes this tool from siblings like tsf_forecast (single model forecast) and tsf_list_models (model names). It states it replaces guesswork about model choice, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Always include SeasonalNaive as the baseline' and 'Start with statistical models on the real horizon before reaching for anything pretrained.' It also explains that the tool produces evidence for model selection. However, it does not explicitly state when not to use this tool (e.g., for final forecasts) or name alternatives, slightly reducing completeness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_describe_seriesARead-onlyIdempotent
Compute the per-series features that decide which model family to try.
Returns length, mean, sd, coefficient of variation, share of zeros, trend strength (R-squared against time), seasonal strength (variance explained by the period means), and lag-1 autocorrelation of the differenced series.
Read it as evidence, not as an answer: high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models (ADIDA, IMAPA, CrostonClassic); high cv with low structure argues for keeping expectations modest. Cheap -- returns in under a second.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and nondestructive nature. The description adds 'Cheap -- returns in under a second' and 'Read it as evidence, not as an answer,' providing behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with five sentences, front-loaded with purpose, and every sentence adds value. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema (not shown but present), the description covers the semantic meaning of the features and how to interpret them. It also provides cost and time estimates, making it self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has detailed descriptions for all three parameters (handle, max_series, response_format). The tool description focuses on output features and usage advice, not parameter details. Since schema coverage is high, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a clear verb-resource pair: 'Compute the per-series features that decide which model family to try.' It lists the specific features computed, distinguishing this diagnostic tool from siblings like tsf_forecast or tsf_load_series.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit decision rules: 'high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models...' and positions the tool as 'evidence, not an answer.' This clearly guides when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_detect_anomaliesARead-onlyIdempotent
Flag historical points that fall outside a cross-validated prediction interval.
The detector model defines what "expected" means, so pick one that fits the series: a weak detector flags its own errors rather than real anomalies. Run tsf_describe_series or tsf_cross_validate first.
Returns flagged counts per series, a capped list of flagged rows, and the parquet path with the full result.
LONG-RUNNING, and the default is the expensive one: leaving n_windows unset refits the model once per observation across the whole history, which takes minutes even for a statistical model. Pass n_windows (e.g. 12) unless you genuinely need every point tested.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses long-running nature and expensive default (n_windows unset refits per observation, taking minutes). Annotations (readOnlyHint, idempotentHint) are consistent; description adds critical performance context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact four paragraphs with clear structure: purpose, prerequisites, output summary, performance warning. No redundant sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given tool complexity (cross-validation, long-running, return includes parquet path and capped list), the description covers prerequisites, output, and performance. Output schema exists, so return values are adequately summarized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond schema by explaining the default slowness of n_windows and the risk of weak models. Schema already has clear descriptions, but description provides crucial usage context for these parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it flags historical points outside a cross-validated prediction interval, distinguishes from sibling tools (tsf_describe_series, tsf_cross_validate, etc.), and warns about weak detectors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises running tsf_describe_series or tsf_cross_validate first, warns against weak detectors, and gives concrete guidance on setting n_windows to avoid slow default behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_reportA
Render every step of the analysis as a report someone can read.
Covers the input and its hash, the features, each cross-validation with its metric table and ranking, the forecasts, any anomaly runs, and the pinned environment -- in the order they happened. HTML is self-contained, with no external stylesheet or script, so it opens correctly years later.
Call it after tsf_export_run at the end of an analysis. Pass a note: the report headlines it as the rationale, and a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.
Report from manifest_path instead of handle to re-render an older run --
it needs nothing but the manifest file.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (all hints false), so the description carries the burden. It discloses that HTML output is 'self-contained, with no external stylesheet or script' and that the report covers steps 'in the order they happened.' However, it does not clarify whether the tool writes a file to disk, returns the report content, or has other side effects. The output schema exists but is not described in the tool description, leaving behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a lead sentence defining the tool, then a list of contents, then usage order, then note advice, then alternative invocation. It is informative without being verbose. Minor inefficiency: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again' is slightly colorful but still earns its place. Could be tighter, but overall effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 optional parameters and an output schema (which handles return value documentation), the description covers its purpose, contents, usage order, and parameter trade-offs. It does not explain what happens if both handle and manifest_path are provided (mutual exclusion handled by schema? not specified). It also assumes the agent knows tsf_export_run was called, which is implied by 'at the end of an analysis.' Overall adequately complete for an AI agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has detailed descriptions for all parameters (note, format, handle, manifest_path), so baseline is 3. The description goes beyond by explaining the semantic purpose of note ('headlines it as the rationale') and the trade-off between handle and manifest_path ('re-renders an older run -- it needs nothing but the manifest file'). This adds actionable context for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description starts with a specific verb ('Render every step of the analysis as a report') and enumerates the exact contents (input hash, features, cross-validation, forecasts, anomalies, pinned environment). This immediately distinguishes it from sibling tools like tsf_export_run (which exports run data) and tsf_forecast (which only forecasts). The purpose is unambiguous and comprehensive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states ordering: 'Call it after tsf_export_run at the end of an analysis.' Provides clear alternatives: 'Report from manifest_path instead of handle to re-render an older run.' Also advises on best practice for the note parameter: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.' This gives the agent concrete when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_runA
Write a JSON manifest of everything done to this handle.
Records the source path and SHA-256, the frequency, every call with its arguments and artifact paths, the pinned package versions, and your note. This is the artifact of record: your prose in the conversation is lost, this file is not. Write the note -- say which model you picked, what the metric table showed, and what you rejected.
Call it at the end of any analysis someone might have to defend or rerun. Writes a file, so it is not read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds valuable context: 'Writes a file, so it is not read-only' and lists everything included in the manifest. It also emphasizes that the note is the only place a reasoning trace survives, which is important behavioral insight. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: starts with the core purpose, then enumerates contents, gives usage advice, and closes with a note about file writing. It is not overly long; every sentence contributes. A slight trim could improve conciseness, but it remains clear and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description correctly omits return value details. It covers when to call, what the manifest contains, and the critical role of the note. The only minor gap is no mention of potential side effects (e.g., overwriting existing files), but overall it is sufficiently complete for a tool with good annotations and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already contains descriptions for both parameters: handle ('Handle whose run log should be pinned to a manifest') and note (detailed explanation of what to write). The tool description adds further guidance for the note, specifically: 'Write the note -- say which model you picked, what the metric table showed, and what you rejected.' This enhances the schema's descriptions, making it clear how to use the note parameter effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Write a JSON manifest of everything done to this handle' clearly states the verb (write a manifest) and the resource (handle). The description further details what the manifest includes (source path, SHA-256, calls, arguments, artifact paths, pinned package versions, note), differentiating it from sibling tools like 'tsf_export_report' or 'tsf_describe_series'. This makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to call it: 'Call it at the end of any analysis someone might have to defend or rerun.' This provides clear context for use. It does not explicitly state when not to use it or name alternatives, but for a specialized export tool the guidance is sufficient and well-placed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_forecastARead-onlyIdempotent
Fit on the full history and forecast h periods ahead with intervals.
Use after tsf_cross_validate has justified the model choice. The full forecast goes to parquet; the response carries the path, the column list, the row count and a small preview. Read the parquet for anything more -- raising max_preview_rows to dump the frame into the conversation is the one thing that reliably ruins a long session.
LONG-RUNNING for foundation models.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the bar is lowered. The description adds valuable context: the full forecast goes to parquet, the response carries path/column list/row count/preview, warns against raising max_preview_rows, and flags 'LONG-RUNNING for foundation models' — all beyond what annotations provide. No contradictions with annotations (readOnlyHint=true is consistent with generating forecasts without mutating data).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: a terse three-sentence explanation that front-loads the core action, then adds usage guidance and behavioral warnings. Every sentence adds essential information without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (forecasting with multiple models and intervals), the output schema exists, so return values don't need elaboration. The description covers the critical workflow (use after cross-validation), output format (parquet with preview), and a key gotcha (don't dump full frame). It lacks explicit error conditions or prerequisite checks, but for the scope, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema provides no parameter descriptions — so the description must compensate. Although the main description does not detail parameters, the parameter `max_preview_rows` receives meaningful context: 'Rows of the forecast to inline... read that instead of raising this.' Other parameters (handle, models, h, level) have descriptions in the schema via the JSON Schema, but since coverage is 0% (likely meaning no separate param list in the description), the main text does not clarify their meaning beyond what the schema already provides. Baseline 3 is appropriate as the description adds some context for max_preview_rows but not for others.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool fits on full history and forecasts h periods ahead with intervals. It uses specific verbs like 'fit' and 'forecast' and explicitly identifies the resource as the time series forecast. However, it does not directly differentiate from siblings like tsf_cross_validate, though the usage guideline addresses that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after tsf_cross_validate has justified the model choice,' providing clear sequencing context and an alternative (cross-validation). It does not mention when not to use it or list specific alternatives for other tasks like anomaly detection, but the primary usage guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_list_modelsARead-onlyIdempotent
Probe which models actually import in this environment.
Call this before cross-validating so you never propose a model that cannot run here. Returns {available, statistical, foundation, unavailable}, where each unavailable entry carries the real reason -- some models are gated on the Python version, not merely absent.
The first call imports TimeCopilot and can take ~30 seconds; later calls are instant.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds important behavioral details beyond that: the first call may take ~30 seconds to import TimeCopilot, later calls are instant; it returns structured output with real reasons for unavailability (including Python version gating). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and structured: first sentence states purpose, second gives usage guidance, third explains return structure, and fourth notes startup latency. Every sentence adds essential information, with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (the agent can rely on structured return type details), the description covers all necessary context: what the tool does, when to use it, its runtime behavior (latency), and the high-level shape of results. No gaps remain for a list/probe tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema contains descriptions for both parameters (family and include_unavailable) that are self-explanatory. The tool description does not add new parameter information—it only repeats that statistical models are cheap and always installed, which already appears in the schema. Schema coverage via inline descriptions is present, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Probe which models actually import in this environment' and specifies the tool's role in preventing proposal of non-running models. It clearly distinguishes itself from siblings like cross_validate by giving a precise pre-check use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contains an explicit directive: 'Call this before cross-validating so you never propose a model that cannot run here.' It also notes that statistical models are always installed, which helps with decision-making. No alternatives are listed, but the use case is crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_load_seriesARead-onlyIdempotent
Read a CSV or Parquet panel from disk and register it under a handle.
Call this first; every other tool takes the handle it returns. The file must be in Nixtla long format (unique_id, ds, y). Returns a compact JSON summary -- series count, inferred frequency, date range, missing values, and the SHA-256 of the source -- and nothing else: the data stays in the server so it never consumes your context.
Read the summary before choosing a horizon. If obs_per_series.min is small, a long horizon or many CV windows will not fit.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description reveals that data stays server-side ('never consumes your context') and that the return is a compact summary with specific fields. It also hints at the side effect of reusing a handle (replacing the panel). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five sentences, each serving a distinct purpose: purpose, ordering, format, return details, caution. Front-loaded with the core action. No superfluous words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's role as the entry point for a time series workflow, the description covers purpose, required file format, return value (with summary contents), and a concrete usage caution. With an output schema present, the lack of detailed return structure is acceptable. The description is complete enough for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides detailed descriptions for all three parameters. The tool description adds the critical constraint that the file must be in 'Nixtla long format (unique_id, ds, y)', which is not in the schema. This adds meaningful value beyond the schema, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action: 'Read a CSV or Parquet panel from disk and register it under a handle.' It distinguishes from sibling tools by explicitly saying 'Call this first; every other tool takes the handle it returns.' This makes the purpose unambiguous and contextually positioned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit ordering ('Call this first'), explains the handle's role in subsequent tools, and gives a practical caution about horizon choices based on the summary output. This equips the agent with clear when-to-use and how-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
tsf_cross_validate - First observed
tsf_describe_series - First observed
tsf_detect_anomalies - First observed
tsf_export_report - First observed
tsf_export_run - First observed
tsf_forecast - First observed
tsf_list_models - First observed
tsf_load_series
TDQS
Scored across 8 tools
Each tool has a distinct and well-defined purpose within the time series forecasting workflow: loading, describing, listing models, cross-validating, forecasting, detecting anomalies, and exporting results. There is no overlap or ambiguity between tools.
All tools follow a consistent pattern: the prefix 'tsf_' followed by a verb (and optional noun), all in snake_case. Examples include tsf_load_series, tsf_describe_series, tsf_cross_validate, and tsf_export_report. The naming is predictable and uniform.
With 8 tools, the server covers a complete analysis pipeline without excess. Each tool is necessary and corresponds to a clear step in the workflow, from data loading to report generation. The count is well-scoped for the domain.
The tool surface covers the full lifecycle of a typical time series analysis: load data, explore features, check available models, cross-validate, forecast, detect anomalies, and export manifests/reports. There are no obvious gaps; the workflow feels self-contained and actionable.
Maintenance
Related MCP Connectors
Probabilistic time-series forecasts from zero-shot foundation models: routed, single or ensembled.
1PredictOracle - 12 forecasting tools: time-series, scenario analysis, risk projections.
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Deterministic time tools for AI agents: timezone conversion, business-day math, cron interpretation.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnable any AI agent to forecast time-series data (e.g., sales, traffic) using Google's TimesFM or a zero-dependency statistical baseline.3Apache 2.0
- AlicenseAqualityDmaintenanceDeterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.17MIT
- AlicenseNot gradedqualityCmaintenanceEnables time-series analysis and forecasting through a structured tool catalogue, including data loading, quality repair, diagnostics, and forecasting with ARIMA, exponential smoothing, Chronos-2, Toto 2.0, and AutoML.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to run TimesFM-3 forecasting workflows locally, including joint multivariate forecasting with known future drivers, backtesting against a baseline, what-if scenario comparison, and historical anomaly detection. It exposes the studio's tools and bundled public and synthetic demo datasets over MCP.Apache 2.0