Skip to main content
Glama

tslab-mcp

Ein MCP-Server, der deterministische Zeitreihenprognosen als Werkzeuge bereitstellt, sodass Ihr Agent die Denkmaschine ist und jede Zahl aus gewöhnlichem, reproduzierbarem Python stammt.

In diesem Paket wird kein LLM aufgerufen. Es ist kein API-Schlüssel erforderlich (es sei denn, Sie fragen nach TimeGPT, das die Nixtla-API aufruft).

Warum

Einige Prognosebibliotheken liefern einen Agenten mit, der Merkmale liest, ein Modell auswählt und das Ergebnis mit einem LLM im Kreislauf erklärt. Wenn Sie einen solchen von Ihrem eigenen Agenten aus aufrufen, verschachteln Sie einen Agenten in einem anderen – zwei Prompts, zwei Rechnungen, zwei Quellen für Nichtdeterminismus und eine undurchsichtige Zwischenschicht, die die Begründung der Modellauswahl nicht überprüfbar macht.

Hier ist die Kontrolle also umgekehrt: Die Prognosebibliothek ist das Werkzeug, und Ihr Agent ist derjenige, der denkt. Er liest die Merkmale, argumentiert für eine Modellfamilie, validiert die Kandidaten kreuzweise und schreibt die Begründung in ein Manifest. Jede Zahl auf dem Weg wird durch einen Bibliotheksaufruf erzeugt, den Sie ohne LLM im Pfad erneut ausführen können.

Diese Trennung setzt sich in der Art und Weise fort, wie das Paket selbst aufgebaut ist. Die Basisinstallation führt elf statistische Modelle aus – AutoARIMA, AutoETS, Theta, CrostonClassic und weitere – über statsforecast: etwa 340 MB, kein PyTorch, und es startet in Sekunden. Ein optionales foundation-Extra fügt die vortrainierten Modelle von TimeCopilot hinzu – Chronos, Moirai, TimesFM, TiRex, Toto und andere – sowie Prophet, für den Fall, dass eine statistische Basislinie nicht ausreicht. Eine Anfrage, die nur statistische Modelle nennt, importiert niemals TimeCopilot oder torch; eine Anfrage, die auch nur ein Foundation-Modell nennt, läuft vollständig über TimeCopilot, das auch die statistischen Modelle mit sich führt. In beiden Fällen meldet tsf_list_models, was tatsächlich installiert ist, bevor Sie sich auf ein Modell festlegen.

Related MCP server: timeseries-mcp

Installation

Erfordert Python 3.10+ (3.13 empfohlen, siehe Python-Version).

uvx tslab-mcp                     # run without installing
uv tool install tslab-mcp         # or install the CLI

Die Basisinstallation führt die elf statistischen Modelle über statsforecast aus: etwa 340 MB, kein PyTorch, und es startet sofort. Für die vortrainierten Foundation-Modelle – Chronos, Moirai, TimesFM, Toto, TiRex – und Prophet fügen Sie das Extra hinzu:

uvx --from 'tslab-mcp[foundation]' tslab-mcp

Das foundation-Extra zieht TimeCopilot nach sich, das torch, transformers und lightning mitbringt: etwa 2 GB bei der ersten Installation, und der erste Werkzeugaufruf, der darauf zugreift, benötigt etwa 30 Sekunden für den Import. Beides sind einmalige Vorgänge, und keiner ist kostenpflichtig, es sei denn, Sie fragen nach einem Modell, das sie benötigt.

Von GitHub

uv und uvx akzeptieren beide eine Git-URL anstelle eines Paketnamens, was den aktuellen main-Zweig installiert, ohne auf ein Release zu warten:

uvx --from git+https://github.com/pedrobtz/tslab-mcp tslab-mcp
uv tool install git+https://github.com/pedrobtz/tslab-mcp        # or install the CLI

# with the foundation extra
uvx --from 'tslab-mcp[foundation] @ git+https://github.com/pedrobtz/tslab-mcp' tslab-mcp

Fixieren Sie einen Ref für alles andere als beiläufige Tests – der Branch-Head kann sich sonst unter Ihnen bewegen. Ein Commit funktioniert heute; ein Versionstag wird ebenfalls funktionieren, sobald einer erstellt ist:

uv tool install "git+https://github.com/pedrobtz/tslab-mcp@136824c1cc2a"

Aus einem Checkout

git clone https://github.com/pedrobtz/tslab-mcp
cd tslab-mcp
uv sync                              # base
uv sync --extra foundation           # with the pretrained models
uv run tslab-mcp

Konfiguration

Fügen Sie den Server zur Konfiguration Ihres MCP-Clients hinzu. Die Datei unterscheidet sich pro Client – oft .mcp.json im Projektstammverzeichnis – aber der Eintrag selbst hat die gleiche Form:

{
  "mcpServers": {
    "tslab": {
      "command": "uvx",
      "args": ["tslab-mcp"],
      "env": {
        "TSLAB_MCP_HOME": "~/.tslab-mcp"
      }
    }
  }
}

TSLAB_MCP_HOME legt fest, wo Artefakte geschrieben werden; der Standardwert ist ~/.tslab-mcp, und Ausgaben von Läufen landen in <home>/runs.

Der Transport ist nur stdio, aus gutem Grund: Ihre Daten werden als sensibel betrachtet und verlassen niemals die Maschine. Der Server tätigt keine ausgehenden Anfragen, außer den Modellgewichts-Downloads, die TimeCopilot selbst für Foundation-Modelle durchführt, und den Nixtla-API-Aufrufen, die TimeGPT tätigt, wenn Sie explizit danach fragen.

GitHub Copilot

Copilot entdeckt MCP-Server aus einer mcp.json-Datei und stellt ihre Werkzeuge im Agentenmodus zur Verfügung – die Werkzeuge erscheinen nicht im Frage- oder Bearbeitungsmodus.

VS Code. Platzieren Sie den Server in .vscode/mcp.json, um ihn mit dem Repository zu teilen, oder führen Sie MCP: Open User Configuration aus der Befehlspalette aus, um ihn in Ihrem eigenen Profil über alle Arbeitsbereiche hinweg zu behalten. Beachten Sie, dass der Schlüssel servers ist, nicht mcpServers:

{
  "servers": {
    "tslab": {
      "type": "stdio",
      "command": "uvx",
      "args": ["tslab-mcp"],
      "env": {
        "TSLAB_MCP_HOME": "${userHome}/.tslab-mcp"
      }
    }
  }
}

Aus einem Checkout heraus verweisen Sie stattdessen auf den Arbeitsbaum:

{
  "servers": {
    "tslab": {
      "type": "stdio",
      "command": "uv",
      "args": ["run", "--directory", "${workspaceFolder}", "tslab-mcp"]
    }
  }
}

Dann: Öffnen Sie Chat, schalten Sie den Modus-Wähler auf Agent und verwenden Sie die Schaltfläche Tools, um zu bestätigen, dass die acht tsf_*-Werkzeuge aufgelistet und aktiviert sind. MCP: List Servers zeigt den Status des Servers und seine Protokolle an, wo ein fehlgeschlagener Start erklärt wird. Copilot begrenzt, wie viele Werkzeuge gleichzeitig aktiv sein können. Wenn Sie also mehrere MCP-Server betreiben, müssen Sie möglicherweise einige abwählen, um alle acht unterzubringen.

Visual Studio. Gleiche JSON-Form, in .mcp.json im Lösungsstammverzeichnis (oder %USERPROFILE%\.mcp.json für alle Lösungen), dann aktivieren Sie die Werkzeuge aus der Copilot Chat-Agentenmodus-Werkzeugauswahl.

JetBrains, Eclipse und Xcode. Öffnen Sie die Copilot Chat-Agentenmodus-Werkzeugauswahl, wählen Sie Edit MCP configuration und fügen Sie denselben servers-Eintrag zu der geöffneten mcp.json hinzu.

Copilot Coding Agent (der Cloud-Agent auf github.com) ist für diesen Server schlecht geeignet: Er führt Ihre MCP-Server in einer flüchtigen GitHub Actions-Umgebung aus, was bedeutet, dass die etwa 2 GB große TimeCopilot-Installation bei jedem Lauf bezahlt werden muss, und er hat keinen Zugriff auf lokale Datendateien. Verwenden Sie ihn stattdessen von Ihrem Editor aus.

Werkzeuge

Werkzeug

Zweck

Rückgabe

tsf_load_series

CSV/Parquet lesen, den unique_id/ds/y-Vertrag validieren, Frequenz ableiten, Handle registrieren

JSON-Zusammenfassung + SHA-256

tsf_describe_series

Pro-Serie-Merkmale zur Auswahl einer Modellfamilie

Markdown-Tabelle oder JSON, zeilenbegrenzt

tsf_list_models

Prüfen, welche Modelle hier tatsächlich importieren

{available, statistical, foundation, unavailable}

tsf_cross_validate

Rollierenden-Ursprung-Vergleich über Modelle hinweg

Metriktabelle, Rangfolge, Parquet-Pfad

tsf_forecast

Anpassen und Prognose mit Vorhersageintervallen

Parquet-Pfad + begrenzte Vorschau

tsf_detect_anomalies

Kreuzvalidierte Intervallmarkierung

Zählungen, begrenzte Flag-Liste, Parquet-Pfad

tsf_export_run

Sitzung in einem wiederholbaren Manifest fixieren

Manifest-Pfad

tsf_export_report

Jeden Schritt als lesbaren Bericht darstellen

HTML- oder Markdown-Pfad

Alles außer den beiden tsf_export_*-Werkzeugen ist als schreibgeschützt markiert; nichts hier löscht, daher ist das Bereinigen von ~/.tslab-mcp/runs Ihre Aufgabe, nicht die des Agenten.

Starten einer Sitzung

Die Werkzeuge erzwingen keine Reihenfolge, daher ist der Eröffnungsprompt das, was acht aufrufbare Funktionen in eine Analyse verwandelt. So etwas funktioniert gut:

Verwenden Sie die tslab-Werkzeuge, um die Serie in /Users/me/data/deposits.csv 12 Monate im Voraus zu prognostizieren.

Arbeiten Sie in dieser Reihenfolge und zeigen Sie Ihre Überlegungen bei jedem Schritt:

  1. Laden Sie die Datei und sagen Sie mir, was Sie gefunden haben – wie viele Serien, welche Frequenz, etwaige Lücken oder fehlende Werte.

  2. Beschreiben Sie die Merkmale und sagen Sie, für welche Modellfamilien sie sprechen und warum.

  3. Überprüfen Sie, welche Modelle tatsächlich installiert sind, bevor Sie welche vorschlagen.

  4. Validieren Sie Ihre Shortlist gegen eine SeasonalNaive-Baseline über 4 Fenster. Vorerst nur statistische Modelle.

  5. Prognostizieren Sie mit dem Gewinner, mit 80%- und 95%-Intervallen.

  6. Exportieren Sie ein Laufmanifest und einen HTML-Bericht und setzen Sie die Begründung der Modellauswahl in die Notiz: was Sie gewählt haben, was die Metriktabelle zeigte und was Sie verworfen haben.

Fassen Sie die Ergebnisse zusammen und geben Sie mir die Parquet-Pfade – fügen Sie keine ganzen DataFrames in den Chat ein.

Vier Dinge in diesem Prompt leisten echte Arbeit:

  • Ein absoluter Pfad. Relative Pfade werden relativ zum Arbeitsverzeichnis des Servers aufgelöst, das Ihr MCP-Client wählt und das Sie im Allgemeinen nicht vorhersagen können.

  • Ein Horizont, der zur Entscheidung passt. h treibt sowohl die Prognose als auch den Umfang der Historie, die jedes CV-Fenster verbraucht; 12 monatliche Schritte sind ein Jahr Planung, kein willkürlicher Standard.

  • „Vorerst nur statistische Modelle.“ Ohne diese Einschränkung könnte ein Agent nach einem Foundation-Modell greifen und mehrere Minuten damit verbringen, Gewichte herunterzuladen, um eine Frage zu beantworten, die AutoETS in Sekunden erledigt hätte. Heben Sie die Einschränkung auf, sobald die billigen Modelle eine Untergrenze gesetzt haben.

  • Die Begründung in der Manifest-Notiz anfordern. Das Chat-Transkript ist wegwerfbar; das Manifest ist der Teil, den jemand erneut ausführen und prüfen kann. Wenn die Begründung nur in der Konversation existiert, ist sie praktisch verloren.

Kürzere Eröffnungen, wenn Sie wissen, was Sie wollen:

Laden Sie /Users/me/data/sales.parquet und beschreiben Sie die Merkmale. Noch nicht prognostizieren – ich möchte zuerst sehen, womit wir es zu tun haben.

Vergleichen Sie SeasonalNaive, AutoETS und AutoARIMA auf dem geladenen deposits-Handle über 6 Fenster bei h=12 und sagen Sie mir dann, ob etwas die Baseline genug schlägt, um die zusätzliche Komplexität wert zu sein.

Rein statistische Aufrufe antworten in Sekunden. Der erste Aufruf, der ein Foundation-Modell nennt, benötigt etwa 30 Sekunden für den Import von TimeCopilot, bevor er etwas anderes tut – diese Pause ist erwartet, kein Hängen, und sie tritt nur auf, wenn das foundation-Extra installiert ist und eine Anfrage tatsächlich nach einem greift.

Eine durchgearbeitete Sitzung

Starten Sie von einer CSV im Nixtla-Langformat:

unique_id,ds,y
branch_01,2018-01-01,1043.2
branch_01,2018-02-01,1102.7
...

1. Laden Sie sie. Das Panel bleibt im Serverprozess; das Handle ist alles, was die Sitzung mit sich führt.

{"handle": "deposits", "n_series": 12, "n_obs": 864, "freq": "MS",
 "start": "2018-01-01T00:00:00", "end": "2023-12-01T00:00:00",
 "obs_per_series": {"min": 72, "median": 72, "max": 72},
 "n_missing_y": 0, "sha256": "9f2c…"}

2. Beschreiben Sie sie. Dies sind die Zahlen, über die Sie nachdenken.

| id        | n  | mean   | cv    | %zero | trend | seasonal | acf1(diff) |
|-----------|----|--------|-------|-------|-------|----------|------------|
| branch_01 | 72 | 1180.4 | 0.112 | 0.0   | 0.83  | 0.62     | -0.31      |

Hohe saisonale Stärke und ein klarer Trend sprechen für AutoETS und AutoARIMA gegenüber einer naiven Baseline; ein hoher %zero-Wert hätte stattdessen für ADIDA oder CrostonClassic gesprochen.

seasonal ist eine STL-Stärke – die saisonale Komponente gemessen an dem, was übrig bleibt, sobald der Trend entfernt ist – daher meldet eine wachsende Serie ihre Saisonalität dennoch ehrlich. Sie trägt ein Rauschniveau von etwa 0,3–0,5: Werte in diesem Band bedeuten „keine Evidenz“, nicht „leicht saisonal“.

3. Überprüfen Sie, was installiert ist mit tsf_list_models, damit Sie niemals ein Modell vorschlagen, das diese Maschine nicht ausführen kann.

4. Validieren Sie die Kandidaten kreuzweise – immer einschließlich SeasonalNaive, da ein Modell, das es nicht schlagen kann, nicht einsatzbereit ist:

{"kind": "cross_validation", "models": ["SeasonalNaive", "AutoETS", "AutoARIMA"],
 "h": 12, "n_windows": 4, "seasonality_used_for_mase": 12,
 "metrics": {"mase": {"SeasonalNaive": 1.0, "AutoETS": 0.71, "AutoARIMA": 0.68}},
 "ranking": {"mase": ["AutoARIMA", "AutoETS", "SeasonalNaive"]},
 "artifact": "~/.tslab-mcp/runs/cv_deposits_3f1a9c02.parquet"}

5. Prognostizieren Sie mit dem Gewinner. Der vollständige Frame geht nach Parquet; die Antwort enthält den Pfad, die Spalten und eine kurze Vorschau.

6. Exportieren Sie den Lauf und den Bericht. Schreiben Sie warum in die Notiz – es ist der einzige Teil Ihrer Überlegungen, der die Konversation überlebt:

{"manifest": "~/.tslab-mcp/runs/manifest_deposits_77b0e415.json", "n_runs": 3,
 "kinds": ["cross_validation", "forecast"]}

Das Manifest enthält den Quellpfad und den Hash, die Frequenz, jeden Aufruf mit seinen Argumenten und Artefaktpfaden, die fixierten Versionen von allem, was tatsächlich installiert ist – statsforecast, pandas und Python immer; TimeCopilot und torch auch, wenn das foundation-Extra dabei ist – und Ihre Notiz. Es reicht aus, um die Zahlen bei gestopptem Server zu reproduzieren.

tsf_export_report verwandelt dasselbe Manifest in etwas, das eine Person liest – Merkmale, Metriktabellen nach Besten geordnet, Prognosen, Anomalien und die Umgebung, in der Reihenfolge ihres Auftretens:

{"report": "~/.tslab-mcp/runs/report_deposits_5c31d0a7.html",
 "format": "html", "n_steps": 3,
 "steps": ["features", "cross_validation", "forecast"]}

Der Bericht ist eine reine Funktion des Manifests: Er liest kein Parquet und ruft kein Modell auf, daher rendert tsf_export_report mit manifest_path einen Lauf von vor Monaten neu, ohne dass etwas geladen ist. Das HTML bettet sein eigenes CSS ein und referenziert kein externes Skript, Stylesheet oder Schriftart, sodass es auch offline korrekt geöffnet wird.

Design

Vier Invarianten und die Gründe für ihre Existenz:

Handles, keine Dataframes.
Ein Kreuzvalidierungsrahmen umfasst
n_series × h × n_windows × n_models Zeilen. Wenn man ihn als Tool-Ergebnis
serialisiert, erschöpft das beim ersten Aufruf den Kontext der Sitzung und
verschlechtert jede spätere Interaktion. Tools übergeben einen Handle und geben
Zusammenfassungen, Aggregate und Dateipfade zurück; jeder Bulk-Pfad ist
begrenzt und meldet, was ausgelassen wurde, damit die Sitzung weiß, dass sie
das Parquet lesen soll, anstatt erneut zu fragen.

Blockierende Arbeiten berühren nie die Event-Schleife.
Das Kreuzvalidieren mehrerer Modelle über eine große Tabelle hinweg bedeutet
Minuten Rechenzeit. Jeder Tool-Body ist ein synchroner Closure, der über
anyio.to_thread.run_sync ausgeführt wird, sodass der stdio-Transport weiterhin
antwortet und der Client den Server nicht mitten im Lauf verwirft.

Die Umgebung wird entdeckt, nicht angenommen.
Modelle werden lazy importiert und geprüft, niemals als vorhanden
vorausgesetzt. tsf_list_models meldet, was hier tatsächlich aufgelöst wurde.
Wenn man also Chronos ohne die zusätzliche Abhängigkeit anfragt, erhält man
eine Nachricht, die das Extra benennt, anstatt zehn Minuten nach Beginn des
Laufs einen Traceback zu bekommen.

Das Backend wird dadurch bestimmt, was man anfragt: Eine Anfrage, deren Modelle
alle statistisch über statsforecast laufen, und nur eine Anfrage, die ein
vorab trainiertes Modell benötigt, greift auf TimeCopilot zu. Statistische Läufe
importieren daher niemals torch, und der Server startet in beiden Fällen
sofort.

statsforecast wird bewusst auf dem Standardwert n_jobs=1 belassen. Sein
Parallelmodus startet Worker-Prozesse, die das Einstiegsmodul erneut
importieren, was innerhalb eines MCP-Servers eher Konflikte und eine
stdout-Gefahr als Geschwindigkeitsvorteile bringt.

Das Manifest ist das Artefakt der Aufzeichnung.
Prosa im Gespräch ist Kommentar. Das Manifest ist das, was jemand in sechs
Monaten erneut ausführt, und das, was ein Prüfer liest, um zu sehen, welche
Modelle verglichen wurden und auf welcher Grundlage.

Python-Version

TimeCopilot macht mehrere Modelle von der Interpreter-Version abhängig, und
unter Python < 3.13 pinnt es tabpfn-time-series, was pandas unter 2.2
festlegt.

Python

Modelle

pandas

3.13

alles außer TabPFN und Sundial

≥ 2.2

3.10–3.12

fügt TabPFN, Sundial hinzu

< 2.2

3.13 ist das empfohlene Ziel. In beiden Fällen meldet tsf_list_models, was
tatsächlich aufgelöst wurde, mit dem Grund für alles, was nicht aufgelöst
werden konnte.

Entwicklung

uv sync --all-groups
uv run pytest                  # fast suite
uv run pytest -m slow          # exercises TimeCopilot; slower, no weight downloads
uv run ruff check src tests
uv run mypy

Überprüfen Sie die Tool-Oberfläche mit dem MCP Inspector:

npx @modelcontextprotocol/inspector uv run tslab-mcp

Lizenz

MIT

Available Tools

8 tools
tsf_cross_validateA
Read-onlyIdempotent

Compare models by rolling-origin cross-validation.

This is the tool that replaces guesswork about model choice: it produces the evidence, you read the table and decide. Always include SeasonalNaive as the baseline -- a model that cannot beat it is not worth deploying.

Returns a per-model metric table aggregated over series and windows, a best-first ranking per metric, and the parquet path holding every per-window prediction. LONG-RUNNING: seconds for statistical models, many minutes for foundation models on a large panel. Start with statistical models on the real horizon before reaching for anything pretrained.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnly, idempotent, non-destructive behavior. The description adds crucial behavioral context: runtime warning ('LONG-RUNNING'), output description (aggregated table, ranking, parquet path), and implicitly that it is safe but compute-intensive. This goes well beyond the annotations and aids agent expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short, purposeful paragraphs. The first sentence immediately states the tool's function. Each subsequent section (usage advice, output details, runtime warning) earns its place with no redundancy. It is front-loaded and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (cross-validation, multiple models, windows, output schema exists), the description covers purpose, usage, baseline recommendation, output contents (aggregated metrics, rankings, prediction parquet), and runtime behavior. It is sufficiently complete for an agent to understand when and how to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The nested input schema (CrossValidateInput) already provides parameter descriptions (e.g., models, metrics). The description adds high-level advice (like horizon matching the decision, baseline recommendation) but no new parameter-level semantics beyond what the schema offers. Schema coverage is effectively high despite the 0% top-level stat, so the description's incremental value here is moderate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource combination ('Compare models by rolling-origin cross-validation') and clearly distinguishes this tool from siblings like tsf_forecast (single model forecast) and tsf_list_models (model names). It states it replaces guesswork about model choice, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: 'Always include SeasonalNaive as the baseline' and 'Start with statistical models on the real horizon before reaching for anything pretrained.' It also explains that the tool produces evidence for model selection. However, it does not explicitly state when not to use this tool (e.g., for final forecasts) or name alternatives, slightly reducing completeness.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_describe_seriesA
Read-onlyIdempotent

Compute the per-series features that decide which model family to try.

Returns length, mean, sd, coefficient of variation, share of zeros, trend strength (R-squared against time), seasonal strength (variance explained by the period means), and lag-1 autocorrelation of the differenced series.

Read it as evidence, not as an answer: high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models (ADIDA, IMAPA, CrostonClassic); high cv with low structure argues for keeping expectations modest. Cheap -- returns in under a second.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and nondestructive nature. The description adds 'Cheap -- returns in under a second' and 'Read it as evidence, not as an answer,' providing behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with five sentences, front-loaded with purpose, and every sentence adds value. No redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the existence of an output schema (not shown but present), the description covers the semantic meaning of the features and how to interpret them. It also provides cost and time estimates, making it self-contained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already has detailed descriptions for all three parameters (handle, max_series, response_format). The tool description focuses on output features and usage advice, not parameter details. Since schema coverage is high, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a clear verb-resource pair: 'Compute the per-series features that decide which model family to try.' It lists the specific features computed, distinguishing this diagnostic tool from siblings like tsf_forecast or tsf_load_series.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit decision rules: 'high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models...' and positions the tool as 'evidence, not an answer.' This clearly guides when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_detect_anomaliesA
Read-onlyIdempotent

Flag historical points that fall outside a cross-validated prediction interval.

The detector model defines what "expected" means, so pick one that fits the series: a weak detector flags its own errors rather than real anomalies. Run tsf_describe_series or tsf_cross_validate first.

Returns flagged counts per series, a capped list of flagged rows, and the parquet path with the full result.

LONG-RUNNING, and the default is the expensive one: leaving n_windows unset refits the model once per observation across the whole history, which takes minutes even for a statistical model. Pass n_windows (e.g. 12) unless you genuinely need every point tested.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses long-running nature and expensive default (n_windows unset refits per observation, taking minutes). Annotations (readOnlyHint, idempotentHint) are consistent; description adds critical performance context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact four paragraphs with clear structure: purpose, prerequisites, output summary, performance warning. No redundant sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given tool complexity (cross-validation, long-running, return includes parquet path and capped list), the description covers prerequisites, output, and performance. Output schema exists, so return values are adequately summarized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Adds meaning beyond schema by explaining the default slowness of n_windows and the risk of weak models. Schema already has clear descriptions, but description provides crucial usage context for these parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it flags historical points outside a cross-validated prediction interval, distinguishes from sibling tools (tsf_describe_series, tsf_cross_validate, etc.), and warns about weak detectors.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises running tsf_describe_series or tsf_cross_validate first, warns against weak detectors, and gives concrete guidance on setting n_windows to avoid slow default behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_export_reportA

Render every step of the analysis as a report someone can read.

Covers the input and its hash, the features, each cross-validation with its metric table and ranking, the forecasts, any anomaly runs, and the pinned environment -- in the order they happened. HTML is self-contained, with no external stylesheet or script, so it opens correctly years later.

Call it after tsf_export_run at the end of an analysis. Pass a note: the report headlines it as the rationale, and a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.

Report from manifest_path instead of handle to re-render an older run -- it needs nothing but the manifest file.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (all hints false), so the description carries the burden. It discloses that HTML output is 'self-contained, with no external stylesheet or script' and that the report covers steps 'in the order they happened.' However, it does not clarify whether the tool writes a file to disk, returns the report content, or has other side effects. The output schema exists but is not described in the tool description, leaving behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a lead sentence defining the tool, then a list of contents, then usage order, then note advice, then alternative invocation. It is informative without being verbose. Minor inefficiency: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again' is slightly colorful but still earns its place. Could be tighter, but overall effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 optional parameters and an output schema (which handles return value documentation), the description covers its purpose, contents, usage order, and parameter trade-offs. It does not explain what happens if both handle and manifest_path are provided (mutual exclusion handled by schema? not specified). It also assumes the agent knows tsf_export_run was called, which is implied by 'at the end of an analysis.' Overall adequately complete for an AI agent to use correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has detailed descriptions for all parameters (note, format, handle, manifest_path), so baseline is 3. The description goes beyond by explaining the semantic purpose of note ('headlines it as the rationale') and the trade-off between handle and manifest_path ('re-renders an older run -- it needs nothing but the manifest file'). This adds actionable context for parameter selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description starts with a specific verb ('Render every step of the analysis as a report') and enumerates the exact contents (input hash, features, cross-validation, forecasts, anomalies, pinned environment). This immediately distinguishes it from sibling tools like tsf_export_run (which exports run data) and tsf_forecast (which only forecasts). The purpose is unambiguous and comprehensive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states ordering: 'Call it after tsf_export_run at the end of an analysis.' Provides clear alternatives: 'Report from manifest_path instead of handle to re-render an older run.' Also advises on best practice for the note parameter: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.' This gives the agent concrete when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_export_runA

Write a JSON manifest of everything done to this handle.

Records the source path and SHA-256, the frequency, every call with its arguments and artifact paths, the pinned package versions, and your note. This is the artifact of record: your prose in the conversation is lost, this file is not. Write the note -- say which model you picked, what the metric table showed, and what you rejected.

Call it at the end of any analysis someone might have to defend or rerun. Writes a file, so it is not read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds valuable context: 'Writes a file, so it is not read-only' and lists everything included in the manifest. It also emphasizes that the note is the only place a reasoning trace survives, which is important behavioral insight. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: starts with the core purpose, then enumerates contents, gives usage advice, and closes with a note about file writing. It is not overly long; every sentence contributes. A slight trim could improve conciseness, but it remains clear and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, the description correctly omits return value details. It covers when to call, what the manifest contains, and the critical role of the note. The only minor gap is no mention of potential side effects (e.g., overwriting existing files), but overall it is sufficiently complete for a tool with good annotations and schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already contains descriptions for both parameters: handle ('Handle whose run log should be pinned to a manifest') and note (detailed explanation of what to write). The tool description adds further guidance for the note, specifically: 'Write the note -- say which model you picked, what the metric table showed, and what you rejected.' This enhances the schema's descriptions, making it clear how to use the note parameter effectively.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence 'Write a JSON manifest of everything done to this handle' clearly states the verb (write a manifest) and the resource (handle). The description further details what the manifest includes (source path, SHA-256, calls, arguments, artifact paths, pinned package versions, note), differentiating it from sibling tools like 'tsf_export_report' or 'tsf_describe_series'. This makes the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells when to call it: 'Call it at the end of any analysis someone might have to defend or rerun.' This provides clear context for use. It does not explicitly state when not to use it or name alternatives, but for a specialized export tool the guidance is sufficient and well-placed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_forecastA
Read-onlyIdempotent

Fit on the full history and forecast h periods ahead with intervals.

Use after tsf_cross_validate has justified the model choice. The full forecast goes to parquet; the response carries the path, the column list, the row count and a small preview. Read the parquet for anything more -- raising max_preview_rows to dump the frame into the conversation is the one thing that reliably ruins a long session.

LONG-RUNNING for foundation models.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the bar is lowered. The description adds valuable context: the full forecast goes to parquet, the response carries path/column list/row count/preview, warns against raising max_preview_rows, and flags 'LONG-RUNNING for foundation models' — all beyond what annotations provide. No contradictions with annotations (readOnlyHint=true is consistent with generating forecasts without mutating data).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise: a terse three-sentence explanation that front-loads the core action, then adds usage guidance and behavioral warnings. Every sentence adds essential information without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (forecasting with multiple models and intervals), the output schema exists, so return values don't need elaboration. The description covers the critical workflow (use after cross-validation), output format (parquet with preview), and a key gotcha (don't dump full frame). It lacks explicit error conditions or prerequisite checks, but for the scope, it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the schema provides no parameter descriptions — so the description must compensate. Although the main description does not detail parameters, the parameter `max_preview_rows` receives meaningful context: 'Rows of the forecast to inline... read that instead of raising this.' Other parameters (handle, models, h, level) have descriptions in the schema via the JSON Schema, but since coverage is 0% (likely meaning no separate param list in the description), the main text does not clarify their meaning beyond what the schema already provides. Baseline 3 is appropriate as the description adds some context for max_preview_rows but not for others.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that the tool fits on full history and forecasts h periods ahead with intervals. It uses specific verbs like 'fit' and 'forecast' and explicitly identifies the resource as the time series forecast. However, it does not directly differentiate from siblings like tsf_cross_validate, though the usage guideline addresses that.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use after tsf_cross_validate has justified the model choice,' providing clear sequencing context and an alternative (cross-validation). It does not mention when not to use it or list specific alternatives for other tasks like anomaly detection, but the primary usage guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_list_modelsA
Read-onlyIdempotent

Probe which models actually import in this environment.

Call this before cross-validating so you never propose a model that cannot run here. Returns {available, statistical, foundation, unavailable}, where each unavailable entry carries the real reason -- some models are gated on the Python version, not merely absent.

The first call imports TimeCopilot and can take ~30 seconds; later calls are instant.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds important behavioral details beyond that: the first call may take ~30 seconds to import TimeCopilot, later calls are instant; it returns structured output with real reasons for unavailability (including Python version gating). No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and structured: first sentence states purpose, second gives usage guidance, third explains return structure, and fourth notes startup latency. Every sentence adds essential information, with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (the agent can rely on structured return type details), the description covers all necessary context: what the tool does, when to use it, its runtime behavior (latency), and the high-level shape of results. No gaps remain for a list/probe tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema contains descriptions for both parameters (family and include_unavailable) that are self-explanatory. The tool description does not add new parameter information—it only repeats that statistical models are cheap and always installed, which already appears in the schema. Schema coverage via inline descriptions is present, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states 'Probe which models actually import in this environment' and specifies the tool's role in preventing proposal of non-running models. It clearly distinguishes itself from siblings like cross_validate by giving a precise pre-check use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description contains an explicit directive: 'Call this before cross-validating so you never propose a model that cannot run here.' It also notes that statistical models are always installed, which helps with decision-making. No alternatives are listed, but the use case is crystal clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tsf_load_seriesA
Read-onlyIdempotent

Read a CSV or Parquet panel from disk and register it under a handle.

Call this first; every other tool takes the handle it returns. The file must be in Nixtla long format (unique_id, ds, y). Returns a compact JSON summary -- series count, inferred frequency, date range, missing values, and the SHA-256 of the source -- and nothing else: the data stays in the server so it never consumes your context.

Read the summary before choosing a horizon. If obs_per_series.min is small, a long horizon or many CV windows will not fit.

ParametersJSON Schema
NameRequiredDescriptionDefault
paramsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (readOnly, idempotent, non-destructive), the description reveals that data stays server-side ('never consumes your context') and that the return is a compact summary with specific fields. It also hints at the side effect of reusing a handle (replacing the panel). No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five sentences, each serving a distinct purpose: purpose, ordering, format, return details, caution. Front-loaded with the core action. No superfluous words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's role as the entry point for a time series workflow, the description covers purpose, required file format, return value (with summary contents), and a concrete usage caution. With an output schema present, the lack of detailed return structure is acceptable. The description is complete enough for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides detailed descriptions for all three parameters. The tool description adds the critical constraint that the file must be in 'Nixtla long format (unique_id, ds, y)', which is not in the schema. This adds meaningful value beyond the schema, justifying a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action: 'Read a CSV or Parquet panel from disk and register it under a handle.' It distinguishes from sibling tools by explicitly saying 'Call this first; every other tool takes the handle it returns.' This makes the purpose unambiguous and contextually positioned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit ordering ('Call this first'), explains the handle's role in subsequent tools, and gives a practical caution about horizon choices based on the summary output. This equips the agent with clear when-to-use and how-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedtsf_cross_validate
    • First observedtsf_describe_series
    • First observedtsf_detect_anomalies
    • First observedtsf_export_report
    • First observedtsf_export_run
    • First observedtsf_forecast
    • First observedtsf_list_models
    • First observedtsf_load_series

TDQS

A4.6/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a distinct and well-defined purpose within the time series forecasting workflow: loading, describing, listing models, cross-validating, forecasting, detecting anomalies, and exporting results. There is no overlap or ambiguity between tools.

Naming Consistency5/5

All tools follow a consistent pattern: the prefix 'tsf_' followed by a verb (and optional noun), all in snake_case. Examples include tsf_load_series, tsf_describe_series, tsf_cross_validate, and tsf_export_report. The naming is predictable and uniform.

Tool Count5/5

With 8 tools, the server covers a complete analysis pipeline without excess. Each tool is necessary and corresponds to a clear step in the workflow, from data loading to report generation. The count is well-scoped for the domain.

Completeness5/5

The tool surface covers the full lifecycle of a typical time series analysis: load data, explore features, check available models, cross-validate, forecast, detect anomalies, and export manifests/reports. There are no obvious gaps; the workflow feels self-contained and actionable.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Deterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.
    17
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables time-series analysis and forecasting through a structured tool catalogue, including data loading, quality repair, diagnostics, and forecasting with ARIMA, exponential smoothing, Chronos-2, Toto 2.0, and AutoML.
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to run TimesFM-3 forecasting workflows locally, including joint multivariate forecasting with known future drivers, backtesting against a baseline, what-if scenario comparison, and historical anomaly detection. It exposes the studio's tools and bundled public and synthetic demo datasets over MCP.
    Apache 2.0