tslab-mcp
tslab-mcp
Un servidor MCP que expone pronósticos de series temporales deterministas como herramientas, para que tu agente sea el motor de razonamiento y cada número provenga de Python ordinario y reproducible.
No se llama a ningún LLM en ningún lugar de este paquete. No se requiere ninguna clave API (a menos que solicites TimeGPT, que llama a la API de Nixtla).
Por qué
Algunas bibliotecas de pronóstico incluyen un agente que lee características, elige un modelo y explica el resultado con un LLM en el bucle. Llamar a uno de esos desde tu propio agente anida un agente dentro de otro agente: dos prompts, dos facturas, dos fuentes de no determinismo y una capa intermedia opaca que hace que la justificación de la selección del modelo no sea auditable.
Por lo tanto, aquí el control está invertido: la biblioteca de pronóstico es la herramienta, y tu agente es el que razona. Lee las características, argumenta a favor de una familia de modelos, valida cruzadamente los candidatos y escribe la justificación en un manifiesto. Cada número en el camino es producido por una llamada a la biblioteca que puedes re-ejecutar sin un LLM en la ruta.
Esa división se traslada a cómo se construye el propio paquete. La instalación base ejecuta once modelos estadísticos — AutoARIMA, AutoETS, Theta, CrostonClassic y amigos — a través de statsforecast: aproximadamente 340 MB, sin PyTorch, y se inicia en segundos. Un extra opcional foundation añade los modelos preentrenados de TimeCopilot — Chronos, Moirai, TimesFM, TiRex, Toto y otros — además de Prophet, para cuando una línea base estadística no es suficiente. Una solicitud que solo nombra modelos estadísticos nunca importa TimeCopilot o torch; una solicitud que nombra incluso un modelo fundacional se ejecuta completamente a través de TimeCopilot, que también incluye los modelos estadísticos. De cualquier manera, tsf_list_models informa qué está realmente instalado antes de que te comprometas con un modelo.
Related MCP server: timeseries-mcp
Instalar
Requiere Python 3.10+ (se recomienda 3.13, consulte Versión de Python).
uvx tslab-mcp # run without installing
uv tool install tslab-mcp # or install the CLILa instalación base ejecuta los once modelos estadísticos a través de statsforecast: aproximadamente 340 MB, sin PyTorch, y se inicia instantáneamente. Para los modelos fundacionales preentrenados — Chronos, Moirai, TimesFM, Toto, TiRex — y Prophet, añade el extra:
uvx --from 'tslab-mcp[foundation]' tslab-mcpEl extra
foundationtrae TimeCopilot, que a su vez trae torch, transformers y lightning: aproximadamente 2 GB en la primera instalación, y la primera llamada a la herramienta que lo toca tarda ~30 segundos en importar. Ambos son únicos, y ninguno es de pago a menos que solicites un modelo que los necesite.
Desde GitHub
uv y uvx aceptan una URL de git en lugar de un nombre de paquete, lo que instala el main actual sin esperar un lanzamiento:
uvx --from git+https://github.com/pedrobtz/tslab-mcp tslab-mcp
uv tool install git+https://github.com/pedrobtz/tslab-mcp # or install the CLI
# with the foundation extra
uvx --from 'tslab-mcp[foundation] @ git+https://github.com/pedrobtz/tslab-mcp' tslab-mcpFija una referencia para cualquier cosa que no sea una prueba casual — la cabeza de la rama puede moverse debajo de ti de otra manera. Un commit funciona hoy; una etiqueta de versión también funcionará una vez que se haya creado:
uv tool install "git+https://github.com/pedrobtz/tslab-mcp@136824c1cc2a"Desde un clon
git clone https://github.com/pedrobtz/tslab-mcp
cd tslab-mcp
uv sync # base
uv sync --extra foundation # with the pretrained models
uv run tslab-mcpConfigurar
Añade el servidor a la configuración de tu cliente MCP. El archivo difiere por cliente — a menudo .mcp.json en la raíz del proyecto — pero la entrada en sí tiene la misma forma:
{
"mcpServers": {
"tslab": {
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "~/.tslab-mcp"
}
}
}
}TSLAB_MCP_HOME establece dónde se escriben los artefactos; por defecto es ~/.tslab-mcp, y las salidas de ejecución se guardan en <home>/runs.
El transporte es solo stdio, por diseño: se asume que tus datos son sensibles y nunca salen de la máquina. El servidor no realiza solicitudes salientes excepto las descargas de pesos de modelo que el propio TimeCopilot realiza para modelos fundacionales, y las llamadas a la API de Nixtla que TimeGPT realiza si lo solicitas específicamente.
GitHub Copilot
Copilot descubre servidores MCP desde un archivo mcp.json y expone sus herramientas en el modo agente — las herramientas no aparecen en el modo de preguntar o editar.
VS Code. Coloca el servidor en .vscode/mcp.json para compartirlo con el repositorio, o ejecuta MCP: Open User Configuration desde la Paleta de Comandos para mantenerlo en tu propio perfil en todos los espacios de trabajo. Ten en cuenta que la clave es servers, no mcpServers:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uvx",
"args": ["tslab-mcp"],
"env": {
"TSLAB_MCP_HOME": "${userHome}/.tslab-mcp"
}
}
}
}Desde un clon, apúntalo al árbol de trabajo en su lugar:
{
"servers": {
"tslab": {
"type": "stdio",
"command": "uv",
"args": ["run", "--directory", "${workspaceFolder}", "tslab-mcp"]
}
}
}Luego: abre Chat, cambia el selector de modo a Agente, y usa el botón Herramientas para confirmar que las ocho herramientas tsf_* están listadas y habilitadas. MCP: List Servers muestra el estado del servidor y sus registros, que es donde se explica un inicio fallido. Copilot limita cuántas herramientas pueden estar activas a la vez, por lo que si ejecutas varios servidores MCP, es posible que debas deseleccionar algunos para que quepan las ocho.
Visual Studio. Misma forma JSON, en .mcp.json en la raíz de la solución (o %USERPROFILE%\.mcp.json para todas las soluciones), luego habilita las herramientas desde el selector de herramientas del modo agente de Copilot Chat.
JetBrains, Eclipse y Xcode. Abre el selector de herramientas del modo agente de Copilot Chat, elige Edit MCP configuration y añade la misma entrada servers al mcp.json que se abre.
Copilot coding agent (el agente en la nube en github.com) es una mala opción para este servidor: ejecuta tus servidores MCP dentro de un entorno efímero de GitHub Actions, lo que significa pagar la instalación de ~2 GB de TimeCopilot en cada ejecución, y no tiene acceso a archivos de datos locales. Úsalo desde tu editor en su lugar.
Herramientas
Herramienta | Propósito | Devuelve |
| Leer CSV/Parquet, validar el contrato | Resumen JSON + SHA-256 |
| Características por serie para elegir una familia de modelos | Tabla Markdown o JSON, limitada por filas |
| Probar qué modelos se importan realmente aquí |
|
| Comparación de origen rodante entre modelos | Tabla de métricas, clasificación, ruta parquet |
| Ajustar y pronosticar con intervalos de predicción | Ruta parquet + vista previa acotada |
| Marcado de intervalos validados cruzadamente | Conteos, lista de banderas acotada, ruta parquet |
| Fijar la sesión a un manifesto re-ejecutable | Ruta del manifiesto |
| Renderizar cada paso como un informe legible | Ruta HTML o Markdown |
Todo excepto las dos herramientas tsf_export_* está marcado como solo lectura; nada aquí elimina, por lo que limpiar ~/.tslab-mcp/runs es tu responsabilidad, no la del agente.
Iniciar una sesión
Las herramientas no imponen un orden, por lo que el prompt inicial es lo que convierte ocho funciones invocables en un análisis. Algo como esto funciona bien:
Usa las herramientas de tslab para pronosticar la serie en
/Users/me/data/deposits.csv, 12 meses adelante.Trabaja en este orden y muestra tu razonamiento en cada paso:
Carga el archivo y dime lo que encontraste — cuántas series, qué frecuencia, si hay huecos o valores faltantes.
Describe las características, y di qué familias de modelos sugieren, y por qué.
Verifica qué modelos están realmente instalados antes de proponer ninguno.
Valida cruzadamente tu lista corta contra una línea base SeasonalNaive en 4 ventanas. Solo modelos estadísticos por ahora.
Pronostica con el ganador, con intervalos del 80% y 95%.
Exporta un manifiesto de ejecución y un informe HTML, y pon la justificación de la selección del modelo en la nota: qué elegiste, qué mostró la tabla de métricas y qué rechazaste.
Resume los resultados y dame las rutas parquet — no pegues marcos completos en el chat.
Cuatro cosas en ese prompt están haciendo trabajo real:
Una ruta absoluta. Las rutas relativas se resuelven contra el directorio de trabajo del servidor, que tu cliente MCP elige y generalmente no puedes predecir.
Un horizonte que coincide con la decisión.
himpulsa tanto el pronóstico como cuánto historial consume cada ventana de CV; 12 pasos mensuales es un año de planificación, no un valor predeterminado arbitrario."Solo modelos estadísticos por ahora." Sin esto, un agente puede recurrir a un modelo fundacional y pasar varios minutos descargando pesos para responder una pregunta que
AutoETShabría resuelto en segundos. Levanta la restricción una vez que los modelos baratos hayan establecido un piso.Pedir la justificación en la nota del manifiesto. La transcripción del chat es desechable; el manifiesto es la parte que alguien puede re-ejecutar y auditar. Si el razonamiento solo existe en la conversación, está efectivamente perdido.
Aperturas más cortas, cuando sabes lo que quieres:
Carga
/Users/me/data/sales.parquety describe las características. No pronostiques todavía — quiero ver con qué estamos tratando primero.
Compara SeasonalNaive, AutoETS y AutoARIMA en el manejador
depositscargado, en 6 ventanas con h=12, luego dime si algo supera la línea base lo suficiente como para que valga la pena la complejidad adicional.
Las llamadas solo estadísticas responden en segundos. La primera llamada que nombra un modelo fundacional tarda ~30 segundos en importar TimeCopilot antes de hacer cualquier otra cosa — esa pausa es esperada, no un bloqueo, y solo ocurre si el extra foundation está instalado y una solicitud realmente recurre a uno.
Una sesión de trabajo
Comienza desde un CSV en formato largo de Nixtla:
unique_id,ds,y
branch_01,2018-01-01,1043.2
branch_01,2018-02-01,1102.7
...1. Cargarlo. El panel permanece en el proceso del servidor; el manejador es todo lo que lleva la sesión.
{"handle": "deposits", "n_series": 12, "n_obs": 864, "freq": "MS",
"start": "2018-01-01T00:00:00", "end": "2023-12-01T00:00:00",
"obs_per_series": {"min": 72, "median": 72, "max": 72},
"n_missing_y": 0, "sha256": "9f2c…"}2. Describirlo. Estos son los números sobre los que razonas.
| id | n | mean | cv | %zero | trend | seasonal | acf1(diff) |
|-----------|----|--------|-------|-------|-------|----------|------------|
| branch_01 | 72 | 1180.4 | 0.112 | 0.0 | 0.83 | 0.62 | -0.31 |Una alta fuerza estacional y una tendencia clara abogan por AutoETS y AutoARIMA sobre una línea base ingenua; un alto %zero habría abogado por ADIDA o CrostonClassic en su lugar.
seasonal es una fuerza STL — el componente estacional medido contra lo que queda una vez que se elimina la tendencia — por lo que una serie creciente aún informa su estacionalidad honestamente. Lleva un piso de ruido de aproximadamente 0.3–0.5: las puntuaciones en esa banda significan "sin evidencia", no "levemente estacional".
3. Verificar qué está instalado con tsf_list_models, para que nunca propongas un modelo que esta máquina no pueda ejecutar.
4. Validar cruzadamente los candidatos — incluyendo siempre SeasonalNaive, ya que un modelo que no puede superarlo no vale la pena implementarlo:
{"kind": "cross_validation", "models": ["SeasonalNaive", "AutoETS", "AutoARIMA"],
"h": 12, "n_windows": 4, "seasonality_used_for_mase": 12,
"metrics": {"mase": {"SeasonalNaive": 1.0, "AutoETS": 0.71, "AutoARIMA": 0.68}},
"ranking": {"mase": ["AutoARIMA", "AutoETS", "SeasonalNaive"]},
"artifact": "~/.tslab-mcp/runs/cv_deposits_3f1a9c02.parquet"}5. Pronosticar con el ganador. El marco completo va a parquet; la respuesta lleva la ruta, las columnas y una vista previa corta.
6. Exportar la ejecución y el informe. Escribe por qué, en la nota — es la única parte de tu razonamiento que sobrevive a la conversación:
{"manifest": "~/.tslab-mcp/runs/manifest_deposits_77b0e415.json", "n_runs": 3,
"kinds": ["cross_validation", "forecast"]}El manifiesto contiene la ruta de origen y el hash, la frecuencia, cada llamada con sus argumentos y rutas de artefactos, las versiones fijadas de lo que sea que esté realmente instalado — statsforecast, pandas y Python siempre; TimeCopilot y torch también si el extra foundation está incluido — y tu nota. Es suficiente para reproducir los números con el servidor detenido.
tsf_export_report convierte ese mismo manifiesto en algo que una persona lee — características, tablas de métricas ordenadas de mejor a peor, pronósticos, anomalías y el entorno, en el orden en que ocurrieron:
{"report": "~/.tslab-mcp/runs/report_deposits_5c31d0a7.html",
"format": "html", "n_steps": 3,
"steps": ["features", "cross_validation", "forecast"]}El informe es una función pura del manifiesto: no lee parquet ni llama a ningún modelo, por lo que tsf_export_report con manifest_path re-renderiza una ejecución de hace meses sin nada cargado. El HTML incrusta su propio CSS y no referencia ningún script, hoja de estilo o fuente externa, por lo que aún se abre correctamente sin conexión.
Diseño
Cuatro invariantes, y las razones por las que existen:
Manejadores, no dataframes. Un marco de validación cruzada tiene
n_series × h × n_windows × n_models filas. Serializarlo como resultado de una
herramienta agota el contexto de la sesión en la primera llamada y empeora cada
turno posterior. Las herramientas toman un manejador y devuelven resúmenes,
agregados y rutas de archivos; cada ruta masiva está limitada e informa lo que
omitió, para que la sesión sepa que debe leer el parquet en lugar de volver a
preguntar.
El trabajo bloqueante nunca toca el bucle de eventos. Validar
cruzadamente varios modelos sobre un panel grande son minutos de CPU. Cada
cuerpo de herramienta es una clausura síncrona enviada a través de
anyio.to_thread.run_sync, por lo que el transporte stdio sigue respondiendo
y el cliente no abandona el servidor a mitad de la ejecución.
El entorno se descubre, no se asume. Los modelos se importan de forma
perezosa y se sondean, nunca se asume que están presentes. tsf_list_models
informa lo que realmente se resolvió aquí, por lo que pedir Chronos sin el
extra devuelve un mensaje nombrando el extra en lugar de un traceback diez
minutos después de iniciada una ejecución.
El backend se elige según lo que pidas: una solicitud cuyos modelos son todos estadísticos se ejecuta a través de statsforecast, y solo una solicitud que necesita un modelo preentrenado recurre a TimeCopilot. Por lo tanto, las ejecuciones estadísticas nunca importan torch, y el servidor se inicia al instante de cualquier manera.
statsforecast se deja deliberadamente en su n_jobs=1 predeterminado. Su modo
paralelo genera procesos trabajadores que reimportan el módulo de entrada, lo
que dentro de un servidor MCP genera contención y un riesgo en stdout en lugar
de velocidad.
El manifiesto es el artefacto de registro. La prosa en la conversación es comentario. El manifiesto es lo que alguien reejecuta en seis meses, y lo que un revisor lee para ver qué modelos se compararon y sobre qué base.
Versión de Python
TimeCopilot condiciona varios modelos a la versión del intérprete, y en
Python < 3.13 fija tabpfn-time-series, lo que limita pandas por debajo de
2.2.
Python | Modelos | pandas |
3.13 | todo excepto | ≥ 2.2 |
3.10–3.12 | añade | < 2.2 |
3.13 es el objetivo recomendado. En cualquier caso, tsf_list_models informa
lo que realmente se resolvió, con la razón de todo lo que no se resolvió.
Desarrollo
uv sync --all-groups
uv run pytest # fast suite
uv run pytest -m slow # exercises TimeCopilot; slower, no weight downloads
uv run ruff check src tests
uv run mypyInspecciona la superficie de las herramientas con el MCP Inspector:
npx @modelcontextprotocol/inspector uv run tslab-mcpLicencia
MIT
Available Tools
8 toolstsf_cross_validateARead-onlyIdempotent
Compare models by rolling-origin cross-validation.
This is the tool that replaces guesswork about model choice: it produces the evidence, you read the table and decide. Always include SeasonalNaive as the baseline -- a model that cannot beat it is not worth deploying.
Returns a per-model metric table aggregated over series and windows, a best-first ranking per metric, and the parquet path holding every per-window prediction. LONG-RUNNING: seconds for statistical models, many minutes for foundation models on a large panel. Start with statistical models on the real horizon before reaching for anything pretrained.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, idempotent, non-destructive behavior. The description adds crucial behavioral context: runtime warning ('LONG-RUNNING'), output description (aggregated table, ranking, parquet path), and implicitly that it is safe but compute-intensive. This goes well beyond the annotations and aids agent expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, purposeful paragraphs. The first sentence immediately states the tool's function. Each subsequent section (usage advice, output details, runtime warning) earns its place with no redundancy. It is front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (cross-validation, multiple models, windows, output schema exists), the description covers purpose, usage, baseline recommendation, output contents (aggregated metrics, rankings, prediction parquet), and runtime behavior. It is sufficiently complete for an agent to understand when and how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The nested input schema (CrossValidateInput) already provides parameter descriptions (e.g., models, metrics). The description adds high-level advice (like horizon matching the decision, baseline recommendation) but no new parameter-level semantics beyond what the schema offers. Schema coverage is effectively high despite the 0% top-level stat, so the description's incremental value here is moderate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource combination ('Compare models by rolling-origin cross-validation') and clearly distinguishes this tool from siblings like tsf_forecast (single model forecast) and tsf_list_models (model names). It states it replaces guesswork about model choice, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Always include SeasonalNaive as the baseline' and 'Start with statistical models on the real horizon before reaching for anything pretrained.' It also explains that the tool produces evidence for model selection. However, it does not explicitly state when not to use this tool (e.g., for final forecasts) or name alternatives, slightly reducing completeness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_describe_seriesARead-onlyIdempotent
Compute the per-series features that decide which model family to try.
Returns length, mean, sd, coefficient of variation, share of zeros, trend strength (R-squared against time), seasonal strength (variance explained by the period means), and lag-1 autocorrelation of the differenced series.
Read it as evidence, not as an answer: high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models (ADIDA, IMAPA, CrostonClassic); high cv with low structure argues for keeping expectations modest. Cheap -- returns in under a second.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and nondestructive nature. The description adds 'Cheap -- returns in under a second' and 'Read it as evidence, not as an answer,' providing behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with five sentences, front-loaded with purpose, and every sentence adds value. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the existence of an output schema (not shown but present), the description covers the semantic meaning of the features and how to interpret them. It also provides cost and time estimates, making it self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has detailed descriptions for all three parameters (handle, max_series, response_format). The tool description focuses on output features and usage advice, not parameter details. Since schema coverage is high, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a clear verb-resource pair: 'Compute the per-series features that decide which model family to try.' It lists the specific features computed, distinguishing this diagnostic tool from siblings like tsf_forecast or tsf_load_series.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit decision rules: 'high seasonal_strength argues for SeasonalNaive or AutoETS; a high pct_zero argues for the intermittent-demand models...' and positions the tool as 'evidence, not an answer.' This clearly guides when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_detect_anomaliesARead-onlyIdempotent
Flag historical points that fall outside a cross-validated prediction interval.
The detector model defines what "expected" means, so pick one that fits the series: a weak detector flags its own errors rather than real anomalies. Run tsf_describe_series or tsf_cross_validate first.
Returns flagged counts per series, a capped list of flagged rows, and the parquet path with the full result.
LONG-RUNNING, and the default is the expensive one: leaving n_windows unset refits the model once per observation across the whole history, which takes minutes even for a statistical model. Pass n_windows (e.g. 12) unless you genuinely need every point tested.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses long-running nature and expensive default (n_windows unset refits per observation, taking minutes). Annotations (readOnlyHint, idempotentHint) are consistent; description adds critical performance context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact four paragraphs with clear structure: purpose, prerequisites, output summary, performance warning. No redundant sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given tool complexity (cross-validation, long-running, return includes parquet path and capped list), the description covers prerequisites, output, and performance. Output schema exists, so return values are adequately summarized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning beyond schema by explaining the default slowness of n_windows and the risk of weak models. Schema already has clear descriptions, but description provides crucial usage context for these parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it flags historical points outside a cross-validated prediction interval, distinguishes from sibling tools (tsf_describe_series, tsf_cross_validate, etc.), and warns about weak detectors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises running tsf_describe_series or tsf_cross_validate first, warns against weak detectors, and gives concrete guidance on setting n_windows to avoid slow default behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_reportA
Render every step of the analysis as a report someone can read.
Covers the input and its hash, the features, each cross-validation with its metric table and ranking, the forecasts, any anomaly runs, and the pinned environment -- in the order they happened. HTML is self-contained, with no external stylesheet or script, so it opens correctly years later.
Call it after tsf_export_run at the end of an analysis. Pass a note: the report headlines it as the rationale, and a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.
Report from manifest_path instead of handle to re-render an older run --
it needs nothing but the manifest file.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (all hints false), so the description carries the burden. It discloses that HTML output is 'self-contained, with no external stylesheet or script' and that the report covers steps 'in the order they happened.' However, it does not clarify whether the tool writes a file to disk, returns the report content, or has other side effects. The output schema exists but is not described in the tool description, leaving behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a lead sentence defining the tool, then a list of contents, then usage order, then note advice, then alternative invocation. It is informative without being verbose. Minor inefficiency: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again' is slightly colorful but still earns its place. Could be tighter, but overall effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 optional parameters and an output schema (which handles return value documentation), the description covers its purpose, contents, usage order, and parameter trade-offs. It does not explain what happens if both handle and manifest_path are provided (mutual exclusion handled by schema? not specified). It also assumes the agent knows tsf_export_run was called, which is implied by 'at the end of an analysis.' Overall adequately complete for an AI agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has detailed descriptions for all parameters (note, format, handle, manifest_path), so baseline is 3. The description goes beyond by explaining the semantic purpose of note ('headlines it as the rationale') and the trade-off between handle and manifest_path ('re-renders an older run -- it needs nothing but the manifest file'). This adds actionable context for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description starts with a specific verb ('Render every step of the analysis as a report') and enumerates the exact contents (input hash, features, cross-validation, forecasts, anomalies, pinned environment). This immediately distinguishes it from sibling tools like tsf_export_run (which exports run data) and tsf_forecast (which only forecasts). The purpose is unambiguous and comprehensive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states ordering: 'Call it after tsf_export_run at the end of an analysis.' Provides clear alternatives: 'Report from manifest_path instead of handle to re-render an older run.' Also advises on best practice for the note parameter: 'a table of numbers without the reasoning is what makes a reviewer ask for the whole thing again.' This gives the agent concrete when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_export_runA
Write a JSON manifest of everything done to this handle.
Records the source path and SHA-256, the frequency, every call with its arguments and artifact paths, the pinned package versions, and your note. This is the artifact of record: your prose in the conversation is lost, this file is not. Write the note -- say which model you picked, what the metric table showed, and what you rejected.
Call it at the end of any analysis someone might have to defend or rerun. Writes a file, so it is not read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false and destructiveHint=false. The description adds valuable context: 'Writes a file, so it is not read-only' and lists everything included in the manifest. It also emphasizes that the note is the only place a reasoning trace survives, which is important behavioral insight. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: starts with the core purpose, then enumerates contents, gives usage advice, and closes with a note about file writing. It is not overly long; every sentence contributes. A slight trim could improve conciseness, but it remains clear and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description correctly omits return value details. It covers when to call, what the manifest contains, and the critical role of the note. The only minor gap is no mention of potential side effects (e.g., overwriting existing files), but overall it is sufficiently complete for a tool with good annotations and schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already contains descriptions for both parameters: handle ('Handle whose run log should be pinned to a manifest') and note (detailed explanation of what to write). The tool description adds further guidance for the note, specifically: 'Write the note -- say which model you picked, what the metric table showed, and what you rejected.' This enhances the schema's descriptions, making it clear how to use the note parameter effectively.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Write a JSON manifest of everything done to this handle' clearly states the verb (write a manifest) and the resource (handle). The description further details what the manifest includes (source path, SHA-256, calls, arguments, artifact paths, pinned package versions, note), differentiating it from sibling tools like 'tsf_export_report' or 'tsf_describe_series'. This makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to call it: 'Call it at the end of any analysis someone might have to defend or rerun.' This provides clear context for use. It does not explicitly state when not to use it or name alternatives, but for a specialized export tool the guidance is sufficient and well-placed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_forecastARead-onlyIdempotent
Fit on the full history and forecast h periods ahead with intervals.
Use after tsf_cross_validate has justified the model choice. The full forecast goes to parquet; the response carries the path, the column list, the row count and a small preview. Read the parquet for anything more -- raising max_preview_rows to dump the frame into the conversation is the one thing that reliably ruins a long session.
LONG-RUNNING for foundation models.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the bar is lowered. The description adds valuable context: the full forecast goes to parquet, the response carries path/column list/row count/preview, warns against raising max_preview_rows, and flags 'LONG-RUNNING for foundation models' — all beyond what annotations provide. No contradictions with annotations (readOnlyHint=true is consistent with generating forecasts without mutating data).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise: a terse three-sentence explanation that front-loads the core action, then adds usage guidance and behavioral warnings. Every sentence adds essential information without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (forecasting with multiple models and intervals), the output schema exists, so return values don't need elaboration. The description covers the critical workflow (use after cross-validation), output format (parquet with preview), and a key gotcha (don't dump full frame). It lacks explicit error conditions or prerequisite checks, but for the scope, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema provides no parameter descriptions — so the description must compensate. Although the main description does not detail parameters, the parameter `max_preview_rows` receives meaningful context: 'Rows of the forecast to inline... read that instead of raising this.' Other parameters (handle, models, h, level) have descriptions in the schema via the JSON Schema, but since coverage is 0% (likely meaning no separate param list in the description), the main text does not clarify their meaning beyond what the schema already provides. Baseline 3 is appropriate as the description adds some context for max_preview_rows but not for others.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool fits on full history and forecasts h periods ahead with intervals. It uses specific verbs like 'fit' and 'forecast' and explicitly identifies the resource as the time series forecast. However, it does not directly differentiate from siblings like tsf_cross_validate, though the usage guideline addresses that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Use after tsf_cross_validate has justified the model choice,' providing clear sequencing context and an alternative (cross-validation). It does not mention when not to use it or list specific alternatives for other tasks like anomaly detection, but the primary usage guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_list_modelsARead-onlyIdempotent
Probe which models actually import in this environment.
Call this before cross-validating so you never propose a model that cannot run here. Returns {available, statistical, foundation, unavailable}, where each unavailable entry carries the real reason -- some models are gated on the Python version, not merely absent.
The first call imports TimeCopilot and can take ~30 seconds; later calls are instant.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as readOnly, idempotent, and non-destructive. The description adds important behavioral details beyond that: the first call may take ~30 seconds to import TimeCopilot, later calls are instant; it returns structured output with real reasons for unavailability (including Python version gating). No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and structured: first sentence states purpose, second gives usage guidance, third explains return structure, and fourth notes startup latency. Every sentence adds essential information, with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists (the agent can rely on structured return type details), the description covers all necessary context: what the tool does, when to use it, its runtime behavior (latency), and the high-level shape of results. No gaps remain for a list/probe tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema contains descriptions for both parameters (family and include_unavailable) that are self-explanatory. The tool description does not add new parameter information—it only repeats that statistical models are cheap and always installed, which already appears in the schema. Schema coverage via inline descriptions is present, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Probe which models actually import in this environment' and specifies the tool's role in preventing proposal of non-running models. It clearly distinguishes itself from siblings like cross_validate by giving a precise pre-check use case.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description contains an explicit directive: 'Call this before cross-validating so you never propose a model that cannot run here.' It also notes that statistical models are always installed, which helps with decision-making. No alternatives are listed, but the use case is crystal clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tsf_load_seriesARead-onlyIdempotent
Read a CSV or Parquet panel from disk and register it under a handle.
Call this first; every other tool takes the handle it returns. The file must be in Nixtla long format (unique_id, ds, y). Returns a compact JSON summary -- series count, inferred frequency, date range, missing values, and the SHA-256 of the source -- and nothing else: the data stays in the server so it never consumes your context.
Read the summary before choosing a horizon. If obs_per_series.min is small, a long horizon or many CV windows will not fit.
| Name | Required | Description | Default |
|---|---|---|---|
| params | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations (readOnly, idempotent, non-destructive), the description reveals that data stays server-side ('never consumes your context') and that the return is a compact summary with specific fields. It also hints at the side effect of reusing a handle (replacing the panel). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Five sentences, each serving a distinct purpose: purpose, ordering, format, return details, caution. Front-loaded with the core action. No superfluous words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's role as the entry point for a time series workflow, the description covers purpose, required file format, return value (with summary contents), and a concrete usage caution. With an output schema present, the lack of detailed return structure is acceptable. The description is complete enough for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides detailed descriptions for all three parameters. The tool description adds the critical constraint that the file must be in 'Nixtla long format (unique_id, ds, y)', which is not in the schema. This adds meaningful value beyond the schema, justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action: 'Read a CSV or Parquet panel from disk and register it under a handle.' It distinguishes from sibling tools by explicitly saying 'Call this first; every other tool takes the handle it returns.' This makes the purpose unambiguous and contextually positioned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit ordering ('Call this first'), explains the handle's role in subsequent tools, and gives a practical caution about horizon choices based on the summary output. This equips the agent with clear when-to-use and how-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
tsf_cross_validate - First observed
tsf_describe_series - First observed
tsf_detect_anomalies - First observed
tsf_export_report - First observed
tsf_export_run - First observed
tsf_forecast - First observed
tsf_list_models - First observed
tsf_load_series
TDQS
Scored across 8 tools
Each tool has a distinct and well-defined purpose within the time series forecasting workflow: loading, describing, listing models, cross-validating, forecasting, detecting anomalies, and exporting results. There is no overlap or ambiguity between tools.
All tools follow a consistent pattern: the prefix 'tsf_' followed by a verb (and optional noun), all in snake_case. Examples include tsf_load_series, tsf_describe_series, tsf_cross_validate, and tsf_export_report. The naming is predictable and uniform.
With 8 tools, the server covers a complete analysis pipeline without excess. Each tool is necessary and corresponds to a clear step in the workflow, from data loading to report generation. The count is well-scoped for the domain.
The tool surface covers the full lifecycle of a typical time series analysis: load data, explore features, check available models, cross-validate, forecast, detect anomalies, and export manifests/reports. There are no obvious gaps; the workflow feels self-contained and actionable.
Maintenance
Related MCP Connectors
Probabilistic time-series forecasts from zero-shot foundation models: routed, single or ensembled.
1PredictOracle - 12 forecasting tools: time-series, scenario analysis, risk projections.
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Deterministic time tools for AI agents: timezone conversion, business-day math, cron interpretation.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnable any AI agent to forecast time-series data (e.g., sales, traffic) using Google's TimesFM or a zero-dependency statistical baseline.3Apache 2.0
- AlicenseAqualityDmaintenanceDeterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.17MIT
- AlicenseNot gradedqualityCmaintenanceEnables time-series analysis and forecasting through a structured tool catalogue, including data loading, quality repair, diagnostics, and forecasting with ARIMA, exponential smoothing, Chronos-2, Toto 2.0, and AutoML.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to run TimesFM-3 forecasting workflows locally, including joint multivariate forecasting with known future drivers, backtesting against a baseline, what-if scenario comparison, and historical anomaly detection. It exposes the studio's tools and bundled public and synthetic demo datasets over MCP.Apache 2.0