Skip to main content
Glama

Minimum backtest length

validate_backtest_length

Minimum backtest length (years) before the best of N independent trials is not expected to reach a target Sharpe by luck; with backtest_years, the most trials those years allow. For planning a search; once it has a result, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
backtest_yearsNoBacktest length in years, for the most trials it allows.
target_sharpe_annualizedNoIn-sample annualized Sharpe you would call a discovery; default 1.
effective_independent_trialsNoIndependent trials tried (backtests, parameter sets, ideas).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • changedInput schema / properties / backtest_years / description
      Previous value: -"Length of the backtest in years, to get the most independent trials it allows."New value: +"Backtest length in years, for the most trials it allows."
    • changedInput schema / properties / effective_independent_trials / description
      Previous value: -"Independent trials (backtests, parameter sets, ideas) tried; gives the minimum backtest length."New value: +"Independent trials tried (backtests, parameter sets, ideas)."
    • changedInput schema / properties / target_sharpe_annualized / description
      Previous value: -"In-sample annualized Sharpe you would take as a discovery; default 1."New value: +"In-sample annualized Sharpe you would call a discovery; default 1."
  2. First observed

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, destructiveHint=false, which is an odd profile for what the description frames as a pure planning calculation; the description does not reconcile that (no side effects, no persisted state, no external calls disclosed). It does add genuine interpretive context that outputs are not admissions or forecasts, which goes beyond the annotations. With annotations already present, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the sibling handoff are front-loaded, which is good, but the opening is a single long clause-stacked sentence that is hard to parse on first read. The closing sentence about thresholds not being admission is useful but borders on editorial and is not clearly tied to what this tool returns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers purpose, planning-time usage, and the handoff tool. For a stateless computation with fully documented parameters, that is close to sufficient; only the reconciliation of the readOnlyHint=false annotation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema with bounds and defaults. The description restates backtest_years ('with backtest_years, the most trials those years allow') without adding syntax or format meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific quantity (minimum backtest length in years) and the condition under which it holds (best of N independent trials not reaching target Sharpe by luck), which is far more specific than the title alone. It is somewhat syntactically dense and reads like a textbook definition rather than a plain 'what this does', and it does not directly contrast itself with siblings like validate_luck_trials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the setting explicitly ('For planning a search') and names the alternative with the condition that selects it ('once it has a result, use validate_deflated_sharpe'). It also adds a when-not caveat: a deflated Sharpe or overfitting probability past any threshold is not admission to anything, so it should not be used as a decision gate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.