Skip to main content
Glama

Validate overfitting (CSCV)

validate_overfitting

Probability of backtest overfitting (0 to 1) by CSCV: how often the in-sample best variant falls below the out-of-sample median. Needs every variant's returns (periods by variants); with summary statistics only, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNoSampling seed; default 42.
matrixYesReturns as fractions, one row per period, one column per variant.
n_splitsNoEven number of blocks, at least 2; default 16.
max_combinationsNoMost splits evaluated, up to 2000 (default).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedInput schema / properties / matrix / description
      Previous value: -"Returns as fractions: one row per period, one column per variant."New value: +"Returns as fractions, one row per period, one column per variant."
  2. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly=false, openWorld=true, destructive=false, which says little about a statistical computation. The description compensates with the required input shape and a meaningful interpretive caveat (threshold crossings are neither admission nor forecast), which is real behavioral guidance beyond structured fields. It stops short of runtime traits like cost, determinism, or failure modes when the matrix is malformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the metric and method, then requirements, then sibling routing, then the interpretation caveat. Three dense sentences with no filler, each carrying distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers inputs, the alternative tool, and result interpretation for a non-trivial statistical procedure. It could say more about assumptions (e.g. minimum period/variant counts) or determinism, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents matrix, seed, n_splits, and max_combinations. The description's 'periods by variants' restates the matrix schema description rather than adding syntax or constraints. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific outcome (probability of backtest overfitting, 0-1) and the method (CSCV), then defines the metric operationally as how often the in-sample best variant falls below the out-of-sample median. It explicitly distinguishes itself from validate_deflated_sharpe, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the input condition that selects this tool (every variant's returns required) and names the alternative for the other case (summary statistics only -> validate_deflated_sharpe). This is an explicit when/alternatives pairing, not implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.