Skip to main content
Glama

Validate deflated Sharpe

validate_deflated_sharpe

Deflated Sharpe ratio: the probability (0 to 1) that the selected strategy's Sharpe beats the best that luck gives across the variants tried, with the probabilistic Sharpe and that luck benchmark. Send the seven statistics or a return series. With every variant's returns use validate_overfitting; luck as a trial count, validate_luck_trials; a multiple-testing haircut, validate_haircut_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
skewNoSkewness of returns; 0 if Normal.
returnsNoPeriodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis.
observationsNoNumber of return observations.
periods_per_yearNoPeriods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly.
non_excess_kurtosisNoKurtosis, not excess kurtosis; 3 if Normal.
observed_sharpe_annualizedNoAnnualized Sharpe as observed.
effective_independent_trialsNoIndependent variants tried before choosing this one.
cross_trial_sharpe_sd_annualizedNoStandard deviation of annualized Sharpe across those trials.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changed
    • changedInput schema / properties / cross_trial_sharpe_sd_annualized / description
      Previous value: -"Standard deviation of the annualized Sharpe across those trials."New value: +"Standard deviation of annualized Sharpe across those trials."
    • changedInput schema / properties / periods_per_year / description
      Previous value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly."
    • changedInput schema / properties / returns / description
      Previous value: -"Periodic returns as fractions (0.01 is 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis fields."New value: +"Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis."
  2. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true and destructiveHint=false, so safety signalling is already covered. The description still adds real context beyond that: it explains the two valid input modes (statistics vs. return series) and warns that the resulting probability is not admission to anything and not a forecast. It does not mention auth, cost, or rate limits, but for a local statistical computation that gap is small.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The whole definition fits in three dense sentences with no filler; the routing alternatives are front-loaded mid-paragraph and the interpretive caveat closes it. It is slightly run-on and could be split, but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a tool with 8 optional parameters and no required inputs, the description covers purpose, the two input modes, sibling alternatives, and interpretation limits. It could go further on which combination of statistics is minimally sufficient, but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline would be 3, but the description adds the either/or input mode ('the seven statistics or a return series') that the schema only implies via the 'replaces' note on returns. This clarifies mutual exclusivity across the 8 all-optional parameters, which is meaningful routing information beyond the field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool computes: the probability (0-1) that the selected strategy's Sharpe beats the best luck produces across the variants tried, plus the probabilistic Sharpe and luck benchmark. It also names the sibling tools it is not (validate_overfitting, validate_luck_trials, validate_haircut_sharpe), so the agent can separate it from adjacent validators. It falls short of 5 only because it opens with a concept definition rather than a direct verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing: send the seven statistics or a return series; use validate_overfitting when every variant's returns are available; validate_luck_trials when luck is expressed as a trial count; validate_haircut_sharpe for a multiple-testing haircut. The alternative tools and the conditions selecting them are spelled out rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.