Skip to main content
Glama

TradingCalc MCP: Options, Forex, Risk Stats, Prediction Markets, On-Chain & Crypto Futures

Deflated Sharpe Ratio

workflow.run_dsr
Read-only

Deflated Sharpe Ratio (Bailey & Lopez de Prado): given how many strategy variants you tried (and how correlated they are), what Sharpe ratio would the best of N clear by luck alone, and does your actual strategy still clear that higher bar? Use when user asks "is my backtested Sharpe ratio real, or did I get lucky trying many variants?" or "how many independent trials does this really represent?". Provide either trial_sharpes[] (the N trials' own observed Sharpe ratios, most rigorous) or n_trials (+ optional avg_correlation to correct for correlated trials via Kish's design effect). Returns: expected_max_sharpe (the luck-alone threshold), dsr (probability your strategy's true Sharpe exceeds it, 0-1), n_trials_effective.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
returnsYesThe candidate strategy's own return series
n_trialsNoNumber of strategy variants tried, if trial_sharpes were not tracked individually
trial_sharpesNoObserved Sharpe ratios of all N trials tried, if tracked. Takes priority over n_trials if both are given.
avg_correlationNoAverage pairwise correlation between trials, 0-1 (default 0 = independent). Only used with n_trials; lowers the effective trial count via Kish's design effect.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safe-read profile (readOnlyHint, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds substantive method context: how correlated trials are handled (Kish's design effect) and the priority rule between trial_sharpes and n_trials. It doesn't state the trial_sharpes-priority caveat is enforced by the tool, but the intent is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads method and the core question, then usage, then inputs, then return values. Efficient for a statistically dense tool, though the run-on sentence in the usage trigger is slightly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description supplies the return fields (expected_max_sharpe, dsr, n_trials_effective) with plain-language meaning and the 0-1 range for dsr. For a four-parameter statistical tool with no output schema, this is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already carries parameter semantics; baseline is 3. The description earns above baseline by explaining the conditional relationship (avg_correlation only meaningful with n_trials, lowers effective count) and asserting trial_sharpes takes priority over n_trials — semantics beyond the field descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific statistical operation (deflating a Sharpe ratio for multiple trials) with the exact question it answers, and cites the method authors. Clearly distinguishable from sibling workflow.run_sharpe_stats, which computes plain Sharpe stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger questions ('is my backtested Sharpe ratio real, or did I get lucky') and an explicit selection rule between the two input modes (trial_sharpes most rigorous, else n_trials + avg_correlation). Nothing left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.