Skip to main content
Glama

Data-snooping tests (SPA, Reality Check, StepM)

validate_reality_check

Data-snooping tests on every variant a search tried: Hansen's SPA p-value that the best beat the benchmark only by luck, White's Reality Check, and the variants Romano-Wolf StepM finds better. Send all variants tried, not only the winners. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
repsNoBootstrap draws; default 2000.
seedNoSampling seed; default 42.
alphaNoFamilywise error for StepM; default 0.05.
matrixNoReturns of every variant the search tried, one row per period, one column per variant.
benchmarkNoBenchmark return per period; default zero.
matrix_fileNoPath to a CSV or JSON with one numeric column per variant, instead of matrix.
block_lengthNoMean bootstrap block in periods; default round(n^(1/3)).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (destructiveHint=false, openWorldHint=true), and the description adds a genuine interpretive caveat beyond them: a deflated Sharpe or overfitting probability above/below a threshold "is not admission to anything and is not a forecast." That tells the agent the output is diagnostic, not a decision — useful context the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what the tool computes and then the input rule. Dense but each clause earns its place; minor run-on structure in the second sentence keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained; all seven parameters are fully described in the schema, and the description supplies the key input expectation (include every variant tried). For a stateless computation tool this is essentially complete, missing only explicit sibling routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters including defaults and bounds, and the baseline is 3. The description only reinforces the matrix requirement ("every variant a search tried"), which the schema already states, so it adds little beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (data-snooping tests) plus the three concrete tests it runs (SPA p-value, White's Reality Check, Romano-Wolf StepM), so the agent knows exactly what computation happens. It only implicitly separates itself from siblings like validate_deflated_sharpe and validate_overfitting via the closing caveat rather than stating routing outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Send all variants tried, not only the winners" is an explicit, actionable usage directive that tells the agent what input to gather, which most siblings don't provide. It stops short of naming when to pick this over validate_deflated_sharpe or validate_overfitting, so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.