Skip to main content
Glama

Money Mind — the judge

Money Mind — Multiple Testing — free allowance, then $0.20

multipletest
Read-only

You tested N strategies and kept the best; is it noise. Use when you searched many strategies or parameter sets and kept the best one. Quoting that winner's solo p-value is the most common way a backtest lies. Give the number of things you tested and the winner's t-statistic; returns the family-wise p-value, the t you actually needed, and the t that pure RUNS NOW: served from a daily free allowance (250 left today), then $0.20 USDC on Base via x402. No account, no API key. Example request: {"n_tested": 1000, "best_t": 3.0}

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
best_tYesexample: 3.0
n_testedYesexample: 1000

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/destructive=false/openWorld=false, so the safety profile is covered. The description adds real context beyond that: output contents, the pricing model (free daily allowance then $0.20 USDC via x402), and that no account or API key is required. The one gap is a truncated sentence ('the t that pure'), which leaves a return value unexplained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core framing is front-loaded and useful, but the description is redundant ('You tested N strategies and kept the best' vs 'Use when you searched many strategies...') and the sentence 'the t that pure' is cut off mid-thought, and pricing/logistics clutter the entry.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the description does explain the key returns (family-wise p-value, needed t). It is nearly complete, losing a point only for the truncated return-value sentence and sparse handling of the example payload's meaning.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions only contain bare examples ('example: 3.0'), so they carry little semantics despite nominal 100% coverage. The description supplies the actual meaning: 'the number of things you tested' (n_tested) and 'the winner's t-statistic' (best_t), adding value the schema does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific statistical operation: correcting a winning backtest for multiple testing, returning the family-wise p-value and the t-stats needed. It's clearly a multiple-testing tool, distinguishable from siblings like samplesize and deflatedsharpe. It doesn't explicitly name a sibling it's not, so not a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Use when you searched many strategies or parameter sets and kept the best one,' plus the motivating failure mode (quoting a winner's solo p-value). Clear adoption context, but no explicit when-not or named-alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources