Skip to main content
Glama

Money Mind — the judge

Money Mind — Judge Nano: is this backtest winner real? — $0.01 per call

judge-nano
Read-only

Is this backtest winner real, or just the best of N tries? The multiple-testing selection correction (family-wise p) in one $0.01 call — verdict NO or PROVISIONAL. Use when your agent has a backtest winner, a research result, or any return series that looks good — and needs the cheapest possible check on whether it is real or the luckiest of N tries. One call runs the selection correction: the family-wise p of the series' own t across your declared n_tested. V PAID: $0.01 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3], "n_tested": 20}

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
returnsYesexample: [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3]
n_testedYesexample: 20

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnly, non-destructive, closed-world). The description adds substantive behavior the annotations cannot: a $0.01 USDC-on-Base x402 payment that must be satisfied before the call, the fact that calling returns a payment challenge, the exact computation performed, and the possible verdicts. It stops short of describing the full response payload, which matters since no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening question is a good hook, but the 'best of N tries' vs 'luckiest of N tries' framing and the 'cheapest possible check' phrasing repeat ideas, and the payment mechanics are jammed into the same breath as the method. It is serviceable but not tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with no output schema, the description should carry the return-value burden. It discloses the verdict vocabulary and the family-wise p output but not the response shape or fields, and does not clarify how an agent should act on 'NO' vs 'PROVISIONAL' or on the p-value itself. Adequate but leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (albeit with thin 'example:' text), so the baseline is 3. The description adds mild meaning by framing n_tested as 'your declared n_tested' (the number of trials involved in selection) and returns as the series under test, but gives no format, length, or scaling guidance for either parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation — the multiple-testing selection correction (family-wise p) on a return series with a declared n_tested — and states the output verdicts (NO or PROVISIONAL). It positions itself as the 'cheapest possible check' within an apparent suite (judge-lite, judge-batch, multipletest exist as siblings), but never names those alternatives, so the differentiation is implicit rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear trigger: use when the agent has a backtest winner, research result, or any return series that looks good and needs a cheap realness check. No when-not conditions or named alternative tools are supplied, so the agent must infer when to escalate to judge-lite/judge-batch.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources