Skip to main content
Glama

Money Mind — the judge

Money Mind — Judge Lite: backtest verdict in three gates — $0.05 per call

judge-lite
Read-only

Is this backtest, strategy or DeFi result real, or the luckiest of N tries? Three gates in one verdict — multiple-testing selection, deflated Sharpe, date-clustered t — built for trading agents. Use when your agent has a backtest, a research result, or any return series that looks good — and needs to know whether it is real or the luckiest of N tries before spending money on it. One call runs the three cheapest high-signal gates: selection correction (family-wise p across your declared n_te PAID: $0.05 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3], "n_tested": 20, "dates": ["2024-06-03", "2024-06-03", "2024-06-04", "2024-06-04", "2024-06-05", "20

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
datesYesexample: ["2024-06-03", "2024-06-03", "2024-06-04", "2024-06-04", "2024-06-05", "2024-06-05", "2024-06-06", "2024-06-06", "2024-0
returnsYesexample: [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3]
n_testedYesexample: 20

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuinely non-redundant behavior: this is a paid call ($0.05 USDC on Base via x402) and you must first call it to receive the payment challenge. That payment/auth flow is valuable context the annotations cannot convey. It does not disclose the response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is corrupted: it cuts off mid-word at 'n_te', splices in the payment line, then dumps a long inline example and truncates at '20'. Structure is poor and forces the reader to reassemble the intent, though the gate list is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should explain what the three-gate verdict returns; it lists the gates but never states the output form. The payment flow is covered, and the input example is present, making it minimally adequate but incomplete for a no-output-schema tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported as 100%, but the parameter 'descriptions' are only truncated examples rather than semantic definitions; the description adds no meaning beyond repeating 'n_tested' and a sample payload. With coverage nominally complete, the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the concrete analysis ('multiple-testing selection, deflated Sharpe, date-clustered t') and frames the verdict question, so an agent can tell what it computes. It is stronger than the bare name 'judge-lite' and distinguishable from siblings like deflatedsharpe or multipletest, though the mid-sentence corruption ('n_te PAID: $0.05 USDC on Base via x402') blurs the statement of purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit triggering context: 'Use when your agent has a backtest, a research result, or any return series that looks good ... before spending money on it.' That is a clear when-to-use condition. It does not name when-not-to-use or point to the single-gate siblings (deflatedsharpe, multipletest) as cheaper alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources