Money Mind — the judge
Server Details
Statistical adjudication for agents: backtest judges, fact-checks. Pay per call in USDC via x402.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 17 tools
The tools cover distinct statistical tests and checks, with clear guidance in descriptions for when to use each (e.g., judge-nano vs judge-lite vs judge-batch). However, there is some overlap between multipletest and judge-nano for single-series selection correction, and the inclusion of web-fetching tools (v1_extract, factcheck) alongside backtest validation may momentarily confuse an agent about scope.
Tool names mix conventions: some are squished compound words (clusteredt, deflatedsharpe, multipletest, abtest) while others use underscores (check_eval_selection, check_price_data) or hyphens (judge-batch, judge-lite). This inconsistency makes the set feel less predictable.
17 tools is on the heavy side for a judging service, with several tools (judge-nano, judge-lite, multipletest, deflatedsharpe) addressing overlapping statistical corrections at different price points. While each has a niche, the volume may overwhelm an agent.
The surface covers a wide range of validation gates: multiple-testing correction, clustering, drawdown, regime shifts, sample size, data quality, and comprehensive audits. Minor gaps exist (e.g., no explicit cross-validation or walk-forward tool), but the core judging needs are well met.
Available Tools
17 toolsabtestMoney Mind — A/B Significance — free allowance, then $0.31BRead-onlyInspect
Did B really beat A, corrected for how many variants you compared. Did B really beat A, corrected for how many variants you compared RUNS NOW: served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402. No account, no API key. Example request: {"n_a": 1200, "conv_a": 96, "n_b": 1180, "conv_b": 130}
| Name | Required | Description | Default |
|---|---|---|---|
| n_a | Yes | example: 1200 | |
| n_b | Yes | example: 1180 | |
| conv_a | Yes | example: 96 | |
| conv_b | Yes | example: 130 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, openWorldHint=false, destructiveHint=false). The description adds real behavioral context beyond them: a metered payment model with a daily free allowance and per-call USDC cost via x402, plus no account or API key required. It stops short of describing what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is duplicated verbatim ('Did B really beat A, corrected for how many variants you compared' appears twice), wasting roughly half the prose. The remaining pricing and example content is front-loaded and useful, but the repetition is a clear structural flaw.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a statistical test with four required numeric inputs and no output schema, the description explains the computation but never indicates what the agent gets back (p-value, lift, significance verdict). Pricing and auth are covered, but the result contract is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, though the per-field descriptions are only 'example: N'. The description echoes a full example request, reinforcing that the four required params are counts (n_a, conv_a, n_b, conv_b) for each variant. This is useful but adds little beyond the schema's own example, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific analysis: whether variant B beat variant A, corrected for the number of comparisons. The verb+resource (significance test on A/B counts) is clear. However, it does not differentiate itself from statistical siblings such as multipletest, samplesize, or clusteredt, which an agent may confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named despite several statistical siblings. The only contextual cues are availability and pricing (free allowance, then x402), which say nothing about when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auditMoney Mind — Full Book Audit — $300.00 per callCRead-onlyInspect
Every gate we have, composed in the right order, over your whole book — including the cross-instrument check that catches one edge wearing many names. Every gate we have, composed in the right order, over your whole book — including the cross-instrument check that catches one edge wearing many names PAID: $300.00 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"bars": {"EURUSD": [{"time": "2024-01-01T00:00:00Z", "open": 1.0995, "high": 1.1001, "low": 1.0986, "close": 1.0992}, {"time": "2024-01-01T01:00:00Z", "open": 1.1002, "high": 1.1012, "low": 1.0996, "close": 1.1006}, {"t
| Name | Required | Description | Default |
|---|---|---|---|
| bars | Yes | example: {"EURUSD": [{"time": "2024-01-01T00:00:00Z", "open": 1.0995, "high": 1.1001, "low": 1.0986, "close": 1.0992}, {"time": " | |
| signals | Yes | example: {"EURUSD": [1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, -1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, -1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, -1, 0, |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description usefully adds non-annotation behavior: the call is paid ($300 USDC on Base via x402) and invoking it returns a payment challenge rather than results. However it says nothing about what the audit returns, runtime cost, or side effects beyond the paywall.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is duplicated verbatim, wasting roughly a third of the description. The remainder mixes payment information with a truncated JSON example that cuts off mid-object, harming scannability and front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining what an audit returns, yet it never does. For a composite, paid operation that subsumes 17 sibling gates, the description omits the result shape, the effect of the payment challenge, and how it relates to the individual gate tools, leaving the agent materially under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters (bars, signals) carry descriptions, so the schema does the heavy lifting and baseline is 3. The description only repeats a truncated request example already present in the schema's examples, adding no new meaning about parameter format or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys scope (a full-book run of every gate in order, plus a cross-instrument dedup check) but does so in metaphorical language rather than a clear verb+resource statement. It implicitly distinguishes itself from siblings like drawdown or multipletest as the 'everything composed' variant, but never states that contrast explicitly. An agent can roughly infer 'run all checks over the whole book' but must guess at the concrete operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage instruction is 'Call it to receive the payment challenge,' which is about payment mechanics, not when to chose this tool over the many sibling gates. There is no guidance on when a full audit is preferable to calling individual checks such as deflatedsharpe or multipletest, nor any prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_eval_selectionDid I really find a better prompt? (free)ARead-onlyInspect
FREE. Use after picking the best of several prompts, models, configs or hyperparameters. Keeping the top scorer out of N is SELECTION: the winner's score is inflated simply by having looked N times, and every eval harness reports it as though one experiment were run. Give the score each variant achieved and it returns how much of the winner's margin the search itself explains, what pure noise would have handed you, and whether the excess is real. Supply n_per_variant to also learn whether your variants can be told apart at all. Call this before concluding an optimisation found something.
| Name | Required | Description | Default |
|---|---|---|---|
| scores | Yes | Score each variant achieved. 3 or more. | |
| n_per_variant | No | Trials behind each score (optional). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only and side-effect free, and the description goes well beyond that: it explains what is returned (share of margin explained by search, what pure noise would give, whether the excess is real) and the conditional behavior when n_per_variant is supplied. This is rich behavioral context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the trigger condition and the mechanism, and most sentences carry information. Minor padding in the 'FREE.' lead and the rhetorical framing keeps it from being maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description bears the burden of explaining returns — and it does, describing the three quantities produced. Combined with the usage trigger and the optional-parameter effect, nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it says n_per_variant reveals whether variants are distinguishable at all, which the schema's 'Trials behind each score' does not convey. The scores parameter's semantics are largely left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific purpose: quantifying how much of a selected winner's margin is explained by the search itself, versus noise. An agent can tell this is a selection-bias correction tool, but the description never names or contrasts with close statistical siblings such as multipletest or deflatedsharpe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering context ('Use after picking the best of several prompts, models, configs or hyperparameters') and an explicit guardrail ('Call this before concluding an optimisation found something'). No alternative tool is named for the case where the user hasn't run a selection procedure, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_price_dataIs this price data real? (free)ARead-onlyInspect
FREE, no payment or account. Audits OHLC bars for the defects that silently invalidate a backtest. Free and indicative feeds often SYNTHESISE the bar's open instead of observing it, which pins |open-close|/range near zero — measured on the same FX pairs, a free feed reads 0.020 against a real broker's 0.464. Anything depending on the open (gap trades, overnight holds, next-bar entries) is invalid on such data. Also detects frozen zero-range bars, duplicate and out-of-order timestamps, and bars whose high/low do not bracket open/close. Returns BROKER-GRADE, SUSPECT or INDICATIVE with the numbers behind it. Use this BEFORE trusting any backtest built on the data.
| Name | Required | Description | Default |
|---|---|---|---|
| bars | Yes | OHLC bars keyed by instrument. Each bar needs time (ISO-8601), open, high, low, close. 50-3000 bars, one instrument, for the free check. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnlyHint, non-destructive, closed-world); the description adds substantial context: it is free with no account, describes the detection heuristic with a concrete numeric example, and names the three output verdicts (BROKER-GRADE, SUSPECT, INDICATIVE). This is well beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with 'FREE, no payment or account' and structured around the problem it solves. Slightly long due to the numeric illustration, but each sentence carries information useful for selection and interpretation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming the three verdict categories, so an agent knows what to expect back. Combined with the detection list and usage timing, it is nearly complete, missing only finer return-shape detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
A single parameter with 100% schema description coverage; the schema already documents the bar shape, keying, and 50-3000 range. The description adds no parameter syntax or format detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audits) on a specific resource (OHLC bars) and enumerates the exact defects it targets (synthesised opens, frozen zero-range bars, duplicate/out-of-order timestamps, non-bracketing high/low). This clearly distinguishes it from the statistical siblings (deflatedsharpe, drawdown, regime), which test strategies rather than data integrity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to run it 'BEFORE trusting any backtest built on the data,' giving clear trigger context. It does not name a sibling alternative or spell out when-not to use it, so it stops short of the 5 bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clusteredtMoney Mind — Clustered t — free allowance, then $0.25ARead-onlyInspect
Your t-stat assumes independence; your trades cluster. Use when your trades are not independent — several in the same day, the same regime, or a correlated basket. A t-stat computed as if every trade were an independent draw is inflated; in this repo the factor was about 1.9x. Give per-trade returns and their dates; returns the naive t, the date-cluster RUNS NOW: served from a daily free allowance (250 left today), then $0.25 USDC on Base via x402. No account, no API key. Example request: {"returns": [0.4, -1.1, 0.8, 1.2, -0.6, 0.3, -0.2, 0.9, -1.4, 0.7, 0.5, -0.3, 1.1, -0.8, 0.2, 0.6, -0.5, 0.4, 0.9, -0.7], "dates": ["2024-01-02", "2024-01-02", "2024-01-03", "2024-01-03", "2024-01-04", "2024-01-05", "202
| Name | Required | Description | Default |
|---|---|---|---|
| dates | Yes | example: ["2024-01-02", "2024-01-02", "2024-01-03", "2024-01-03", "2024-01-04", "2024-01-05", "2024-01-05", "2024-01-08", "2024-0 | |
| returns | Yes | example: [0.4, -1.1, 0.8, 1.2, -0.6, 0.3, -0.2, 0.9, -1.4, 0.7, 0.5, -0.3, 1.1, -0.8, 0.2, 0.6, -0.5, 0.4, 0.9, -0.7] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false). The description adds meaningful context beyond them: the pricing model (250/day free allowance, then $0.25 USDC on Base via x402), the no-account/no-API-key access model, and the ~1.9x inflation factor observed. The only gap is that the intended return contents ('the naive t, the date-cluster...') are cut off.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The rationale is well front-loaded, but the body is a run-on where the pricing sentence is jammed mid-clause ('the date-cluster RUNS NOW: served from a daily free allowance...') and the example request is truncated. The core message survives, but the structure and truncation hurt readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter statistical tool with no output schema, the description should explain what comes back; instead the return-value clause is interrupted before listing the clustered statistic or any p-value/CI. Inputs, trigger conditions, and pricing are covered, but the output story is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With two required parameters at 100% schema coverage, the schema already carries the burden, and its 'descriptions' are merely truncated examples. The description adds a little meaning by phrasing the inputs as 'per-trade returns and their dates', implying they are paired observations, but adds no format or alignment detail beyond the schema examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific computation (clustered t-statistic that corrects for date clustering), names the problem it addresses, and contrasts the naive vs. clustered result. It is clearly distinguishable from statistical siblings like abtest, multipletest, and deflatedsharpe, though it never names those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit when-to-use trigger — trades that are not independent, several in the same day, same regime, or a correlated basket — which is concrete and actionable. It stops short of naming a competing tool to use when trades ARE independent, so routing is implied rather than fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
correlationMoney Mind — Spurious Correlation — free allowance, then $0.31CRead-onlyInspect
Real relationship, or two autocorrelated series drifting together. Real relationship, or two autocorrelated series drifting together RUNS NOW: served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402. No account, no API key. Example request: {"x": [0.4, -1.1, 0.8, 1.2, -0.6, 0.3, -0.2, 0.9, -1.4, 0.7, 0.5, -0.3], "y": [0.3, -0.9, 0.6, 1.0, -0.7, 0.1, -0.4, 0.8, -1.2, 0.5, 0.6, -0.1]}
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | example: [0.4, -1.1, 0.8, 1.2, -0.6, 0.3, -0.2, 0.9, -1.4, 0.7, 0.5, -0.3] | |
| y | Yes | example: [0.3, -0.9, 0.6, 1.0, -0.7, 0.1, -0.4, 0.8, -1.2, 0.5, 0.6, -0.1] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive, closed-world), so the bar is lower. The description adds genuinely new behavioral context: the two-tier payment model (daily free allowance, then $0.31 USDC on Base via x402) and that no account or API key is required. That is exactly the kind of auth/cost context annotations cannot express, though it still omits anything about computation or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is duplicated verbatim ("Real relationship, or two autocorrelated series drifting together." twice), wasting the most prominent position. The example request in the description also duplicates the input schema's example, and payment text is interleaved before any statement of what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining return values, and it says nothing about what comes back (correlation coefficient, p-value, spuriousness verdict). It also omits length/type constraints on the series, leaving the agent unable to predict or interpret results beyond billing details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is reported at 100%, so the baseline is 3. The description adds only an example invocation, and the schema 'descriptions' are themselves just repeated example arrays, so neither source explains what x and y represent, that they are equal-length numeric series, or any constraints. No meaning beyond the schema is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description frames the tool rhetorically ("Real relationship, or two autocorrelated series drifting together") rather than stating a concrete verb+resource such as "test whether the correlation between two series is spurious." The name 'correlation' plus the framing imply the purpose, but an agent gets no clean statement of what is computed or returned. It also does nothing to distinguish itself from statistical siblings like deflatedsharpe, multipletest, or clusteredt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-tool guidance at all. The only operational text is pricing/allowance mechanics (free 250/day, then $0.31), which tells the agent about billing, not about when this tool is the right choice over its many statistical siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deflatedsharpeMoney Mind — Deflated Sharpe — free allowance, then $0.31CRead-onlyInspect
Is this Sharpe real, or the best of N tries? Deflated for the search you ran. Is this Sharpe real, or the best of N tries? Deflated for the search you ran RUNS NOW: served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402. No account, no API key. Example request: {"returns": [0.012, -0.004, 0.021, 0.003, -0.011, 0.017, 0.006, -0.008, 0.014, 0.002, -0.006, 0.019, 0.001, -0.013, 0.009, 0.004], "n_trials": 40}
| Name | Required | Description | Default |
|---|---|---|---|
| returns | Yes | example: [0.012, -0.004, 0.021, 0.003, -0.011, 0.017, 0.006, -0.008, 0.014, 0.002, -0.006, 0.019, 0.001, -0.013, 0.009, 0.004] | |
| n_trials | Yes | example: 40 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=false and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful operational context beyond that: a daily free allowance (250 remaining), a $0.31 USDC payment via x402 on Base, and no account or API key required. This billing/auth disclosure is exactly the kind of value structured fields cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is duplicated word-for-word, and the run-together text 'the search you ran RUNS NOW' merges two unrelated thoughts with no punctuation or break. The pricing paragraph is front-loaded as a wall of text rather than separated from the conceptual explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining results, yet it never says whether the tool returns a deflated Sharpe value, a p-value, or a verdict. It does cover cost, auth and a sample input, which makes it usable, but an agent still cannot anticipate the response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the two parameters (returns, n_trials) already carry example values in the schema. The description repeats the same example request but adds no new meaning about units, minimum series length, or how n_trials should be counted. Baseline 3 applies when the schema does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description frames the tool as answering 'Is this Sharpe real, or the best of N tries?' adjusted for the number of trials, which conveys the concept of a deflated Sharpe ratio. However, the verb+resource statement is implicit and the key sentence is duplicated verbatim, so it never states plainly what it computes or returns. It also does not distinguish itself from siblings like multipletest or check_eval_selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use or when-not-to-use guidance and never names an alternative among the many statistics siblings. The only conditional content is commercial (free allowance exhausted, then paid), not about selecting this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drawdownMoney Mind — Drawdown Check — free allowance, then $0.31CRead-onlyInspect
Is this drawdown a broken strategy or a normal bad patch, vs your own shuffled returns. Is this drawdown a broken strategy or a normal bad patch, vs your own shuffled returns RUNS NOW: served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402. No account, no API key. Example request: {"returns": [0.9, -1.0, 0.4, -1.0, 1.6, -1.0, 0.7, 0.3, -1.0, 2.1, -1.0, -1.0, 0.8, 1.4, -1.0, 0.5, -1.0, 1.9, 0.2, -1.0, 1.1, -1.0, 0.6, 1.3]}
| Name | Required | Description | Default |
|---|---|---|---|
| returns | Yes | example: [0.9, -1.0, 0.4, -1.0, 1.6, -1.0, 0.7, 0.3, -1.0, 2.1, -1.0, -1.0, 0.8, 1.4, -1.0, 0.5, -1.0, 1.9, 0.2, -1.0, 1.1, -1.0, |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/openWorldHint/destructiveHint, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: no account or API key required, a 250-call daily free allowance, and a $0.31 USDC-on-Base x402 charge thereafter. It does not disclose computation traits (e.g., required sample size, rerandomization count), which keeps it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is duplicated verbatim and then mangled into '…vs your own shuffled returns RUNS NOW:', which reads as a copy-paste artifact. Pricing and an inline JSON example are appended without structure, so the description is noisy rather than front-loaded; only the fact that it is short keeps this above 1.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with a high-coverage schema, the description covers access/payment and gives an example payload. What it omits is the output: with no output schema, the agent has no idea what comes back (a p-value, a verdict label, a distribution?) or what input constraints apply, so it is only marginally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single 'returns' parameter, so the schema already carries the burden. The description supplies a worked example request, which is mildly helpful, but it never explains what 'returns' means (per-period decimal returns? percentages?) or any constraints such as minimum series length. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description frames the tool as a question ('Is this drawdown a broken strategy or a normal bad patch') rather than stating a specific verb+resource like 'test whether an observed drawdown is statistically significant against shuffled return series'. The 'vs your own shuffled returns' clause hints at the method, but the framing is sales copy-style and offers no differentiation from siblings such as deflatedsharpe, multipletest, or samplesize, which perform adjacent statistical checks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to reach for this tool versus the many statistical siblings. The only usage-adjacent content is billing mechanics ('free allowance, then $0.31'), which tells the agent nothing about when the tool is the right choice or what inputs make it meaningful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
factcheckMoney Mind — Fact Check: verify a claim against sources — $0.03 per callARead-onlyInspect
Verify a claim against 2-4 source URLs: cited passages, cross-source agreement, verdict: consistent | conflicting | insufficient_evidence. Evidence, not truth. Use when your agent needs to verify a claim against independent sources before acting on it. Give a claim plus 2-4 source URLs; each source is extracted (three-fetch cross-check, confidence-scored), the checkable figures are compared across sources, and you get the cited passages plus a verdict: con PAID: $0.03 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"claim": "Acme Corp revenue was $4.2 billion in 2025", "urls": ["https://example.com", "https://example.org"]}
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | example: ["https://example.com", "https://example.org"] | |
| claim | Yes | example: "Acme Corp revenue was $4.2 billion in 2025" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld/non-destructive, so the safety profile is covered. The description adds substantive behavior beyond that: the three-fetch cross-check with confidence scoring, cross-source figure comparison, and crucially the paid model ($0.03 USDC on Base via x402) with the payment-challenge handshake, which is essential context for invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core (claim + 2-4 URLs in, passages + verdict out) is front-loaded, but the text is a run-on with a garbled/truncated segment ('...verdict: con PAID: $0.03 USDC...') where sentences collide, hurting readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param tool with no output schema, the description covers inputs, processing, return values, verdict vocabulary, and the payment model. The only shortfall is the missing explanation of the output field shape, which is partly implied by the listed verdict values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the schema 'descriptions' are only bare examples, whereas the description adds a real constraint (2-4 URLs, not arbitrary) and a full example request payload, giving meaning the schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (verify) and resource (a claim against 2-4 source URLs) and enumerates the outputs: cited passages, cross-source agreement, and a verdict enum. It is clearly distinguishable from statistical siblings like correlation or abtest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states the context: 'Use when your agent needs to verify a claim against independent sources before acting on it.' It also tells the caller the payment flow ('Call it to receive the payment challenge'). It does not name any sibling alternative, but none of the siblings share this function.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
judge-batchMoney Mind — Judge Batch: backtest triage — $0.01 per callARead-onlyInspect
Triage up to 1,000 backtests in one call: the best-of-N multiple-testing selection correction over every return series — $0.01 per series, survivors ranked best-first. Use when your agent has a FARM of backtests, strategies, or research results — dozens to a thousand — and needs triage, not one-by-one verdicts. One call runs the selection correction over every series: each gets NO (killed), PROVISIONAL (survived this gate), or ERROR (malformed). Results ranked sur PAID: $0.01 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"series": [{"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3], "n_tested": 20, "label": "weak"}, {"returns": [2.0, 2.1, 1.9, 2.2, 2.0, 2.1, 1.8, 2.0,
| Name | Required | Description | Default |
|---|---|---|---|
| series | Yes | example: [{"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3] |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds important behavior beyond annotations: paid call at $0.01 per series, USDC on Base via x402, payment challenge, and per-series outcomes of NO, PROVISIONAL, or ERROR ranked best-first.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and usage, but it contains a malformed/truncated sentence ('Results ranked sur PAID:'), repeats payment information, and ends with an incomplete example request. The structure is messy and several sentences do not earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a paid, batch-oriented tool with no output schema, the description does disclose return categories and ranking, plus payment mechanics. However, it is truncated and does not fully cover parameter fields or output format details, leaving gaps an agent must infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With one parameter and high schema coverage, the baseline is 3. The description adds some semantics, such as up to 1,000 series and per-series payment, and the example shows returns, n_tested, and label, but it does not explain fields like n_tested or the expected returns format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: triage up to 1,000 backtests using best-of-N multiple-testing selection correction over every return series. It distinguishes the batch triage use case from one-by-one verdicts, though it does not explicitly name sibling alternatives such as judge-lite or judge-nano.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use the tool: when an agent has a farm of backtests, strategies, or research results numbering dozens to a thousand and needs triage rather than individual verdicts. It does not state when-not-to-use or name a specific alternative sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
judge-liteMoney Mind — Judge Lite: backtest verdict in three gates — $0.05 per callARead-onlyInspect
Is this backtest, strategy or DeFi result real, or the luckiest of N tries? Three gates in one verdict — multiple-testing selection, deflated Sharpe, date-clustered t — built for trading agents. Use when your agent has a backtest, a research result, or any return series that looks good — and needs to know whether it is real or the luckiest of N tries before spending money on it. One call runs the three cheapest high-signal gates: selection correction (family-wise p across your declared n_te PAID: $0.05 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3], "n_tested": 20, "dates": ["2024-06-03", "2024-06-03", "2024-06-04", "2024-06-04", "2024-06-05", "20
| Name | Required | Description | Default |
|---|---|---|---|
| dates | Yes | example: ["2024-06-03", "2024-06-03", "2024-06-04", "2024-06-04", "2024-06-05", "2024-06-05", "2024-06-06", "2024-06-06", "2024-0 | |
| returns | Yes | example: [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3] | |
| n_tested | Yes | example: 20 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds genuinely non-redundant behavior: this is a paid call ($0.05 USDC on Base via x402) and you must first call it to receive the payment challenge. That payment/auth flow is valuable context the annotations cannot convey. It does not disclose the response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is corrupted: it cuts off mid-word at 'n_te', splices in the payment line, then dumps a long inline example and truncates at '20'. Structure is poor and forces the reader to reassemble the intent, though the gate list is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should explain what the three-gate verdict returns; it lists the gates but never states the output form. The payment flow is covered, and the input example is present, making it minimally adequate but incomplete for a no-output-schema tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 100%, but the parameter 'descriptions' are only truncated examples rather than semantic definitions; the description adds no meaning beyond repeating 'n_tested' and a sample payload. With coverage nominally complete, the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the concrete analysis ('multiple-testing selection, deflated Sharpe, date-clustered t') and frames the verdict question, so an agent can tell what it computes. It is stronger than the bare name 'judge-lite' and distinguishable from siblings like deflatedsharpe or multipletest, though the mid-sentence corruption ('n_te PAID: $0.05 USDC on Base via x402') blurs the statement of purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit triggering context: 'Use when your agent has a backtest, a research result, or any return series that looks good ... before spending money on it.' That is a clear when-to-use condition. It does not name when-not-to-use or point to the single-gate siblings (deflatedsharpe, multipletest) as cheaper alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
judge-nanoMoney Mind — Judge Nano: is this backtest winner real? — $0.01 per callARead-onlyInspect
Is this backtest winner real, or just the best of N tries? The multiple-testing selection correction (family-wise p) in one $0.01 call — verdict NO or PROVISIONAL. Use when your agent has a backtest winner, a research result, or any return series that looks good — and needs the cheapest possible check on whether it is real or the luckiest of N tries. One call runs the selection correction: the family-wise p of the series' own t across your declared n_tested. V PAID: $0.01 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3], "n_tested": 20}
| Name | Required | Description | Default |
|---|---|---|---|
| returns | Yes | example: [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3] | |
| n_tested | Yes | example: 20 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly, non-destructive, closed-world). The description adds substantive behavior the annotations cannot: a $0.01 USDC-on-Base x402 payment that must be satisfied before the call, the fact that calling returns a payment challenge, the exact computation performed, and the possible verdicts. It stops short of describing the full response payload, which matters since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening question is a good hook, but the 'best of N tries' vs 'luckiest of N tries' framing and the 'cheapest possible check' phrasing repeat ideas, and the payment mechanics are jammed into the same breath as the method. It is serviceable but not tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, the description should carry the return-value burden. It discloses the verdict vocabulary and the family-wise p output but not the response shape or fields, and does not clarify how an agent should act on 'NO' vs 'PROVISIONAL' or on the p-value itself. Adequate but leaves real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (albeit with thin 'example:' text), so the baseline is 3. The description adds mild meaning by framing n_tested as 'your declared n_tested' (the number of trials involved in selection) and returns as the series under test, but gives no format, length, or scaling guidance for either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation — the multiple-testing selection correction (family-wise p) on a return series with a declared n_tested — and states the output verdicts (NO or PROVISIONAL). It positions itself as the 'cheapest possible check' within an apparent suite (judge-lite, judge-batch, multipletest exist as siblings), but never names those alternatives, so the differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear trigger: use when the agent has a backtest winner, research result, or any return series that looks good and needs a cheap realness check. No when-not conditions or named alternative tools are supplied, so the agent must infer when to escalate to judge-lite/judge-batch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
multipletestMoney Mind — Multiple Testing — free allowance, then $0.20ARead-onlyInspect
You tested N strategies and kept the best; is it noise. Use when you searched many strategies or parameter sets and kept the best one. Quoting that winner's solo p-value is the most common way a backtest lies. Give the number of things you tested and the winner's t-statistic; returns the family-wise p-value, the t you actually needed, and the t that pure RUNS NOW: served from a daily free allowance (250 left today), then $0.20 USDC on Base via x402. No account, no API key. Example request: {"n_tested": 1000, "best_t": 3.0}
| Name | Required | Description | Default |
|---|---|---|---|
| best_t | Yes | example: 3.0 | |
| n_tested | Yes | example: 1000 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive=false/openWorld=false, so the safety profile is covered. The description adds real context beyond that: output contents, the pricing model (free daily allowance then $0.20 USDC via x402), and that no account or API key is required. The one gap is a truncated sentence ('the t that pure'), which leaves a return value unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core framing is front-loaded and useful, but the description is redundant ('You tested N strategies and kept the best' vs 'Use when you searched many strategies...') and the sentence 'the t that pure' is cut off mid-thought, and pricing/logistics clutter the entry.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description does explain the key returns (family-wise p-value, needed t). It is nearly complete, losing a point only for the truncated return-value sentence and sparse handling of the example payload's meaning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions only contain bare examples ('example: 3.0'), so they carry little semantics despite nominal 100% coverage. The description supplies the actual meaning: 'the number of things you tested' (n_tested) and 'the winner's t-statistic' (best_t), adding value the schema does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific statistical operation: correcting a winning backtest for multiple testing, returning the family-wise p-value and the t-stats needed. It's clearly a multiple-testing tool, distinguishable from siblings like samplesize and deflatedsharpe. It doesn't explicitly name a sibling it's not, so not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger: 'Use when you searched many strategies or parameter sets and kept the best one,' plus the motivating failure mode (quoting a winner's solo p-value). Clear adoption context, but no explicit when-not or named-alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refereeMoney Mind — Referee — $5.00 per callARead-onlyInspect
11-gate adjudication of a trading signal series. Use when you have OHLC bars and a -1/0/1 signal series and need to know whether a backtest result is real or an artefact. Applies 11 gates most backtests fail: matched random-entry drift control, date-clustered standard errors, entry at the next bar's open, events-not-fills, absolute (not just exces PAID: $5.00 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"bars": {"EURUSD": [{"time": "2024-01-01T00:00:00Z", "open": 1.0995, "high": 1.1001, "low": 1.0986, "close": 1.0992}, {"time": "2024-01-01T01:00:00Z", "open": 1.1002, "high": 1.1012, "low": 1.0996, "close": 1.1006}, {"t
| Name | Required | Description | Default |
|---|---|---|---|
| rr | Yes | example: 2.0 | |
| bars | Yes | example: {"EURUSD": [{"time": "2024-01-01T00:00:00Z", "open": 1.0995, "high": 1.1001, "low": 1.0986, "close": 1.0992}, {"time": " | |
| horizon | Yes | example: 5 | |
| signals | Yes | example: {"EURUSD": [1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, -1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, -1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, -1, 0, | |
| cost_bps | Yes | example: 10 | |
| stop_atr | Yes | example: 1.5 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnlyHint=true, destructiveHint=false, openWorldHint=false), and the description adds genuinely new operational context: a $5.00 USDC-on-Base x402 payment and a 'call it to receive the payment challenge' first step. The gate enumeration is useful but is cut off mid-list ('absolute (not just exces'), so the full behavioral picture is incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and gates are front-loaded well, but the text is visibly truncated mid-word and a trailing 'Example request: {...}' block is both cut off and redundant with the schema examples, wasting space. The splice point where payment info interrupts the gate list also harms readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex statistical validator with no output schema, the description partially covers the method (listing ~5 of 11 gates before truncation) and the payment flow, but leaves the return/verdict format and the remaining gates unexplained. Not adequate as a full specification, though not empty either.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported at 100%, but each 'description' is merely an example value ('example: 2.0', 'example: 10'), so the schema itself adds little semantic meaning. The description's example request duplicates those same values and does not explain units, ordering, or the relationship between signals and bars. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: '11-gate adjudication of a trading signal series.' An agent knows this validates a signal series rather than running a backtest itself. However, it never names or contrasts a sibling (e.g., deflatedsharpe, clusteredt, multipletest, abtest) that could also be used for backtest validation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger condition: use when you have OHLC bars plus a -1/0/1 signal series and need to know whether a backtest result is real or an artefact. No explicit when-not or named alternative is provided, so routing among the many sibling validation tools is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
regimeMoney Mind — Regime Break — free allowance, then $0.31CRead-onlyInspect
Did the process change, or is recent weakness inside normal variation. Did the process change, or is recent weakness inside normal variation RUNS NOW: served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402. No account, no API key. Example request: {"returns": [0.9, -1.0, 0.4, -1.0, 1.6, -1.0, 0.7, 0.3, -1.0, 2.1, -1.0, -1.0, 0.8, 1.4, -1.0, 0.5, -1.0, 1.9, 0.2, -1.0, 1.1, -1.0, 0.6, 1.3, -0.8, 1.5]}
| Name | Required | Description | Default |
|---|---|---|---|
| returns | Yes | example: [0.9, -1.0, 0.4, -1.0, 1.6, -1.0, 0.7, 0.3, -1.0, 2.1, -1.0, -1.0, 0.8, 1.4, -1.0, 0.5, -1.0, 1.9, 0.2, -1.0, 1.1, -1.0, |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful non-schema behavior: it is served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402, with no account or API key required. It says nothing about latency, output shape, or what happens when the allowance is exhausted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening question is duplicated verbatim in back-to-back sentences, which is pure waste and makes the description look machine-generated. The example request also duplicates the schema's example verbatim. Pricing is front-loaded before any functional statement of what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a statistical test with no output schema, the description should at least indicate what is returned (a break statistic, p-value, regime labels, etc.). It instead spends its length on pricing and a duplicated question, leaving an agent unable to predict the response or judge whether the result answers its question.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single required parameter (returns), so the schema carries most of the burden; its property description is truncated mid-example, so the description's full example request is a modest useful addition. It confirms the expected input form (a numeric return series) but adds no semantics such as minimum length, units, or whether the series must be chronological.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a rhetorical question ('Did the process change, or is recent weakness inside normal variation') rather than a statement of what the tool computes or returns. The name 'regime' plus the title 'Regime Break' gives a hint, but the text never says what a call produces or what a 'regime break' result looks like, and it does nothing to distinguish this from siblings like clusteredt, drawdown, or correlation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no named alternative among the many analysis siblings. The only implicit signal is the example request, which shows a returns series as input, implying it applies to a return-stream rather than, say, a price series or an A/B setup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
samplesizeMoney Mind — Sample Size — free allowance, then $0.31CRead-onlyInspect
How many observations before this test can answer anything — ask before you spend. How many observations before this test can answer anything — ask before you spend RUNS NOW: served from a daily free allowance (250 left today), then $0.31 USDC on Base via x402. No account, no API key. Example request: {"baseline_rate": 0.08, "min_detectable_effect": 0.2, "power": 0.8, "alpha": 0.05}
| Name | Required | Description | Default |
|---|---|---|---|
| alpha | Yes | example: 0.05 | |
| power | Yes | example: 0.8 | |
| baseline_rate | Yes | example: 0.08 | |
| min_detectable_effect | Yes | example: 0.2 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, correctly signalling a safe, self-contained computation. The description usefully adds the payment/allowance model and the no-account/no-API-key fact, which is real context, but says nothing about what the computation returns or its assumptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is duplicated verbatim, wasting the front-loaded position on a repeat. Pricing and example content is jammed together without clear separation, making the definition noisy rather than tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining the return value — but it never states whether it returns per-arm or total N, or what power/baseline assumptions produce. For a statistical tool with four required numeric parameters, this leaves a significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description only echoes the same example values the schema already lists and never explains what baseline_rate (baseline conversion rate?) or min_detectable_effect (relative vs absolute?) actually mean, so it adds no real semantic value over the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The collocated name and title ('Sample Size') plus the illustrative parameter set make it inferable that this computes a required sample size for a test. However, the description's actual prose ('How many observations before this test can answer anything') never names the method (power analysis / two-proportion test) and gives no differentiation from statistical siblings like abtest, clusteredt, or multipletest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'ask before you spend' hints at pre-experiment use, and the pricing line explains cost. But there is no statement of when to use this versus the many sibling statistical tools, no prerequisites, and no exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
v1_extractMoney Mind — Web Extraction, cross-checked 3 ways — free allowance, then $0.02ARead-onlyInspect
Fetch a web page three independent ways (rendered, static, second-render cross-check), confidence-scored, with a tamper-evident attestation. Refuses protected, login and CAPTCHA pages without charging. Use when your agent needs web content it cannot fetch itself — JavaScript-rendered pages, static pages, or paginated listings — and needs to show its principal the evidence. Every extraction runs three independent fetches (rendered, static, second-render cross-check), is confidence-scored, and is wr RUNS NOW: served from a daily free allowance (25 left today), then $0.02 USDC on Base via x402. No account, no API key. Example request: {"url": "https://example.com", "extraction_target": "the page title and the body text"}
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | example: "https://example.com" | |
| extraction_target | Yes | example: "the page title and the body text" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint, openWorldHint, and destructiveHint already in the annotations, the description adds valuable context: three independent fetches, confidence scoring, tamper-evident attestation, refusal of protected/login/CAPTCHA pages without charging, and cost/auth details (daily free allowance then $0.02 USDC, no API key). The malformed fragment 'and is wr RUNS NOW:' slightly hurts clarity, but the added behavioral disclosure is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description repeats that extraction runs three independent fetches after already stating it in the opening sentence. More seriously, it contains the broken fragment 'and is wr RUNS NOW:' that disrupts the flow and makes the text feel patched or malformed rather than front-loaded and clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, usage context, refusal behavior, pricing, authentication model, and the key output concepts (confidence-scored, tamper-evident attestation). However, with no output schema, it could say more about the actual response format, so it is complete enough to invoke but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the schema property descriptions are only examples rather than semantic definitions. The description repeats the example request and does not add meaning beyond what the schema already provides, so the baseline of 3 is appropriate for high schema coverage with no supplementary parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names a specific verb and resource: fetching a web page via three independent extraction methods, with confidence scoring and attestation. It is easy to distinguish from the analytical sibling tools, but it does not explicitly differentiate itself from any sibling or name an alternative, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use the tool — when an agent needs web content it cannot fetch itself, including JavaScript-rendered pages, static pages, and paginated listings — and when it will refuse (protected, login, CAPTCHA pages). No alternative tools are named, so it misses the top score's explicit comparison against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
17 tool updates
- First observed
abtest - First observed
audit - First observed
check_eval_selection - First observed
check_price_data - First observed
clusteredt - First observed
correlation - First observed
deflatedsharpe - First observed
drawdown - First observed
factcheck - First observed
judge-batch - First observed
judge-lite - First observed
judge-nano - First observed
multipletest - First observed
referee - First observed
regime - First observed
samplesize - First observed
v1_extract
Related MCP Connectors
Adversarial verification for AI agents - pay an independent skeptic per verdict in USDC via x402.
Pay-per-call DeFi and macro intel for AI agents. x402 USDC tools via streamable HTTP /api/mcp.
Pay-per-call ($0.005–$0.03 USDC) market, on-chain, and prediction-market data API for AI agents and trading bots via the x402 protocol — no signup, no API key. Exposed as a remote MCP server with 15 tools: pre-trade token security (honeypot/liquidity checks), kimchi premium, funding rate APR, DEX slippage, Polymarket arbitrage & liquidity audits, and Hyperliquid HIP-4 prediction-market odds.
Market data and web intelligence for AI agents, paid per call in USDC on Base via x402.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables agents to submit claims, plans, or drafts for evaluation by a panel of frontier models from multiple vendors, receiving a structured consensus verdict with dissent, paid per call via x402 using USDC on Base without needing vendor accounts.1-
- FlicenseAqualityDmaintenancePay-per-call tools for AI agents including trust checks, due diligence, market data, and human-verified approvals, settled in USDC on Base via the x402 protocol.16-
- AlicenseAqualityCmaintenanceEnables MCP-capable AI agents to make pay-per-call Solana data requests (snapshots, risk checks, token reports, health, balances, transactions, trending pairs) with automatic USDC settlement over x402.5MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to obtain an independent review of their drafts, returning a pass/fail verdict with specific issues and suggested fixes, plus a signed receipt. Payments are made per check over x402, with no account or API key needed.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.