Regime
Server Details
Checks crypto trading claims on real data: leverage, drawdown, stop-loss, seasonality, luck.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 12 tools
Each tool targets a distinct type of analysis (DCA, missing days, seasonality, coin comparison, copy trading, diversification, drawdown, leaderboard rank, leverage, overfitting, stop-loss, strategy lookup). The descriptions explicitly cross-reference related tools and clarify when to use which, eliminating overlap confusion.
All tool names use snake_case and are descriptive noun phrases. While eight end with '_check' and four do not, the naming convention is still consistent and predictable; no mixed styles or vague verbs.
Twelve tools is well within the ideal range for a focused crypto claim-checking server. Each tool earns its place by covering a specific common claim or analysis, and none appear redundant.
The surface covers a wide range of common trading claims (DCA, best days, seasonality, drawdown, leverage, stop-loss, diversification, copy trading, overfitting, strategy backtests). Minor gaps exist such as a dedicated take-profit check or explicit market-regime detection, but core workflows are well supported.
Available Tools
12 toolsaveraging_in_checkAveraging in checkARead-onlyIdempotentInspect
Whether spreading the entry would have helped, measured. One amount into a coin over a window (90, 365 or 730 days): all of it on day one, or equal instalments weekly or monthly. Both start the same day, move the same money and are valued on the same day, started on every day of the history. You get how often spreading ended ahead, the middle result of each, and the worst start day of each — which is what spreading actually buys and the half nobody shows. Spot tape, no fees, dead coins included.
Use for dollar-cost averaging versus lump sum on one coin. For claims about particular days or months use best_days_check or calendar_check; for the fall after a single buy, drawdown_check.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | Ticker or full pair. An unknown one comes back with the list we hold and the ones left out. | |
| days | Yes | The window, in days: 90, 365 or 730. | |
| instalments | Yes | How often a slice goes in: weekly or monthly. Defaults to weekly. | weekly |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| model | No | The assumptions, stated: cash not yet in earns nothing (against spreading), no fees (in favour of spreading). |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| starts | No | Start days tested. Fewer than 120 and we refuse to give a percentage. |
| instalments | No | How many slices fit in the window. |
| spreading_won | No | Start days on which spreading it out ended ahead. A count, not a rate: «100%» is only said when it was all of them. |
| p05_spread_pct | No | The same for spreading. |
| p05_lump_sum_pct | No | Fifth percentile of starts, lump sum: one worst day is an anecdote. |
| worst_spread_pct | No | The same for spreading. This pair is the whole argument for averaging in, and it is the one that never appears next to the advice. |
| median_spread_pct | No | Middle start, spread out. |
| spreading_won_pct | No | The same as a share of starts. |
| median_edge_points | No | Percentage points spreading made over lump sum in the middle case. Negative means lump sum won. |
| worst_lump_sum_pct | No | The single worst start day, all in on day one. |
| median_lump_sum_pct | No | Middle start, all in on day one. |
| share_in_profit_spread_pct | No | The same for spreading. |
| share_in_profit_lump_sum_pct | No | Share of starts that ended up, lump sum. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already confirm this is a read-only, idempotent, non-destructive check. The description adds substantial behavioral context: spot tape only, no fees, dead coins included, both strategies start the same day and use the same money, and the result includes how often spreading won, median outcomes, and worst start days.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core comparison and then moves into methodology and routing guidance. It is longer than strictly necessary and includes colloquial phrases like 'the half nobody shows' and 'Spot tape,' but most sentences carry operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With full schema coverage, annotations, and an output schema, the remaining burden is usage and behavioral context. The description supplies both, including methodology assumptions and what the result conveys, so an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the coin, days enum, and instalments enum/default. The description restates the same information in prose but does not add constraints or syntax beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific comparison: whether spreading an entry would have helped, using one amount into one coin over a fixed window, either lump sum on day one or equal weekly/monthly instalments. It also distinguishes itself from siblings like best_days_check, calendar_check, and drawdown_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use this for dollar-cost averaging versus lump sum on one coin, then names the alternatives for adjacent questions: best_days_check or calendar_check for particular days/months, and drawdown_check for the fall after a single buy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
best_days_checkBest days checkARead-onlyIdempotentInspect
The 'miss the ten best days and you get nothing' claim, measured on both sides. The same window of a coin lived four ways: all of it, without its N best days, without its N worst days, and without either, from every start day of the history. You also get how many of those best days landed within a week of a worst one, which is what decides whether the claim means anything. Backtest on spot daily candles, dead coins included.
Use when someone argues for staying invested because missing the best days ruins returns, or for timing the market to dodge the worst ones. Which weekday or month does best is calendar_check.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | Ticker or full pair. An unknown one comes back with the list we hold and the ones left out. | |
| days | Yes | Window in days: 90, 365 or 730. | |
| best_days | Yes | How many days are missed: 1, 5, 10 or 20. Defaults to 10, the number in the claim. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| page | No | The page with these exact numbers already in. |
| error | No | no_data when ok is false. |
| model | No | What missing a day means here, and what is not charged. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| history | No | From, to, days held and which tape. |
| windows | No | Start days tested. Fewer than 120 and we refuse to give a percentage. |
| median_pct | No | The middle window, lived whole. The alternative, always. |
| biggest_days | No | The biggest up days and down days of the whole history, with their dates, so the clustering can be checked rather than believed. |
| share_in_profit_pct | No | Windows that ended in profit, lived whole. |
| cost_of_missing_best_pp | No | Percentage points the best days were worth. |
| median_without_best_pct | No | The same window without its best days. This is the half of the claim you get shown. |
| gift_of_missing_worst_pp | No | Percentage points dodging the worst days would have paid. |
| median_without_worst_pct | No | The same window without its WORST days. This is the half nobody shows, and it is just as big. |
| median_without_either_pct | No | Without both groups: usually back near the first number, after dodging the days that supposedly decided everything. |
| next_to_means_within_days | No | What «next to» means, in days. |
| winners_turned_losers_pct | No | Of the windows that ended in profit, the share that end in loss once their best days are removed. |
| share_in_profit_without_best_pct | No | The same, without the best days. |
| best_days_next_to_a_worst_day_pct | No | Share of the best days that fell within a few days of one of the worst. High means you cannot dodge one without dodging the other. |
| share_in_profit_without_worst_pct | No | The same, without the worst days. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered; the description adds real behavioral context beyond that — data domain (spot daily candles), survivorship handling (dead coins included), the full-history rolling start-day sweep, and a derived output (count of best days within a week of a worst one). It stops short of failure behavior or cost/latency characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two paragraphs, both earning their place, front-loaded with the core computation before usage guidance. Slightly dense, but no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out; the description nonetheless frames what the four scenarios yield and when the tool is applicable. For a read-only analytical tool with fully covered params, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and enums are already documented, so the baseline is 3. The description explains what the analysis does conceptually with best_days and the window, but adds no format, range, or syntax detail beyond what the schema already carries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States precisely what it computes: the 'miss the ten best days' claim, measured four ways (all, minus best N, minus worst N, minus both) across every start day. It explicitly names the sibling calendar_check and what that sibling covers instead, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit trigger conditions ('when someone argues for staying invested because missing the best days ruins returns, or for timing the market to dodge the worst ones') and names the alternative tool (calendar_check) for the adjacent question.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calendar_checkCalendar checkARead-onlyIdempotentInspect
'Uptober'. 'Mondays dip'. 'Sell in May'. Any calendar claim on 34 coins, measured against chance. With seven days one has to come first, so the answer is a permutation test: the SAME returns dealt out at random hundreds of times, and how often chance alone produces a bucket that good. Plus how many times the bucket really happened - October is 279 days of bitcoin but nine Octobers - and the round-trip cost. Nine years of daily candles.
Tests whether one day of the week, month of the year or hour of the day really beats the rest for a coin. Use for seasonality claims (Uptober, Monday dips, Sell in May). For the claim about missing the market's best days use best_days_check; for when to spread an entry, averaging_in_check.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | Ticker or full pair. An unknown one comes back with the list we hold and the ones left out. | |
| calendar | Yes | day_of_week, month_of_year or hour_of_day. The hourly one exists for the 16 coins we hold hourly candles for. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| page | No | The page with these exact numbers already in. |
| error | No | no_data when ok is false. |
| model | No | The test, stated, including why a monthly p-value flatters itself. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| buckets | No | Every bucket: its name, observations, times occurred, mean, median and share of up days. |
| history | No | From, to, days held and which tape. |
| shuffles | No | How many shuffled worlds were tested. |
| best_bucket | No | The best bucket by average move, named. |
| observations | No | Moves measured across all buckets. |
| worst_bucket | No | The other end. |
| best_mean_pct | No | Its average move. On its own, this is the advertisement. |
| worst_mean_pct | No | Its average move. |
| edge_covers_cost | No | Whether the average move of the best bucket is bigger than that cost. Usually it is not. |
| chance_best_p95_pct | No | What chance reaches in one shuffle out of twenty. |
| round_trip_cost_pct | No | What entering and exiting once costs, so the edge can be compared with it. |
| beats_chance_at_5pct | No | Whether that share is under 0.05. Across all 34 coins, about 5% of these come back true by chance alone, so one true answer on its own is not a finding. |
| chance_best_mean_pct | No | What the BEST bucket of a shuffled world averages. Anything below this is less impressive than nothing. |
| chance_matches_it_share | No | The p-value: share of shuffles whose best bucket matched or beat the real one. High means no pattern. |
| best_bucket_happened_times | No | How many times that bucket has actually occurred. For months this is years, not days: nine Octobers is nine observations however many candles they hold, and it is the number these claims never show. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so safety is covered. The description adds genuine methodological context: a permutation test with hundreds of randomized shuffles, the reported round-trip cost, and sample-size caveats such as 'nine Octobers'. This is real disclosure beyond the structured fields, though it does not describe output shape (which the output schema handles).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The operative routing sentence and alternatives are front-loaded and tight. The opening flavor paragraph ('Uptober', 'Mondays dip', sample-size asides) is stylized and longer than strictly necessary, but it does establish the seasonality framing before the definitions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytical tool with annotations covering safety and an output schema covering returns, the description supplies the method, data scope (nine years of daily candles, 34 coins) and sibling routing. Complete enough to invoke correctly; only the verbose framing keeps it from being maximally efficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are already documented, including the enum values and the hourly limitation. The prose restates the three calendar buckets but adds no syntax or format detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific analytical action ('Tests whether one day of the week, month of the year or hour of the day really beats the rest for a coin') with a clear resource, and names the sibling tools it is not. An agent can distinguish it from best_days_check and averaging_in_check without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it for seasonality claims and gives concrete examples (Uptober, Monday dips, Sell in May). It then routes the agent away with named alternatives: best_days_check for the best-days claim, averaging_in_check for entry spreading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
coin_comparisonCoin comparisonARead-onlyIdempotentInspect
Two or three coins side by side, every check at once: for each, how often a buy fell 30% or more before the window was out and where the middle one ended, how often a leveraged month ended with nothing, whether spreading the entry won, how often a stop-loss sold a buy that ended in profit, and how much it moves with bitcoin. Nothing is ranked and no coin is called better. Spot and perpetual tapes, fees charged.
Use when two or three coins are being compared or chosen between. For one coin or one question the specific check gives more detail: drawdown_check, leverage_survival, averaging_in_check, stop_loss_check, diversification_check.
| Name | Required | Description | Default |
|---|---|---|---|
| days | Yes | 90, 365 (default) or 730. | |
| coins | Yes | Two or three tickers, comma-separated: BTC,ETH,SOL. Kept in the order sent. | |
| leverage | Yes | For the leverage row. Default 25. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| rows | No | Per coin: the drawdown, leverage, averaging-in and stop-loss answers, and its correlation with bitcoin (null for bitcoin itself). |
| error | No | no_data when ok is false. |
| model | No | The method, stated. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public files these numbers come from. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), yet the description adds real behavioral context beyond them: results are unranked and no coin is declared better, both spot and perpetual tapes are used, and fees are charged. Those are exactly the traits an agent needs to interpret the output and set user expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core scope ('Two or three coins side by side, every check at once') before the check enumeration, and the usage paragraph follows cleanly. The check list is long but each item is load-bearing information about output content, so little is wasted; slightly dense for a single sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value formatting is handled structurally, and the description still covers scope, the checks performed, the non-ranking behavior, fee and tape assumptions, and routing to single-coin alternatives. Nothing an agent needs in order to select and call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% — ticker format, comma separation, order preservation, the 90/365/730 enum with default, and the leverage default are all documented in the schema itself. The description only alludes to 'the window' and 'the leverage row', adding no syntax or format detail. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (two or three coins side by side, every check at once) and then enumerates exactly which checks are run: drawdown, leveraged month survival, entry spreading, stop-loss, bitcoin correlation. It also states what it is not — nothing is ranked and no coin is called better — which an agent cannot infer from the name or siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes use to 'two or three coins being compared or chosen between', and names the alternatives for the other case: 'For one coin or one question the specific check gives more detail: drawdown_check, leverage_survival, averaging_in_check, stop_loss_check, diversification_check.' Both the when and the when-not, with named siblings, are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
copy_trading_checkCopy trading checkARead-onlyIdempotentInspect
Whether copying the top traders works, measured. Accounts in the top 10% or 1% of Hyperliquid one month, by return or by dollar profit: how many were in the top again the next month, next to chance, how many fell to the bottom instead, and how many made money the month you would have copied them for, next to every active account. 4,000 accounts sampled at random, not today's leaders; about 30 month pairs.
Use when someone proposes copying top traders, a leaderboard or signal leaders. Takes no account address. To place one specific return on the leaderboard use leaderboard_rank_check; to ask whether luck explains a record, overfitting_odds.
| Name | Required | Description | Default |
|---|---|---|---|
| top_pct | Yes | 10 (default) or 1. 0.1 is read as 10. | |
| ranked_by | Yes | return (default) or profit. | return |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| model | No | The method, stated. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| repeated | No | Of those, times it was in the top again next month. |
| chance_pct | No | What picking accounts at random gives. The alternative, always. |
| repeated_pct | No | The same as a share. |
| times_in_top | No | Times an account finished a month in the top group. |
| stopped_trading | No | Counted as not repeating. |
| fell_to_bottom_pct | No | Share that fell to the bottom group instead: the size control. |
| in_profit_next_month_pct | No | Share that made money the next month. |
| in_profit_next_month_all_accounts_pct | No | The same for every active account. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds genuinely useful scope context the annotations cannot: no account address is required, the sample is 4,000 random accounts rather than today's leaders, and the horizon is ~30 month pairs. It stops short of describing return format or pagination, but the output schema likely covers results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Structure is correct — what it measures first, then the use-when and alternatives in a clearly separated paragraph. The opening sentence and the run-on enumeration of metrics are dense and slightly redundant given an output schema exists, but each clause carries real information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a statistical-analysis tool with a full annotation set and an output schema, the description supplies everything the agent needs: the exact question answered, the data scope, the no-address constraint, the trigger condition, and the two sibling alternatives. Return values are correctly left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both params are enums with descriptions, so the baseline is 3. The description modestly enriches meaning by spelling out the semantics behind the choices ('top 10% or 1%', 'by return or by dollar profit'), but adds no format or edge-case detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific analytic question (does copying top traders persist?) and the exact population measured (Hyperliquid top 10%/1% by return or dollar profit, 4,000 random accounts, ~30 month pairs). It explicitly distinguishes itself from leaderboard_rank_check and overfitting_odds, so an agent can route correctly without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete trigger ('Use when someone proposes copying top traders, a leaderboard or signal leaders') and names two alternatives with the conditions that select them (a single return -> leaderboard_rank_check; is luck the explanation -> overfitting_odds). When-to-use, when-not, and alternatives are all present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diversification_checkDiversification checkARead-onlyIdempotentInspect
Diversification check: how many independent bets a basket of coins really is. Give the coins (and optionally the window: 90, 365 or 730 days): you get the average correlation between the pairs on real daily returns, the number of independent bets it works out to, each coin's correlation with bitcoin — and what the equal-weight basket did on the days bitcoin closed 3% or more down: how often it fell too, by how much, and its worst such day. Dead coins included.
Use when a basket of two or more coins is called diversified or hedged. For one coin use drawdown_check; for two or three coins side by side on every check, coin_comparison.
| Name | Required | Description | Default |
|---|---|---|---|
| days | Yes | Window ending on the last day of the data: 90, 365 or 730. Defaults to 365. | |
| coins | Yes | Tickers or full pairs, e.g. ["BTC", "ETH", "SOL"]. A comma-separated string also works. One we do not hold comes back with the list. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| model | No | The assumptions, stated. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| btc_bad_days | No | The days bitcoin closed 3% or more down: how many, on what share of them the basket fell too, its median and worst move. Diversification that vanishes on those days was never there. |
| independent_bets | No | N / (1 + (N-1)*avg_corr). Five coins that move as one are one bet; five that ignore each other are five. |
| all_coins_we_hold | No | What every coin we hold, at equal weights, adds up to — the ceiling, for comparison. |
| correlation_with_btc | No | Each coin's correlation with bitcoin in the window. |
| average_pair_correlation | No | Pearson correlation of daily log returns, averaged over every pair in the basket. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: the window is optional with a default, and 'Dead coins included' tells the agent how delisted assets are handled — a meaningful data-coverage disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The output set is front-loaded and the routing rule follows; every sentence carries information. It is denser than strictly necessary — the long output enumeration in sentence one could be trimmed — but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, yet the description previews them usefully and covers the trigger, alternatives, and dead-coin handling. An agent has enough to call it correctly; only minor gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented with examples, enum values, and the missing-coin behavior. The description restates the window choices (90/365/730) without adding syntax or meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb/resource ('how many independent bets a basket of coins really is') and concretely enumerates the outputs: average pairwise correlation, independent-bet count, per-coin correlation with bitcoin, and basket behavior on bitcoin-down days. It clearly differentiates itself from siblings like drawdown_check and coin_comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the trigger condition ('when a basket of two or more coins is called diversified or hedged') and routes to alternatives: 'For one coin use drawdown_check; for two or three coins side by side on every check, coin_comparison.' Both the when and the when-not are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
drawdown_checkDrawdown checkARead-onlyIdempotentInspect
Drawdown check: what it cost to collect the return you were shown. The coin is bought at the close of every day of its history and held for the period you name: you get the drawdown from the entry before the period was out, the share of starts that fell 30% and 50%, the days spent under water, how it ended — and, among the starts that ended in profit, the drawdown they sat through first. Spot tape, no fees, coins that died included.
Use when a return is quoted for holding a coin ("BTC did +150% in a year") and you want the fall endured on the way. With leverage use leverage_survival; to test a stop that cuts the fall, stop_loss_check.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | Ticker or full pair. An unknown one comes back with the list we hold and the ones left out. | |
| days | Yes | How long it is held: 30, 90, 365 or 730. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| model | No | The assumptions, stated. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| starts | No | Start days tested. Fewer than 120 and we refuse to give a percentage. |
| median_return_pct | No | How the middle start ended. |
| share_fell_30_pct | No | Share of starts that fell 30% or more from entry first. |
| share_fell_50_pct | No | Same, 50% or more. |
| asset_max_drawdown | No | The coin's own biggest top-to-bottom fall, with dates and days to recover, for comparison. |
| share_in_profit_pct | No | Share of starts that ended up. |
| median_worst_fall_pct | No | The middle start's worst fall from its own entry price within the hold. |
| median_days_under_water | No | Days of the hold spent below the entry price, middle start. |
| share_never_recovered_pct | No | Starts that never closed back at their entry price within the hold. |
| median_worst_fall_among_winners_pct | No | Among the starts that ended in profit, the median fall they sat through first. The price of the number you were shown. A return without this is advertising. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuinely useful scope context beyond them: 'Spot tape, no fees, coins that died included' tells the agent what data is and isn't in the result, plus it enumerates the computed metrics. It stops short of discussing rate limits or auth, but those are marginal for a local read-only computation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded and free of filler sentences, but the first paragraph enumerates output metrics ('share of starts that fell 30% and 50%, days spent under water...') that the existing output schema already delivers, and the prose is dense and stylized rather than tight. Some of the length duplicates structured data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with a full schema, complete enum coverage, annotations, and an output schema, the description supplies the remaining needed context: what is computed, the trigger condition, the data-scope caveats, and the alternatives. An agent has everything required to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the enum of allowed hold lengths (30/90/365/730) is documented in the schema itself. The description only restates the hold concept, adding no format or syntax detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete computation on a specific resource: buy at every daily close across a coin's history, hold for the named period, and return drawdown statistics. It is immediately distinguishable from siblings like leverage_survival and stop_loss_check, which it names explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('when a return is quoted for holding a coin ... and you want the fall endured on the way') and names two alternatives with the conditions that select them: leverage_survival for leverage and stop_loss_check for stop-cut falls. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leaderboard_rank_checkLeaderboard rank checkARead-onlyIdempotentInspect
How much company a return has. Send a return (+340%) and a window (day, week, month or all time) and get how many accounts on Hyperliquid's whole public leaderboard did the same or better, out of how many traded, its percentile, and the median account next to it. Every row of the exchange's own table, refreshed daily; the denominator is the accounts that traded, not the ones that sat idle.
Use when a trader or an ad quotes a return ("+340% this month") and you want to know how rare it is. Hyperliquid accounts only. Whether last month's top accounts stay on top is copy_trading_check; whether luck explains a win rate, overfitting_odds.
| Name | Required | Description | Default |
|---|---|---|---|
| window | Yes | day, week, month (default) or allTime. | month |
| return_pct | Yes | In percent: 340 means +340%. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| model | No | The method, stated. |
| one_in | No | One account in how many. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| percentile | No | Share of accounts below it. |
| same_or_better | No | Accounts at that return or above. |
| median_return_pct | No | The middle account, same window. |
| share_in_profit_pct | No | Accounts in profit, same window. |
| accounts_that_traded | No | The denominator. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower, yet the description adds real context: the data is the exchange's full table refreshed daily, and the denominator is accounts that traded rather than idle ones. It does not discuss rate limits or latency, but for a read-only lookup that is a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and scoped tightly to two sentences, and the sibling routing is compact. The opening phrase "How much company a return has" is slightly mannered, but it does not obscure the meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not needed, and the description covers the remaining essentials: trigger conditions, scope restriction to Hyperliquid accounts, data freshness, and denominator definition. An agent has everything required to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented in the schema, including the enum and the percent convention, so the baseline is 3. The description restates 'return (+340%)' and the window options but adds no format or edge-case detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (check how a return ranks against Hyperliquid's public leaderboard) with the exact output it produces: count doing the same or better, denominator, percentile, and median account. It also names the siblings it is not (copy_trading_check, overfitting_odds), so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use trigger ("when a trader or an ad quotes a return and you want to know how rare it is") plus scope limits ("Hyperliquid accounts only") and two named alternatives with their distinct questions. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leverage_survivalLeverage survivalARead-onlyIdempotentInspect
What a leveraged trade actually did, opened on every single day of the history instead of the one day that worked. Give coin, side, leverage and holding period: you get the share of those days that ended liquidated, the median outcome, the best day, and what the same coin did with no leverage at all. Real perpetual tape, with the funding that was actually paid charged daily and eating into margin. Coins that blew up included.
Use when the question involves leverage, perpetuals or liquidation. For the fall an unleveraged holder sits through use drawdown_check; for two or three coins at once, coin_comparison.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | Ticker or full pair. An unknown one comes back with the list we hold and the ones we do not. | |
| days | Yes | How long it is held: 1, 7, 30 or 90. | |
| side | No | long or short. Defaults to long. | long |
| leverage | Yes | 2, 3, 5, 10, 20, 25, 50 or 100. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| model | No | The assumptions, stated: liquidation at exactly 1/leverage with no maintenance margin, funding charged daily, the wick decides not the close. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| attempts | No | Starting days tested. A position still open when the history ends is not counted. |
| best_pct | No | The best day. It is the one you were shown, and it is not hidden here. |
| liquidated_pct | No | Share of them that ended with nothing left. |
| median_return_pct | No | The middle outcome, on the margin. |
| median_days_to_liquidation | No | How fast, when it happened. |
| median_return_unlevered_pct | No | What the same coin did with no leverage over the same days. A result without this is advertising. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, it discloses key behavior: real perpetual tape, daily funding charged into margin, inclusion of coins that blew up, and the counterfactual every-day simulation rather than cherry-picked dates. This gives the agent material context for trusting and interpreting results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core simulation concept, then adds data realism and edge-case behavior, then routing guidance. Despite being longer than a minimal definition, every sentence carries useful information for this complex tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be fully documented, yet the description still summarizes key outputs. Combined with annotations covering safety and the description covering methodology, input use, edge cases, and sibling routing, it is complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already defines coin, days, side, and leverage, including enums and defaults. The description names these inputs, but does not add syntax, constraints, or meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description concretely states that this tool simulates opening a leveraged trade on every historical day and reports outcomes like liquidation share, median, best day, and unleveraged comparison. It clearly distinguishes itself from siblings by naming drawdown_check and coin_comparison for adjacent questions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use it when the question involves leverage, perpetuals, or liquidation, and it routes the agent to drawdown_check for unleveraged falls and coin_comparison for multi-coin comparisons. Both when-to-use and alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
overfitting_oddsOverfitting oddsARead-onlyIdempotentInspect
Overfitting check: how many attempts plain chance needs to produce the track record you were shown. Give the wins, the losses and how many versions were tried: you get the exact binomial odds of one try reaching it, the odds once somebody shows you the best of N, and how many tries would make it an even bet. Arithmetic only — no market data, no model, no opinion. A 62% win rate over 100 trades is one thing on the first try, nothing on the fiftieth.
Use when you are shown a win rate, a backtest or a signal channel's record and need to know whether luck explains it. Needs no coin. For what a named rule did on real prices use strategy_grid_lookup; to place a return among real accounts, leaderboard_rank_check.
| Name | Required | Description | Default |
|---|---|---|---|
| wins | Yes | Winning trades. | |
| tries | No | How many versions were tried before this one was shown to you. Defaults to 1, which is almost never true. | |
| losses | No | Losing trades. Give this or trades. | |
| trades | Yes | Total trades. | |
| baseline | No | Probability a single trade wins by chance. Defaults to 0.5. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| error | No | no_data when ok is false. |
| reason | No | Why, in one sentence, and what we do have instead. |
| reading | No | The same three numbers in one sentence. |
| odds_one_try | No | P(at least this many wins) for a single attempt. Exact binomial, not simulated. |
| win_rate_pct | No | The win rate. |
| odds_best_of_tries | No | Same result, once you take the best of the attempts made. |
| tries_for_even_odds | No | Attempts needed for chance alone to reach it half the time. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds real behavioral context beyond them: 'Arithmetic only — no market data, no model, no opinion' tells the agent this is a pure deterministic calculation needing no external inputs, and it notes no coin is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The operational content is front-loaded and the second paragraph cleanly separates the when-to-use rule from the sibling routing. It is a touch prose-heavy, but every clause (pure arithmetic, no coin, the 62% illustration) carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a pure-arithmetic tool with an output schema, full-annotation coverage and complete schema descriptions, the definition supplies everything needed: trigger, routing, input expectations and the deterministic nature of the computation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description reinforces the semantics of 'tries' with the worked example ('62% over 100 trades is one thing on the first try, nothing on the fiftieth'), clarifying that the parameter dominates the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: computing the binomial odds that chance alone explains a shown track record from wins, losses and tries. It explicitly distinguishes itself from siblings by naming strategy_grid_lookup and leaderboard_rank_check as the tools for different questions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('Use when you are shown a win rate, a backtest or a signal channel's record and need to know whether luck explains it') and routes the agent elsewhere for the two adjacent cases by naming the alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stop_loss_checkStop loss checkARead-onlyIdempotentInspect
Whether a stop-loss would have helped, measured. The same buy of a coin, held 30, 90, 365 or 730 days, with a stop 5, 10, 15, 20 or 30% below the entry and without one, started on every day of the history. The stop fires on the day's low, not its close, and fees are charged to both. You get how often it fired, how often it sold a buy that would have ended in profit, and the worst case of each. Spot tape, dead coins included.
Use when asked whether a stop-loss protects a spot buy-and-hold. For leveraged positions, where liquidation is the stop, use leverage_survival; for a rule with a stop and a take-profit, strategy_grid_lookup.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | Ticker or full pair. An unknown one comes back with the list we hold and the ones left out. | |
| days | Yes | Holding period in days: 30, 90, 365 or 730. | |
| stop | Yes | How far below the entry, in percent: 5, 10, 15, 20 or 30. 0.05 is read as 5. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| buys | No | Start days tested. Fewer than 120 and we refuse to give a percentage. |
| error | No | no_data when ok is false. |
| model | No | The assumptions, stated: the low triggers, fills at the stop or the gap open, out until the end, same costs both ways. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| stop_fired | No | Buys on which the stop fired. A count. |
| stop_fired_pct | No | The same as a share of buys. |
| p05_with_stop_pct | No | Fifth percentile, with the stop: what the stop actually buys. |
| stop_beat_holding | No | Buys on which the stop ended strictly ahead of holding. |
| mean_with_stop_pct | No | Average buy, with the stop. |
| median_days_to_fire | No | Among the buys it fired on. Null when it never fired. |
| worst_with_stop_pct | No | The single worst buy, with the stop. Worse than the stop itself only when a day gapped through. |
| median_with_stop_pct | No | Middle buy, with the stop. |
| p05_without_stop_pct | No | The same, holding. |
| mean_without_stop_pct | No | Average buy, holding. Skewed by a few huge winners. |
| stop_beat_holding_pct | No | The same as a share of buys. |
| worst_without_stop_pct | No | The same, holding. |
| median_without_stop_pct | No | Middle buy, just holding. The alternative, always. |
| share_in_profit_with_stop_pct | No | Share of buys that ended up, stop. |
| share_in_profit_without_stop_pct | No | The same, holding. |
| fired_but_would_have_ended_in_profit | No | Buys the stop sold that, held to the end, would have ended in profit after fees. The half nobody shows. |
| fired_but_would_have_ended_in_profit_pct | No | The same as a share of ALL buys. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description goes well beyond them: the stop fires on the day's low rather than the close, fees are charged to both paths, the simulation starts on every day of history, spot tape only with dead coins included, and it names the returned metrics (fire frequency, false-positive sells, worst case). That is substantive methodology disclosure an agent needs to interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the outcome question, then methodology, then routing — every sentence carries information and nothing is padding. The prose is dense and slightly comma-chained in the middle, but it stays within a compact paragraph rather than sprawling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the read-only/idempotent profile and an output schema presumably documenting the returned metrics, the description only needs to supply methodology and routing — and it does both. Nothing required to invoke this correctly (coin, days, stop semantics, when to prefer a sibling) is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with enums fully documented, so the schema carries the baseline load. The description adds genuine meaning beyond it: the stop trigger is defined on the low (not the close) and fees apply to both the stop and no-stop paths, which changes how a result should be read. It also restates the enum ranges, which is redundant with the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line specifies the exact analysis ('Whether a stop-loss would have helped, measured') and the body enumerates the simulation design: buy-and-hold across 30/90/365/730 days, stops at 5-30%, run on every day of history. It explicitly separates itself from leverage_survival and strategy_grid_lookup by naming them. An agent can distinguish it from every sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a positive trigger ('Use when asked whether a stop-loss protects a spot buy-and-hold') and two explicit exclusions with the correct routing: leveraged positions go to leverage_survival, stop+take-profit rules go to strategy_grid_lookup. This is textbook when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
strategy_grid_lookupStrategy grid lookupARead-onlyIdempotentInspect
Backtest lookup: what one exact version of a strategy actually did. Buy-the-dip and the moving-average cross, every combination of their knobs, computed over nine years of real prices with the exchange's fees charged both ways — sixteen coins, three of which died. Give the coin and the numbers and you get the return, the trades, the worst drawdown, and what buying and holding did over the same window. Read from a file you can download.
Use when someone quotes a specific dip-buying or moving-average rule ("buy 10% dips, take 5%") and you want its real backtest. For leverage use leverage_survival; for a stop on a plain buy-and-hold, stop_loss_check; to judge a win rate you were shown, overfitting_odds.
| Name | Required | Description | Default |
|---|---|---|---|
| coin | Yes | BTC, ETH, SOL, DOGE, PEPE, SRM… ticker or full pair. Ask with an unknown one and the answer lists the sixteen we hold. | |
| drop | No | dip only: how far below its recent high to buy, in %. | |
| fast | No | cross only: fast average. | |
| slow | No | cross only: slow average. | |
| stop | No | dip only: stop loss, in %. | |
| hours | No | dip only: give up after this many hours. | |
| family | No | dip (buy when it falls) or cross (moving average crossover). Defaults to dip. | dip |
| target | No | dip only: take profit, in %. | |
| candles | No | cross only: 1d or 4h. |
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | No | false when we do not hold that data. Never a zero standing in for an answer. |
| rule | No | The rule in plain English. |
| error | No | no_data when ok is false. |
| reason | No | Why, in one sentence, and what we do have instead. |
| source | No | The public file these numbers come from. |
| trades | No | How many trades it took. |
| return_pct | No | What it returned, after fees. |
| win_rate_pct | No | Share of winners. |
| buy_and_hold_pct | No | What buying the coin and leaving it alone returned over the same window. A result without this is advertising. |
| max_drawdown_pct | No | Worst peak-to-trough fall. |
| versions_in_family | No | How many versions of this idea exist. |
| versions_beating_buy_and_hold | No | How many of them beat doing nothing. Often zero. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false, and destructiveHint=false, covering the safety profile. The description adds context beyond annotations: data is computed over nine years of real prices with exchange fees charged both ways, across sixteen coins (three of which died), and results are read from a downloadable file. It also previews return values (return, trades, worst drawdown, buy-and-hold), though an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence. The description is slightly verbose with colorful phrasing ('knobs', 'sixteen coins, three of which died') but each sentence carries relevant information about scope, returns, data source, and routing. It could be tightened, but it remains structured and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with an output schema and rich annotations, the description covers purpose, data coverage, outputs, and alternatives. It does not detail every parameter or edge case, but the output schema and 100% schema coverage handle those. It is largely complete for correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the input schema already documents all nine parameters, including enum values and defaults. The description only broadly references 'the coin and the numbers' and 'every combination of their knobs', adding no syntax or format details beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Backtest lookup') and resource: what one exact version of a strategy actually did. Distinguishes from siblings by naming leverage_survival, stop_loss_check, and overfitting_odds for alternative use cases. An agent can immediately tell this is the tool for backtesting a specific dip-buying or moving-average rule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use: when someone quotes a specific dip-buying or moving-average rule and you want its real backtest. Names three alternatives and the conditions that select them: leverage_survival for leverage, stop_loss_check for a stop on plain buy-and-hold, overfitting_odds to judge a win rate. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
- Changed
averaging_in_check3 fields changed- added
Input schema / properties / days / enumAdded value: +[ + 90, + 365, + 730 +] - added
Input schema / properties / instalments / defaultAdded value: +"weekly" - added
Input schema / properties / instalments / enumAdded value: +[ + "weekly", + "monthly" +]
- Changed
best_days_check3 fields changed- added
Input schema / properties / best_days / defaultAdded value: +10 - added
Input schema / properties / best_days / enumAdded value: +[ + 1, + 5, + 10, + 20 +] - added
Input schema / properties / days / enumAdded value: +[ + 90, + 365, + 730 +]
- Changed
calendar_check1 field changed- added
Input schema / properties / calendar / enumAdded value: +[ + "day_of_week", + "month_of_year", + "hour_of_day" +]
- Changed
coin_comparison3 fields changed- added
Input schema / properties / days / defaultAdded value: +365 - added
Input schema / properties / days / enumAdded value: +[ + 90, + 365, + 730 +] - added
Input schema / properties / leverage / defaultAdded value: +25
- Changed
copy_trading_check4 fields changed- added
Input schema / properties / ranked_by / defaultAdded value: +"return" - added
Input schema / properties / ranked_by / enumAdded value: +[ + "return", + "profit" +] - added
Input schema / properties / top_pct / defaultAdded value: +10 - added
Input schema / properties / top_pct / enumAdded value: +[ + 10, + 1 +]
- Changed
diversification_check2 fields changed- added
Input schema / properties / days / defaultAdded value: +365 - added
Input schema / properties / days / enumAdded value: +[ + 90, + 365, + 730 +]
- Changed
drawdown_check1 field changed- added
Input schema / properties / days / enumAdded value: +[ + 30, + 90, + 365, + 730 +]
- Changed
leaderboard_rank_check2 fields changed- added
Input schema / properties / window / defaultAdded value: +"month" - added
Input schema / properties / window / enumAdded value: +[ + "day", + "week", + "month", + "allTime" +]
- Changed
leverage_survival4 fields changed- added
Input schema / properties / days / enumAdded value: +[ + 1, + 7, + 30, + 90 +] - added
Input schema / properties / leverage / enumAdded value: +[ + 2, + 3, + 5, + 10, + 20, + 25, + 50, + 100 +] - added
Input schema / properties / side / defaultAdded value: +"long" - added
Input schema / properties / side / enumAdded value: +[ + "long", + "short" +]
- Changed
overfitting_odds2 fields changed- added
Input schema / properties / baseline / defaultAdded value: +0.5 - added
Input schema / properties / tries / defaultAdded value: +1
- Changed
stop_loss_check2 fields changed- added
Input schema / properties / days / enumAdded value: +[ + 30, + 90, + 365, + 730 +] - added
Input schema / properties / stop / enumAdded value: +[ + 5, + 10, + 15, + 20, + 30 +]
- Changed
strategy_grid_lookup3 fields changed- added
Input schema / properties / candles / enumAdded value: +[ + "1d", + "4h" +] - added
Input schema / properties / family / defaultAdded value: +"dip" - added
Input schema / properties / family / enumAdded value: +[ + "dip", + "cross" +]
12 tool updates
- First observed
averaging_in_check - First observed
best_days_check - First observed
calendar_check - First observed
coin_comparison - First observed
copy_trading_check - First observed
diversification_check - First observed
drawdown_check - First observed
leaderboard_rank_check - First observed
leverage_survival - First observed
overfitting_odds - First observed
stop_loss_check - First observed
strategy_grid_lookup
Related MCP Connectors
Read-only crypto strategy backtest & verification verdicts. No trading tools; every result is cited.
- mcpOAuthcom.market-graphs
Praxis: published trading strategies re-run and audited — verdicts, claimed vs measured stats.
Is a trading strategy's track record luck? Free luck check, plus what a full strategy audit checks.
Crypto backtesting & Bitcoin cycle analytics. Point-in-time, DSR-corrected, look-ahead-aware.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceChecks whether a trading backtest survives its own statistics: deflated Sharpe, multiple-testing correction against a best-of-N-noise benchmark, minimum track record length, and fill realism. Takes no market data and no API keys, and cannot recommend a trade — it only reports that a result is weaker than claimed or not yet provable.MIT
- AlicenseNot gradedqualityCmaintenanceEnables statistical validation of whether a trading backtest’s edge is real using bootstrap-resampling and reshuffling Monte Carlo methods, including confidence intervals, drawdown-path percentiles, expected value calculations, and prop-firm challenge pass-probability simulation. It also compares multiple win-rate/risk-reward geometries by simulated pass rate.MIT
- AlicenseNot gradedqualityBmaintenanceProvides tools to research crypto trading strategies via backtesting, walk-forward validation, and paper trading, with a deflated-Sharpe overfitting check. Enables natural-language-driven analysis and interpretation of strategy performance.3Apache 2.0
- AlicenseAqualityAmaintenanceMost trading signals are noise. AlphaAssay puts them on trial — deflated Sharpe, out-of-sample, leakage forensics — and returns signed pass/fail verdicts anyone can verify. Methodology audits, not investment advice.17Apache 2.0
Glama MCP Gateway
Add one secure layer between your agents and this server.