Skip to main content
Glama

Backtesting Arena

Server Details

Crypto backtesting & Bitcoin cycle analytics. Point-in-time, DSR-corrected, look-ahead-aware.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
100.0% over 47 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
Schoasch/skill-backtesting-arena
GitHub Stars
0
Server Listing
Backtesting Arena

TDQS

A4/5.0

Scored across 49 tools

Disambiguation3/5

Most tools have clearly distinct purposes (snapshot vs history vs subscription vs backtest), but with 49 tools there is real overlap: arena_get_gem_score vs arena_get_gem_scores (singular/plural on different resources), arena_get_signal_status vs arena_get_signal_context vs arena_get_signal_events, and a cluster of regime/volatility readers (arena_get_pulse, arena_get_cycle, arena_get_bullmarket_ampel, arena_get_volatility_phases/insights/recommendations) that a hurried agent could confuse. The descriptions are unusually detailed and do resolve most cases, so this is 'some overlap but descriptions help' rather than ambiguous throughout.

Naming Consistency4/5

A clear arena_ prefix plus verb_noun convention dominates (arena_get_*, arena_list_*, arena_run_*, arena_subscribe_*, arena_share_*), all snake_case. Deviations exist: validate_strategy drops the arena_ prefix entirely, and a handful use non-verb-first or verb-last forms (arena_status, arena_batch, arena_cancel_subscription, arena_dip_decision, arena_is_distinguishable). Mostly consistent with minor deviations.

Tool Count2/5

49 tools is at the threshold of 'too many' for a coherent set, well beyond a well-scoped surface. The server mitigates this with arena_batch (batched reads) and arena_call_extended (a gateway that folds per-metric history/niche tools into one line), which shows the authors know the count is a problem, but the browsable surface is still very heavy and selection cost is high.

Completeness4/5

Coverage is unusually broad: backtest create/read/list/trades, universe and job lifecycle, filters/regime/robustness analysis, signal validation, market/on-chain/options data, alert subscriptions, and knowledge objects. Minor gaps remain (e.g. no listable delete/prune for saved backtests, no visible report-request tool alongside arena_get_report_status), but arena_call_extended provides an escape hatch for extended operations.

Available Tools

49 tools
arena_batchBatch — several snapshots in one callAInspect

Several market snapshots in ONE call instead of one roundtrip per read. Batchable reads (14): spot_price, pulse, cycle, fear_greed, funding_rate, macro_regime, iv_snapshot, etf_flows, stablecoin_supply, mayer_multiple, onchain_latest, max_pain, altcoin_season, bullmarket_ampel. Pass 1-6 queries; each returns its result OR a structured error (partial success — one failing query does not abort the rest). Each query consumes one rate-limit unit: the batch saves roundtrips, not quota. Payloads, tier gates and source attribution are identical to the single tools; per-query args match the single tool's parameters (e.g. {tool: "iv_snapshot", args: {currency: "BTC"}}). For history tools, backtests or anything not in the list, call the single tool. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
queriesYes1-6 queries, executed in order.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses partial success semantics (each query returns a result or a structured error, one failure doesn't abort the rest), the rate-limit accounting (batch saves roundtrips, not quota), and that payloads, tier gates, and source attribution are identical to the single tools. This gives the agent a precise model of side effects and error handling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds information: purpose, allowed tools, failure behavior, rate-limit nuance, payload equivalence, and an exclusion. It is front-loaded with the core value proposition. While it is a longer description, it avoids filler and is well-structured for an agent to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description adequately covers all essential details: the allowed set of reads, query count limits (1-6, enforced by schema), behavior on partial failure, rate-limit implications, and how to route any request outside the batch to a single tool. The agent has enough to call the tool correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes the queries array and the tool enum, with 100% coverage. The description adds value by providing an explicit example ({tool: 'iv_snapshot', args: {currency: 'BTC'}}), clarifying that args map to the single tool's parameters, and noting that args should be omitted when the underlying tool takes none. This goes beyond the schema's generic 'Args of the underlying single tool' text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool batches multiple market snapshot reads into one call, and it enumerates the 14 specific supported reads (spot_price, pulse, cycle, etc.). This distinguishes it from the sibling single-read tools (arena_get_spot_price, arena_get_cycle, etc.) and leaves no ambiguity about which tool to use for batching.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when the batch tool is appropriate (multiple reads of the listed types) and when it is not: 'For history tools, backtests or anything not in the list, call the single tool.' It also explains the practical benefit (saves roundtrips, not quota) and gives a concrete example of query construction, so the agent knows exactly how to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_call_extendedCall an extended toolAInspect

Gateway to the EXTENDED tools of this server — listed here in one line each instead of individually, to keep the tool list short. Use it when no listed tool fits: per-metric daily history series, subscriptions, chart images, niche primitives. mode="call" runs the tool with arguments (same result, auth, tier and rate limits as calling it directly); mode="describe" returns its full description and parameters first, if the arguments are unclear.

  • arena_cancel_subscription: Stop this alert?

  • arena_check_subscription_updates: Has anything I subscribed to fired?

  • arena_cross_series: Did two market series move together, and what did BTC do next when they agreed or diverged?

  • arena_dca_scenario: Should I invest all at once or spread it out (DCA)?

  • arena_dip_scenario: Where would I add on a dip, and when is the thesis wrong?

  • arena_get_altcoin_season_history: Has capital been rotating into or out of altcoins?

  • arena_get_backtest_trades: Which trades did that backtest actually take?

  • arena_get_btc_macro_correlations: What does Bitcoin actually move with?

  • arena_get_carry_monitor: Which part of a week's spot-BTC-ETF inflows coincided with a short build-up by CME Leveraged Funds — i.e.

  • arena_get_chart: Renders one of the named platform series as a PNG line chart and returns it as an MCP image content block, plus a JSON …

  • arena_get_cost_basis_spread: Is the market in profit or at a loss?

  • arena_get_cycle_history: How has the cycle score moved over time?

  • arena_get_drift_log: Do two independent providers still agree on the same on-chain quantity?

  • arena_get_filter_insights: Do entry filters help, and which ones?

  • arena_get_funding_rate_history: How has leverage positioning shifted over time?

  • arena_get_gem_score: How does this altcoin score?

  • arena_get_gem_validation: Did the screener picks actually beat BTC?

  • arena_get_halvings: When were the halvings, and what followed?

  • arena_get_hash_ribbons: Are miners capitulating?

  • arena_get_kimchi_premium: Is Korean spot trading BTC at a premium?

  • arena_get_ma_distance_history: How far above or below its moving averages did price stand back then?

  • arena_get_mayer_multiple: Is BTC stretched against its 200-day average?

  • arena_get_mayer_multiple_history: How stretched has BTC been against its 200-day average?

  • arena_get_ontology_term: What does this term mean here, exactly?

  • arena_get_pattern_scan: Is a classic chart formation (descending/ascending wedge, symmetric triangle, range) forming on a crypto pair right now…

  • arena_get_platform_activity: What are people backtesting right now?

  • arena_get_portfolio_correlation: Is my crypto portfolio actually diversified?

  • arena_get_pulse_history: How did market heat get to where it is?

  • arena_get_reference_models: How much does a Bitcoin valuation-model line depend on WHEN it was computed?

  • arena_get_report_status: Is my report ready?

  • arena_get_shared_backtest: What is in this shared backtest link?

  • arena_get_signal_events: When did which classic top/bottom signal actually flip?

  • arena_get_signal_status: Is this strategy signalling buy or sell right now?

  • arena_get_taker_imbalance: Are taker buys or taker sells dominating?

  • arena_get_trend_channels: Where does price sit inside its trend channel?

  • arena_get_volatility_insights: Does this strategy work better in calm or wild markets?

  • arena_get_volatility_phases: Is this pair calm or wild right now?

  • arena_get_volatility_recommendations: Which strategies suit the current volatility phase?

  • arena_get_volume_profile: WHERE was Bitcoin actually traded, as opposed to where it turned?

  • arena_get_winners: What are the strongest backtest results on the platform?

  • arena_list_subscriptions: Which alerts do I have running?

  • arena_quote_report: What would a custom report cost?

  • arena_share_grid_backtest: Want a public link for a grid result?

  • arena_subscribe_bullmarket_stage: Notify me when the bull-market stage count changes?

  • arena_subscribe_cycle_changes: Notify me when the cycle band changes?

  • arena_subscribe_pulse_changes: Notify me when market heat crosses a threshold?

  • arena_subscribe_signal_alerts: Notify me when this signal flips?

  • arena_suggest_grid_range: Which price range should my grid bot use?

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo"call" runs the tool; "describe" returns its description and input schemacall
toolYesName of the extended tool to run
argumentsNoArguments for the tool, exactly as the tool itself takes them (mode="call")

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the burden and does well: it discloses that mode="call" has 'same result, auth, tier and rate limits as calling it directly' and that describe is a safe preflight. It does not, however, warn that the catalog itself includes destructive/state-changing members (arena_cancel_subscription, subscribe_*), nor does it note any caveats about the truncated one-line summaries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the gateway purpose and mode semantics before the catalog, so the important content comes first. The 48-entry catalog is long but is the functional payload; however several entries are visibly truncated ('i.e.', trailing ellipses), which blunts their value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a meta-dispatch tool with no output schema and no annotations, the description covers purpose, routing, mode behavior, and parity of auth/limits. It leaves the return shape of mode="call" only implied ('same result ... as calling it directly'), which is acceptable since the underlying tools define it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains what mode="call" versus mode="describe" actually do at runtime and that `arguments` must be 'exactly as the tool itself takes them', which is not stated in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific role: a gateway that exposes 'EXTENDED' tools one line each rather than as separate entries, and explains the reason (keeping the tool list short). An agent can immediately distinguish this from the ~50 flat sibling tools like arena_get_pulse or arena_get_cycle.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the selection condition: 'Use it when no listed tool fits', and then sub-routes internally with mode="call" vs mode="describe", recommending describe 'if the arguments are unclear'. That is a complete when-to-use and how-to-decide story.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_compare_strategiesCompare 2-5 StrategiesAInspect

Which of these strategies performed best on the same data? Run 2–5 strategies against the SAME pair, interval and date range and return per-strategy metrics plus a comparison summary (best by CAGR, best by win-rate, worst by drawdown). Use this when the user asks which of several strategies fits a market — it holds the pair, interval and requested date range fixed, which a series of separate arena_run_backtest calls does not guarantee. What it does NOT equalize is the EVALUATION window: a strategy with a long warmup starts trading later, so compare actual_date_from across the runs and check result.benchmark before ranking by CAGR. For one strategy across many pairs use arena_run_universe_backtest instead. Caveat worth passing on: comparing N strategies and reporting the winner IS multiple testing — the winner’s edge is upward-biased. arena_get_robustness_field puts a counted N on that. Sequential, expect 10–50s. Per-day quota: Pro=20, Power=200. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCrypto pair symbol, e.g. BTCUSDT — the same pair for every strategy.
capitalNoStarting capital in quote currency. Default 10000. Affects absolute figures only, not CAGR or win-rate.
date_toNoEnd date, YYYY-MM-DD. Default: today.
filtersNoOptional entry filters (Pro+). Each one only ever REMOVES entries — filters never create trades. Omit for the unfiltered baseline.
intervalYesCandle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
date_fromYesStart date, YYYY-MM-DD. Earlier than the pair listing is clamped to the first available candle.
asset_typeYesAsset class. Use 'crypto' unless you are explicitly backtesting a tokenized real-world asset. Note: tokenized stocks/ETFs/gold trade AS crypto pairs (e.g. spybUSDT, qqqbUSDT) — there is no separate stocks/forex backtest surface; non-crypto asset classes were retired.
strategiesYesThe 2–5 strategies to compare, each with optional own params.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden — and meets it thoroughly. It reveals the non-equalized evaluation window (long-warmup strategies start trading later; compare actual_date_from and check result.benchmark before ranking by CAGR), the multiple-testing upward bias of reporting a winner, sequential execution with 10–50s latency, per-day quotas (Pro=20, Power=200), and the Pro-tier API requirement. This is far beyond a generic 'runs a comparison' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but disciplined — the opening question front-loads the purpose and each subsequent sentence adds a distinct fact: same-data guarantee, sibling differentiation, evaluation-window caveat, multiple-testing caveat, latency, quota, tier. The tail stacks several caveats and operational notes in quick succession, which is slightly heavy, but there is no redundant filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no output schema, no annotations, and a family of 75+ siblings, this is remarkably complete: purpose, scoping guarantees, the metric categories returned (best by CAGR, best by win-rate, worst by drawdown), behavioral caveats, latency, quota, tier, and sibling routing all appear. The high-level return shape compensates adequately for the missing output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds genuine meaning on top: it establishes the core guarantee that pair, interval and date range are held constant across all strategies, and warns that date_from clamping interacts with warmup so the actual data windows differ across runs. It correctly avoids restating per-parameter docs the schema already covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a concrete question ('Which of these strategies performed best on the same data?') and states the verb-resource pair: run 2–5 strategies against the same pair/interval/date range and return per-strategy metrics plus a comparison summary. It explicitly distinguishes itself from a series of separate arena_run_backtest calls and from arena_run_universe_backtest (one strategy across many pairs), so an agent can select it correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger: 'Use this when the user asks which of several strategies fits a market.' It names the alternative it competes with (a sequence of arena_run_backtest calls) and explains why this tool is superior (it holds pair/interval/date range fixed, which separate runs do not guarantee). It also routes the inverse case to arena_run_universe_backtest and the multiple-testing follow-up to arena_get_robustness_field — full when-to-use and when-not-to-use coverage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_dip_decisionDip Decision — buy now or wait?AInspect

Buy now or wait for the dip? Decision-math over the user's OWN assumptions (target/dip prices, probabilities, capital). Two modes: "compare" = expected value of Buy-Now vs Wait vs Split + the breakeven dip probability (prices as MULTIPLES of today); "allocate" = the risk-adjusted (Kelly / risk-aversion γ) optimal fraction to deploy now vs reserve for the dip (ABSOLUTE prices). Ask the user for the missing inputs, then call. Returns scenario numbers and which option wins on expected value — NOT a buy/sell recommendation. For the full interactive version (incl. leverage & Elliott-wave planning) point the user to https://tradingstrategies.work/analyse/dip-decision. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNo'compare' (default): EV of buy-now vs wait vs split + breakeven dip probability. 'allocate': risk-adjusted optimal deploy-now fraction under γ.compare
compareNoRequired when mode='compare'.
allocateNoRequired when mode='allocate'.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It explicitly states it returns scenario numbers and the winner on expected value, and explicitly denies being a buy/sell recommendation. It also notes that calculations are over the user's own assumptions (not live data) and mentions a free tier, adding transparency. It doesn't cover edge cases like inconsistent probabilities, but this is acceptable for a math tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (~200 words) but well-structured: it leads with a clear question and purpose, then details modes, clarifies the non-recommendation aspect, and ends with a practical URL. While not as terse as ideal, every sentence carries information relevant to correct invocation, and the structure is logical and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (two modes, many parameters, no output schema), the description covers the essential items: what each mode does, the input types (multiples vs absolute), the non-recommendation nature, and guidance to ask for missing inputs. It lacks a precise description of return format, but the stated 'scenario numbers and which option wins' is adequate for an agent to interpret results. The URL fallback also compensates for missing interactivity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema by explaining that 'compare' prices are multiples of today and 'allocate' uses absolute prices, and by describing the risk-aversion parameter in a practical way. This enhances understanding without being redundant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: decision-math over the user's own assumptions to decide buy-now vs wait, with two explicit modes ('compare' and 'allocate'). It distinguishes itself from siblings by detailing what each mode computes (expected value vs risk-adjusted allocation) and explicitly says it is not a buy/sell recommendation, making its role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives practical guidance: 'Ask the user for the missing inputs, then call' sets a clear precondition. It also provides a fallback alternative for the full interactive version via URL, clarifying when to direct users elsewhere instead of calling. It does not explicitly compare to sibling tools, but the guidance is sufficient for typical use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_altcoin_seasonGet Altcoin Season SnapshotAInspect

Is it altcoin season? Daily Altcoin-Season indicator (v7 Native-Filter methodology). Returns BTC-Dominance, Alt-Dominance, 4 Layer-1 signals (USDT.D, USDC.D, BTC-DOM, ETH-DOM), overall color (red/amber/green) + Top-50 CoinGecko snapshot. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the data fields returned, the update frequency ('Daily'), and a 'Free tier' caveat suggesting possible limitations. It does not explicitly state that it is read-only, but the verb 'get' and the nature of a snapshot make it implicit. The methodology mention adds valuable context. Overall, it provides solid transparency beyond trivial details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the core question ('Is it altcoin season?') and then compresses the key details into a list. There is no fluff, and every phrase adds value—methodology, data fields, color scheme, and tier note. It is appropriately concise and immediately scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description must convey what the call returns. It lists all major fields: BTC-Dominance, Alt-Dominance, 4 Layer-1 signals, overall color, and a Top-50 snapshot. It also notes the daily update cycle and methodology. While the exact structure of the Top-50 snapshot is not detailed, the description is sufficient for an agent to understand what to expect and to call the tool correctly. Minor lack of formatting details keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing for the description to clarify. Per the rubric, a baseline of 4 applies when there are no parameters. The description does not need to add parameter semantics, and it correctly omits them. It adds no redundancy with the schema, which is empty and trivially 100% covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it reports the current altcoin season status using a specific methodology (v7 Native-Filter). It lists exact outputs (BTC-Dominance, Alt-Dominance, Layer-1 signals, color, and Top-50 snapshot). The word 'Daily' and 'snapshot' differentiate it from the sibling 'arena_get_altcoin_season_history', establishing it as the current-state version. This is a specific verb+resource with clear distinguishing scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for current altcoin season state through 'Daily' and 'snapshot', but it never explicitly states when to use this tool versus alternatives like the history version or other snapshot indicators. No exclusion conditions or alternative names are provided. An agent would infer the usage from context, but explicit guidance is missing, which is a gap given the large sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_asset_snapshotAsset Snapshot — one coin, one callAInspect

Where does this coin stand? ONE call per Binance USDT pair instead of six: last daily close, 7/30/90/365-day returns, relative strength vs BTC and vs ETH on the same horizons (with the MEASURED base rate next to it — the median altcoin loses against Bitcoin, so a positive number is a description, not an edge), the F6 trend state vs BTC, ATH/drawdown/days-since-ATH on the available exchange history (ath_scope says which), SMA200 distance, a parabolic state (in a parabolic run now? last run? plus what followed such runs per exit rule, from knowledge object parabolic_base_rate), realized 30d volatility and ATR%, liquidity from our own daily Binance universe measurement (24h-volume rank today vs 30 days ago, 30d mean/median volume, band), tokenomics ratios from the gem screener (Pro+, CoinGecko ratios only), derivatives (BTC only so far) and a data_quality block: history span, candle count, missing days, coverage %, source/stitch, listing status (delisted pairs are flagged) and a mechanical A/B/C grade whose rule travels in the payload. Every source can fail independently — sources_used / sources_unavailable make the basis auditable. Works for any Binance USDT pair, not just BTC/ETH/SOL. detail: 'full' adds the raw BTC/ETH benchmark returns behind the relative numbers. Descriptive, no signal. For the strategy-side question ("should I take this entry?") use arena_get_signal_context. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesBinance USDT pair, e.g. 'SOLUSDT'. Case-insensitive.
detailNo'standard' (default). 'full' adds the raw BTC/ETH benchmark returns used for the relative numbers.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond basic operation to reveal failure modes ('Every source can fail independently'), data-quality caveats ('history span, candle count, missing days... delisted pairs are flagged'), the mechanical A/B/C grade rule that travels in the payload, and an interpretive caution ('the median altcoin loses against Bitcoin, so a positive number is a description, not an edge'). This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, very dense paragraph with a long enumeration of metrics, conditions, and caveats. It lacks visual structure (bullets, sections) and is considerably longer than needed for an agent to parse quickly. While every sentence carries useful information, the lack of organization and sheer length hinder readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's high complexity and the absence of an output schema, the description covers the major return components comprehensively (returns, trend, ATH, volatility, liquidity, tokenomics, derivatives, data quality, sources used/unavailable). It does not specify the exact response format, but it still gives an agent enough context to predict what the tool will return and how to interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description adds no new semantic detail beyond what the schema provides (pair case-insensitivity, any Binance USDT pair, detail full adds raw returns). While it reinforces the pair scope, it doesn't add meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description answers 'Where does this coin stand?' and enumerates a specific set of metrics (returns, relative strength, trend, drawdown, volatility, liquidity, tokenomics, data quality), making the resource and scope clear. It also explicitly differentiates from arena_get_signal_context ('For the strategy-side question...'), which distinguishes it from the closest sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states when to use the tool (descriptive snapshot for any Binance USDT pair) and when not to use it ('Descriptive, no signal', and 'For the strategy-side question... use arena_get_signal_context'). It also identifies an alternative tool by name, which is exactly what high-quality usage guidance should do.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_backtestGet Backtest DetailAInspect

What exactly did that backtest do? Returns the full record of ONE backtest run by id: strategy, pair, interval, date range, parameters, filters and the aggregate metrics (CAGR, total return, win-rate, max drawdown, trade count, Buy & Hold comparison, net-of-fees figures). Only your own runs (admins may read others). Get ids from arena_list_backtests; for the individual trades add arena_get_backtest_trades; to create a new run use arena_run_backtest. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesUUID of the backtest run.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses access control (own runs only, admin override) and the API Pro tier requirement, both behavioral traits. However, it doesn't mention error behavior (e.g., 404 for nonexistent id) or response format details, but for a simple get-by-id operation the key caveats are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a purpose question, then compresses a full field list, access rule, sibling pointers, and tier note into three sentences. All sentences carry useful information with no filler, though the density might be slightly high for quick scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one param and no output schema, the description covers what it returns, how to get the id, alternatives, access control, and tier requirement. Nothing essential is missing for an agent to call it correctly; even edge cases like permission and tier are addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (id is described as 'UUID of the backtest run'), so the baseline is 3. The description adds value by instructing where to obtain the id (arena_list_backtests) and reinforcing that the id identifies a single run, going beyond the schema's terse field description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete question that frames the tool's purpose, then states precisely: 'Returns the full record of ONE backtest run by id' and enumerates the exact contents (strategy, pair, interval, metrics). It explicitly differentiates from siblings by naming arena_get_backtest_trades and arena_run_backtest, so an agent can distinguish it without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing: 'Get ids from arena_list_backtests', 'for the individual trades add arena_get_backtest_trades', and 'to create a new run use arena_run_backtest'. It also clarifies access scope ('Only your own runs (admins may read others)') and tier requirement ('[API Pro tier]'), giving an agent clear when-to-use and when-not-to-use signals.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_btc_market_structureGet BTC Market StructureAInspect

Is the trend up or down, and how fresh is the flip? Daily Bitcoin market structure from 1000-bar Phantomflow adaptation (BTCUSDT 1d). Returns current_trend (up/down/sideways), last trend change timestamp, counts of waves + fractals, last-5 fractals on each side (up = pivot highs, down = pivot lows), and trend_context: previous trend + its duration, flip_age_days, and a descriptive historical flip base rate over the SAME 1000 bars (total flips, share reverted within 5 bars, median trend duration) — a fresh same-day flip is the least settled observation — the base rate tells you how often such flips reverted historically, so you can weight the current one yourself. Educational analysis of price action. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It discloses the underlying timeframe, the 1000-bar sample, the trend_context fields, and even the interpretive caveat that a fresh same-day flip is the least settled observation. The 'Educational analysis' tag further signals that this is analytical, not financial advice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but densely packed with essential information: inputs, timeframe, all returned fields, and interpretive guidance. It is front-loaded with the core question and purpose. Slightly run-on in places, but no sentence is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema and no annotations, the description fully compensates by enumerating every major return component: current_trend, timestamp, wave/fractal counts, last-5 fractals per side, and the entire trend_context object with its base-rate stats. The agent can understand what call will yield and how to interpret the freshest flip.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description provides all necessary context about what the tool outputs without needing to explain parameter behavior. There is nothing missing here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete question ('Is the trend up or down, and how fresh is the flip?') and names the exact resource: daily Bitcoin market structure from a 1000-bar Phantomflow adaptation. It clearly identifies the tool as a trend-structure getter, distinct from the many other arena_get_* indicator tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context around when this is relevant: for daily BTCUSDT trend direction, flip freshness, and historical flip behavior. It stops short of explicitly naming alternatives or stating when not to use it, but the data source and 'educational analysis' framing give the agent enough context to select it from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_bullmarket_ampelGet Bullmarket Ampel SnapshotAInspect

Is this still a bull market? Bitcoin Bullmarket-Ampel current state (0-5 active stages). Returns active_count, a stages[] breakdown (each stage with key, label, active and since = first day of its current state; null when the state predates the 400-day lookup) and stage_history — per day active_count PLUS all five per-stage booleans, so which stage flipped when is readable directly (history_days 1-365, default 30). Higher count = more bull-market signals firing. Stages evaluate weekly 20W/50W-MA conditions. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
history_daysNoDays of stage_history to return (1-365, default 30). Each row carries active_count plus all five per-stage booleans, so stage flips are readable per day instead of only via the derived `since` of the current run.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the burden, and it discloses substantial behavior: output shape (active_count, stages[], stage_history), the null `since` edge case for pre-400-day states, per-day booleans, weekly 20W/50W-MA evaluation, and free-tier status. It doesn't explicitly address rate limits or side effects, but the read-only nature is strongly implied by 'Returns'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but mostly front-loaded: it leads with the purpose, then return structure, parameter behavior, interpretation, and cadence. A few long parentheticals make it parse slightly harder, but no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must explain return values, and it does: active_count, stages[] fields including the null `since` semantics, and stage_history as per-day booleans. The parameter bounds, default, evaluation frequency, and free-tier context are also present, so an agent has enough to invoke and interpret the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and the schema's parameter description already contains the same detail about history_days bounds, default, and per-day booleans. The tool description repeats that content without adding new meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the exact question it answers and names a specific resource: the Bitcoin Bullmarket-Ampel. It clearly defines scope (current state 0-5 plus history) and the unique 'stages' concept, distinguishing it from generic market/correlation siblings without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The rhetorical 'Is this still a bull market?' and 'Higher count = more bull-market signals firing' imply the use case: assessing bull-market strength via the Ampel stages. However, it does not state when to prefer this over related siblings (e.g., cycle, market structure, macro regime), leaving the selection logic implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_cycleGet Crypto Cycle Snapshot (BTC / ETH / SOL)AInspect

Crypto cycle position — where are we in the cycle? Default BTC: point-in-time 9-indicator aggregation (Pi-Cycle Top & Bottom, Mayer Multiple, weekly RSI, 200-week-MA distance, halving position, Fear & Greed, BTC-dominance trend, mining-difficulty trend — weights in indicator_scores; components without input are excluded and weights renormalized, see indicator_coverage). Includes an ath block (E32): ATH on UTC daily-close basis with ath_date, days_since_ath and drawdown_from_ath_pct vs BOTH the scoring price and the live spot. Pass asset=ETH or asset=SOL for a per-coin cycle read built from the transferable price-derived indicators (Mayer, weekly-RSI, 200-week-MA distance) with renormalized weights; BTC-native indicators (halving, dominance, mining, F&G, Pi-Cycle) are returned as not_applicable rather than faked. All return raw + Z-Score, signal enum, and a percentiles block ranking each indicator against that asset’s own history. The signal enum is a FIXED SCORE-BAND LABEL (<25 accumulation · 25–45 recovery · 45–60 expansion · 60–75 distribution · ≥75 overheated), not an independent market-phase detection: the 45–60 band is the neutral middle, so a mid-band score reads "expansion" even in a drawdown market — the label describes the score band, not the market. BTC additionally returns highlights[] (rule-based markers for currently unusual indicator values — descriptive, versioned ruleset; empty array = nothing unusual) and price_context (price at scoring time vs live spot with drift % — the scores rest on the scoring-time price). Point-in-time scored — not reconstructable from a generic price API. The volatility series itself is arena_get_volatility_history; this tool carries the regime context around it. score_fields_note explains the four score fields: z_score/z_adj_score are the composite standardized against its own history and mapped back onto the 0-100 scale, NOT statistical z-values; halving_context.ath_days_after_halving puts the observed cycle high next to days_since_halving. Related: arena_get_historical_analog (what followed states like this one), arena_get_bullmarket_ampel, arena_get_pulse. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
assetNoWhich asset’s cycle. Default BTC. ETH/SOL return a price-derived cycle read with not_applicable fields for BTC-native indicators.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it excels. It discloses the aggregation logic (9 indicators, weights, renormalization when components are missing), the ATH block details, the per-coin behavior for ETH/SOL (returning not_applicable for BTC-native indicators), the signal enum being a fixed score-band label rather than market-phase detection, the highlights and price_context blocks, and the fact that scores are point-in-time and not reconstructable. It even explains the z-score fields and the halving_context field. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but every sentence carries meaningful information. It is structured with a clear lead-in (core purpose), then detailed blocks for ATH, per-coin behavior, signal enum caveat, highlights, price_context, and related tools. The most critical facts (default asset, indicator aggregation, signal meaning) are front-loaded. It is dense but not redundant; it earns its length, though it could be tightened slightly without losing value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of the tool and the lack of an output schema, the description is remarkably complete. It explains what is returned (raw + Z-Score, signal enum, percentiles, ATH block, highlights, price_context), the meaning of the signal enum, the scoring methodology, and the relationship to other tools. It also covers edge cases (excluded indicators, not_applicable fields) and the non-reconstructability of scores. There is no missing information that would prevent correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers the single parameter (asset) with an enum and a brief description. The tool description adds substantial meaning: it explains the default value (BTC), the different behavior for ETH/SOL (price-derived indicators vs. not_applicable for BTC-native ones), and the implications of choosing each asset. This goes well beyond the schema's 'Which asset’s cycle' and provides context that helps the agent select the correct parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states what the tool does: 'Crypto cycle position — where are we in the cycle?' and specifies the default asset (BTC), the 9-indicator aggregation, and the optional per-coin read for ETH/SOL. It distinguishes from siblings like arena_get_cycle_history and arena_get_historical_analog by focusing on the current cycle snapshot rather than history or analog states. The verb is specific ('get'), the resource is unambiguous, and the scope (point-in-time aggregation) is explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names related tools and their purpose: 'Related: arena_get_historical_analog (what followed states like this one), arena_get_bullmarket_ampel, arena_get_pulse.' It also clarifies a key distinction: 'The volatility series itself is arena_get_volatility_history; this tool carries the regime context around it.' This tells the agent when to use this tool vs. alternatives, and what it does not cover. The guidance is direct and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_edge_reportsGet Edge Library — Filter Effect ReportsAInspect

Which entry filter carries a real edge? Platform-wide aggregated analysis: how each Pro+ entry filter (200 WMA, ATR low/high/expansion, Altcoin Season, Bullmarket confirm/strict) affects strategy CAGR — baseline vs. filtered, asset-equal-weighted (per-asset medians over param-deduplicated runs, then the median across assets — no single asset's run grid can dominate an arm). delta_cagr is the median of PER-ASSET deltas over MATCHED assets only (present in both arms) — so it usually differs from filtered_cagr − baseline_cagr; pairs_matched/pairs_filtered and the baseline pairs count declare the basis. Verdicts come from the effect's 90% paired-bootstrap interval (delta_ci_low/delta_ci_high), not the point estimate: helps (whole interval > +1pp) / hurts (< −1pp) / neutral (inside ±1pp) / insufficient_evidence (runs disagree) / insufficient_data (fewer than 30 runs per arm or fewer than 10 matched assets). Below the gate, derived fields (delta_*, dsr, dsr_pass) are null; every gated null carries its reason (dsr_pass_reason, *_net_reason); the envelope evidence block declares the gate's referent and threshold machine-readably. Response is GROUPED by strategy: envelope fields (market, computed_at, n_trials) once, per strategy one baseline block {cagr, net_cagr, sharpe} plus filter cells; filter cells with zero runs are folded into filters_without_data. A full market is a few hundred cells — use limit/offset (strategies per page) plus the truncated flag for partial reads. Filters evaluated in isolation (no stacking); net values are median CAGR after per-side trading costs (verdict/delta stay gross). [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoStrategies per page (1–100). Omit for all.
marketYesMarket to analyze (crypto or tokenized).
offsetNoStrategies to skip (paging).
verdictNoFilter by verdict. Default 'all'. Note 'insufficient_evidence' is NOT the same as 'insufficient_data': the former has enough runs but they disagree (the effect's 90% interval straddles the ±1pp line), the latter simply lacks runs.
strategyNoRestrict to a single strategy key (e.g. golden_cross). Omit for all strategies.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does this exceptionally well: explains the non-trivial delta_cagr computation, the bootstrap-based verdict logic, the gating and null reasons, the response structure, pagination behavior, and even that filters are evaluated in isolation with net vs gross distinction. Nothing is hidden; the tool's quirks (e.g., delta differs from filtered−baseline) are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries substantive information needed for correct invocation given the tool's complexity and lack of output schema. It front-loads the core question and then cascades into details logically. While it could be restructured with sections, it avoids fluff and repetition. The length is justified by the information density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, no output schema, and complex gating/grouping behavior, the description is thoroughly complete. It explains the response envelope, per-strategy blocks, folding of empty cells, pagination with truncation, and the free-tier note. An agent has everything needed to call the tool correctly and interpret results without additional schemas.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes further by explaining the verdict parameter's nuanced meanings (already partially in schema but reinforced), the paging semantics (strategies per page), and how the strategy param restricts to a single key. It adds context about the response grouping that clarifies parameter effects, though some concrete param guidance (e.g., exact values for market) is already in schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear question and then states the tool's purpose precisely: 'Platform-wide aggregated analysis: how each Pro+ entry filter affects strategy CAGR'. It mentions specific filter names and the exact metric (delta_cagr) and distinguishes its aggregated, platform-wide scope from the many sibling get_* tools that are per-strategy or narrower. The verb 'get edge reports' is clear and the description leaves no ambiguity about what resource is returned.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives detailed context about the tool's aggregation semantics, gating rules, and response grouping, but it never explicitly says when to choose this tool over alternatives like arena_get_strategy_filter_effect or arena_get_filter_insights. It implies it is the platform-wide edge analysis, but does not name alternatives or exclusion criteria. An agent would need to infer the differentiation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_etf_flowsGet Spot-ETF Net-Flow Trend (BTC / ETH / SOL)AInspect

Spot-ETF net flows (USD millions) — is the flow impulse turning or accelerating? The summary only gives point-in-time deltas; this exposes the trend: 30d/90d net flow, a direction label (inflows/outflows/flat) and a daily series (every US trading day: cumulative inflow + that day's net flow; resolution states points and spacing) so direction and speed are visible, not just a single delta. Read impulse for what the flow is doing — it has four states (accelerating / decelerating / reversal / flat) and is the field to quote. Two neighbouring fields measure different things and are easy to confuse: acceleration_usd_m is the signed difference last-30d minus prior-30d and gets LARGE precisely when the flow reverses, while the older boolean accelerating requires the same direction AND a bigger magnitude — so a swing from outflows to inflows shows a big positive acceleration_usd_m together with accelerating: false, which is correct and reads like a contradiction. impulse reports that case as 'reversal'. When impulse is 'reversal', reversal_recovered_pct says how much of the preceding counter-move has actually come back, with its denominator in reversal_basis_usd_m — quote it alongside, because a reversal in direction is not yet a reversal in the stock. Both are null otherwise. source_inconsistencies lists the days on which the source's own daily net flow and the change of its cumulative disagree (BTC: 2 of ~700 days, 2026-04-23/24 — the two-day sum matches, the day split does not); cumulative deltas and the 30d/90d windows are unaffected, per-day sums over those dates are. NYSE closing days that the source still lists come back with net_flow_usd_m null and market_closed: true (a closed market is not a zero flow). availability states when a day's flow was first known (measured from our own ingest; the source has no publication time) — use it for point-in-time work. For long windows pass format: 'columns' (parallel arrays, far smaller). Default BTC; pass asset=ETH or asset=SOL. Source SoSoValue. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoLength of the returned daily series in days (every US trading day in the window). Default 365, clamped 7–1095.
assetNoWhich spot-ETF flows. Default BTC.
formatNoDefault 'rows' (series[] of objects). 'columns' returns series_columns instead — parallel arrays (date, cum_net_inflow_usd_m, net_flow_usd_m, market_closed) with each field name once; use it for long windows, it is much smaller.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly: it clarifies the apparent contradiction between acceleration_usd_m and the boolean accelerating, defines impulse states, explains that reversal_recovered_pct/reversal_basis_usd_m are null otherwise, documents source_inconsistencies dates, market_closed:true on NYSE holidays (closed market != zero flow), and availability timing. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and result shape, and nearly every sentence carries unique field semantics. It is dense and formatted as one long block, which slightly hurts scannability, but little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain returns — and it does comprehensively: series contents, resolution, direction labels, null semantics, and provenance caveats. An agent has everything needed to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning: it ties format:'columns' to long-window efficiency, implies the 30d/90d windows that days spans, and notes the default asset. It goes beyond restating the schema even if the schema already covers most fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource: it exposes the spot-ETF net-flow TREND (30d/90d net flow, direction label, daily series) rather than point-in-time deltas, and explicitly contrasts itself with 'the summary' that only gives deltas. An agent can distinguish it from the many arena siblings without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete selection guidance: pass format:'columns' for long windows, default asset is BTC with ETH/SOL alternatives, and use availability for point-in-time work. It stops short of naming a competing sibling or stating when NOT to use it, so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_fear_greedGet Fear & Greed IndexAInspect

How fearful or greedy is the market right now? Crypto Fear & Greed Index (alternative.me). Returns the current value (0-100) and classification (extreme fear / fear / neutral / greed / extreme greed) as their own fields, plus history — the last 90 daily readings by default, so you can see whether today is a move or a plateau. The window is capped in SIZE but free in POSITION: end_date moves it anywhere in the history since 2018 (e.g. end_date=2025-10-06 reads the sentiment around the October 2025 top), and the range block states requested / granted / available days with the reason — a short series here is a window, not a young index. On Pro and Elite two Arena-derived blocks add what the upstream index does not publish: cadence (how far smoothed sentiment has travelled versus ~90 days ago) and tempo (how FAST the index is moving — 7d and 30d change ranked as a rolling percentile against three years of same-direction moves, not a fixed threshold; rank compares with its own history, not with "normal"). On Free both blocks are present but their values are null with a stated reason. For the regime around a reading use arena_get_cycle; for what followed comparable sentiment states use arena_get_historical_analog(preset="deep_fear"). [Free tier · cadence/tempo Pro+]

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoHow many daily readings to return (1-365, default 90). The full history since 2018 is deliberately not offered in one response — it is ~3,100 points and does not fit a tool response. The cap limits window SIZE, not position: combine with end_date to read any window since 2018.
end_dateNoLast day of the window (YYYY-MM-DD, inclusive). Positions the window anywhere in the history since 2018-02 — e.g. end_date=2025-10-06 answers "what was sentiment at the October 2025 top". Omit for a window ending today. value/classification/as_of describe the LAST day of the window; cadence/tempo (Pro+) compute on the history up to end_date only, never on later data.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and meets it: it discloses the default 90-day window, the size/position semantics of days vs end_date, the range block's requested/granted/available behavior, tier-based nulls for cadence/tempo on Free, and how cadence/tempo are computed. It also warns against misreading a short series as a young index.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core answer and then walks through output, window semantics, tier behavior, and siblings in a logical order. A few rhetorical flourishes (e.g., 'a move or a plateau') are not strictly necessary, but the density is high and every substantive claim supports correct use.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description explains the shape of the response (value, classification, history, range, cadence, tempo) and the tier differences on top of the schema's parameter detail. It also covers edge semantics such as end_date positioning and Free-tier nulls, making it complete enough to invoke correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents days and end_date exceptionally well, including the 2018-02 position, the last-day semantics, and the cap-not-position rule. The description restates the same mental model but adds little literally new about parameter syntax or meaning, so the high-coverage baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description names the specific resource (Crypto Fear & Greed Index from alternative.me), the exact output fields (value, classification, history, cadence, tempo), and how it differs from adjacent tools by naming arena_get_cycle and arena_get_historical_analog. The 'get' verb is backed by concrete return content, so an agent knows immediately what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description not only states when to use this tool ('to see whether today is a move or a plateau') but explicitly routes to alternatives: 'For the regime around a reading use arena_get_cycle; for what followed comparable sentiment states use arena_get_historical_analog(preset="deep_fear")'. This is explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_funding_rateGet Funding Rate SnapshotAInspect

Are longs or shorts paying right now? Latest BTC perpetual funding rate, averaged across up to 4 exchanges (Binance, Bybit, OKX, Deribit; 8h settlement cadence). Returns value, 30d moving average and Z-Score. Positive = longs pay shorts (bullish bias), negative = shorts pay longs (bearish bias). Read coverage before comparing values across dates: it says how many exchanges stand behind that day (4 = full average, 1 = a single exchange), and a day-over-day move can be a change in composition rather than in the market; venues_present/venues_missing name the exchanges. A value of exactly 0.0001 (0.01 % per 8h) on many days is the exchanges' base-rate clamp on USDT perpetuals, not a cap in our pipeline: it means "no premium beyond the base rate", and values above it are real market readings. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and does so well: it explains the sign convention, the exchange-averaging basis, the composition caveat via coverage/venues_present/venues_missing, and the 0.0001 base-rate clamp that could otherwise be misread as a capped value. It also flags the free-tier status.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the question it answers, then progressively deeper interpretive detail. Dense but nearly every sentence adds decision-relevant meaning; the base-rate-clamp passage is long yet guards against a real misreading, so only mild trimming is warranted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description names the returned components (value, 30d moving average, Z-Score) and the diagnostic fields, covering everything needed to both call and interpret the result. Complete for a zero-parameter snapshot tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline is 4 and there is nothing for the description to compensate for. The field-level guidance it provides (coverage, venues_present/venues_missing) concerns return values rather than inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource and scope: the latest BTC perpetual funding rate averaged across up to four named exchanges on an 8h cadence, plus 30d MA and Z-Score. An agent can distinguish it from siblings like arena_get_spot_price or arena_get_iv_snapshot without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Opens with the decision it answers ('Are longs or shorts paying right now?') and gives a clear interpretive caveat: read `coverage` before comparing values across dates because composition shifts can masquerade as market moves. It never names a sibling alternative or an explicit when-not, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_gem_scoresGet Altcoin Screener RankingsAInspect

Altcoin screener ranking — which altcoins look strong right now? Today's CoinGecko Top-200 minus stablecoins and tokenized fiat, scored by a composite of 3 factor groups: Mean-Reversion (A), Tokenomics (B), Market-Structure (C). Each score carries plain (one sentence: rank with its base scored_total, composite, factor groups) and the response carries scored_total. Backtest-validated factors, not a hype list. Limit gated by tier: Free top-10, Pro top-50, Power up to 200 (the full scored set). [Free tier, daily refresh]

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoNumber of coins to return (tier-capped)
from_rankNoStart from this rank (default 1)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses the exact universe, the composite scoring model, the response shape including the `plain` summary and `scored_total`, tier-based limit gating, and daily refresh behavior. This is strong transparency for a read-only screener.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense but every sentence earns its place: scope, scoring model, output format, validation credibility, tier limits, and refresh cadence. It is front-loaded with the most important question and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description is remarkably complete for calling the tool correctly. It explains what the tool returns, how scores are composed, which top coins are included, and how the `limit` parameter interacts with the user's tier. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful parameter semantics by explaining that `limit` is tier-gated (Free top-10, Pro top-50, Power up to 200), which is not in the schema. It does not add much about `from_rank`, but the schema already documents it clearly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it is an altcoin screener ranking of today's CoinGecko Top-200, scored by a composite of three named factor groups. The title and phrasing clearly separate it from sibling tools like arena_get_gem_score, which implies single-coin scoring, and arena_get_winners.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use it — 'which altcoins look strong right now?' — and sets expectations with 'Backtest-validated factors, not a hype list.' It does not explicitly name alternatives or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_historical_analogHistorical analog — conditional forward returnsAInspect

What happened historically after the Bitcoin cycle looked like this? Conditional forward-return distribution for a named preset cycle state — over N DISTINCT historical episodes matching that state (matched_episodes), returns median/IQR/positive-share forward returns (30/90/180/365d) with per-horizon n, small-n warnings, point-in-time integrity and an evidence block that names which field its sample-size gate checked (gate_applies_to), against which threshold, over which data window. A distribution with its sample size. Not obtainable from web search or public market-data APIs — requires point-in-time indicator history and look-ahead-free episode matching. Presets: cycle_bottom_cluster (Cycle bottom cluster), cycle_top_cluster (Cycle top cluster), deep_fear (Deep fear), euphoria (Euphoria), quiet_volatility (Quiet volatility regime). The response opens with "preset_definition" (machine-readable condition set) plus current_state_matches (does the state hold TODAY?) and last_matching_date. Some presets carry a "study_finding" field — a state already investigated, with a NULL result where that is what the study found. EVERY preset returns "vs_unconditional_drift": the raw forward median contains the asset's contemporaneous drift; the drift and excess columns separate the two, and the excess can be negative while the raw median is positive. For quiet_volatility, vol_rank_threshold (fixed steps 5/10/20/50) asks the stricter "UNUSUALLY quiet" question the null study left open, and condition_on_direction conditions episodes on the sign of the first post-anchor move over direction_window_days (default 5) — both mark study_finding_applies=false, and horizons within direction_window_days are suppressed as circular. Also works for asset=ETH/SOL (F2 cycle history), but only price-derived presets (cycle_bottom_cluster, cycle_top_cluster) — fear-greed and volatility presets are BTC-only. Related: arena_get_volatility_history (the series behind the volatility preset), arena_get_cycle (the current state to compare against), arena_dip_scenario (composes this base rate into a tranche structure). [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
assetNoWhich asset’s cycle history. Default BTC. ETH/SOL support only price-derived presets (cycle_bottom_cluster, cycle_top_cluster).
presetYesNamed ex-ante cycle-state condition set. One of: cycle_bottom_cluster, cycle_top_cluster, deep_fear, euphoria, quiet_volatility.
forward_horizonsNoForward-return horizons in days. Default [30, 90, 180, 365] — except for quiet_volatility, which defaults to the horizons its study actually tested ([30, 90, 180]); anything beyond that is flagged as outside the protocol.
vol_rank_thresholdNoquiet_volatility only. Reference threshold as a FIXED step: 50 (default, below trailing median — the studied definition) or 5/10/20 (unusually quiet: RV30 below its trailing Nth percentile). Any value other than 50 sets study_finding_applies=false — the null study covered only the default.
direction_window_daysNoClassification window for condition_on_direction (default 5). Only meaningful together with condition_on_direction.
condition_on_directionNoquiet_volatility only. Condition episodes on the direction of the FIRST post-anchor move (sign of the direction_window_days-day return). Horizons <= direction_window_days are suppressed as circular. Sets study_finding_applies=false.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full disclosure responsibility, and it delivers. It transparently discloses point-in-time integrity, look-ahead-free episode matching, small-n warnings, the 'study_finding' field with NULL results, the vs_unconditional_drift separation and its interpretation (excess can be negative while raw median is positive), circularity suppression for horizons within direction_window_days, and the fact that study_finding_applies=false when non-default parameters are used. It also clarifies that quiet_volatility's vol_rank_threshold has fixed steps and different semantics. No contradictions with any structured metadata exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but exceptionally information-dense and front-loaded. It opens with the core purpose, then systematically covers output fields, parameter nuances, asset limitations, and related tools. While it could be broken into bullet points for scannability, every sentence carries unique information and none is filler. The length is justified by the tool's complexity (6 parameters, 5 presets, multiple conditional behaviors). It loses a point only for being a single dense paragraph that might overwhelm an agent scanning quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no output schema and no annotations, the description must compensate for both. It does so comprehensively: it explains the return structure (median/IQR, positive-share, per-horizon n, small-n warnings, evidence block with gate_applies_to), the preset definitions and their semantics, asset support restrictions, the special quiet_volatility parameters and their implications, and the relationship to unconditional drift. It even notes which presets carry study_finding fields. For an agent to call this tool correctly and interpret results, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though the input schema covers 100% of parameters with descriptions, the tool description adds substantial meaning beyond the schema. It explains the conceptual meaning of presets (e.g., 'cycle_bottom_cluster' as a named condition set), the special default behavior for forward_horizons under quiet_volatility, the semantic difference between vol_rank_threshold values (50 vs 5/10/20) and the 'unusually quiet' interpretation, and the directional conditioning semantics. The description enriches the schema with protocol-level nuance (e.g., horizons beyond study are flagged as outside protocol) that is not present in the schema fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete question ('What happened historically after the Bitcoin cycle looked like this?') and defines the tool as a 'Conditional forward-return distribution for a named preset cycle state.' It specifies the verb (returns a distribution), the resource (historical analog episodes), and the distinguishing feature (point-in-time, look-ahead-free matching). It also names three sibling tools and when to use them, clearly differentiating itself from the broader family of arena_get_* dataset tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states explicit usage context: it is not obtainable from web search or public APIs and requires point-in-time history and look-ahead-free episode matching. It names related tools (arena_get_volatility_history, arena_get_cycle, arena_dip_scenario) and states when each is appropriate. It also spells out asset restrictions (ETH/SOL only for price-derived presets) and parameter-specific usage (vol_rank_threshold, condition_on_direction only for quiet_volatility). No ambiguity remains about when to invoke this tool versus alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_indicator_snapshotGet Indicator Snapshot with Historical Percentile RanksAInspect

What do the classic indicators read right now? Current RSI(14), MACD(12/26/9), Bollinger(20,2), ATR(14) and OBV for a pair — each with a PERCENTILE RANK against that indicator's own history on that pair, plus the observation count — the rank turns a raw reading into a placement. ATR comes as a percentage of price so it is comparable across time, and OBV as a 30-bar slope normalised by that window's volume (raw cumulative OBV would mostly rank how long the series has existed). Where the reading sits in an extreme AND a study on this platform has tested that exact state, the payload carries the study verdict — including a null result: a Bollinger squeeze returns the quiet_volatility finding that tight bands did NOT carry an edge. Below 500 bars (1d) / 150 (1w) the raw values still come but percentile is null with a reason, rather than a rounded number from too small a sample. Set interval to '1w' for the weekly view. On the 1d view the payload also carries rsi_14_weekly (weekly RSI with its own rank) — for BTCUSDT this is the SAME series as arena_get_cycle rsi_weekly, measured character-identical (its source_note carries the measurement). It also carries rsi_14_4w (RSI-14 on 28-day candles, with bars — few, so crossing counts stay small; its source_note states what followed crossings). state (oversold/neutral/overbought) names where a reading sits on its own scale. Related: arena_get_trend_channels (structure), arena_get_historical_analog (did a condition like this one ever pay?), arena_get_volatility_history (the volatility series behind ATR). [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairNoPair, e.g. "BTCUSDT" (default), "ETHUSDT", "PAXGUSDT".
intervalNoDefault '1d'. '1w' computes every indicator on weekly bars.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: percentile rank methodology, ATR normalization, OBV slope normalization, null percentile thresholds, additional weekly/4w fields, state naming, and the null-result verdict behavior. This is far beyond a generic read snapshot.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but information-dense; every sentence adds a distinct behavioral or interpretive fact. Front-loaded with the core question and indicator list, then caveats, then related tools. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param read tool with no output schema, this is unusually complete: it covers what is returned, how metrics are computed, when percentile is null, what extra fields appear, and related tools. An agent can call it correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema already documents pair and interval 100%. Description adds meaning by explaining interval='1w' changes all indicators to weekly bars, and gives example pairs. It also explains the threshold behavior tied to interval, so it adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete question and enumerates the exact indicators (RSI(14), MACD(12/26/9), Bollinger(20,2), ATR(14), OBV) plus percentile ranks, making the resource and scope unmistakable. It also distinguishes itself from sibling tools by naming related tools and their different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly lists related tools with parentheticals explaining what each is for, and tells the user to set interval='1w' for the weekly view. It doesn't provide strict when-not-to-use rules, but the context is clear enough to route an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_iv_snapshotGet Deribit IV SnapshotAInspect

What is the options market pricing in? Latest Deribit volatility snapshot for BTC or ETH. Returns DVOL (30d vol index), constant-maturity ATM implied vol (30/60/90/180d via options chain), 30d realized vol, and vol_risk_premium_30d, which is the TRAILING spread: ATM implied vol (30d, from the options chain — not DVOL) minus the realised volatility of the PAST 30 days. It answers "are options priced expensively right now?". Set include_implied=true to additionally get the FORWARD premium in an implied block: DVOL(t) minus the realised volatility of the FOLLOWING 30 days, which answers the different question "did the expectation actually materialise?". These two are NOT interchangeable — measured 2026-08 they carried OPPOSITE signs on 17.3% (BTC) / 30.5% (ETH) of paired days. The forward field is spelled out as vol_risk_premium_forward_30d so the two cannot be confused. The most recent 30 days carry premium_complete=false and no premium value at all, because their forward window has not closed yet; they are excluded from every aggregate. Source: Deribit DVOL Index. History: BTC from 2021-04-01, ETH from 2022-02-15. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
currencyYesCurrency to fetch IV snapshot for
include_impliedNoDefault false (response unchanged). When true, adds an `implied` block with the FORWARD volatility risk premium, its percentile and the historical base rate.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully bears the burden of behavioral disclosure. It reveals the exact fields returned, the semantics of both trailing and forward risk premiums (including their non-interchangeability and sign opposition), what include_implied adds, the edge case of recent 30 days with premium_complete=false, and source/history dates. This is comprehensive and goes far beyond a bare 'get snapshot' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but each sentence contributes: purpose, returned fields, flag behavior, edge cases, source, and history. Use of bolding, backticks, and code-style names aids scannability. It is longer than typical but not bloated; it earns its length. A minor deduction for not front-loading the most common usage (snapshot) vs. the detailed metric explanations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description covers all essential returning fields and their meanings, the optional block, data availability caveats, and historical start dates. It tells the agent exactly what to expect in the response (e.g., implicit JSON structure with an 'implied' block). Nothing critical is missing for correct invocation or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (both params described), so baseline is 3. The description adds substantial meaning: it explains that include_implied appends a forward premium block, clarifies the difference between trailing and forward premiums, details the implied block fields, and notes the exclusion of incomplete recent data. For currency, it reinforces the BTC/ETH scope. This far exceeds merely restating schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns the latest Deribit volatility snapshot for BTC or ETH, enumerating specific metrics (DVOL, implied vols, realized vol, premiums). It distinguishes itself from siblings like arena_get_volatility_history by emphasizing 'latest snapshot' and answering 'are options priced expensively right now?'. The verb+resource is explicit and differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides a clear use case ('are options priced expensively right now?') and explains when to set include_implied (to answer a different forward-looking question). It also warns about the non-interchangeability of trailing and forward premiums and the incomplete recent 30 days. However, it does not explicitly name alternative tools or state when to prefer history/insights siblings over this snapshot tool, leaving some ambiguity for an agent choosing among many volatility tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_job_statusGet Async Job StatusAInspect

Is my universe backtest finished? Polls an async job by job_id (created via arena_run_universe_backtest). Returns status (pending/running/completed/failed), progress_pct, pairs_completed, and once completed: the full result (summary + per-pair results). [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesUUID job_id returned by arena_run_universe_backtest.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it delivers: it discloses the status lifecycle (pending/running/completed/failed), the intermediate progress fields (progress_pct, pairs_completed), and the transition to full results upon completion. It does not address polling cadence or rate limits for a repeated-poll operation, which would strengthen it further.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the question, states the verb, names the provenance, and enumerates return fields with no filler. The opening conversational fragment is slightly unusual for a formal definition but works well for orientation; otherwise it is tightly packed with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter polling tool with no output schema, the description covers purpose, usage provenance, the parameter, and the return shape across lifecycle states. Minor gaps are the absence of error/status semantics (e.g., handling of failed jobs) and no recommended polling interval, but the core is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already describes job_id as the UUID returned by arena_run_universe_backtest. The description's reference to job_id largely restates what the schema provides, adding minimal beyond-schema meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Polls') and resource (async job by job_id) and orients the reader with the opening question 'Is my universe backtest finished?'. It names the provenance tool (arena_run_universe_backtest), which clearly distinguishes this polling endpoint from sibling sync-retrieval tools like arena_get_backtest and arena_get_report_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the usage context clearly: it polls an async job created via arena_run_universe_backtest, and it notes the [Free tier] cost signal. It does not explicitly name alternatives or state when not to use it, but the 'created via' provenance effectively separates it from synchronous backtest retrieval siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_key_levelsGet BTC Key Levels (S/R clusters + indicator levels)AInspect

Which price levels matter above and below spot? Reproducible Bitcoin structural levels on BOTH sides of spot, in TWO distinct provenance classes. (1) resistance/support: swing-pivot clusters — where past pivot highs+lows cluster into price zones (touch-count, band, last-touch date, signed distance), resistance above spot, support below, nearest-first. (2) indicator_levels.above / .below: named indicator STANDS as marks — 200-day & 200-week simple moving averages, short-term-holder cost basis, Pi-Cycle legs — each carrying its source, formula and as_of date. The two classes are kept separate on purpose: pivots are where price REACTED before, indicator levels are where an indicator STANDS now. Both are measured price clusters: they say where trading has concentrated, not where anyone defends a level. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral burden. It adds meaningful context: levels are 'reproducible', 'measured price clusters' not 'defended levels', and notes a free tier. It discloses that the two classes are kept separate on purpose. This goes beyond a generic 'get levels' description, though it does not mention rate limits or response format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat long but well-structured with a front-loaded question and clear numbered classes. Every sentence adds distinct information (provenance, purpose, free tier). No redundancy, but slightly verbose for a zero-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description is complete. It explains what levels are included, how they are derived, the separation of classes, and the free tier limitation. Nothing essential is missing for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the baseline is 4. The description correctly focuses on the output content rather than parameter details, which is appropriate. No further parameter clarification is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('BTC Key Levels') and immediately clarifies it returns structural levels in two provenance classes (pivot clusters and indicator levels). It clearly distinguishes this from the many sibling tools by enumerating what is included, without needing to compare to others.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: it answers 'Which price levels matter above and below spot?' and explains the two classes. However, it does not explicitly name alternative tools or state when not to use this tool. Given the large sibling list, more explicit routing (e.g., 'for indicator-only values use arena_get_indicator_snapshot') would improve clarity, but the context is understandable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_knowledgeGet Knowledge ObjectAInspect

What does the platform know about this subject? Fetch a versioned, explainable Knowledge Object by type + subject (e.g. type='market_regime', subject='GLOBAL'). Returns the current published envelope: payload, explanation (factors + weights + confidence), provenance (inputs + params), ontology binding, compute version. ONE tool covers ALL knowledge types. Set include_graph=true to also walk the knowledge graph: resolved outbound edges (what this object is derived_from / references) + inbound edges (what derives from / references it), each with api_path + seo_slug so you can follow them. [Free tier; per-object access additionally gated by min_tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYesKnowledge object type, e.g. 'market_regime'.
as_ofNoSpecific date YYYY-MM-DD. Omit for latest.
subjectYesSubject ref, e.g. 'GLOBAL', 'BTC'.
include_graphNoIf true, attach the resolved edge neighbourhood (outbound + inbound) for graph traversal.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses the exact return envelope (payload, explanation, provenance, ontology binding, compute version), notes the objects are versioned and published, and explains the include_graph extension with edge details (api_path + seo_slug). It also mentions per-object access gating by min_tier. These are substantive behavioral disclosures beyond a simple 'fetch' description, though it omits error or rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-organized: starts with a framing question, then purpose, response envelope, scope ('ONE tool'), graph extension, and access note. The rhetorical opener adds flavor but every following sentence delivers concrete information. It is slightly long but efficiently structured with key details front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Comprehensive for a tool with no output schema. It specifies the full return envelope fields, the graph traversal behavior with edge attributes, and the access gating condition. The scope clarification and examples ensure an agent can call it correctly without needing an output schema. All essential information for successful invocation is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds practical value by giving example values for type and subject, clarifying that as_of defaults to the latest when omitted, and explaining the effect of include_graph on graph traversal. These clarifications go beyond the schema's brief property descriptions, making the parameters more actionable for an agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Fetch' and resource 'Knowledge Object' by type and subject, with concrete examples (type='market_regime', subject='GLOBAL'). The phrase 'ONE tool covers ALL knowledge types' clearly differentiates it from the many specific arena_get_* siblings, establishing that this is the general knowledge retrieval tool. The description is unambiguous and immediately scopes the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context by stating it covers ALL knowledge types, which implies it is the general-purpose alternative to the numerous type-specific getters (e.g., arena_get_macro_regime, arena_get_btc_market_structure). It gives practical examples and explains the optional include_graph for graph traversal. However, it does not explicitly enumerate when-not-to-use cases or name specific alternatives, so the guidance is strong but not fully exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_macro_regimeGet Macro Regime SnapshotAInspect

What is the macro backdrop doing? Daily Macro Regime snapshot from 18 components in 6 tiers (Liquidity 30%, Financial Conditions 20%, Risk Appetite 15%, Crypto Liquidity 10%, Business Cycle 15%, Inflation/Real Rates 10%). FRED-sourced. Returns composite_score (0-100), regime_label (risk_off/neutral/risk_on_leaning/risk_on), cycle_phase_label (contraction/early_expansion/mid_expansion/late_expansion), matrix_quadrant (sweet_spot/late_cycle_warning/crisis/recovery), tier_scores (6 sub-scores), components (flat key/value of all 18), plus stale_components_detail dating each stale input (last_good_date + age_days + discontinued flag for series the upstream has retired for good) so freshness is quantified, not a vague caveat. Two component keys mean something narrower than their name suggests, so read them carefully: vix_score is the derived 0-100 score (a value of 71 means VIX around 18.6), NOT the VIX index level — the raw Cboe level is not redistributed over this channel; and broad_dollar_index is FRED DTWEXBGS (Broad USD Index, Jan 2006 = 100), NOT the ICE DXY, so readings near 120 are normal. fed_funds_rate is FRED FEDFUNDS, the monthly AVERAGE effective rate (lags; not the daily DFF, not the target range); global_m2_yoy is NOT M2 but the YoY change of G3 central-bank balance sheets (Fed+ECB+BoJ in USD) — it can fall while US M2 hits a record. component_notes carries these definitions in the payload. The former names vix and dxy were removed on 2026-09-01 after their announced deprecation window; consumer_confidence_value went with them (OECD retired the series, frozen since 2024-01-01, never weighted in the composite). [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly. It discloses FRED sourcing, stale-component freshness quantification, discontinued flags, and detailed semantic caveats for component keys like vix_score, broad_dollar_index, fed_funds_rate, and global_m2_yoy. It even documents removed legacy names, giving agents strong behavioral grounding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but front-loaded with the core snapshot summary and output fields. The caveats are dense and mostly necessary given the absence of an output schema, though a few details (deprecation history, 'Free tier') could be trimmed without losing essential guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema and no annotations, the description must fully specify return values and interpretation. It does: composite_score, regime_label, cycle_phase_label, matrix_quadrant, tier_scores, components, stale_components_detail, and component_notes are all described. It also explains freshness quantification and key naming traps, making it complete for a zero-parameter snapshot tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so the baseline is 4. The description adds no parameter-level semantics because there are none to document; it instead uses the space to explain output semantics, which is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific question ('What is the macro backdrop doing?') and names a concrete resource: a Daily Macro Regime snapshot built from 18 components in 6 tiers. It clearly distinguishes this from sibling tools by focusing on the composite macro regime output rather than single indicators or strategies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use case is clear: retrieve the current macro backdrop as a daily snapshot with composite scores, regime labels, and tier breakdowns. It does not explicitly name alternative tools or state when not to use it, so it stops short of a 5, but the context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_max_painGet Deribit BTC Max Pain (latest + upcoming)AInspect

What happened at the last Deribit expiry? Max pain and how spot settled against it: max_pain_strike, spot_at_expiry, %-diff, put_call_ratio, notional. Plus up to 10 upcoming expiries, each with current live max-pain level, days_to_expiry, open_interest_contracts and open_notional_usd. Field semantics: days_to_expiry is floored at 0 and cannot separate "expires later today" from "already settled" — settles_at (full ISO timestamp) and hours_to_settlement (SIGNED; negative = settled but not yet finalized) carry that distinction. settlement_time_utc names the settlement time where evidenced against the exchange (08:00:00Z for DERIBIT_BTC); where not evidenced, all three timing fields are null. open_interest_contracts (upcoming: latest daily snapshot) and total_contracts (settled: last snapshot BEFORE expiry) are the SAME measurement at different observation times; contracts_as_of names the snapshot. total_notional_usd is computed against the SETTLEMENT spot and never changes; open_notional_usd uses the CURRENT spot and moves with spot (notional_spot/notional_spot_date name the reference). oi_available distinguishes "null" from "not collected". Expiry flags NEST rather than partition (quarterly ⊂ monthly ⊂ weekly ⊂ daily): filter on the booleans, read expiry_type as the label — only it separates a Friday expiry from a mid-week one. All flags are calendar-derived, so upcoming expiries carry them too. spot_at_expiry is the exchange settlement price: for DERIBIT_BTC the Deribit delivery price (30-min index TWAP before 08:00 UTC — rows before 2026-08-31 were recomputed from that series; they had carried the BTCUSDT daily close, 16 h later), for IBIT the ETF close of the expiry day. Pass market to switch venue (DERIBIT_BTC default, IBIT). include_strike_ladder=true adds, per expiry, open interest per 2.5 % price band around spot (±25 %, calls/puts, absolute contracts) with day-over-day delta — a stock, not a side: no hedge direction follows from it. include_gex=true (DERIBIT_BTC only) adds per expiry a gex block plus gex_totals across the book — Black-Scholes gamma notional per band from LIVE Deribit mark IV (gex_data_as_of names the fetch, a different observation time than the snapshot fields); the dealer sign is an ASSUMPTION, both conventions published side by side; zero_gamma_level flips only under the SqueezeMetrics convention (short-all has no zero crossing by construction, its null is structural — zero_gamma_level.note says so). With ladder or GEX, pass expiry_date to get ONE expiry instead of all upcoming ones — the full ten-expiry response with both is large. Cron collects daily 02:00 UTC from Deribit Public API. Related: arena_get_max_pain_history (base rates + daily snapshots of open expiries), arena_get_iv_snapshot (implied vol for the same expiries). [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
marketNoOptions market. Only DERIBIT_BTC is served: IBIT (BlackRock spot-ETF options) is still collected daily but no longer delivered — the chain comes from an unlicensed source, so it cannot be redistributed (2026-09-24).
expiry_dateNoOptional YYYY-MM-DD of ONE open expiry: upcoming[] (and its strike_ladder/gex) is restricted to it, which keeps ladder/GEX responses small; gex_totals still covers the whole book. An expiry that is not open returns invalid_input listing the open ones. Omit for all upcoming expiries.
include_gexNoDefault false (response unchanged). DERIBIT_BTC only. When true, each upcoming expiry carries a `gex` block plus `gex_totals` across the whole book: Black-Scholes gamma notional (USD per 1 % spot move) per 2.5 % band from LIVE Deribit mark IV per strike (gex_data_as_of names the fetch, ~10 min cache — a different observation time than the 02:00 UTC snapshot fields). The dealer SIGN is an assumption, not a measurement: both conventions are published side by side (assuming_dealers_short_all, assuming_squeezemetrics_convention); where they disagree, the data does not know the answer. zero_gamma_level flips only under the SqueezeMetrics convention — short-all is <= 0 everywhere and has no zero crossing by construction (its null is structural; zero_gamma_level.note says so). Tau floor 2 h near expiry (tau_clamped flags it); instruments without usable IV are excluded and counted.
include_strike_ladderNoDefault false (response unchanged). When true, every expiry carries a `strike_ladder`: open interest per 2.5 % price band around the snapshot spot (±25 %, calls/puts separate, absolute contracts, share_pct), below_range/above_range sums, max_pain_recomputed (cross-check against the stored level) and `delta` vs the previous day's snapshot on the same band grid (null with delta_reason when there is none). OI is a stock, not a side — no hedge direction follows; the note travels with the response.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discharges it unusually well: it discloses the collection cadence ('Cron collects daily 02:00 UTC from Deribit Public API'), the ~10 min GEX cache and that gex_data_as_of is a different observation time than the snapshot fields, the dealer-sign ASSUMPTION with both conventions published, the structural (not missing) null in short-all zero_gamma_level, and the historical recomputation of pre-2026-08-31 rows. These are exactly the provenance/assumption caveats an agent would otherwise misread.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The answer is front-loaded (the core question leads), which is good, but the passage runs several hundred words and re-explains the GEX and strike-ladder parameters in nearly the same detail the input schema already provides, so those sentences do not fully earn their place. Density is high but so is duplication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema and this is a rich options-analytics payload, yet the description defines the key fields (max_pain_strike, spot_at_expiry per venue, open_interest_contracts vs total_contracts, total_notional_usd vs open_notional_usd, oi_available, nested expiry flags, ladder and GEX blocks) and their observation-time distinctions. An agent has enough to interpret results correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters in near-identical terms, establishing a baseline of 3. The description adds limited marginal value (rationale that ladder/GEX responses get large, notes about nesting expiry flags), and the statement 'Pass `market` to switch venue (DERIBIT_BTC default, IBIT)' actually conflicts with the schema enum, which permits only DERIBIT_BTC and states IBIT is no longer delivered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific framing ('What happened at the last Deribit expiry? Max pain and how spot settled against it') and enumerates the concrete return fields (max_pain_strike, spot_at_expiry, put_call_ratio, notional) plus up to 10 upcoming expiries, so the resource and scope are unambiguous. It also names two siblings (arena_get_max_pain_history, arena_get_iv_snapshot) with their distinct roles, letting an agent route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes explicitly: history tool for 'base rates + daily snapshots', iv_snapshot for 'implied vol for the same expiries', and it tells the agent to pass expiry_date when combining ladder/GEX because 'the full ten-expiry response with both is large'. What's missing is any negative guidance on when NOT to call this (e.g. vs live spot/options tools), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_max_pain_historyGet Deribit BTC Max Pain HistoryAInspect

Does max pain actually pull price to the strike? Settled Deribit BTC options expiries with the max-pain level we compute per expiry, for measuring the convergence question: does spot drift toward the max-pain level as expiry approaches? Each row: expiry_date, max_pain_strike, spot_at_expiry, %-diff, P/C ratio, notional, expiry-type flags. The mandatory base_rates block answers the convergence question PER expiry class (n, median |diff|, shares within 1%/2%, max, sample_adequate at n>=30) — the pooled median mixes tiny daily expiries with large quarterlies, which is what the per-class split separates. Filter with expiry_type / min_contracts / snapshot_expiry_date instead of post-processing the full row set. With include_open_snapshots=true it adds the daily observation series of still-open expiries — that series starts 2026-05-28, is not backfillable, and its per-expiry depth is thin, so check open_snapshot_coverage before computing anything from it. Days auto-capped by tier: Pro 365d, Power 3650d. Max-pain levels are our own aggregation across the option chain; the chain itself is not redistributed. Source: Deribit. Related: arena_get_max_pain (current + upcoming), arena_get_iv_snapshot. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoDays back from today (default 90, capped by tier).
marketNoOptions market. Only DERIBIT_BTC is served: IBIT (BlackRock spot-ETF options) is still collected daily but no longer delivered — the chain comes from an unlicensed source, so it cannot be redistributed (2026-09-24).
expiry_typeNoFilter expiries AND open_snapshots to one expiry class (label = highest level reached; the nesting booleans stay untouched). base_rates are always computed BEFORE this filter.
min_contractsNoOnly finalized expiries with total_contracts >= this (rows with unknown contracts drop out when set).
snapshot_expiry_dateNoReduce open_snapshots[] to exactly this expiry date (YYYY-MM-DD). Only meaningful with include_open_snapshots=true.
include_open_snapshotsNoDefault false. When true, adds open_snapshots[] (daily observations of not-yet-expired contracts) plus open_snapshot_coverage. Omit for the unchanged response.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly: it discloses that max-pain levels are its own aggregation (chain not redistributed), that IBIT is no longer delivered due to licensing, that open snapshots start on a fixed date and are not backfillable, that days are tier-capped, and that base_rates are computed before expiry_type filtering. No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: purpose, data contents, base_rates rationale, filtering advice, open-snapshot caveats, tier caps, source, and related tools. It's front-loaded with the core question and logically organized, though slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with no output schema and no annotations, the description is fairly complete: it explains the row contents, the base_rates block, the open_snapshots series and its caveats, and the tier caps. It could explicitly list the response fields or pagination, but the main behavioral aspects are covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers all 6 parameters with detailed descriptions (e.g., market explains IBIT exclusion, expiry_type explains filtering). The description adds extra semantics: explains the base_rates block, that base_rates are computed before the filter, that snapshot_expiry_date is only meaningful with include_open_snapshots=true, and the tier caps on days. Adds value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific research question ('Does max pain actually pull price to the strike?') and states the exact resource: settled Deribit BTC options expiries with computed max-pain levels. It clearly differentiates from siblings by naming arena_get_max_pain (current + upcoming) and arena_get_iv_snapshot as related, and by describing its unique historical/convergence focus.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on how to filter (use expiry_type/min_contracts/snapshot_expiry_date instead of post-processing), explains why per-class base_rates are preferred over pooled medians, and warns to check open_snapshot_coverage before computing anything. Names related tools for current/upcoming data, though it doesn't explicitly state 'use this for X, not Y' with direct contrast.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_onchain_historyGet On-Chain Series Historical ValuesAInspect

How has this on-chain metric moved over time? Returns the full TIME SERIES of one on-chain metric from the Bitcoin Research Kit — date/value pairs in ascending order, with history back to 2009 for most series. Use it for trend and percentile work; for the single current reading call arena_get_onchain_latest, and to discover valid series_ids call arena_list_onchain_series. Values are as-reported: on-chain metrics can be revised retroactively, so this is not a point-in-time vintage. Range capped by tier — the response carries a range block (requested_days, granted_days, clamped, clamp_reason, tier), so a clamped window announces itself instead of silently looking like the full history. [Free 30d / Pro 365d / Power unlimited]

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoDays back from today (clamped by tier).
series_idYesBRK series id, e.g. 'mvrv'.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that values are as-reported and can be revised retroactively (not point-in-time), and it transparently explains tier-based clamping via the 'range' block. It could add more on response format or error behavior, but the key caveats are covered thoroughly. The lack of annotations is mitigated by this strong prose.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place. The description leads with intent, explains the data, routes to alternatives, and flags caveats. It is dense but not bloated, and it front-loads the core function. The structure is a model of clarity without superfluous wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a time-series tool with no output schema, the description covers the essential contextual needs: what is returned, how to use it, alternatives, data revision caveat, and tier clamping behavior. It omits explicit mention of pagination or exact number of data points, but the range block and history depth are explained. Minor gaps but overall complete enough for an agent to call confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are documented. The description adds value by explaining the 'days' parameter's interaction with tier limits (Free 30d / Pro 365d / Power unlimited) and by mentioning the 'range' block that reports clamping. It also gives an example series_id ('mvrv'). This goes beyond the schema's basic descriptions, enriching the semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the full time series of one on-chain metric as date/value pairs, with history back to 2009. It distinguishes itself from siblings by explicitly naming arena_get_onchain_latest for single readings and arena_list_onchain_series for discovering series IDs. The verb 'returns' and resource 'time series' are specific and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance: 'for the single current reading call arena_get_onchain_latest, and to discover valid series_ids call arena_list_onchain_series.' It also states the intended use case ('trend and percentile work'). This is a textbook example of alternating tool routing with clear conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_onchain_latestGet On-Chain Series Latest ValueAInspect

What does this on-chain metric read right now? Returns the most recent value of ONE on-chain series from the Bitcoin Research Kit as { series_id, metric_name, date, value }. Cheapest way to answer "what is X right now" (MVRV, SOPR, realized price, hash rate, …). Discover valid series_ids with arena_list_onchain_series; for the history behind the number use arena_get_onchain_history. A single reading has no context — pair it with the series percentile before calling any level high or low. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
series_idYesBRK series id, e.g. 'mvrv', 'sopr', 'realized_price'.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool returns a single reading, notes it has no context ('A single reading has no context') and advises pairing with the series percentile, and hints at cost with '[Free tier]' and 'Cheapest way'. While it doesn't explicitly say 'read-only' or cover error behavior, the description provides substantial operational context beyond the bare function.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and output format, then adds practical guidance (alternatives, context warning, free tier) in a compact set of sentences. Every sentence earns its place: the opening question captures intent, the return format is explicit, and the caveat prevents misuse. It is dense but not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (1 parameter, no output schema), and the description covers everything an agent needs: the return shape, how to discover valid inputs, the relationship to history, and a usage pitfall (single reading lacks context). Since there is no output schema, the explicit field list is essential and provided. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers the parameter with examples ('mvrv', 'sopr', 'realized_price'), so baseline is 3. The description adds value by explaining that series_ids come from a separate listing tool and by reiterating the output fields, which clarifies how to interpret the response. This goes beyond merely restating the schema and compensates slightly for the lack of an output schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Returns the most recent value of ONE on-chain series') with a concrete output format and examples ('MVRV, SOPR, realized price, hash rate'). It clearly distinguishes itself from arena_get_onchain_history by contrast, and names the discovery tool arena_list_onchain_series. The purpose is unambiguous and immediately separates it from dozens of siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly frames the tool as 'Cheapest way to answer "what is X right now"' and gives direct alternatives: 'for the history behind the number use arena_get_onchain_history' and 'Discover valid series_ids with arena_list_onchain_series'. This tells the agent exactly when to use this tool versus its siblings, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_pulseGet Arena Pulse TodayAInspect

How hot is the Bitcoin market today? Daily 0-100 heat score for the Bitcoin market, aggregated from 8 components (BTC-Cycle, F&G, Altcoin-Season, Bullmarket-Ampel, Funding-Rate, Hash-Ribbons, Mayer-Multiple, MVRV-Z). Returns score, band label, color, 7d/30d delta, verdict, components breakdown, plus score_percentile ranking today’s score against its own history (e.g. 42 = 44th percentile — how hot/cold vs history, not just the raw number). score_semantics says which value you hold: the daily snapshot frozen once a day by the cron, or — before that cron has run for today — a live preliminary that still moves and whose percentile/deltas compare against frozen snapshots. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Given no annotations are present, the description carries the full responsibility of explaining behavior, and it does: it discloses the frozen vs. live preliminary state, the cron-based daily freeze, and the percentile comparison base. It wisely doesn't include what happens on failure or the exact caching, but the most important behavioral nuance is there.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences but avoids wasted words: it leads with the core question, enumerates components and outputs, and then adds the crucial interpretation detail. It is slightly long, but for a helper with no output schema every sentence provides useful knowledge that would otherwise have to be discovered at runtime.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, the description is the sole guide to the return shape: it lists all fields and explains score_percentile with a concrete example. It also covers the cron/frozen context and the [Free tier] access notion. The only missing piece is a note about error behavior when no data exists, but otherwise an agent can call and interpret the response correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With zero parameters, schema coverage is trivially 100%, so there is nothing missing for the agent to clarify. The description mentions score_semantics, which is a value of the returned object rather than a parameter, and this does not confuse the param attributes. Baseline for a 0-param tool is 4, and no deduction is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it provides a daily 0-100 heat score for the Bitcoin market, lists the 8 component indicators, and names all the return fields (score, band, color, delta, verdict, components, percentile). This makes it easy to tell apart from the many sister get_* tools in the sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys the context that this is a current-day snapshot and introduces the cron/live semantics, but it never spells out when to prefer it over alternatives such as arena_get_pulse_history or other single-indicator getters. There is clear implied usage but no explicit when-not/exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_robustness_fieldRobustness Field — plateau vs. spike + Deflated Sharpe with a counted NAInspect

Is this backtest result real, or a lucky cell? Assess one backtest result against its neighborhood instead of trusting a single "+X% CAGR" cell. Given a (strategy, interval, pair) and YOUR result (user_cagr, optional user_sharpe), returns: the cross-asset distribution of the SAME strategy+interval across every pair the backtest factory ran it on (median, IQR, positive-share, your percentile), a plateau/spike/fragile/mixed verdict, and — where Sharpe coverage allows — a Deflated Sharpe threshold whose N is COUNTED (the number of neighbor assets IS the testing family), not guessed. Honest small-n handling: fewer than 15 neighbors → "insufficient", no DSR-N claimed. Set axis="parameter" for the secondary, always-anecdotal view (the few parameter settings tested on this exact pair). Read-only over result aggregates, look-ahead free. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
axisNoNeighborhood axis. 'cross_asset' (default, dense, carries the verdict + DSR-N) or 'parameter' (secondary, always anecdotal — the parameter settings tested on this one pair).cross_asset
pairYesTrading pair of your cell, e.g. 'BTCUSDT'.
paramsNoOptional: numeric strategy parameters of your cell. Only numeric params define the neighborhood; matched per pair where the factory ran them.
intervalYesCandle interval, e.g. '1d', '1w', '1M'.
strategyYesStrategy key, e.g. 'rsi_sma'.
user_cagrYesYour result: CAGR in percent (e.g. 41 for +41%) — the cell being assessed.
asset_typeNoAsset class filter (default 'crypto').
user_sharpeNoOptional: your annualized Sharpe (result_sharpe scale). Used for the counted-N Deflated Sharpe where neighbor coverage allows.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully owns behavioral disclosure. It reveals that the tool is read-only and look-ahead free, states the honest small-n handling (no DSR-N claimed under 15 neighbors), and explicitly labels the parameter axis as 'always anecdotal'. It also explains the counted-N logic for DSR, providing a complete picture with no contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries value: the hook, the input, the output, the caveat about small-n, and the axis distinction. It is front-loaded with the core question and adds specifics in a natural progression, with zero filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 params, nested object, no output schema), the description is remarkably complete: it explains the output shape (distribution metrics, verdict, DSR threshold), the edge case handling, and the semantic difference between axes. An agent can confidently call this tool without additional inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is already documented. The description adds meaning by explaining the conceptual role of params ('Only numeric params define the neighborhood'), how user_sharpe feeds into the counted-N DSR, and what axis values mean. This goes beyond a bare schema listing, though it doesn't re-explain each field individually.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a direct, specific question ('Is this backtest result real, or a lucky cell?') and a clear verb ('Assess one backtest result against its neighborhood'). It names the exact resource (a strategy+interval+pair cell) and what it returns (distribution, verdict, DSR), making it unmistakably distinct from sibling tools like arena_get_backtest or arena_is_distinguishable, even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly contrasts this tool with a naïve approach ('instead of trusting a single "+X% CAGR" cell'), giving a clear when-to-use. It also provides conditional guidance: 'Set axis="parameter" for the secondary, always-anecdotal view' and explains the small-n rule ('fewer than 15 neighbors → "insufficient"'). This is actionable and non-ambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_signal_contextSignal Context — filters vs. today, in one callAInspect

Should I take this entry? Answers it for one (strategy, pair, interval) in ONE call instead of seven. Aligns what each entry filter historically did to this strategy (arena_get_strategy_filter_effect) with where that filter stands TODAY (bull-market gauge, altcoin-season signal, volatility phase, 200-week trend for BTC): filters[].blocks_this_entry says which filter would sit this entry out, with the measured worst-loss / return deltas next to it. Adds the current signal state (anticipated is always false — before candle close there is no signal), an edge_vs_benchmark block gated by the MEASURED noise floor (a gap below the floor is a measurement artifact, not a finding), a contradictions block (e.g. Pulse risk-off while the macro regime reads risk-on — reported, never resolved), and measured invalidation zones (pivot clusters, 200-week SMA; BTC only). detail: 'headline' (default) returns the statement, three key numbers and only the decisive filters; 'full' adds every variant, the raw pulse/macro/filter-effect blocks. Every source can fail independently — sources_used / sources_unavailable make the basis auditable; the answer never silently narrows. Returns a plain-language statement with its confidence and the reason for that confidence — state it, do not hedge it further; the payload carries its own scope note. Compose further with arena_get_strategy_performance_by_regime (WHEN has this worked) and arena_is_distinguishable. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesPair, e.g. 'BTCUSDT'. Case-insensitive.
detailNo'headline' (default): statement + key numbers + decisive filters. 'full': every measured variant plus the raw source blocks.
intervalNoDefault '1w'. Candle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyYesStrategy key, e.g. 'rsi_sma'. See arena_list_strategies.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and handles it exceptionally. It discloses that 'anticipated is always false' before candle close, that edge_vs_benchmark is gated by the measured noise floor with artifacts flagged, that contradictions are reported but never resolved, that each source can fail independently with auditable sources_used/sources_unavailable, and that the answer never silently narrows. This is comprehensive disclosure beyond what any schema could provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. It front-loads the core question, then logically organizes the response blocks, caveats, error handling, and composition guidance. Dense with actionable information but never verbose or repetitive, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, this description is remarkably complete. It explains the response structure (filters[], edge_vs_benchmark, contradictions, invalidation zones), the detail modes, source auditability, failure independence, and the confidence/reason block. Unusual edge cases like the pre-candle-close signal state and the noise-floor artifact are explicitly covered. Nothing an agent needs to call and interpret the result correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already provides rich descriptions for all four parameters, including the alignment caveat for multi-day intervals and the detail enum values. The description reinforces these meanings but doesn't add new parameter-level semantics beyond what the schema already documents, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the exact question it answers ('Should I take this entry?') and states the precise scope (one strategy/pair/interval in one call instead of seven). It explicitly differentiates from sibling arena_get_strategy_filter_effect by showing how it aligns historical filter effect with current signal state. The purpose is unambiguous and distinctive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: it's a one-call replacement for a seven-call flow, and it names two companion tools for composition (arena_get_strategy_performance_by_regime and arena_is_distinguishable). However, it doesn't explicitly state when NOT to use this tool or list mutually exclusive alternatives, so it's a slight step below the highest bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_spot_priceGet BTC/ETH/SOL Spot PriceAInspect

Current BTC, ETH and SOL spot price — what is Bitcoin (or ETH/SOL) worth right now? Live USDT-quoted last price plus 24h change %, high and low from Binance. Use this to anchor the connector’s own analytics (cycle, historical-analog, gem scores) with the current market price instead of switching to web search mid-analysis. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the data source (Binance), freshness ('Live'), and that it is on the free tier, implying accessibility. While it doesn't mention rate limits, authentication, or response structure, for a no-parameter read tool this is reasonable. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no fluff. The core purpose is stated first, followed by specific data details and a practical usage note. The '[Free tier]' tag is a useful addition without verbosity. It is well-organized and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of a no-param get with no output schema, the description adequately conveys what is returned (prices, change, high/low), the source, and the intended use case. It does not explicitly state that all three assets are returned simultaneously, but the wording implies it. Minor ambiguity aside, it is complete enough for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the schema provides no semantic detail; per the rubric, the baseline for 0 params is 4. The description doesn't need to explain parameters and does not attempt to, so it maintains the baseline without additional value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool retrieves the current BTC, ETH, and SOL spot price, and specifies the exact data points: USDT-quoted last price, 24h change %, high, and low, sourced from Binance. The phrase 'what is Bitcoin (or ETH/SOL) worth right now?' makes the purpose immediately clear, and it is distinct from siblings like funding rate or fear-greed tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description directly instructs to use this tool to anchor the connector's analytics (cycle, historical-analog, gem scores) with current market price, and explicitly says 'instead of switching to web search mid-analysis.' This provides clear context and an alternative (web search), though it doesn't name specific sibling tools, which is acceptable given the large sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_stablecoin_supplyGet Stablecoin Supply TrendAInspect

Aggregate stablecoin supply (crypto-liquidity proxy) — is the liquidity impulse turning or accelerating? macro_regime only gives the 30d delta; this exposes the trend: current supply, 30d/90d change (USD + %) plus a daily time series (days, default 365; resolution states points and spacing) so direction and speed are visible, not just a single delta. Read impulse for what the supply change is doing — four states (accelerating / decelerating / reversal / flat). The neighbouring acceleration_usd is the signed difference last-30d minus prior-30d and gets LARGE exactly when the trend reverses, while the older boolean accelerating requires the same direction AND a bigger magnitude; a reversal therefore shows a big acceleration_usd next to accelerating: false. Source DefiLlama peggedUSD. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoLength of the returned daily series in days. Default 365, clamped 7–1095.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does well: it discloses the data source (DefiLlama peggedUSD), the free-tier status, and the structure of the returned trend (current supply plus 30d/90d change and a daily series). It explains tricky field semantics such as what `impulse` states mean, how `acceleration_usd` behaves on reversal, and why it can be large while `accelerating` is false. Return format is described at concept level, though exact response shape is left undefined since no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single paragraph is front-loaded with the tool's purpose and then elaborates on field interpretation. It is dense and information-rich without being repetitive; a minor deduction for packing several concepts (trend, impulse, acceleration_usd, accelerating) into a long sentence that could be split for scanability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given one optional parameter, no annotations, no output schema, and a complex set of sibling tools, the description supplies enough conceptual guidance about returned fields and their interpretation to call and use the tool correctly. It remains slightly short of fully describing the response shape, but no output schema exists, so the conceptual description is reasonable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter exists and schema coverage is 100%, so baseline would be 3. The description goes beyond the schema by restating the default (365), the control over both points and spacing via `resolution`, and the effective meaning of `days` for the series length. That is real added value over the schema's clamp range.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Aggregate stablecoin supply') and explicitly differentiates itself from the sibling 'macro_regime,' which 'only gives the 30d delta.' This distinguishes scope and intent clearly. It stops just short of 5 because the routing language names one sibling but doesn't provide a positive selection rule across the broader toolset.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear context for use - determining whether the liquidity impulse is turning or accelerating - and explicitly contrasts the tool with arena_get_macro_regime as the alternative for a single 30d delta. There is no explicit 'when not to use,' but the comparative framing is strong enough to route an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_sth_cost_basisGet BTC Short-Term-Holder Cost Basis (latest)AInspect

What did recent buyers pay on average — and how far is spot from that? Latest BTC short-term-holder cost basis (realized price of coins younger than ~155 days, BRK brk_sth_realized_price), derived STH-MVRV (spot ÷ STH cost basis), an in_loss flag, plus ±1σ/±2σ bands: basis × exp(±k·σ), σ of ln(price ÷ basis) over a 730-day ROLLING window (sigma_method/sigma_window_days travel in the payload; similar construction to public STH band charts, own convention — not a rebuild). band_zone names the state (above/below basis, beyond ±2σ); sth_mvrv_percentile is the rolling 730d rank. Measured band coverage (2026-08-25, full history): 32.1% of days outside ±1σ (near the Gaussian 31.7%), 7.4% outside ±2σ (wider than the Gaussian 4.6% — fat tails); the bands are descriptive geometry (the measured coverage above tells you how literally to take them). On-chain context you weigh with the percentile field. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does this well: it discloses the 730-day rolling sigma window, the 'own convention' construction, the band coverage caveat, and the free-tier status. There is minor ambiguity about whether sigma_method/sigma_window_days are response fields or request payload fields, but overall the behavior is unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description opens with a plain-language hook and then layers definitions, output fields, and interpretive caveats in a sensible order. It is dense and somewhat long, but the measured-coverage statistics and convention caveat earn their place by helping the agent interpret the bands.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-argument latest-value tool, the description is nearly complete: it names the returned metrics, explains band construction, and tells the agent how literally to read the bands. With no output schema, explicit units and exact response field names would be the only meaningful additions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing for the description to document and the baseline is 4. The mention of sigma_method/sigma_window_days 'traveling in the payload' is slightly confusing against the empty schema, but it does not harm parameter understanding because no parameters exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact resource (BTC short-term-holder cost basis) and the operative action (get latest), then enumerates the derived metrics: STH-MVRV, in_loss flag, sigma bands, band_zone, and percentile. This makes it clearly distinct from the large sibling list, especially with the explicit 'own convention — not a rebuild' note.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: an agent needing the latest STH cost basis and how far spot is from it can infer this tool fits. However, there is no explicit 'when not to use' or a pointer to sibling alternatives such as arena_get_cost_basis_spread or arena_get_onchain_latest, so routing guidance is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_strategy_filter_effectGet Strategy Filter Effect Snapshot (per Asset)AInspect

What would each entry filter have changed for this strategy? Per-(strategy, asset, interval) filter-effect analysis. Returns baseline-stats (no filters) + each observed filter-variant's stats with cagr_delta / drawdown_delta / win_rate_delta vs the time-overlap-matched baseline + best_by_cagr pick (null with best_by_cagr_reason when every variant is low_data or none beats the baseline — no pick below the data gate) + not_applicable_filters list (e.g. altcoin_season excluded on BTC-pair). Baseline and each variant carry their aggregation window (from/to + avg_run_years) — CAGR is time-normalized, so identical trade sets over different windows legitimately produce different CAGR. Based on REAL backtest aggregations — not theoretical 2^5 permutations. Use this to answer 'Which filters would improve my backtest for X on Y?'. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
assetYesPair / symbol (e.g. BTCUSDT). Case-insensitive.
intervalNoDefault '1w'. Candle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyYesStrategy key (see arena_get_strategies).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the baseline+variant structure, delta metrics, the best_by_cagr null-and-reason logic, the data gate, the not_applicable_filters list, and the time-normalization/window caveat. It omits auth or rate-limit context, but the behavioral semantics of the call are unusually rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense with parentheticals but front-loaded with the guiding question, and each clause carries signal (deltas, gate logic, window caveat). Slightly heavy but nearly every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a genuinely complex return shape, the description explains the baseline/variant/delta/pick structure thoroughly enough that an agent knows what to expect. Nothing essential for correct invocation appears missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents asset, interval (including the alignment caveat) and strategy. The description adds little parameter-level meaning beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a concrete question ('What would each entry filter have changed for this strategy?') and immediately narrows scope to per-(strategy, asset, interval) filter-effect analysis, distinguishing it from sibling analytics like arena_get_filter_insights and arena_get_strategy_insights. An agent can tell what resource this operates on without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit use case: 'Use this to answer "Which filters would improve my backtest for X on Y?"'. This gives clear context for selection, but does not name a when-not condition or a specific alternative sibling tool for adjacent questions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_strategy_insightsGet Strategy Insights Matrix or DetailAInspect

Which strategy and interval combinations actually performed? Aggregated backtest performance per (strategy × interval) cell. If strategy AND interval provided, returns detail with per-asset breakdown + param variants. Otherwise returns the matrix. Free tier is limited to the same strategies that are free in the backtester itself (rsi_sma, golden_cross, rsi_ob_os, bnh_fixed, dca_reference, dca_reference_v2); the response then carries plan_capped: true plus plan_cap_note, so a short matrix is never mistaken for a thin database. Detail mode on a Pro-only strategy returns 403 rather than a silently empty answer. API Pro and Power receive every cell. Counts: runCount = deduplicated runs above the trade floor that carry the averages, inertRuns = 0-trade runs of the same cell counted IN ADDITION, runs_total = both. avgWinRate averages only runs with a rated trade (0-trade runs and open single positions store 0, which is not a hit rate); avgBuyholdCagr/beatsBuyhold are shown only when >= 50 % of the cell's assets carry the strategy-window benchmark (bhAssetCoverage) — the envelope carries benchmark_definition, win_rate_definition and counts_definition. Matrix mode also carries summary (cells_total, cells_with_benchmark, cells_beating_bh, cells_beating_bh_with_negative_cagr, cells_without_cagr) counted over the cells actually returned — quote those instead of counting cells yourself. [Free: 6 strategies / Pro+: full]

ParametersJSON Schema
NameRequiredDescriptionDefault
intervalNoDetail mode: interval. Candle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
min_runsNoMatrix mode: minimum runs per cell. Default 5.
strategyNoDetail mode: strategy key (used together with `interval`).
asset_typeNoRestrict to one asset class.
assets_modeNo'top10' restricts to top-10 pairs by run-count.
ref_strategyNoBenchmark reference. Default 'bh'.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral burden and does so: free-tier strategy caps plus `plan_capped`/`plan_cap_note`, 403 rather than silent empty detail on Pro-only strategies, and precise definitions for `runCount`/`inertRuns`/`runs_total` and `avgWinRate`. It also discloses the >=50% `bhAssetCoverage` gate before benchmark fields are shown, which is exactly the kind of non-obvious behavior an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long, but front-loaded with the purpose and mode logic before the definitional detail, and each clause (counts, win-rate definition, benchmark gate, free tier) earns its place given there is no output schema. Slightly dense for a single paragraph.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, zero-annotation, no-output-schema tool, the description covers mode selection, error behavior, tier limits, return-field semantics and the embedded definition envelope. An agent has everything needed to call it correctly and interpret the response without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds mode semantics the schema does not: `strategy`+`interval` jointly select detail mode, `min_runs` is matrix-only, and it names the free strategies explicitly. The `ref_strategy`, `asset_type` and `assets_mode` params get no added meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening question and the phrase 'Aggregated backtest performance per (strategy × interval) cell' give a specific verb and resource, and the matrix/detail duality is spelled out. It does not, however, explicitly differentiate itself from close siblings such as arena_get_strategy_performance or arena_compare_strategies, so an agent must infer the split from the schema shape.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Routing is well specified internally: supplying both `strategy` and `interval` yields detail (with a 403 for Pro-only strategies), otherwise the matrix, and `min_runs` is scoped to matrix mode. There is no guidance on when to prefer this over the sibling performance/compare tools, which keeps it below 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_strategy_performanceGet Strategy Performance Snapshot (per Asset)AInspect

How did this exact strategy, asset and interval perform? Aggregated backtest performance for ONE specific (strategy, asset, interval) combination. Returns run_count, avg_cagr, avg_win_rate, avg_drawdown, effective_years, vs_buy_hold comparison (beats_buy_hold, cagr_delta) and an evidence block declaring the gate machine-readably (gate_applies_to: stats.run_count, threshold 5 runs, benchmark value, aggregation data window). For multi-strategy overview use arena_get_strategy_insights. Use this to answer 'How does strategy X perform on asset Y?'. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
assetYesCrypto pair / symbol (e.g. BTCUSDT, ETHUSDT). Case-insensitive.
intervalNoDefault '1w'. Candle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyYesStrategy key (e.g. rsi_sma, golden_cross). See arena_get_strategies for valid keys.
asset_typeNoOptional asset class filter to disambiguate (e.g. when same pair-name exists in two classes).
ref_strategyNoBenchmark reference. Default 'bh' (Buy & Hold).

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It discloses not just the return fields but also the machine-readable `evidence` block, the gate threshold (5 runs), and a critical behavioral nuance: the interval alignment issue for 2d/3d candles and its measurable impact on CAGR (6.66 pp on average). This is rich, honest behavioral disclosure beyond any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: the opening question orients the agent, the return list is compact, the sibling pointer is explicit, and the interval caveat is crucial. It is longer than the two-sentence gold standard but not bloated; a 4 reflects the tight packing rather than fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no output schema, the description provides a complete picture: it names the output fields and their sematics (including the evidence block with gate_applies_to, threshold, benchmark, aggregation window), states the defaults, and covers the alignment caveat that could otherwise lead to misinterpretation. An agent has everything needed to invoke and interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by specifying defaults for `interval` and `ref_strategy` (not in the schema) and by explaining the alignment-dependent behavior of multi-day intervals — semantic nuance that materially affects result interpretation. It stops short of fully explaining `asset_type` disambiguation beyond what the schema already says, hence 4 rather than 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a precise resource ('strategy performance snapshot per asset'), and clearly delimits scope ('ONE specific (strategy, asset, interval) combination'). It lists the exact returned fields and explicitly distinguishes this from the multi-strategy sibling `arena_get_strategy_insights`. No ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use ('Use this to answer ...') and names the alternative for multi-strategy overview (`arena_get_strategy_insights`). This is enough for an agent to choose correctly, and the contrast with the sibling is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_strategy_performance_by_regimeGet Regime-Aware Strategy PerformanceAInspect

In which macro regime has this strategy worked? Historical backtest performance for ONE (strategy, asset, interval) combination SPLIT BY macro market regime (sweet_spot / late_cycle_warning / crisis / recovery — classified at each trade's entry date), PLUS the CURRENT live regime so you can align the buckets yourself. You get the per-regime numbers to weigh directly (per-bucket verdicts live in the per-cell tools, where the pool is stable). Each regime bucket returns trades, trades_per_config (trade counts pool ALL parameter-variant configs — see config_count), win_rate, avg_pnl_pct (per-trade return, not annualized), reward_risk_ratio (per-trade mean/stddev, NOT annualized Sharpe), share_of_time_pct (calendar-day-weighted — each regime observation counts the days until the next one, so the mixed weekly/daily cadence of the regime history does not skew the share) and a rating. The benchmark block anchors the payload with the combination's buy-and-hold CAGR (identical to arena_get_strategy_performance vs_buy_hold — without that anchor, regime avg_pnl_pct is a trajectory, not an excess). For a decision-grade view compose with arena_get_strategy_filter_effect and arena_is_distinguishable. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
assetYesCrypto pair / symbol (e.g. BTCUSDT, ETHUSDT). Case-insensitive.
intervalNoDefault '1w'. Candle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyYesStrategy key (e.g. rsi_sma, golden_cross). See arena_get_strategies.
asset_typeNoOptional asset class filter to disambiguate identical pair-names.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well. It discloses that regime classification happens at trade entry date, that trades_per_config counts all parameter-variant configs, that avg_pnl_pct is not annualized, reward_risk_ratio is not annualized Sharpe, and that share_of_time_pct is calendar-day-weighted to avoid cadence skew. It also explains the benchmark anchor's purpose. Minor gaps remain around the rating scale and live regime presentation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and informative, but it is long and structured as a single block, which reduces scannability. Many parentheticals and qualifiers are valuable, but some could be trimmed or front-loaded more effectively. It earns a mid-range score because every sentence carries meaningful content, but formatting and pacing are not optimal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description compensates by enumerating most returned fields and their semantics, explaining the benchmark block, and addressing the large interval caveat. It also references sibling tools for composition. Missing details such as rating scale or live regime format are minor. Overall, the description is robust enough for an agent to call and interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaningful context beyond the schema, especially for the interval parameter: it warns that multi-day candle alignment can move CAGR by 6.66 pp on average and that 1d/2d/3d behaved as one block. This enriches interpretation of the interval parameter. The description doesn't add much for asset or asset_type, but the schema already covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: it retrieves historical backtest performance split by macro regime for ONE (strategy, asset, interval) combination. It also distinguishes itself from related tools by clarifying that per-bucket verdicts live elsewhere and that the benchmark block matches arena_get_strategy_performance. This provides clear differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: when you need per-regime numbers to weigh directly and align with the current live regime. It also suggests composing with arena_get_strategy_filter_effect and arena_is_distinguishable for a decision-grade view. It lacks explicit when-not-to-use instructions, but the guidance is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_universeGet Universe DetailAInspect

Which pairs are in this universe? Returns one pair universe in full: its id, label, selection rule and the complete list of pairs it currently contains. Use it to see what you are about to test BEFORE handing a universe_id to arena_run_universe_backtest, or to resolve a universe into explicit pairs. For the list of available universes call arena_list_universes. Without as_of the universe reflects the CURRENT membership (CoinGecko market-cap rank) — a backtest over it carries survivorship bias for the earlier years; the pit block in the payload says so. With as_of (YYYY-MM-DD, >= 2026-07-14) it returns the membership as MEASURED on that day from our own daily record of Binance USDT spot, ranked by 24h quote volume (not market cap) — coins delisted since are included, coins listed later are not. Point-in-time universes are recorded forward-only; earlier dates are refused, not reconstructed. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
as_ofNoOptional YYYY-MM-DD. Point-in-time membership on that day (recorded since 2026-07-14, volume-ranked).
universe_idYesUniverse id, e.g. 'crypto-top-50'.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the burden of behavioral disclosure. It explains the survivorship bias without as_of, the forward-only nature of point-in-time universes, the rejection of earlier dates, and the difference in ranking methodology (market cap vs 24h quote volume). This is exemplary transparency for a tool with subtle behavioral nuances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

While long, every sentence earns its place. The first sentence states the core purpose; subsequent sentences cover usage guidance, the as_of semantics, the bias warning, and the free tier. Information is front-loaded, and the structure flows logically from purpose to parameters to caveats. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has inherent complexity (current vs point-in-time membership, survivorship bias, forward-only recording), and the description addresses all of it. It tells the agent exactly what is returned and how to interpret the as_of parameter. No output schema exists, so the description's explicit statement of the return fields ('id, label, selection rule and the complete list of pairs') covers that gap. Nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema description coverage is 100%, the description adds substantial meaning beyond the schema. For as_of, it clarifies the format, the minimum date (2026-07-14), the definition of 'measured on that day', and the consequence of using it (no survivorship bias). For universe_id, it implies the format via the example. This goes well beyond the schema's terse descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Returns') and resource ('one pair universe in full') and enumerates the exact fields (id, label, selection rule, complete list of pairs). It clearly distinguishes itself from arena_list_universes and arena_run_universe_backtest by naming them as siblings with different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells when to use the tool: 'to see what you are about to test BEFORE handing a universe_id to arena_run_universe_backtest, or to resolve a universe into explicit pairs.' It also names the alternative for listing universes ('call arena_list_universes'). No ambiguity remains.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_get_volatility_historyGet BTC Volatility History (RV + ATR%)AInspect

How volatile has Bitcoin been? Daily Bitcoin volatility time series: realized volatility (30d & 90d, √252-annualized, close-to-close) and ATR% (Wilder EMA-14, captures intraday range + gaps), on the same scale. Ranks come in two flavours answering different questions: rvRank/atrPctAnnRank expand from the start of history and are look-ahead-free, but they include BTC's structural volatility decline; rvRankRolling/atrPctAnnRankRolling rank against a trailing 2-year window, which removes that trend from the comparison. History reaches back to 2009 via a stitched pre-Binance close series; ATR is null before the Binance era because no daily high/low exists that far back (see meta.coverage). Use from/to for a specific window instead of pulling everything and discarding it, and granularity/fields to keep long ranges affordable. For long ranges pass format: "columns" (one array per field — the largest saving) together with schema_version: "2026-08" (rounds floats; opt-in until the default flips 2026-11-01), fields: "minimal" and meta: "minimal" — every response carries a size block with chars_before/chars_after/saved_pct measuring the saving for YOUR call. Free tier: last 365 days. Related: arena_get_volatility_phases (current phase per pair), arena_get_iv_snapshot (implied vs. this realized — same RV method, but its realized_vol_30d is computed at snapshot time BEFORE that date has traded, so on fresh breakout days the two can differ; this series uses completed closes and is the one to trust for finished days), arena_get_cycle (regime context). [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoISO date (YYYY-MM-DD), inclusive. End of the window. Defaults to the latest bar.
daysNoNumber of most recent days to return. Free tier capped at 365; API Pro unlimited. Ignored when from/to are given.
fromNoISO date (YYYY-MM-DD), inclusive. Start of the window. Free tier still only sees the last 365 days.
metaNoDefault full. 'minimal' drops params/params_hash/warmup, which are only useful on the first call.
fieldsNoDefault full. 'minimal' returns date, close, rv, rvRank, rvRankRolling, atrPctAnnRank, atrPctAnnRankRolling only — measured saving 18–20 % of characters (full-history series, 2026-07-31; the `size` block in the response has the figure for your actual call), not a fifth of the size. Combine with granularity or a from/to window for a real reduction; dropping fields alone saves less than it looks.
formatNoDefault 'rows' (series[] of objects). 'columns' returns series_columns instead — one array per field (date, close, rv, …) with each field name written once; the largest single saving for long ranges. Combines with fields, granularity and schema_version (the size block then measures the combined saving).
granularityNoDefault daily. weekly/monthly keep the LAST observation of each period (a state, not an average).
schema_versionNoDefault '2026-07' (unchanged output). '2026-08' rounds floats to 2 decimals (ranks 1) and reports the saving. Default flips 2026-11-01.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so richly: it discloses the free-tier 365-day cap, that ATR is null before the Binance era, that history reaches 2009 via a stitched series, that rvRank/atrPctAnnRank are look-ahead-free but trend-biased while the Rolling variants are not, and that the size block measures savings. These are non-obvious behavioral traits an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening question ('How volatile has Bitcoin been?') front-loads the purpose, and each subsequent sentence carries distinct information (rank semantics, coverage caveats, affordability knobs, related tools). It is dense and long, but not padded — though the savings/format guidance could be tightened for an agent skimming it.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, yet the description names the returned fields, explains the two rank families and their trade-offs, documents data-coverage limits (ATR null pre-Binance, 2009 stitching, free-tier window), and points to the meta.coverage/size blocks for self-description. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3; the description goes beyond it by explaining the rank-flavour distinction behind rvRankRolling, why `format: "columns"` produces the largest saving, how fields/granularity/format combine in the size block, and that schema_version 2026-08 is opt-in with a default flip date. It adds real interpretive value without re-documenting the enum values already in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Daily Bitcoin volatility time series') and enumerates exactly what is returned: realized volatility (30d/90d, √252-annualized) and ATR% (Wilder EMA-14). It also names the differentiating sibling (arena_get_iv_snapshot) and explains the methodological distinction, so an agent can separate it from related tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing is provided: use `from`/`to` rather than pulling everything, use `granularity`/`fields` for affordability, and pass `format: "columns"` + `schema_version: "2026-08"` for long ranges. It also states the free-tier limit and directs to sibling tools for the questions this one does not answer (arena_get_volatility_phases, arena_get_cycle).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_is_distinguishableIs this difference real — or smaller than the measurement noise?AInspect

Do these two CAGR figures actually differ? Check before ranking them. Pass the two values as a and b (gross CAGR in percent, same basis) plus axes — which arbitrary choices went into them — and the tool returns whether their gap clears the MEASURED noise floor of those choices, along with the floor itself, the dominant axis, and the probe + date it was measured on. axes accepts: grid_phase (how a multi-day candle grid is aligned to the Unix epoch; exists only on 2d/3d), parameter_choice (neighbouring parameter settings — by far the largest axis), window_edges (shifting the start date), pair_selection (which pairs made it into the universe). Pass ALL axes that genuinely varied; the floor is their maximum, not their sum. Optionally set interval to the candle interval so the floor can be sharpened where an axis was measured per interval — passing grid_phase together with a non-multi-day interval is a hard error, because that axis does not exist there. label_a and label_b are optional display names for the two values and are echoed back inside the explanation, so a multi-way comparison stays readable. Read-only, no market data touched. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
aYesFirst value — gross CAGR in percent (e.g. 33.1 for +33.1%).
bYesSecond value, same unit and same basis as a.
axesYesWhich arbitrary choices differ between a and b. Pass every one that genuinely varied — omitting an axis makes the answer look more certain than it is.
label_aNoOptional name for a, echoed in the explanation.
label_bNoOptional name for b, echoed in the explanation.
intervalNoCandle interval, if known (e.g. '1d', '2d', '3d', '1w'). Sharpens the floor where an axis was measured per interval. Passing grid_phase with a non-multi-day interval is an error, not a rounding detail — that axis does not exist there.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses it is read-only ('Read-only, no market data touched'), describes the return contents (whether gap clears noise floor, the floor itself, dominant axis, probe and date), explains the consequence of omitting an axis (answer looks more certain than it is), and documents the hard error for grid_phase with non-multi-day intervals. This goes well beyond a minimal behavioral statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place. It front-loads the core question, then methodically explains the return, the axes, the aggregation, the optional interval, the labels, and the safety guarantee. No fluff or repetition; structured with clear logical flow from purpose to parameters to behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, multiple axes, error conditions, no output schema), the description is complete. It explains the return structure, the error condition, the aggregation rule, and every parameter's role. An agent has everything needed to call it correctly and interpret the result, even without an output schema. The only minor omission is the exact JSON return shape, but the description's narrative covers the components.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds substantial meaning: it explains the axes enum in detail (including that parameter_choice is the largest axis), clarifies the floor aggregation rule, defines the error condition for interval/grid_phase, and explains that labels are display names echoed back. This transforms the schema's raw definitions into practical, operational guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a crisp question — 'Do these two CAGR figures actually differ?' — and states the verb ('Check') and the resource (two CAGR values). It distinguishes itself from sibling tools by being the only one that performs a noise-floor significance test on CAGR differences, rather than retrieving data or running strategies. The phrase 'before ranking them' positions it clearly in the analysis workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'Check before ranking them.' It instructs the caller to pass all axes that genuinely varied and clarifies the floor is their maximum, not sum. It also warns about the hard error when combining grid_phase with a non-multi-day interval. Though it doesn't name alternative tools, the context makes the intended use unambiguous and the conditions precise.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_list_backtestsList Your BacktestsAInspect

Which backtests have I run? Lists the backtest runs belonging to the authenticated user — newest first, with id, strategy, pair, interval, date range and headline metrics per run. Use it to find a run_id, then call arena_get_backtest for its detail or arena_get_backtest_trades for the individual trades. Only your OWN runs; for the public cross-user leaderboard use arena_get_winners. Paginated via limit + offset. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairNoFilter by pair symbol, e.g. BTCUSDT. Omit for all.
limitNoPage size, max 100, default 50.
offsetNoRows to skip for paging; default 0.
intervalNoFilter by candle interval; omit for all. Candle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyNoFilter by strategy key, e.g. 'rsi_sma'. Omit for all.
asset_typeNoFilter by asset class, e.g. 'crypto'. Omit for all.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavior disclosure. It discloses the operation is a read-only list ('Lists the backtest runs'), scopes results to the authenticated user, describes ordering ('newest first'), and states the tier ('API Pro'). It does not explicitly say there are no side effects, but 'list' implies a read. It could have noted that no mutation occurs or described auth failure behavior, but overall it provides sufficient behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no filler. The opening question 'Which backtests have I run?' immediately orients the agent, followed by the core purpose, usage flow, scope, and pagination note. Every clause earns its place, and the structure is front-loaded with the most critical information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

As a list tool with no output schema, the description compensates by listing the fields returned (id, strategy, pair, interval, date range, headline metrics), pagination parameters, scope, and related tools. It does not describe error or empty-result behavior, but that is a minor gap given the tool's simplicity and the absence of annotations. The mention of the API tier and the authenticated-user scope adds context. Overall, it is near-complete for its purpose.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with all six parameters already described in the input schema (e.g., pair filter, limit max 100, offset default 0, interval enum, strategy, asset_type). The tool description adds only 'Paginated via limit + offset', which reinforces but does not extend beyond schema. Since coverage is high, the baseline is 3 and the description adds no new parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and the resource 'backtest runs', specifies the scope 'belonging to the authenticated user', and lists the fields returned (id, strategy, pair, interval, date range, headline metrics). It explicitly differentiates from sibling tools by naming arena_get_backtest for detail and arena_get_winners for the public leaderboard, so an agent can distinguish it without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives direct usage guidance: 'Use it to find a run_id, then call arena_get_backtest for its detail or arena_get_backtest_trades for the individual trades.' It also states when not to use it ('Only your OWN runs; for the public cross-user leaderboard use arena_get_winners') and mentions pagination via limit+offset. This is explicit and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_list_knowledgeList Knowledge Objects (catalog)AInspect

What knowledge objects exist here? Discover what Knowledge Objects exist: lists all published types + their subjects (with min_tier, api_path, seo_slug, latest as_of). Use this BEFORE arena_get_knowledge to learn valid type/subject pairs instead of guessing. New types appear automatically. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It states it's a listing operation (implying read-only) and notes the free tier, giving cost context. It does not explicitly declare side effects or rate limits, but the nature of a zero-parameter list operation and the detailed output fields make behavior sufficiently clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, followed by usage guidance and a note about free tier. It is a bit conversational ('What knowledge objects exist here?') but each sentence earns its place, with no redundancy. Length is appropriate for the information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless list tool, the description is complete: it states what is returned, when to use it, and cost implications. It does not describe the exact response format or pagination, but these are less critical for a simple listing. The absence of an output schema is mitigated by listing the returned fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is trivially 100%. The baseline for 0 params is 4. The description adds value by explaining the output contents (types, subjects, and the specific fields), which compensates for the lack of an output schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists all published knowledge types and their subjects, with specific fields (min_tier, api_path, seo_slug, latest as_of). It distinguishes itself from arena_get_knowledge by positioning itself as a discovery step, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Contains an explicit directive: 'Use this BEFORE arena_get_knowledge to learn valid type/subject pairs instead of guessing.' It also notes that new types appear automatically, implying the tool should be re-run to stay current. This clearly tells the agent when and why to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_list_onchain_seriesList Available BRK On-Chain SeriesAInspect

Which on-chain series are available? Lists all 69 available Bitcoin Research Kit (BRK) on-chain series across the groups pilot, sentiment, mining, supply, cointime, activity, liquidity (e.g. MVRV, NUPL, SOPR, Realized-Price, Mayer, Puell, STH/LTH SOPR, Hash-Ribbons). Returns id + label + group. Use the id with arena_get_onchain_latest / _history. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden, and it does reasonably well: it discloses the cardinality (69 series), the grouping taxonomy, the exact return shape (id + label + group), and the pricing tier ([Free tier]). It does not state whether the list is static or updates over time, which is a minor gap for a no-parameter read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core question and answer, followed by tier and return shape in compact sentences. The parenthetical example list (MVRV, NUPL, SOPR, ...) is long but earns its place by letting the agent recognize series names without a second call.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless catalog with no output schema, the description supplies everything needed: what is returned (id, label, group), how many items, the group taxonomy, and the follow-up tools. Nothing required for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which sets the baseline at 4. The description correctly implies no filtering is possible and instead points to the consumer tools that accept the returned ids.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a direct question that names the resource ('on-chain series'), then states it lists all 69 BRK on-chain series with the enumeration groups. An agent can immediately distinguish this catalog tool from arena_get_onchain_latest / _history, which consume the ids this tool returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent downstream: 'Use the id with arena_get_onchain_latest / _history.' That tells the agent this is the entry point before fetching data, but it does not say when NOT to use it or contrast against other list_* siblings (list_strategies, list_universes, list_backtests).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_list_strategiesList Available Trading StrategiesAInspect

Which strategies can I backtest here? Lists all backtest strategies (key, label, plan, supported asset classes, primary indicators). Filterable by asset class and plan. Use this before calling arena_run_backtest to discover valid strategy names. Entries deprecated for an asset class stay listed (historical results depend on them) and carry deprecated_for + deprecation {since, reason} — do NOT call arena_run_backtest or validate_strategy for those combinations, they return 400. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
langNoLocalized names/taglines. Default 'en'.
planNoFilter to strategies of this plan tier.
asset_classNoFilter to strategies supporting this asset class (crypto or tokenized_rwa).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses that deprecated entries intentionally stay listed, that they carry deprecated_for and deprecation metadata, and that invalid combinations fail with 400. This exposes important non-obvious behavior. It doesn't mention response shape, pagination, or auth, but the key traps are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with a natural question, then the core behavior, then the critical deprecation caveat. The deprecation sentence is long but packs essential warning details. Overall it earns its place without unnecessary fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple optional-parameter list tool with no output schema, the description covers what the tool returns, what filters exist, how it relates to arena_run_backtest, and the important deprecation edge case. A small example of output shape or a note about lang default would push it higher, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds only minimal parameter semantics by mentioning filterability by asset class and plan; it doesn't add new meaning for lang or any details beyond the schema. No penalty is warranted because the schema already documents each parameter well.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a concrete user question and states the exact function: 'Lists all backtest strategies' with the specific fields returned (key, label, plan, supported asset classes, primary indicators). It also frames itself as a discovery tool for arena_run_backtest, which distinguishes it from the many other list/get tools in its sibling set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit: 'Use this before calling arena_run_backtest to discover valid strategy names.' It also gives a clear negative directive — do not call arena_run_backtest or validate_strategy for deprecated asset-class combinations, and states the consequence (they return 400). This is strong, unambiguous routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_list_universesList Asset UniversesAInspect

Which asset universes can I test against? Lists all crypto asset universes (BTC, top-10 crypto, top-50 crypto, etc.) — the underlying pair-sets used by custom-report and universe-backtest endpoints. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It states the core action (lists all universes) and clarifies what universes are (underlying pair-sets), which is helpful. However, it does not mention side effects, auth requirements, error cases, or output format. The '[Free tier]' hint is a minor behavioral note. For a read-only list tool, this is adequate but not rich in disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus a bracketed note, with no fluff. The front-loaded question immediately signals the tool's purpose, followed by a precise definition and usage context. Every word contributes meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only list tool, the description covers the essential aspects: what it does, what it returns conceptually (a list of universes), and why it matters (used by specific endpoints). It does not detail the exact output structure, but given no output schema, that is acceptable. It is sufficiently complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the baseline is 4 as per rubric. The description adds value by giving examples of expected universe names (BTC, top-10, etc.), which helps an agent anticipate the output, but there are no parameter semantics to elaborate on.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Lists'), names the exact resource ('all crypto asset universes'), and provides concrete examples (BTC, top-10, top-50). The opening question 'Which asset universes can I test against?' directly addresses an agent's selection need and clearly differentiates this from other list tools (e.g., arena_list_backtests) by focusing on universes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful context on when to use it: it lists the pair-sets used by 'custom-report and universe-backtest endpoints', implying the tool is needed when selecting a universe for those operations. It does not explicitly name alternatives like arena_get_universe, but the context is clear enough for appropriate invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_run_backtestRun a New BacktestAInspect

How would this strategy have performed? Run ONE strategy on ONE pair over a date range and get the full result: CAGR, total return, max drawdown, win-rate, trade count, Buy & Hold comparison, net-of-fees figures, and a run_id for later retrieval. Synchronous, typically 3–10s. Use this when the user wants a concrete result for a specific setup. For several strategies side by side use arena_compare_strategies; for many pairs at once use arena_run_universe_backtest; to judge whether an EXISTING result is trustworthy rather than produce a new one, use validate_strategy or arena_get_robustness_field. Filters are optional and only remove entries; run once without them for the baseline. Read result.benchmark before comparing cagr to buyhold_cagr: warmup or a late listing can shorten the strategy window, and matches_strategy_window:false means the two figures are annualized over DIFFERENT periods — in that case benchmark.strategy_window carries the like-for-like buy-and-hold (its cagr_delta_pp is benchmark-minus-benchmark, defined in cagr_delta_pp_definition; strategy vs like-for-like benchmark is strategy_minus_window_benchmark_pp) over the window the strategy actually traded, and THAT is the one to compare against. Per-day quota: Pro=50, Power=500. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCrypto pair symbol, e.g. BTCUSDT, ETHUSDT, SOLUSDT.
paramsNoStrategy-specific parameters, e.g. { rsi_period: 14 }. Omit to use the audited defaults — changing them without a reason is how overfitting starts.
capitalNoStarting capital in quote currency. Default 10000. Affects absolute figures only, not CAGR or win-rate.
date_toNoEnd date, YYYY-MM-DD. Default: today.
filtersNoOptional entry filters (Pro+). Each one only ever REMOVES entries — filters never create trades. Omit for the unfiltered baseline.
intervalYesCandle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyYesStrategy key — use arena_list_strategies to find valid keys.
date_fromYesStart date, YYYY-MM-DD. Earlier than the pair listing is clamped to the first available candle.
asset_typeYesAsset class. Use 'crypto' unless you are explicitly backtesting a tokenized real-world asset. Note: tokenized stocks/ETFs/gold trade AS crypto pairs (e.g. spybUSDT, qqqbUSDT) — there is no separate stocks/forex backtest surface; non-crypto asset classes were retired.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: synchronous with 3-10s latency, per-day quota, filters only remove entries, and a detailed benchmark-window caveat that prevents misreading cagr versus buyhold_cagr. It even warns that matches_strategy_window:false means different annualization periods.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place: outcome first, then latency, then routing, then filters, then the critical benchmark caveat. It front-loads the decision-relevant facts and does not pad.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with nested filters and no annotations or output schema, the description is remarkably complete: it lists the result metrics, explains how to retrieve later via run_id, warns about benchmark windows, and states quota limits. The absence of an output schema is compensated by the explicit mention of key result fields and the comparison caveat.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already explains every parameter, which sets the baseline at 3. The main description adds only mild reinforcement ('filters are optional and only remove entries') rather than new per-parameter semantics; most of its added value concerns result interpretation, not parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line frames the user question and then states exactly what the tool does: run ONE strategy on ONE pair over a date range and return the full result, including CAGR, drawdown, win-rate, and a run_id. It also names the sibling tools it is not (compare_strategies, run_universe_backtest), so there is no ambiguity with the large sibling list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit 'Use this when' condition, then names concrete alternatives for adjacent use cases: several strategies, many pairs, and judging existing results. This is exactly the routing guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_run_grid_backtestRun a Grid-Trading BacktestAInspect

Would a grid bot have made money here? Simulate a GRID BOT (buy-low / sell-high ladder inside a fixed price range) on historical candles. Returns final value, return %, CAGR, trade count, fees paid and a Buy & Hold comparison — plus zerlegung (spot runs): decomposition splits the result into ladder P&L from completed buy→sell cycles, allocation P&L of the starting coins, open grid buys and fees (identity_check_usd must be ~0); benchmarks anchors buy-and-hold, the never-touched starting split (static_allocation, with coin_share_start) and an arithmetic 50/50 at the entry price — grid_vs_static_pp is the number that says whether the ladder added anything over just holding the split; fee_economics gives the break-even spacing (2 × fee) and flags below_breakeven. Note: buyhold_return/outperformance keep their legacy anchor (first→last candle close); benchmarks anchor at entry. This is a different machine from the strategy backtester: grid bots earn from oscillation inside a range, not from trend — for signal-based strategies use arena_run_backtest instead. The result depends heavily on the range you choose (low_price / high_price); a range the price left early makes the bot idle (by default the grid pauses outside the range and resumes when price returns; set stop_on_range_exit to end the run at the first close outside it instead, selling all coins there), so treat range choice as part of the hypothesis, not a detail — arena_suggest_grid_range proposes a defensible range. Each run is saved to your account (the returned id is the run_id); publish a public snapshot page with arena_share_grid_backtest. grid_mode picks neutral (default) or long. Optional leverage (2/3/5, grid_mode long only, Pro) with funding_mode (conservative default / historical BTCUSDT / none): simulates an isolated-margin futures long grid — margin = total_investment, the grid trades margin × leverage, funding accrues daily on the open position, liquidation is checked per candle at the low. It simulates, it does not recommend: the result can be a total loss of the margin. Free tier limited to BTCUSDT/ETHUSDT. Per-day quota: Free=5, Pro=50, Power=500. [Free / Pro / Power tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCrypto pair symbol, e.g. BTCUSDT. Free tier: BTCUSDT or ETHUSDT only.
end_dateYesSimulation end, YYYY-MM-DD.
fee_rateYesPer-trade fee fraction, e.g. 0.001 for 0.1% (Binance spot taker).
leverageNoOptional, default 1 (spot grid, unchanged). 2/3/5 = isolated-margin long grid (grid_mode must be long; Pro). Adds liquidated, liquidation_time/price, funding_cost_usd and max_notional_exposure to the result; final_value/total_return are then on the margin.
grid_modeNo'neutral' (default): buys coins for every level above the entry at the start (the coin share follows the entry's position in the range — measured 2–83 %, NOT a fixed 50/50; see benchmarks.static_allocation.coin_share_start), then buys and sells around the entry. 'long': starts 100% in cash, buys dips below the entry, sells on recovery — required for leverage.
grid_typeYesLevel spacing: 'arithmetic' = equal price steps, 'geometric' = equal percentage steps (usually the better fit for crypto).
low_priceYesLower bound of the grid range, in quote currency. Below it the bot is fully invested and stops buying.
grid_countYesNumber of grid levels between low_price and high_price (2–200). More levels = more, smaller trades = more fees.
high_priceYesUpper bound of the grid range, in quote currency. Above it the bot is fully in cash and stops selling. Must exceed low_price.
start_dateYesSimulation start, YYYY-MM-DD.
entry_priceNoOptional price at which the bot starts; default is the OPEN of the first candle (must lie inside low_price..high_price).
funding_modeNoOnly with leverage > 1. 'conservative' (default): flat 0.05%/day on the open position. 'historical': recorded daily average of three exchanges, BTCUSDT from 2019-09-08 only — otherwise falls back to conservative and flags funding_fell_back_to_conservative. 'none': no funding (optimistic).
stop_loss_priceNoOptional: liquidate the whole grid and stop once price falls to this level.
total_investmentYesCapital in USDT spread across the grid; min 100.
take_profit_priceNoOptional: liquidate the whole grid and stop once price rises to this level.
stop_on_range_exitNoOptional (default false): stop once a candle CLOSES outside low_price..high_price — sell all coins at that close, stop_reason range_exit_high/low. Default: the grid pauses outside the range and resumes when price returns.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load and does so: it discloses the returned fields, the zerlegung sub-reports, that runs persist to the account with a run_id, tier/quota limits, the simulation-not-advice caveat ('can be a total loss of the margin'), and range-exit pausing semantics with stop_on_range_exit as the override.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long, but the length is justified by a 16-parameter tool with no output schema, and the value proposition is front-loaded in the first sentence. A few clauses (tier tags, T&C-style caveats) sit near the end and could be trimmed, keeping it short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema or annotations, the description compensates fully: it enumerates return values and the decomposition/benchmarks/fee_economics breakdown, explains the legacy vs entry-price anchoring discrepancy, and covers leverage, funding, quotas and persistence. Nothing needed to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema descriptions are already rich, so the baseline is 3. The description still adds real meaning beyond the schema for leverage ('margin = total_investment, the grid trades margin × leverage, funding accrues daily ... liquidation is checked per candle at the low') and reinforces the low_price/high_price range dependency.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening states a concrete verb and resource ('Simulate a GRID BOT ... on historical candles') and defines the mechanism (buy-low/sell-high ladder inside a fixed range). It explicitly distinguishes itself from the sibling strategy backtester ('This is a different machine ... for signal-based strategies use arena_run_backtest instead'), so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the alternative (arena_run_backtest) and the condition selecting it, points to arena_suggest_grid_range when range choice is unresolved, and to arena_share_grid_backtest for publishing. It also frames when the grid archetype applies ('earn from oscillation inside a range, not from trend') and warns that range choice is part of the hypothesis.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_run_universe_backtestRun Backtest on a Pair Universe (async)AInspect

Does this strategy hold up across a whole universe? Runs it against every pair in the universe. Pair cap depends on your API tier: Pro 50, Power 250 — Power covers crypto-top-250 in ONE job, and a single job keeps the ranking on one pair set (merging results across different pair sets measures pair selection, not strategy quality). THIS CALL IS ASYNCHRONOUS AND RETURNS NOTHING BUT A job_id: the result is NOT in this response. You MUST poll arena_get_job_status until status is 'completed'; estimated_seconds in the create-response says how long to budget. Provide either universe_id (call arena_list_universes) OR explicit pairs[]. Benchmarks bnh_fixed and dca_reference_v2 are accepted here — run one of them over the SAME universe and interval alongside: an excess over buy-and-hold is only readable next to the buy-and-hold value itself, which can be negative. beats_bh_count compares each pair's cagr against the LIKE-FOR-LIKE buy-and-hold — the benchmark measured over the window the strategy actually traded, not from the requested start. A long warmup or a pair listed after date_from shifts that start, and comparing across two different windows is a handicap, not a benchmark. The old pairing is still reported as beats_bh_count_requested_window, and pairs_with_window_offset says on how many pairs the two can differ at all; per-pair, buyhold_cagr_strategy_window and benchmark_matches_window carry the same distinction. PERSISTENCE: universe results live ONLY in the job response (api_jobs.result). They are deliberately not written to backtest_runs, so they carry no filter_binding and no coin-denominated history, and you will not find them later via arena_list_backtests — copy what you need out of the job result. Per-day quota: Pro=5, Power=50. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
pairsNoExplicit pair list. Hard schema limit 250; the effective cap is your tier (Pro 50, Power 250). Use instead of universe_id.
paramsNoStrategy-specific parameters applied to EVERY pair in the universe. Omit for audited defaults.
capitalNoStarting capital in quote currency. Default 10000. Affects absolute figures only, not CAGR or win-rate.
date_toNoEnd date, YYYY-MM-DD. Default: today.
filtersNoOptional entry filters (Pro+). Each one only ever REMOVES entries — filters never create trades. Omit for the unfiltered baseline.
intervalYesCandle interval: '1d' daily, '2d'/'3d' multi-day, '1w' weekly, '1M' monthly. Multi-day candles (2d/3d) are anchored to the Unix epoch, so one of n possible alignments is used. Measured on our own corpus, the choice of alignment alone moves CAGR by 6.66 pp on average (max 12.30). Treat differences below that as not distinguishable — 1d/2d/3d behaved as one block in our tests, not a ranking.
strategyYesStrategy key — call arena_list_strategies.
date_fromYesStart date, YYYY-MM-DD. Earlier than the pair listing is clamped to the first available candle.
universe_idNoPre-curated universe — call arena_list_universes for valid IDs. Capped by tier (Pro 50, Power 250); a larger universe is rejected rather than silently truncated.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does: it declares the call asynchronous and that the response contains only a job_id, describes polling and estimated_seconds, tier-dependent pair caps and per-day quotas, and the persistence caveat that results live only in api_jobs.result and are not written to backtest_runs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the highest-value facts (async, job_id only, must poll) before drilling into benchmark-window nuance. The passages on beats_bh_count vs beats_bh_count_requested_window and the window-offset caveats are somewhat repetitive, but each sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, it fully covers the async contract, how to obtain results, quotas, tier limits, and result persistence, including which sibling tools will NOT surface the results later. Nothing an agent needs to invoke and consume this call correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: tiers gate the effective pair cap vs the schema's 250 hard limit, and it frames universe_id vs pairs[] as mutually exclusive sources.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a concrete question and states the exact operation: run the strategy against every pair in the universe. It is clearly distinguishable from arena_run_backtest (single pair) and arena_run_grid_backtest (parameter sweep) without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: provide either universe_id (via arena_list_universes) or explicit pairs[], plus the mandate to poll arena_get_job_status until 'completed'. It also prescribes pairing with bnh_fixed/dca_reference_v2 over the same universe and interval, which is genuine when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

arena_statusConnection, Auth & Quota StatusAInspect

Am I connected, and what can this key do? Returns auth status (key kind: oauth connector or bearer API key, tier), server version, current UTC time, and the rate-limit state (hour/day used, remaining, reset) WITHOUT consuming extra quota beyond this call itself. Call this first when other tools fail: it separates auth problems (reconnect), tier problems (upgrade) and rate limits (wait) from real outages. [Free tier]

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosure. It clearly states that the tool returns auth status, tier, server version, time, and rate-limit state, and explicitly notes that it does not consume extra quota beyond the call itself. This covers the key behavioral traits an agent needs to know for a status probe. It does not explicitly say it never modifies anything, but the read-only nature is strongly implied by the diagnostic framing and the listed outputs, so the slight gap is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is about 90 words and is well-structured: it opens with the core question, then lists outputs, then gives usage guidance and the quota note. Every sentence contributes new information, and the most important guidance ('call this first when other tools fail') is front-loaded near the middle rather than buried at the end. It is slightly longer than strictly necessary, but the density of actionable detail justifies the length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has no parameters, no output schema, and no annotations, the description needs to fully explain what the agent will get and when to use it. It does exactly that: it enumerates all the returned data categories, explains how to interpret them for troubleshooting, and explicitly promises no extra quota consumption. The sibling set is large and all getter-like, so this description successfully carves out a unique diagnostic role. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so schema coverage is trivially 100%. The description goes beyond the empty schema by explaining what the returned fields mean (key kind, tier, time, rate-limit state) and how they map to troubleshooting actions. With no parameters to document, a baseline of 4 is appropriate; the description adds semantic value about the output context even though it cannot describe input semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a customer-facing question ('Am I connected, and what can this key do?') and then enumerates exactly what it returns: auth status (key kind and tier), server version, current UTC time, and rate-limit state. This is a specific verb+resource with no ambiguity. It also distinguishes itself from the many sibling getter tools by positioning itself as the diagnostic endpoint to call when other tools fail, which clearly sets it apart from the rest of the arena_* family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit, actionable guidance: 'Call this first when other tools fail' and then tells the agent how to interpret results by separating auth problems (reconnect), tier problems (upgrade), and rate limits (wait) from real outages. It also states that the call itself does not consume extra quota, which is a practical constraint. This is textbook when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_strategyValidate a strategy/signal (honest backtest)AInspect

Does this strategy survive an honest test? Backtest a trading strategy honestly — look-ahead-aware validation with Deflated-Sharpe-Ratio / multiple-testing correction (Bailey & López de Prado). Returns an EVIDENCE verdict (insufficient_evidence | anecdote | failed_oos | passed_oos) plus metrics, flags and caveats — NOT a buy/sell recommendation. Call this before acting on a strategy or signal list. Accepts a named catalog strategy (type=rules), a timestamped BUY/SELL signal list (signal_list), or a timestamped trade list (trade_list). Checks: realistic next-bar fills (look-ahead/optimism), net of cost, out-of-sample split, and a hard 30-round-trip sample gate (under 30 is always "anecdote"). Not reproducible via generic backtest tools that ignore overfitting. [API Pro tier]

ParametersJSON Schema
NameRequiredDescriptionDefault
oosNoHow the claim is tested out-of-sample. Omit for the default split — the out-of-sample part is what separates a finding from a fit.
costsNoTrading costs. Default 10 bps (crypto) / 5 bps (else) — a gross-only claim usually shrinks once these apply.
marketYesWhich market the claim is about — prices are re-fetched from here, not taken from you.
windowYesPeriod over which the claim is checked.
strategyYesThe claim being validated — supply exactly one of: a catalog strategy (type=rules), your signals (type=signal_list) or your finished trades (type=trade_list).

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It fully discloses the return format (EVIDENCE verdict plus metrics, flags, caveats), the checks performed (look-ahead-aware, net of cost, out-of-sample split, 30-round-trip gate), and explicitly states it is NOT a recommendation. It also mentions the 'API Pro tier' access constraint, which is not covered elsewhere. This is exceptional transparency that goes well beyond a typical description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet comprehensive, covering purpose, inputs, checks, output, and access in about 150 words. It is front-loaded with the core question and key details, with every sentence adding value. While the density might benefit from bullet points, it is well-structured and appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema is thorough (100% coverage), and the description explains the tool's behavior, returns an EVIDENCE verdict, and lists the checks performed. It does not provide an output schema, but it names the possible verdict values. It also mentions the Pro tier. For a tool with nested objects and multiple input types, this description, combined with the detailed schema, is complete enough for an agent to know when and how to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter is already well documented. The description adds context about the input types (rules, signal_list, trade_list) and the overall purpose, but does not add detailed parameter-specific guidance beyond what the schema provides. The 'API Pro tier' note is access-related, not parameter semantics. Given the high schema coverage, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear question and states the tool's function: 'Backtest a trading strategy honestly' with a specific methodology (Deflated-Sharpe-Ratio / multiple-testing correction). It lists the three input types (rules, signal_list, trade_list), returns an EVIDENCE verdict, and explicitly states it is NOT a buy/sell recommendation. It also distinguishes itself from generic backtest tools, making its purpose unambiguous and well differentiated among the many arena_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description advises 'Call this before acting on a strategy or signal list' and notes it is 'not reproducible via generic backtest tools that ignore overfitting,' giving clear when-to-use context. However, it does not explicitly name sibling tools such as arena_run_backtest or arena_get_backtest, nor does it specify when not to use it (e.g., if you only need a simple backtest). This leaves some ambiguity but still provides strong usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updates
    • Changedarena_call_extended1 field changed
      • changedInput schema / properties / tool / enum
        Previous value: -[
        -  "arena_cancel_subscription",
        -  "arena_check_subscription_updates",
        -  "arena_cross_series",
        -  "arena_dca_scenario",
        -  "arena_dip_scenario",
        -  "arena_get_altcoin_season_history",
        -  "arena_get_backtest_trades",
        -  "arena_get_btc_macro_correlations",
        -  "arena_get_carry_monitor",
        -  "arena_get_chart",
        -  "arena_get_cost_basis_spread",
        -  "arena_get_cycle_history",
        -  "arena_get_drift_log",
        -  "arena_get_filter_insights",
        -  "arena_get_funding_rate_history",
        -  "arena_get_gem_score",
        -  "arena_get_gem_validation",
        -  "arena_get_halvings",
        -  "arena_get_hash_ribbons",
        -  "arena_get_kimchi_premium",
        -  "arena_get_ma_distance_history",
        -  "arena_get_mayer_multiple",
        -  "arena_get_mayer_multiple_history",
        -  "arena_get_ontology_term",
        -  "arena_get_platform_activity",
        -  "arena_get_portfolio_correlation",
        -  "arena_get_pulse_history",
        -  "arena_get_reference_models",
        -  "arena_get_report_status",
        -  "arena_get_shared_backtest",
        -  "arena_get_signal_events",
        -  "arena_get_signal_status",
        -  "arena_get_taker_imbalance",
        -  "arena_get_trend_channels",
        -  "arena_get_volatility_insights",
        -  "arena_get_volatility_phases",
        -  "arena_get_volatility_recommendations",
        -  "arena_get_volume_profile",
        -  "arena_get_winners",
        -  "arena_list_subscriptions",
        -  "arena_quote_report",
        -  "arena_share_grid_backtest",
        -  "arena_subscribe_bullmarket_stage",
        -  "arena_subscribe_cycle_changes",
        -  "arena_subscribe_pulse_changes",
        -  "arena_subscribe_signal_alerts",
        -  "arena_suggest_grid_range"
        -]New value: +[
        +  "arena_cancel_subscription",
        +  "arena_check_subscription_updates",
        +  "arena_cross_series",
        +  "arena_dca_scenario",
        +  "arena_dip_scenario",
        +  "arena_get_altcoin_season_history",
        +  "arena_get_backtest_trades",
        +  "arena_get_btc_macro_correlations",
        +  "arena_get_carry_monitor",
        +  "arena_get_chart",
        +  "arena_get_cost_basis_spread",
        +  "arena_get_cycle_history",
        +  "arena_get_drift_log",
        +  "arena_get_filter_insights",
        +  "arena_get_funding_rate_history",
        +  "arena_get_gem_score",
        +  "arena_get_gem_validation",
        +  "arena_get_halvings",
        +  "arena_get_hash_ribbons",
        +  "arena_get_kimchi_premium",
        +  "arena_get_ma_distance_history",
        +  "arena_get_mayer_multiple",
        +  "arena_get_mayer_multiple_history",
        +  "arena_get_ontology_term",
        +  "arena_get_pattern_scan",
        +  "arena_get_platform_activity",
        +  "arena_get_portfolio_correlation",
        +  "arena_get_pulse_history",
        +  "arena_get_reference_models",
        +  "arena_get_report_status",
        +  "arena_get_shared_backtest",
        +  "arena_get_signal_events",
        +  "arena_get_signal_status",
        +  "arena_get_taker_imbalance",
        +  "arena_get_trend_channels",
        +  "arena_get_volatility_insights",
        +  "arena_get_volatility_phases",
        +  "arena_get_volatility_recommendations",
        +  "arena_get_volume_profile",
        +  "arena_get_winners",
        +  "arena_list_subscriptions",
        +  "arena_quote_report",
        +  "arena_share_grid_backtest",
        +  "arena_subscribe_bullmarket_stage",
        +  "arena_subscribe_cycle_changes",
        +  "arena_subscribe_pulse_changes",
        +  "arena_subscribe_signal_alerts",
        +  "arena_suggest_grid_range"
        +]
    • Changedarena_run_grid_backtest1 field changed
      • addedInput schema / properties / stop_on_range_exit
        Added value: +{
        +  "description": "Optional (default false): stop once a candle CLOSES outside low_price..high_price — sell all coins at that close, stop_reason range_exit_high/low. Default: the grid pauses outside the range and resumes when price returns.",
        +  "type": "boolean"
        +}
  2. 3 tool updates
    • Changedarena_get_etf_flows1 field changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "Default 'rows' (series[] of objects). 'columns' returns series_columns instead — parallel arrays (date, cum_net_inflow_usd_m, net_flow_usd_m, market_closed) with each field name once; use it for long windows, it is much smaller.",
        +  "enum": [
        +    "rows",
        +    "columns"
        +  ],
        +  "type": "string"
        +}
    • Changedarena_get_max_pain1 field changed
      • addedInput schema / properties / expiry_date
        Added value: +{
        +  "description": "Optional YYYY-MM-DD of ONE open expiry: upcoming[] (and its strike_ladder/gex) is restricted to it, which keeps ladder/GEX responses small; gex_totals still covers the whole book. An expiry that is not open returns invalid_input listing the open ones. Omit for all upcoming expiries.",
        +  "pattern": "^\\d{4}-\\d{2}-\\d{2}$",
        +  "type": "string"
        +}
    • Changedarena_get_volatility_history1 field changed
      • addedInput schema / properties / format
        Added value: +{
        +  "description": "Default 'rows' (series[] of objects). 'columns' returns series_columns instead — one array per field (date, close, rv, …) with each field name written once; the largest single saving for long ranges. Combines with fields, granularity and schema_version (the size block then measures the combined saving).",
        +  "enum": [
        +    "rows",
        +    "columns"
        +  ],
        +  "type": "string"
        +}
  3. 1 tool update
    • Changedarena_call_extended1 field changed
      • changedInput schema / properties / tool / enum
        Previous value: -[
        -  "arena_cancel_subscription",
        -  "arena_check_subscription_updates",
        -  "arena_cross_series",
        -  "arena_dip_scenario",
        -  "arena_get_altcoin_season_history",
        -  "arena_get_backtest_trades",
        -  "arena_get_btc_macro_correlations",
        -  "arena_get_chart",
        -  "arena_get_cost_basis_spread",
        -  "arena_get_cycle_history",
        -  "arena_get_drift_log",
        -  "arena_get_filter_insights",
        -  "arena_get_funding_rate_history",
        -  "arena_get_gem_score",
        -  "arena_get_gem_validation",
        -  "arena_get_halvings",
        -  "arena_get_hash_ribbons",
        -  "arena_get_kimchi_premium",
        -  "arena_get_ma_distance_history",
        -  "arena_get_mayer_multiple",
        -  "arena_get_mayer_multiple_history",
        -  "arena_get_ontology_term",
        -  "arena_get_platform_activity",
        -  "arena_get_pulse_history",
        -  "arena_get_reference_models",
        -  "arena_get_report_status",
        -  "arena_get_shared_backtest",
        -  "arena_get_signal_events",
        -  "arena_get_signal_status",
        -  "arena_get_taker_imbalance",
        -  "arena_get_trend_channels",
        -  "arena_get_volatility_insights",
        -  "arena_get_volatility_phases",
        -  "arena_get_volatility_recommendations",
        -  "arena_get_volume_profile",
        -  "arena_get_winners",
        -  "arena_list_subscriptions",
        -  "arena_quote_report",
        -  "arena_share_grid_backtest",
        -  "arena_subscribe_bullmarket_stage",
        -  "arena_subscribe_cycle_changes",
        -  "arena_subscribe_pulse_changes",
        -  "arena_subscribe_signal_alerts",
        -  "arena_suggest_grid_range"
        -]New value: +[
        +  "arena_cancel_subscription",
        +  "arena_check_subscription_updates",
        +  "arena_cross_series",
        +  "arena_dca_scenario",
        +  "arena_dip_scenario",
        +  "arena_get_altcoin_season_history",
        +  "arena_get_backtest_trades",
        +  "arena_get_btc_macro_correlations",
        +  "arena_get_carry_monitor",
        +  "arena_get_chart",
        +  "arena_get_cost_basis_spread",
        +  "arena_get_cycle_history",
        +  "arena_get_drift_log",
        +  "arena_get_filter_insights",
        +  "arena_get_funding_rate_history",
        +  "arena_get_gem_score",
        +  "arena_get_gem_validation",
        +  "arena_get_halvings",
        +  "arena_get_hash_ribbons",
        +  "arena_get_kimchi_premium",
        +  "arena_get_ma_distance_history",
        +  "arena_get_mayer_multiple",
        +  "arena_get_mayer_multiple_history",
        +  "arena_get_ontology_term",
        +  "arena_get_platform_activity",
        +  "arena_get_portfolio_correlation",
        +  "arena_get_pulse_history",
        +  "arena_get_reference_models",
        +  "arena_get_report_status",
        +  "arena_get_shared_backtest",
        +  "arena_get_signal_events",
        +  "arena_get_signal_status",
        +  "arena_get_taker_imbalance",
        +  "arena_get_trend_channels",
        +  "arena_get_volatility_insights",
        +  "arena_get_volatility_phases",
        +  "arena_get_volatility_recommendations",
        +  "arena_get_volume_profile",
        +  "arena_get_winners",
        +  "arena_list_subscriptions",
        +  "arena_quote_report",
        +  "arena_share_grid_backtest",
        +  "arena_subscribe_bullmarket_stage",
        +  "arena_subscribe_cycle_changes",
        +  "arena_subscribe_pulse_changes",
        +  "arena_subscribe_signal_alerts",
        +  "arena_suggest_grid_range"
        +]
  4. 1 tool update
    • Changedarena_run_grid_backtest2 fields changed
      • changedInput schema / properties / entry_price / description
        Previous value: -"Optional price at which the bot starts; default is the first close in the range."New value: +"Optional price at which the bot starts; default is the OPEN of the first candle (must lie inside low_price..high_price)."
      • changedInput schema / properties / grid_mode / description
        Previous value: -"'neutral' (default): starts half in coins, buys and sells around the entry. 'long': starts 100% in cash, buys dips below the entry, sells on recovery — required for leverage."New value: +"'neutral' (default): buys coins for every level above the entry at the start (the coin share follows the entry's position in the range — measured 2–83 %, NOT a fixed 50/50; see benchmarks.static_allocation.coin_share_start), then buys and sells around the entry. 'long': starts 100% in cash, buys dips below the entry, sells on recovery — required for leverage."
  5. 3 tool updates
    • Changedarena_call_extended1 field changed
      • changedInput schema / properties / tool / enum
        Previous value: -[
        -  "arena_cancel_subscription",
        -  "arena_check_subscription_updates",
        -  "arena_dip_scenario",
        -  "arena_get_altcoin_season_history",
        -  "arena_get_backtest_trades",
        -  "arena_get_btc_macro_correlations",
        -  "arena_get_chart",
        -  "arena_get_cost_basis_spread",
        -  "arena_get_cycle_history",
        -  "arena_get_drift_log",
        -  "arena_get_filter_insights",
        -  "arena_get_funding_rate_history",
        -  "arena_get_gem_score",
        -  "arena_get_gem_validation",
        -  "arena_get_halvings",
        -  "arena_get_hash_ribbons",
        -  "arena_get_kimchi_premium",
        -  "arena_get_ma_distance_history",
        -  "arena_get_mayer_multiple",
        -  "arena_get_mayer_multiple_history",
        -  "arena_get_ontology_term",
        -  "arena_get_platform_activity",
        -  "arena_get_pulse_history",
        -  "arena_get_report_status",
        -  "arena_get_shared_backtest",
        -  "arena_get_signal_events",
        -  "arena_get_signal_status",
        -  "arena_get_taker_imbalance",
        -  "arena_get_trend_channels",
        -  "arena_get_volatility_insights",
        -  "arena_get_volatility_phases",
        -  "arena_get_volatility_recommendations",
        -  "arena_get_volume_profile",
        -  "arena_get_winners",
        -  "arena_list_subscriptions",
        -  "arena_quote_report",
        -  "arena_share_grid_backtest",
        -  "arena_subscribe_bullmarket_stage",
        -  "arena_subscribe_cycle_changes",
        -  "arena_subscribe_pulse_changes",
        -  "arena_subscribe_signal_alerts",
        -  "arena_suggest_grid_range"
        -]New value: +[
        +  "arena_cancel_subscription",
        +  "arena_check_subscription_updates",
        +  "arena_cross_series",
        +  "arena_dip_scenario",
        +  "arena_get_altcoin_season_history",
        +  "arena_get_backtest_trades",
        +  "arena_get_btc_macro_correlations",
        +  "arena_get_chart",
        +  "arena_get_cost_basis_spread",
        +  "arena_get_cycle_history",
        +  "arena_get_drift_log",
        +  "arena_get_filter_insights",
        +  "arena_get_funding_rate_history",
        +  "arena_get_gem_score",
        +  "arena_get_gem_validation",
        +  "arena_get_halvings",
        +  "arena_get_hash_ribbons",
        +  "arena_get_kimchi_premium",
        +  "arena_get_ma_distance_history",
        +  "arena_get_mayer_multiple",
        +  "arena_get_mayer_multiple_history",
        +  "arena_get_ontology_term",
        +  "arena_get_platform_activity",
        +  "arena_get_pulse_history",
        +  "arena_get_reference_models",
        +  "arena_get_report_status",
        +  "arena_get_shared_backtest",
        +  "arena_get_signal_events",
        +  "arena_get_signal_status",
        +  "arena_get_taker_imbalance",
        +  "arena_get_trend_channels",
        +  "arena_get_volatility_insights",
        +  "arena_get_volatility_phases",
        +  "arena_get_volatility_recommendations",
        +  "arena_get_volume_profile",
        +  "arena_get_winners",
        +  "arena_list_subscriptions",
        +  "arena_quote_report",
        +  "arena_share_grid_backtest",
        +  "arena_subscribe_bullmarket_stage",
        +  "arena_subscribe_cycle_changes",
        +  "arena_subscribe_pulse_changes",
        +  "arena_subscribe_signal_alerts",
        +  "arena_suggest_grid_range"
        +]
    • Changedarena_get_etf_flows1 field changed
      • changedInput schema / properties / days / description
        Previous value: -"Length of the returned cumulative series in days. Default 365, clamped 90–1095."New value: +"Length of the returned daily series in days (every US trading day in the window). Default 365, clamped 7–1095."
    • Changedarena_get_stablecoin_supply3 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / days
        Added value: +{
        +  "description": "Length of the returned daily series in days. Default 365, clamped 7–1095.",
        +  "type": "integer"
        +}
  6. 2 tool updates
    • Addedarena_call_extended
    • Addedarena_get_max_pain_history
  7. 1 tool update
    • Changedarena_get_max_pain2 fields changed
      • changedInput schema / properties / market / description
        Previous value: -"Options market: 'DERIBIT_BTC' (default) or 'IBIT' (BlackRock spot-ETF options, collected since 2026-08-24; settlement-timing fields are null until evidenced)."New value: +"Options market. Only DERIBIT_BTC is served: IBIT (BlackRock spot-ETF options) is still collected daily but no longer delivered — the chain comes from an unlicensed source, so it cannot be redistributed (2026-09-24)."
      • changedInput schema / properties / market / enum
        Previous value: -[
        -  "DERIBIT_BTC",
        -  "IBIT"
        -]New value: +[
        +  "DERIBIT_BTC"
        +]
  8. 45 tool updates
    • Removedarena_cancel_subscription
    • Removedarena_check_subscription_updates
    • Removedarena_dip_scenario
    • Removedarena_get_altcoin_season_history
    • Addedarena_get_asset_snapshot
    • Removedarena_get_backtest_trades
    • Removedarena_get_btc_macro_correlations
    • Removedarena_get_chart
    • Removedarena_get_cost_basis_spread
    • Removedarena_get_cycle_history
    • Removedarena_get_drift_log
    • Removedarena_get_filter_insights
    • Removedarena_get_funding_rate_history
    • Removedarena_get_gem_score
    • Removedarena_get_gem_validation
    • Removedarena_get_halvings
    • Removedarena_get_hash_ribbons
    • Removedarena_get_kimchi_premium
    • Removedarena_get_ma_distance_history
    • Removedarena_get_max_pain_history
    • Removedarena_get_mayer_multiple
    • Removedarena_get_mayer_multiple_history
    • Removedarena_get_ontology_term
    • Removedarena_get_platform_activity
    • Removedarena_get_pulse_history
    • Removedarena_get_report_status
    • Removedarena_get_sentiment
    • Removedarena_get_shared_backtest
    • Removedarena_get_signal_events
    • Removedarena_get_signal_status
    • Removedarena_get_taker_imbalance
    • Removedarena_get_trend_channels
    • Changedarena_get_universe2 fields changed
      • addedInput schema / properties / as_of
        Added value: +{
        +  "description": "Optional YYYY-MM-DD. Point-in-time membership on that day (recorded since 2026-07-14, volume-ranked).",
        +  "type": "string"
        +}
      • changedInput schema / properties / universe_id / description
        Previous value: -"Universe id, e.g. 'top-10-crypto'."New value: +"Universe id, e.g. 'crypto-top-50'."
    • Removedarena_get_volatility_insights
    • Removedarena_get_volatility_phases
    • Removedarena_get_volatility_recommendations
    • Removedarena_get_winners
    • Removedarena_list_subscriptions
    • Removedarena_quote_report
    • Removedarena_share_grid_backtest
    • Removedarena_subscribe_bullmarket_stage
    • Removedarena_subscribe_cycle_changes
    • Removedarena_subscribe_pulse_changes
    • Removedarena_subscribe_signal_alerts
    • Removedarena_suggest_grid_range
  9. 1 tool update
    • Changedarena_get_cycle_history1 field changed
      • addedInput schema / properties / view
        Added value: +{
        +  "description": "Default 'scores' (daily composite score rows). 'timeline' = band per indicator per week + consensus. 'flips' = band-change log + consensus.",
        +  "enum": [
        +    "scores",
        +    "timeline",
        +    "flips"
        +  ],
        +  "type": "string"
        +}
  10. 1 tool update
    • Changedarena_run_grid_backtest3 fields changed
      • addedInput schema / properties / funding_mode
        Added value: +{
        +  "description": "Only with leverage > 1. 'conservative' (default): flat 0.05%/day on the open position. 'historical': recorded daily average of three exchanges, BTCUSDT from 2019-09-08 only — otherwise falls back to conservative and flags funding_fell_back_to_conservative. 'none': no funding (optimistic).",
        +  "enum": [
        +    "none",
        +    "conservative",
        +    "historical"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / grid_mode
        Added value: +{
        +  "description": "'neutral' (default): starts half in coins, buys and sells around the entry. 'long': starts 100% in cash, buys dips below the entry, sells on recovery — required for leverage.",
        +  "enum": [
        +    "neutral",
        +    "long"
        +  ],
        +  "type": "string"
        +}
      • addedInput schema / properties / leverage
        Added value: +{
        +  "description": "Optional, default 1 (spot grid, unchanged). 2/3/5 = isolated-margin long grid (grid_mode must be long; Pro). Adds liquidated, liquidation_time/price, funding_cost_usd and max_notional_exposure to the result; final_value/total_return are then on the margin.",
        +  "enum": [
        +    1,
        +    2,
        +    3,
        +    5
        +  ],
        +  "type": "number"
        +}
  11. 1 tool update
    • Changedarena_get_max_pain1 field changed
      • changedInput schema / properties / include_gex / description
        Previous value: -"Default false (response unchanged). DERIBIT_BTC only. When true, each upcoming expiry carries a `gex` block plus `gex_totals` across the whole book: Black-Scholes gamma notional (USD per 1 % spot move) per 2.5 % band from LIVE Deribit mark IV per strike (gex_data_as_of names the fetch, ~10 min cache — a different observation time than the 02:00 UTC snapshot fields). The dealer SIGN is an assumption, not a measurement: both conventions are published side by side (assuming_dealers_short_all, assuming_squeezemetrics_convention) with a zero_gamma_level each; where they disagree, the data does not know the answer. Tau floor 2 h near expiry (tau_clamped flags it); instruments without usable IV are excluded and counted."New value: +"Default false (response unchanged). DERIBIT_BTC only. When true, each upcoming expiry carries a `gex` block plus `gex_totals` across the whole book: Black-Scholes gamma notional (USD per 1 % spot move) per 2.5 % band from LIVE Deribit mark IV per strike (gex_data_as_of names the fetch, ~10 min cache — a different observation time than the 02:00 UTC snapshot fields). The dealer SIGN is an assumption, not a measurement: both conventions are published side by side (assuming_dealers_short_all, assuming_squeezemetrics_convention); where they disagree, the data does not know the answer. zero_gamma_level flips only under the SqueezeMetrics convention — short-all is <= 0 everywhere and has no zero crossing by construction (its null is structural; zero_gamma_level.note says so). Tau floor 2 h near expiry (tau_clamped flags it); instruments without usable IV are excluded and counted."

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Local-first backtesting engine with built-in overfitting detection (PBO, deflated Sharpe, bootstrap CI, walk-forward) and a native MCP server for AI agents to validate trading strategies.
    4
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides tools to research crypto trading strategies via backtesting, walk-forward validation, and paper trading, with a deflated-Sharpe overfitting check. Enables natural-language-driven analysis and interpretation of strategy performance.
    2
    Apache 2.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.