Skip to main content
Glama

TradingCalc MCP: Options, Forex, Risk Stats, Prediction Markets, On-Chain & Crypto Futures

Server Details

Deterministic options, forex, risk, on-chain & futures math. 75 tools. Not AI estimates.

Ownership verified
Status
Healthy
Uptime
100.0% over 47 days
Last Tested
Transport
Streamable HTTP · MCP 2024-11-05
URL
Repository
SKalinin909/tradingcalc-mcp
GitHub Stars
2
Server Listing
TradingCalc MCP Server

TDQS

A3.9/5.0

Scored across 76 tools

Disambiguation3/5

Many tools have overlapping purposes, such as primitive.average_entry vs workflow.run_dca_entry vs workflow.run_forex_average_entry, and primitive.funding_arb vs workflow.run_funding_arbitrage vs workflow.run_carry_trade. The unusually detailed descriptions explicitly disambiguate when to use each, which keeps boundaries mostly clear, but the large number of near-duplicate and live/manual pairs still creates substantial misselection risk.

Naming Consistency5/5

Tool names follow a highly consistent namespace.snake_case pattern: primitive.*, system.*, and workflow.run_* across all 76 tools. The convention is predictable and readable throughout, with no mixing of camelCase or chaotic verb styles.

Tool Count1/5

At 76 tools, the surface is far beyond the typical well-scoped range and exceeds the 50+ threshold for extreme mismatch. Although the server spans multiple domains (options, forex, risk stats, prediction markets, on-chain, futures), many tools are live/manual duplicates or narrow variants, making the count excessive even for a broad calculator.

Completeness5/5

The tool surface covers an extensive range of trading calculators across options pricing, forex, futures PnL/position sizing, risk statistics, prediction markets, and on-chain checks. No obvious gaps in the stated domains are apparent; the server appears complete for its purpose.

Available Tools

76 tools
primitive.average_entryAverage Entry PriceA
Read-only
Inspect

Calculate the weighted average entry price from multiple buy/sell fills (DCA): the bare number only, no breakeven or per-fill breakdown. Use when user asks only "what's my average entry?" and wants just that figure. For breakeven and a per-level summary too, use workflow.run_dca_entry instead. Returns: averagePrice, totalSize, totalCost.

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYes
symbolYesTrading pair symbol, e.g. BTCUSDT
contractTypeNolinear = USDT-margined, average is the arithmetic mean (default). inverse = coin-margined, average is the harmonic mean (fill quantity is USD notional).
exchangeCodeNoExchange identifier (optional)

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered; the description adds useful scope disclosure beyond that — it returns only the bare average figure and deliberately omits breakeven and per-fill breakdown. It also enumerates the returned fields, which annotations do not. It stops short of any caveats about missing/invalid fills, hence not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, with the core purpose front-loaded before the routing alternative and the return shape. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description enumerates the return fields (averagePrice, totalSize, totalCost), and the nested input is documented by the schema. Between description, schema, and annotations, an agent has everything needed to call this pure-computation tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents symbol, contractType (including the arithmetic vs harmonic mean distinction) and the nested fill fields. The description only adds that fills are 'buy/sell' DCA fills and implies price+quantity weighting; it never mentions contractType, which is the parameter with the most consequential semantics. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('calculate the weighted average entry price from multiple buy/sell fills') and immediately scopes it as the bare number. It distinguishes itself from workflow.run_dca_entry by name, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use condition (user asks only "what's my average entry?" and wants just that figure) and names the alternative tool plus the condition that selects it (breakeven and per-level summary too → workflow.run_dca_entry). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

primitive.hedge_ratioHedge RatioA
Read-only
Inspect

Calculate the short perpetual futures position size needed to hedge a spot holding. Use when user asks "how much should I short to hedge my BTC?" or "what margin do I need for a 100% hedge?". Returns: hedgeNotional, requiredMargin, estimatedFundingCost.

ParametersJSON Schema
NameRequiredDescriptionDefault
leverageNoLeverage on the perp short. Default 1.
spotSizeYesSpot position value in USDT
hedgeRatioNoPercentage of spot to hedge, e.g. 100 for full hedge, 50 for half. Default 100.
fundingRatePctNoCurrent 8h funding rate as percentage, e.g. 0.01. Used for cost estimate.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the description's main added value is the disclosure of the three return fields (hedgeNotional, requiredMargin, estimatedFundingCost) — important since no output schema exists. It doesn't note assumptions such as default leverage or funding-rate behavior, so it is not fully rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences: purpose first, usage triggers second, return values last. No filler, and the most decision-relevant content is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-purpose calculation tool with no output schema, the description supplies the return shape and usage triggers, and annotations cover safety. It omits behavioral edge cases (e.g., what happens with omitted optional params or negative/zero spotSize), so it is nearly but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter (spotSize, leverage, hedgeRatio, fundingRatePct) has an inline description with units and defaults. The description adds no syntactic or semantic detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Calculate) and a precise resource: the short perpetual futures position size needed to hedge a spot holding. This is clearly distinguishable from generic siblings like run_position_sizing or run_liquidation_safety, which do not compute a hedge against spot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger phrasing ("how much should I short to hedge my BTC?", "what margin do I need for a 100% hedge?"), which tells the agent the user intents this tool serves. It stops short of naming when not to use it or pointing to an alternative sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

system.pubkeySigning Public KeyA
Read-only
Inspect

Return the ECDSA P-256 public key (PEM + JWK) and canonical signing format used to sign tool responses, so results can be verified offline without calling back to TradingCalc. Every tools/call result includes a signed second content block when signing is configured; also available at GET /api/mcp/pubkey.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: every tools/call result carries a signed second content block when signing is configured, which tells the agent where signatures come from and when they exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the concrete return payload before the usage rationale. Dense but every clause (PEM + JWK, canonical format, offline verification, signed second block, HTTP mirror) carries information; only the HTTP endpoint mention is arguably secondary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey the return value, and it does (PEM + JWK key plus canonical signing format). With annotations covering the safety profile and no parameters to document, an agent has enough to call and use it correctly; a brief note on key rotation or cache lifetime would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so baseline is 4 per the rubric. The description correctly implies no inputs are needed and instead describes the returned key formats, which is the relevant information here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: returns the ECDSA P-256 public key (PEM + JWK) plus the canonical signing format. An agent can immediately distinguish it from siblings like system.verify and the many workflow.run_* tools, since it is the only key-material provider.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives the purpose-driven context for use: verifying signed tool responses offline without calling back to TradingCalc, and notes the alternate access path GET /api/mcp/pubkey. It does not explicitly name system.verify as a complementary/alternative tool, so it stops short of a full when/when-not routing statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

system.verifyVerify CalculationsA
Read-only
Inspect

Run the full regression suite: 43 canonical test vectors (linear and inverse/coin-margined) across all 12 calculators, and return a pass/fail report with counts and timestamp. Useful as a health check before relying on results in production workflows.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so safety is covered. The description adds real behavioral context beyond that: exactly what is exercised (43 vectors across linear and inverse/coin-margined calculators) and what comes back (pass/fail with counts and timestamp).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and scope, then the use case. The parenthetical 'linear and inverse/coin-margined' is genuine scope information, not filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless, read-only tool with no output schema, the description covers everything an agent needs: what runs, what it validates, what is returned (pass/fail, counts, timestamp), and when it is worth calling. An agent can invoke it blind and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so there is nothing to document and the baseline is 4. The description correctly implies the tool is parameterless, running the fixed suite unconditionally.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('run the full regression suite') plus concrete scope (43 test vectors, 12 calculators) and the return shape. This is unmistakably the meta/health-check tool, cleanly distinguishable from every workflow.run_* calculator sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a clear usage context: 'a health check before relying on results in production workflows.' It does not name an explicit alternative or when-not-to-use condition, but the sibling set makes the role unambiguous, so guidance is strong without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_average_downAverage DownA
Read-only
Inspect

Should-I-average-down check: for an already-open position, compares adding more at a worse price against the always-available alternative of buying the same final total size fresh at today's price. Returns the new blended average entry, liquidation price and breakeven (before vs. after), margin_added (the actual cash/margin required for the add at this leverage, not the full notional), and three risk figures at your stop-loss: existing_risk (what you already risk, before adding), pyramid_risk (what you'd risk after adding), and clean_entry_risk (what a fresh entry at add_price for the same total size would risk). risk_penalty_pct is how much MORE than that fresh-entry alternative you're risking - for any genuine average-down (add_price worse than existing_entry_price) pyramid_risk is provably always greater than clean_entry_risk. liquidation_before_stop is true when the new liquidation price sits at or beyond your own stop-loss, meaning the exchange would force-close the position before the stop-loss ever triggers. Use when user asks "should I add to this losing position?" or "what does averaging down actually cost me here?".

ParametersJSON Schema
NameRequiredDescriptionDefault
mmrNoMaintenance margin rate (default 0.005)
sideYes
add_sizeYesAdditional size being considered, same unit convention as existing_size
leverageYesLeverage multiplier
add_priceYesProposed price to add at - also treated as today's current price for the fresh-entry comparison
stop_lossYesStop-loss price (must be below add_price for a long, above it for a short - the position should already be closed otherwise)
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. Risk figures come back denominated in the base coin for inverse.
fee_open_pctNoOpen fee rate (default 0.0002)
existing_sizeYesSize already held: base-asset quantity for linear, USD notional (contracts) for inverse
fee_close_pctNoClose fee rate (default 0.0005)
existing_entry_priceYesEntry price of the position already held

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/non-destructive behavior, but the description adds substantial behavioral context: it explains the outputs, the meaning of risk_penalty_pct, the provable relationship between pyramid_risk and clean_entry_risk, and the liquidation_before_stop condition. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and the dense output explanations are relevant for a tool with no output schema. It is longer than ideal, but nearly every clause adds decision-useful semantics rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter analytical workflow with no output schema, the description carries the return-value burden well: it names and explains the blended average, liquidation/breakeven, margin_added, risk figures, risk_penalty_pct, and liquidation_before_stop. Enough context is provided to call and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 91%, so parameter semantics are already well documented. The description adds only marginal context for add_price, stop_loss, and existing_entry_price, mostly repeating what the schema already says.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact decision-analysis purpose: comparing an average-down add against a fresh entry for the same final size. It specifies the inputs it reasons over and the metrics it produces, which clearly separates it from generic entry/sizing siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit user-phrasing triggers ("should I add to this losing position?" and "what does averaging down actually cost me here?") and scopes the tool to already-open positions. It does not explicitly name sibling alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_black_scholesBlack-Scholes Option Price and GreeksA
Read-only
Inspect

Theoretical European option price and Greeks (delta, gamma, theta, vega, rho) from Black-Scholes, given manual spot/strike/days-to-expiry/volatility/risk-free-rate inputs: no live data fetch. The live variant is workflow.run_black_scholes_live, which applies when checking a real Deribit BTC/ETH instrument, since that variant also reports how far the instrument's actual quoted price sits from what this formula implies. Prices are USD-denominated (the universal convention); callPriceCoin/putPriceCoin additionally divide by spot to match Deribit's own coin-settled quoting convention. Use when user asks "what should this option be worth at X% IV?" or wants raw Greeks for a hypothetical. Returns: callPriceUsd/putPriceUsd, callPriceCoin/putPriceCoin, deltaCall/deltaPut, gamma, vegaPerPct (per 1 vol point), thetaCallPerDay/thetaPutPerDay, rhoCallPerPct/rhoPutPerPct (per 1 rate point).

ParametersJSON Schema
NameRequiredDescriptionDefault
spotYesUnderlying spot price, USD
strikeYesStrike price, USD
daysToExpiryYesCalendar days until expiry (can be fractional)
volatilityPctYesAnnualized implied volatility in percentage points, e.g. 60 for 60%
riskFreeRatePctNoRisk-free rate in percentage points. Default 0: standard crypto-options convention.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond that: no live data fetch (pure manual-input computation) and the coin-denomination convention that divides price by spot. It stops short of noting edge cases (e.g. zero/negative time to expiry, vol validation), so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is dense and front-loaded — the core purpose, the sibling routing, and the unit convention all appear before the return list. The final sentence is a long output enumeration, but since there is no output schema it earns its place; the prose is slightly heavy overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description usefully enumerates return fields and clarifies per-unit conventions (vegaPerPct per 1 vol point, theta per day, rho per 1 rate point) and USD vs coin quoting. Combined with annotations covering safety, an agent has enough to call it correctly; minor gaps remain around input validation/edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (spot, strike, daysToExpiry, volatilityPct, riskFreeRatePct) is already documented, including volatility units and the risk-free default of 0. The description restates the input set rather than adding new syntax or constraint detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific method (Black-Scholes), the exact resource (theoretical European option price and Greeks), and enumerates the output set. It explicitly distinguishes itself from run_black_scholes_live, so an agent can route between the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the concrete trigger ("what should this option be worth at X% IV?", raw Greeks for a hypothetical), the alternative (run_black_scholes_live), and the condition that selects the alternative (checking a real Deribit BTC/ETH instrument with a quoted price). When-to-use and when-not are both covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_black_scholes_liveBlack-Scholes Option Price and Greeks (Live)A
Read-only
Inspect

Black-Scholes theoretical price and Greeks for a REAL, live Deribit BTC/ETH option instrument: pulls that instrument's own spot, strike, days to expiry, and implied volatility from Deribit, then reports how far Deribit's actual quoted mark price sits from what Black-Scholes implies at that IV (priceDiscrepancyPct). This is the trust-check tool: "is this exchange's quoted price consistent with its own volatility assumption?", not an estimate, a live formula cross-check. Use when user gives a specific instrument name (e.g. "BTC-27FEB27-90000-C") and asks "is this option fairly priced?" or "what are the Greeks on this contract?". Returns everything workflow.run_black_scholes does, plus instrumentName, currency, optionType, deribitMarkPriceCoin, deribitMarkIvPct, bsPriceCoin, priceDiscrepancyPct, available (false + error if the instrument name doesn't resolve).

ParametersJSON Schema
NameRequiredDescriptionDefault
instrumentNameYesExact Deribit instrument name, e.g. BTC-27FEB27-90000-C. Get one from Deribit's options chain; this tool does not browse the chain, it prices one named instrument.
riskFreeRatePctNoRisk-free rate in percentage points. Default 0: standard crypto-options convention.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint false, openWorldHint). The description goes beyond them by disclosing data provenance (spot/strike/DTE/IV pulled live from Deribit), the semantic intent of the discrepancy metric, and the failure mode ('available: false + error if the instrument name doesn't resolve'). It stops short of stating rate limits or freshness/latency caveats for the live pull.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and the trust-check framing, and every clause carries information. The trailing enumeration of return fields is a long run-on list, but given there is no output schema it is mostly earning its place rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by naming the returned fields (instrumentName, currency, optionType, deribitMarkPriceCoin, deribitMarkIvPct, bsPriceCoin, priceDiscrepancyPct, available) and the error path. Combined with clear usage triggers and a 100%-covered schema, an agent has everything needed to call and interpret it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the parameter documentation burden is already met; the description adds context beyond that by giving a concrete instrument format and clarifying that the name must come from an external options chain. It does not elaborate on riskFreeRatePct semantics, but the schema's default-0 convention note covers that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (Black-Scholes theoretical price and Greeks for a live Deribit instrument) and explicitly positions itself against its nearest sibling: 'Returns everything workflow.run_black_scholes does, plus...'. It also names the shape of the answer (priceDiscrepancyPct trust-check), so an agent can distinguish it from both the static BS tool and adjacent option workflows like run_implied_volatility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions and quoted user phrasings ('is this option fairly priced?', 'what are the Greeks on this contract?') plus the required precondition that the user supply a specific instrument name. It also states what this tool does NOT do ('this tool does not browse the chain, it prices one named instrument'), which routes the agent away from misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_bonding_curveBonding CurveA
Read-only
Inspect

Pump.fun-style bonding curve calculator: exact tokens received for a buy, price impact, and graduation progress. Pure constant-product math (Uniswap V2 style) using pump.fun's official virtual-reserve constants: no live lookup needed, works for any token still on the curve (not yet graduated to a real AMM pool). Use when user asks "how many tokens do I get buying X SOL on this curve?" or "will this buy graduate the token?". Returns: tokensOut, priceImpactPct, progressPctBefore/After, willGraduate, partialFill (true if the buy exceeds remaining curve capacity).

ParametersJSON Schema
NameRequiredDescriptionDefault
solToSpendYesSOL amount for this buy
solRaisedSoFarYesSOL already raised on the curve so far (0 for a brand-new token)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false. The description adds real value beyond that: no live lookup is required (pure constant-product math with pump.fun's official virtual-reserve constants) and it discloses the partialFill edge case when a buy exceeds remaining capacity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core identity and mechanics before the usage examples and return list. The return-field enumeration is dense but every element earns its place given there is no output schema; slightly list-heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description enumerates the return values (tokensOut, priceImpactPct, progressPctBefore/After, willGraduate, partialFill) and explains partialFill's meaning. Inputs, scope, and outputs are all covered for a two-parameter pure-math tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters are documented there, so baseline 3 applies. The description clarifies the model context (virtual-reserve constants) but adds no format, range, or unit detail for solToSpend or solRaisedSoFar beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Pump.fun-style bonding curve calculator') with sub-capabilities named: exact tokens out, price impact, graduation progress. It clearly distinguishes itself from the many AMM/pool siblings by scoping to tokens 'still on the curve (not yet graduated to a real AMM pool)'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrasing ('how many tokens do I get buying X SOL on this curve?', 'will this buy graduate the token?') plus an explicit scope boundary: works for curve tokens, not graduated AMM pools. An agent can route to it without inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_breakeven_planningBreakeven PlanningA
Read-only
Inspect

Calculate the break-even exit price that covers all trading fees: this alone, nothing else. Use when user asks only "what price do I need to just break even?" and nothing more. If the user also gave a stop/target or wants a full trade-safety check, use workflow.run_risk_reward or workflow.run_pre_trade_check instead; both already include this breakeven figure plus more. Returns: breakevenPrice, totalFees.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
sizeBaseYesPosition size: base asset qty for linear, USD contracts for inverse
entryPriceYesEntry price (positive)
feeOpenPctNoOpening fee fraction, default 0.0002
feeClosePctNoClosing fee fraction, default 0.0005
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. For inverse, totalFees is returned in the base coin.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe, read-only, closed-world computation, so the safety profile is covered. The description adds genuine value by disclosing the returned fields (breakevenPrice, totalFees), which matters because there is no output schema. It does not discuss fee defaults or numeric behavior, but that is largely delegated to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then the usage rule, then the alternative routing, then the returns. Every sentence earns its place and nothing is redundant with structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the return fields. Combined with explicit scope boundaries and sibling routing, an agent has everything needed to select and invoke the tool correctly for a 6-parameter computation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 83%, so the schema already documents side, sizeBase, entryPrice, and the fee/contract-type parameters well. The description adds no syntax, units, or default information beyond what the schema provides ('covers all trading fees' only hints at the fee params). Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('calculate the break-even exit price that covers all trading fees') and immediately constrains scope ('this alone, nothing else'). It also names the sibling tools it must not be confused with, so an agent can distinguish it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('user asks only what price do I need to just break even') and when-not-to-use ('if the user also gave a stop/target or wants a full trade-safety check'), naming workflow.run_risk_reward and workflow.run_pre_trade_check as the alternatives. Routing is fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_breakout_acceptanceMarket Profile: Breakout AcceptanceA
Read-only
Inspect

Market Profile breakout acceptance: did price accept (hold) beyond the value area / range, or reject back inside (fakeout)? Optional buy/sell delta. Use for "did the break above VAH get accepted?". Returns: state, accepted (boolean), direction, confidence, key_levels (VAH/VAL/VPOC), scenario_framing, invalidation level.

ParametersJSON Schema
NameRequiredDescriptionDefault
venueYesExchange to fetch candles from when candles[] not supplied
candlesNoOptional OHLCV for the session; omit to fetch from venue (reproducible + 0 COGS when supplied)
timeframeNoCandle timeframe (default 15m)
instrumentYesSymbol, e.g. BTCUSDT
prev_candlesNoOptional OHLCV for the previous session
session_dateYesSession date YYYY-MM-DD (UTC)
include_deltaNoInclude buy/sell delta analysis (default true)
value_area_ruleNoValue-area fraction 0.5–0.9 (default 0.70)

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description then adds genuinely new context: the return payload (state, accepted, direction, confidence, key_levels with VAH/VAL/VPOC, scenario_framing, invalidation level) and the delta option. It does not discuss cost/reproducibility tradeoffs of the candle-fetch path, which is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is front-loaded with the concept and the accept-vs-fakeout framing, then the example, then the return contract. Every clause carries information; the return-field list is dense but functional. Slight compression could improve it, but there is no wasted filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analytical tool with a fully documented input schema and no output schema, the description covers purpose, invocation context, and return shape, which is what an agent needs to call it. The absence of a stated alternative for overlapping 'breakout' questions is the only notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all eight parameters, including enums and defaults. The description only echoes the optional buy/sell delta concept, adding little syntax or format meaning beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific analytical concept (Market Profile breakout acceptance) and defines it precisely: did price hold beyond the value area/range or reject back inside (fakeout). This is far more specific than a generic verb+resource and clearly separates it from unrelated siblings like run_open_analysis. It stops short of naming a sibling it could be confused with, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The quoted example question ("did the break above VAH get accepted?") implies the triggering scenario, which is useful. However there is no explicit when-to-use vs when-not, no prerequisites, and no pointer to an alternative workflow for related questions (e.g., session structure). Usage is implied, not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_carry_tradeCarry TradeA
Read-only
Inspect

Delta-neutral carry trade (funding arbitrage) analysis, with a profitable/marginal/loss verdict on top of the same math primitive.funding_arb uses. Compared with primitive.funding_arb, this one adds output for the case where a plain-English verdict is wanted, not just the raw numbers. Use when user asks "is this carry trade worth it?": long on exchange A, short on exchange B, collect the funding rate spread. Returns: netYieldPct, grossProfit, netProfit, breakevenDays, verdict (profitable/marginal/loss).

ParametersJSON Schema
NameRequiredDescriptionDefault
notionalYesPosition notional in USDT
hold_daysYesHold duration in days
interval_hoursNoFunding interval: 1 or 8 hours (default 8)
transfer_fee_pctNoOne-way transfer fee % (default 0.1)
funding_rate_longYesFunding rate on long exchange per interval (decimal)
funding_rate_shortYesFunding rate on short exchange per interval (decimal)

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description usefully confirms this is a deterministic computation layered on an existing primitive and enumerates the computed outputs, though it says nothing about edge cases (e.g., zero/negative spread handling).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and trigger, but the 'compared with primitive.funding_arb' clause restates the first sentence, the math-primitive reference is an awkward sentence fragment, and the Returns list partly duplicates what the schema/verdict already implies.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly fills the gap by naming the returned fields (netYieldPct, grossProfit, netProfit, breakevenDays, verdict). It omits any note on expected input units beyond the schema or on what drives a 'marginal' classification, but is otherwise sufficient to call the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters (including the interval_hours enum and fee defaults) are already documented in the schema. The description only gestures at the funding-rate spread concept and adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific analysis (delta-neutral carry trade / funding arbitrage) plus its distinguishing output (profitable/marginal/loss verdict), and explicitly names the sibling whose math it reuses. An agent can tell it apart from primitive.funding_arb and from workflow.run_funding_arbitrage without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger ('use when user asks "is this carry trade worth it?"') and a discriminator against primitive.funding_arb (plain-English verdict vs raw numbers). The named alternative, however, is not in the sibling list shown, so the routing hint is slightly off-target and there is no explicit when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_cointegrationCointegration TestA
Read-only
Inspect

Engle-Granger two-step cointegration test for a pair of price series: do they share a long-run equilibrium relationship (their spread is stationary/mean-reverting)? The standard pairs-trading signal test. Returns the cointegrating regression's hedge ratio (beta) and an ADF t-statistic on the residuals, compared against 1%/5%/10% critical values. Use when user asks "are these two assets cointegrated?" or "is this a valid pairs trade?". Returns: alpha, beta (hedge ratio), t_stat, critical_values, cointegrated (booleans at each significance level).

ParametersJSON Schema
NameRequiredDescriptionDefault
xYesSecond price series, same length and dates as y
yYesFirst price series (the dependent variable in the cointegrating regression)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description adds genuine behavioral context beyond that: the two-step method, that residual ADF stats are compared to 1/5/10% critical values, and the exact returned fields (alpha, beta, t_stat, critical_values, cointegrated).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the method and the question it answers, then the return payload. Slightly verbose with the rhetorical parenthetical, but every clause carries information an agent would want.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description responsibly enumerates the returned fields and the interpretation threshold logic, and the required parameters are fully documented in the schema. An agent can call and interpret this tool without further reference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with only two parameters, and the schema descriptions already state the role of y (dependent variable) and x (second series, same length/dates). The description adds no format or ordering guidance beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific statistical procedure (Engle-Granger two-step cointegration test) applied to a specific resource (a pair of price series), and frames the exact question it answers (shared long-run equilibrium / stationary spread). This distinguishes it cleanly from siblings like run_forex_correlation and primitive.hedge_ratio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrasing: 'are these two assets cointegrated?' and 'is this a valid pairs trade?' — a clear usage context. It does not, however, name sibling alternatives or state when not to use it (e.g. vs a plain correlation tool), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_compound_fundingCompound FundingA
Read-only
Inspect

Project capital growth from reinvesting perpetual futures funding income (compounding carry). Use when user asks "how much will I make compounding 0.01% funding for 90 days?" or "what's my APY on this carry position?". Returns: finalCapital, totalEarned, apy, growthTable.

ParametersJSON Schema
NameRequiredDescriptionDefault
reinvestPctNoPercentage of earnings reinvested each interval. 100 = full compounding, 0 = no reinvestment. Default 100.
durationDaysYesNumber of days to project
intervalHoursNoFunding interval: 8 (standard) or 1 (Hyperliquid)
fundingRatePctYesFunding rate per interval as percentage, e.g. 0.01 for 0.01%
initialCapitalYesStarting capital in USDT

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds value beyond that by naming the return payload (finalCapital, totalEarned, apy, growthTable), which matters since no output schema exists. It does not mention rounding, interval assumptions, or error behavior, so it is short of excellent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, front-loaded with the operation and followed by usage triggers and return keys. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only projection tool with no output schema and a fully documented 5-param schema, the description covers purpose, triggers, and the four return keys, which is what an agent needs to invoke it correctly. Minor gaps remain around assumptions (e.g. that funding rate is held constant across intervals), preventing a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema, including the 8/1 interval enum and the reinvestPct default. The description only echoes units through its example ('0.01% funding for 90 days'), adding no real semantics beyond the schema. Baseline 3 applies when the schema carries the load.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Project capital growth from reinvesting perpetual futures funding income (compounding carry)'), which cleanly separates it from nearby siblings like run_carry_trade, run_funding_cost, and run_funding_arbitrage. An agent can identify the operation without reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides two concrete trigger utterances ('how much will I make compounding 0.01% funding for 90 days?' and 'what's my APY on this carry position?') that map directly to the intended use case, which is strong context. It stops short of naming when-not-to-use or pointing at an alternative sibling such as run_funding_breakeven, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_covered_call_protective_putCovered Call and Protective PutA
Read-only
Inspect

Covered call (long the coin + short a call against it, for yield) or protective put (long the coin + long a put, for downside insurance) on a Deribit BTC/ETH position. Returns the standard annualized-yield metric (premium ÷ 1 coin, annualized by 365/daysToExpiry) up front: that number doesn't depend on any price scenario. Also returns the USD value of the combined position at a given scenario price: covered call caps upside at strike + premium×scenarioPrice (the premium's own coin-denominated value still scales with price, unlike a textbook USD-settled cap); protective put floors value at strike×(1−premium), which is the true minimum across every possible settlement price, not just an approximation. Use when user asks "what annualized yield do I get selling covered calls on my BTC?" or "how much does insuring my BTC with a put cost me?". Returns: staticYieldPct, annualizedYieldPct, valueAtScenarioUsd, breakevenPrice (covered_call only), floorValueUsd (protective_put only), vsHoldingUsd (vs. just holding the coin).

ParametersJSON Schema
NameRequiredDescriptionDefault
strikeYesOption strike, USD
currencyNoUnderlying coin. Default BTC.
quantityYesCoin units held / contracts (1:1 covered)
strategyYes
spotEntryYesPrice you acquired/value the underlying coin at, USD; used for the covered-call breakeven vs. cost basis
premiumCoinYesPremium received (covered_call) or paid (protective_put) per contract, in the base coin
daysToExpiryYesCalendar days until expiry; used to annualize the yield/cost
scenarioPriceYesUnderlying price in USD to evaluate the combined position's value at

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this read-only, non-destructive and closed-world, but the description adds real behavioral context beyond them: the yield metric is price-scenario independent, covered-call upside is capped by a coin-denominated premium that still scales with price, and the protective-put floor is the true minimum across all settlement prices. That is substantive disclosure the annotations do not carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the strategy definition and the key yield metric before the scenario math, so the agent gets the gist immediately. The dense parenthetical formulas are informative but make it longer than strictly necessary, keeping it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the return fields (staticYieldPct, annualizedYieldPct, valueAtScenarioUsd, breakevenPrice, floorValueUsd, vsHoldingUsd). All 7 required params are documented, and the two strategy modes are explained, so an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (88%), so the baseline is 3, but the description adds meaning the schema text lacks — it explains that the annualization uses 365/daysToExpiry, that spotEntry drives the covered-call breakeven, and that premium is coin-denominated and scales with scenarioPrice. This lifts it slightly above the schema-only baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource — computing covered-call yield / protective-put protection value for a Deribit BTC/ETH position — and explicitly distinguishes the two strategies it handles. An agent can tell it apart from siblings like run_options_payoff or run_spread_payoff from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete trigger phrases ("what annualized yield do I get selling covered calls on my BTC?", "how much does insuring my BTC with a put cost me?"), which clearly establishes when to reach for it. It does not, however, name alternative tools or state when-not-to-use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_cross_venue_arbitrageCross Venue ArbitrageA
Read-only
Inspect

Checks 2-5 quotes for the SAME real-world binary bet across venues (or manually-supplied probabilities) for a guaranteed, direction-independent arbitrage: buy "Yes" at whichever venue quotes it cheapest, buy "No" at whichever venue quotes "Yes" most expensively (its own "No" price is assumed to be 1 minus its own "Yes" price, the standard complementary-binary convention). Reports isArbitrage, the cost to lock in $1 of guaranteed payout, guaranteed profit and ROI for a given stake, and which side to buy where. feePctPerLeg is an optional per-leg trading-fee rate that can and does erase a real-looking spread - reported honestly via isArbitrage rather than always showing a positive number. Use when user asks "can I arbitrage this bet across venues?" or "is there a risk-free profit here?". The caller asserts the venues quote the same bet; this tool does not verify that.

ParametersJSON Schema
NameRequiredDescriptionDefault
quotesNoLive mode: one entry per venue, 2-5 total. Provide this OR manualProbabilitiesPct, not both.
stakeUsdYesTotal capital to deploy across both legs
feePctPerLegNoOptional per-leg trading-fee rate, e.g. 0.02 for 2% (default 0)
manualProbabilitiesPctNoManual mode: 2-5 probabilities (0.01-99.99) already known for the same bet across different venues, when you don't want a live fetch. Provide this OR quotes, not both.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare read-only/no-destruct/open-world; the description adds genuinely load-bearing behavior: the complementary-binary assumption for the 'No' price, that fees can erase a spread and this is reported honestly via isArbitrage, and the critical caveat that the tool does not verify the venues quote the same bet. That last point is exactly the kind of limitation an agent must know before invoking.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then the convention assumption, then the return fields, then the trigger phrase and the caveat. Dense but every clause carries information; the parenthetical about the 'No' convention is long but necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return burden and does so: it names isArbitrage, cost per $1 payout, guaranteed profit, ROI, and which side to buy where. Combined with the unverified-same-bet caveat, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: feePctPerLeg's ability to flip isArbitrage to false, and the semantic intent of manualProbabilitiesPct (bypassing a live fetch). quotes/stakeUsd are largely restated from the schema, keeping this short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('checks 2-5 quotes ... for a guaranteed arbitrage') and pins the domain precisely: same real-world binary bet across prediction-market venues. It is unmistakably distinct from neighbors like run_prediction_market_edge or run_funding_arbitrage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes on user intent ('Use when user asks "can I arbitrage this bet across venues?"'), and clarifies the two mutually exclusive input modes (live quotes vs manual probabilities). It lacks an explicit when-not, but the scope is clear enough that an agent can pick it confidently.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_dca_entryDCA EntryA
Read-only
Inspect

DCA entry planner: weighted average entry price, breakeven, and per-level contribution from multiple fill prices and sizes. Compared with primitive.average_entry, this one adds output for the case where breakeven or the per-level breakdown is also wanted, not just the bare average. Use when user bought at several prices and asks "what's my average entry?" or "where is my DCA breakeven?". Returns: averageEntry, breakeven, per-level summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
entriesYes
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. Each fill's size is USD notional (contracts) for inverse; averageEntry is then the harmonic mean of fill prices, not the arithmetic mean.
fee_open_pctNoOpen fee rate (default 0.0002)
fee_close_pctNoClose fee rate (default 0.0005)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so the safety profile is covered. The description adds the return shape, which is useful, but says nothing about auth needs, compute cost, or failure modes for an inverse-contract edge case it hints at. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the capability, then differentiation, then trigger phrases and returns. The middle sentence is slightly wordy ('output for the case where... is also wanted') but every sentence carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Read-only, no output schema, yet the description enumerates the return fields (averageEntry, breakeven, per-level summary), which compensates for the missing schema. The main residual gap is the undocumented required 'side' parameter and edge-case behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60% and the description adds no parameter meaning at all. It never mentions the required 'side' (long/short) parameter, nor contractType or the fee rates, and 'multiple fill prices and sizes' merely restates 'entries'. One point of clear guidance was available here (e.g. when side flips breakeven) and is absent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('DCA entry planner') and enumerates the exact outputs (weighted average entry price, breakeven, per-level contribution). It explicitly contrasts itself with the sibling primitive.average_entry, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit selection rule ('use when breakeven or the per-level breakdown is also wanted, not just the bare average') and names the alternative tool plus the condition that selects it. Also supplies concrete trigger utterances for when to invoke.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_dsrDeflated Sharpe RatioA
Read-only
Inspect

Deflated Sharpe Ratio (Bailey & Lopez de Prado): given how many strategy variants you tried (and how correlated they are), what Sharpe ratio would the best of N clear by luck alone, and does your actual strategy still clear that higher bar? Use when user asks "is my backtested Sharpe ratio real, or did I get lucky trying many variants?" or "how many independent trials does this really represent?". Provide either trial_sharpes[] (the N trials' own observed Sharpe ratios, most rigorous) or n_trials (+ optional avg_correlation to correct for correlated trials via Kish's design effect). Returns: expected_max_sharpe (the luck-alone threshold), dsr (probability your strategy's true Sharpe exceeds it, 0-1), n_trials_effective.

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsYesThe candidate strategy's own return series
n_trialsNoNumber of strategy variants tried, if trial_sharpes were not tracked individually
trial_sharpesNoObserved Sharpe ratios of all N trials tried, if tracked. Takes priority over n_trials if both are given.
avg_correlationNoAverage pairwise correlation between trials, 0-1 (default 0 = independent). Only used with n_trials; lowers the effective trial count via Kish's design effect.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safe-read profile (readOnlyHint, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds substantive method context: how correlated trials are handled (Kish's design effect) and the priority rule between trial_sharpes and n_trials. It doesn't state the trial_sharpes-priority caveat is enforced by the tool, but the intent is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads method and the core question, then usage, then inputs, then return values. Efficient for a statistically dense tool, though the run-on sentence in the usage trigger is slightly heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description supplies the return fields (expected_max_sharpe, dsr, n_trials_effective) with plain-language meaning and the 0-1 range for dsr. For a four-parameter statistical tool with no output schema, this is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already carries parameter semantics; baseline is 3. The description earns above baseline by explaining the conditional relationship (avg_correlation only meaningful with n_trials, lowers effective count) and asserting trial_sharpes takes priority over n_trials — semantics beyond the field descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific statistical operation (deflating a Sharpe ratio for multiple trials) with the exact question it answers, and cites the method authors. Clearly distinguishable from sibling workflow.run_sharpe_stats, which computes plain Sharpe stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit trigger questions ('is my backtested Sharpe ratio real, or did I get lucky') and an explicit selection rule between the two input modes (trial_sharpes most rigorous, else n_trials + avg_correlation). Nothing left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_evt_tail_riskEVT Tail RiskA
Read-only
Inspect

Extreme Value Theory tail risk (Peaks-Over-Threshold): fits a Generalized Pareto Distribution to the losses beyond a high threshold via Grimshaw's (1993) profile-likelihood MLE, then extrapolates VaR/Expected Shortfall at the requested confidence, without assuming a normal distribution. Use when user asks "what's my tail VaR without assuming normality?" or "how fat is my loss tail, really?". Complements workflow.run_var_cvar (parametric, normal-distribution VaR/CVaR) for exactly the fat-tailed-return case that assumption understates. threshold_percentile (default 90) sets which percentile of the loss distribution (losses = -returns) becomes the threshold u; confidence must be deep enough into the fitted tail (1-confidence < the threshold's own exceedance rate) or the call throws. Returns: threshold, n_exceedances, exceedance_rate, xi (GPD shape: 0=exponential tail, >0=heavy/fat tail, <0=bounded tail; values <= -1 are excluded from the fit domain as a known non-regular/unbounded-likelihood case, Smith 1985), beta (GPD scale), xi_asymptotically_normal (false when xi<=-0.5: the fit is still valid but the usual MLE confidence-interval theory doesn't apply, per that same Smith 1985 result), var, es (null when xi>=1, where Expected Shortfall is mathematically undefined).

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsYesReturn series, one value per period, at least 100 values (POT needs real sample depth in the tail)
confidenceNoVaR/ES confidence level, 0-1 exclusive. Default 0.99. Must satisfy 1-confidence < the threshold's own exceedance rate.
threshold_percentileNoPercentile (0-100 exclusive) of the loss distribution used as the POT threshold u. Default 90.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnly, non-destructive), so the burden is lower, yet the description still discloses real behavioral traits: the confidence-vs-threshold feasibility constraint that triggers an exception, and edge-case semantics for xi, xi_asymptotically_normal, and es becoming null when xi>=1. It omits runtime/cost characteristics, but the error and domain conditions are unusually well documented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation and use case, which is good, but the trailing 'Returns:' block is a long, dense enumeration of six fields with heavy parenthetical citations. It is information-dense but borders on over-specification and would be better absorbed by an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter, no-output-schema tool, this covers method, use case, sibling differentiation, parameter meaning, error conditions, and return fields — an agent has what it needs to call it correctly. Not perfect only because return semantics are embedded in prose rather than a schema, and expected input format conventions (e.g., ordered series, sampling frequency) are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents all three parameters, so the baseline is 3. The description goes beyond by explaining the domain meaning of threshold_percentile (percentile of the loss distribution, losses = -returns) and the cross-parameter constraint on confidence, adding interpretation rather than restating format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb and resource (fits a GPD to tail losses and extrapolates VaR/ES) and distinguishes itself from the sibling workflow.run_var_cvar by contrasting the distributional assumptions. An agent can route between the two without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit triggering user phrasings ('what's my tail VaR without assuming normality?') and names the alternative tool (workflow.run_var_cvar) together with the condition that selects it — the fat-tailed case that normal parametric VaR understates. When-to-use and when-to-prefer-the-sibling are both present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_exit_targetExit TargetA
Read-only
Inspect

Calculate the exact exit price needed to hit a target PnL or ROE percentage. Use when user asks "at what price do I take profit to make $500?" or "where should I set TP for 20% ROE?". Returns: targetExitPrice.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
leverageYesLeverage multiplier
sizeBaseYesPosition size in base asset
entryPriceYesEntry price
feeOpenPctNoOpening fee fraction, default 0.0002
targetModeYes"pnl" = target in USDT, "roe" = target in %
feeClosePctNoClosing fee fraction, default 0.0005
targetValueYesTarget value (USDT/coin for pnl mode, or %)
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. For inverse, pnl-mode targetValue and outputs are in the base coin.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so the safety profile is covered. The description adds only the return field, and does not disclose fee-handling behavior (feeOpenPct/feeClosePct defaults) that meaningfully affects results. Adequate but not rich context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, usage examples second, return value last. Zero waste; every sentence earns its place, and the return declaration compensates for the missing output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-param calculation tool with no output schema, the description covers purpose, triggering conditions, and the returned field, while annotations carry safety and the schema carries parameters. Only nuanced behavior (fee effects on the result) is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 89%, so the schema already documents nearly all parameters, including the targetMode and contractType enums and fee defaults. The description adds no syntax or semantic detail beyond what the schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: 'Calculate the exact exit price needed to hit a target PnL or ROE percentage.' This is clearly distinct in intent from breakeven/scale-out siblings, but no sibling is named explicitly to route the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete user-question triggers ('at what price do I take profit to make $500?', 'where should I set TP for 20% ROE?'), which clearly signal when to use it. No when-not guidance or named alternatives (e.g. breakeven planning) are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_average_entryForex Average EntryA
Read-only
Inspect

Size-weighted average entry price across multiple forex fills: plain arithmetic mean, since forex has no coin-margined analog requiring the harmonic mean the crypto average_entry tool uses for inverse contracts. Use when user asks "what's my average entry after these fills?". Returns: totalUnits, totalCost, avgEntry.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
fillsYesFills to average, each with a price and a size in base-currency units

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, non-destructive, and non-open-world behavior. The description adds meaningful context beyond that: it specifies the calculation method (plain arithmetic mean) and lists the exact return fields (totalUnits, totalCost, avgEntry). It does not address edge cases like empty fills, but for a simple averaging read operation the added computation and return detail is substantial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core operation, follows with the rationale that differentiates it from the crypto sibling, then gives the trigger and return fields. Every clause earns its place and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read-only averaging tool with full schema coverage and annotations, the description is complete: it explains the computation, provides a usage trigger, and enumerates the return values. No output schema exists, so specifying the return fields fills the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents both 'pair' (BASE/QUOTE format) and 'fills' (price and units). The description does not add further parameter syntax, constraints, or formatting details, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Size-weighted average entry price across multiple forex fills.' It explicitly distinguishes itself from the sibling crypto average_entry tool by contrasting arithmetic and harmonic means for inverse contracts, so an agent can tell them apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger phrase: 'Use when user asks what's my average entry after these fills?'. It also implicitly sets boundaries by noting forex has no coin-margined analog, pointing to the crypto average_entry tool for that case. However, it does not explicitly state when not to use this tool or name the alternative as a direct replacement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_breakevenForex BreakevenA
Read-only
Inspect

Breakeven price for a forex position accounting for spread and round-trip commission, in pips and in price. Commission is quoted per standard lot (100,000 units) and expressed in the pair's own quote currency; because both commission and pip value scale with lot size, the commission-in-pips figure is independent of position size by construction. Use when user asks "where's my true breakeven after spread and commission?". Returns: pipSize, commissionPips, totalCostPips, breakevenPrice.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
sideYes
entryPriceYes
spreadPipsYesSpread at entry, in pips
commissionPerLotRoundTripNoRound-trip commission per standard lot, in the pair's own quote currency. Default 0 (pure-spread broker model).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a read-only, non-destructive, closed-world calculation, and the description adds real behavioral context beyond that: commission is per standard lot in the quote currency and the commission-in-pips result is independent of position size by construction. It also previews the return shape. Does not discuss error cases or assumptions on spread handling, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the definition, then the domain nuance, then a 'Returns:' list. The commission-scaling sentence is dense but earns its place by explaining why the pip figure is size-independent. Slightly overpacked but no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description compensates by naming the returned fields (pipSize, commissionPips, totalCostPips, breakevenPrice). Inputs are adequately framed. Complete enough to call correctly, though edge cases (e.g. missing commission default behavior) are only implied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 60% schema coverage, the description usefully explains the semantics of commissionPerLotRoundTrip (per standard lot, pair's quote currency, default of 0 implied) and how it interacts with lot size. It adds meaning over the bare schema. Schema already documents pair, spreadPips, and the side enum, so it does not need to repeat those.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: computes forex breakeven price accounting for spread and round-trip commission, quantified in pips and price. This clearly separates it from generic computation siblings, though it never names a specific sibling (e.g. run_forex_pip_value or run_breakeven_planning) to disambiguate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit trigger: 'Use when user asks where's my true breakeven after spread and commission?'. This is clear usage context, but there are no exclusions or stated alternatives among the many run_forex_* / breakeven siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_correlationForex CorrelationA
Read-only
Inspect

Correlation coefficient and minimum-variance hedge ratio between two price series of matching length, computed on daily % returns (not raw price levels, which would give spuriously high correlation between two unrelated but both-trending series). hedgeRatio follows Hull's standard futures-hedging formula: cov(returns1,returns2)/var(returns2), "how many units of series 2 per unit of series 1 minimizes the combined position's variance." No live data fetch: supply the two price series directly. The live variant is workflow.run_forex_correlation_live, which has both series fetched automatically for two named pairs. Use when user already has two price series and wants their statistical relationship. Returns: n, correlation (-1 to 1), hedgeRatio.

ParametersJSON Schema
NameRequiredDescriptionDefault
rates1YesFirst price series, chronological order, one price per date
rates2YesSecond price series, chronological order, same length and same dates as rates1

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, openWorldHint=false, so safety is covered. The description adds genuinely non-obvious behavior: computation is on daily % returns rather than raw levels, the hedgeRatio uses Hull's cov/var formula, and no live fetch occurs. It stops short of 5 only because it doesn't discuss edge cases (e.g., zero variance, minimum series length enforcement) that an agent invoking with borderline inputs would want.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation and the key no-fetch constraint, and the return values are summarized last. The parenthetical justification of % returns and the quoted formula explanation are slightly verbose, but each sentence carries information an agent would otherwise have to guess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a pure-computation tool with no output schema, the description supplies the algorithmic method, the input shape requirement, the alternative sibling, and an explicit 'Returns: n, correlation, hedgeRatio' summary. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema by clarifying that the series are transformed to daily % returns before computing and that both must be of 'matching length' — a constraint the schema states but the description explains the reason for. It does not document the 5-item minimum beyond what the schema declares.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource (correlation coefficient and minimum-variance hedge ratio between two price series) and immediately distinguishes itself from workflow.run_forex_correlation_live by naming the sibling and the condition that selects it ('No live data fetch: supply the two price series directly'). An agent can route correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the use condition ('Use when user already has two price series and wants their statistical relationship') and names the alternative (run_forex_correlation_live) with its contrasting capability ('both series fetched automatically for two named pairs'). This is the when/when-not/alternatives pattern at full strength.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_correlation_liveForex Correlation (Live)A
Read-only
Inspect

Correlation coefficient and minimum-variance hedge ratio between two forex pairs, fetched live: historical daily rates for both pairs over the given lookback window (frankfurter.app's daily ECB reference time series, business days only), aligned by matching date, computed on daily % returns. hedgeRatio follows Hull's standard futures-hedging formula: "how many units of pair2 per unit of pair1 minimizes the combined position's variance." Use when user asks "how correlated are EUR/USD and GBP/USD?" or "what hedge ratio should I use between these two pairs?". Returns: n (overlapping trading days used), correlation (-1 to 1), hedgeRatio, available (false + error if either pair has no data, or too few dates overlap).

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNoCalendar days to look back, 14-365. Default 30. Business-day-only data means fewer actual points than this number.
pair1YesFirst currency pair in BASE/QUOTE format, e.g. "EUR/USD"
pair2YesSecond currency pair in BASE/QUOTE format, e.g. "GBP/USD"

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover safety (readOnly/openWorld/not destructive); the description carries substantial extra behavior: ECB reference source, business-days-only data, date-aligned series, computation on daily % returns, the specific Hull formula behind hedgeRatio, and explicit failure semantics ('available: false + error if either pair has no data, or too few dates overlap').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded: purpose, then data/compute method, then when-to-use, then returns. Dense but nearly every sentence earns its place; the inlined Hull formula quotation is the only slightly verbose element.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description fully documents the return contract (n, correlation, hedgeRatio, available + error conditions), the data window semantics, and failure modes. Nothing material for a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so baseline is 3, but the description adds real meaning beyond the schema: pairs are aligned by matching date, days is calendar days that yield fewer business-day points, and the required BASE/QUOTE format is reinforced. Adds value on top of an already-documented schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (correlation coefficient and minimum-variance hedge ratio between two forex pairs) and the data provenance (live, frankfurter.app ECB daily series). It is clearly separable from its non-live sibling run_forex_correlation and from primitive.hedge_ratio.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger questions ('how correlated are EUR/USD and GBP/USD?', 'what hedge ratio should I use?') and explains the live-fetch context. It does not explicitly name the alternative to use for historical/static data, which is the one clear gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_currency_converterForex Currency ConverterA
Read-only
Inspect

Converts an amount between currencies using a manually supplied rate: no live data fetch. The live variant is workflow.run_forex_currency_converter_live, which has the rate fetched automatically. Use when user already knows the exact rate they want applied. Returns: converted.

ParametersJSON Schema
NameRequiredDescriptionDefault
rateYesExchange rate to apply (1 unit of source currency = rate units of target currency)
amountYesAmount in the source currency

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds the key behavioral trait that this is a pure, deterministic computation with no network fetch and a caller-supplied rate. It loses a point because 'Returns: converted' is an empty restatement rather than any useful detail about the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core behavior and the sibling routing. Every sentence carries weight except the trailing 'Returns: converted', which is filler and slightly dilutes an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, non-destructive, deterministic conversion with no output schema, the description covers the mode of operation and when to pick it. The only gap is that the return shape is not meaningfully described, though for a simple converter that is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (rate, amount) are documented in the schema with directional semantics for the rate. The description adds only the 'manually supplied' framing, which does not extend the parameter meaning beyond what the schema already supplies, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Converts) and resource (amount between currencies) plus the defining constraint: a manually supplied rate, no live fetch. It explicitly names the sibling it is not (workflow.run_forex_currency_converter_live), so an agent can separate the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit selection condition ('Use when user already knows the exact rate they want applied') and names the alternative tool that should be chosen otherwise. This is exactly the routing guidance a sibling-heavy namespace needs, since many forex siblings exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_currency_converter_liveForex Currency Converter (Live)A
Read-only
Inspect

Converts an amount between currencies using a live FX rate: TrueFX for its 10 quoted majors (genuinely live tick data), frankfurter.app daily ECB reference rate as the fallback for every other currency pair (30 total). Use when user asks "what's $X worth in EUR?" or any currency conversion where the rate itself isn't already known. Returns: converted, rate, source (truefx/frankfurter/identity), asOf, available (false + error if no rate could be found for that pair).

ParametersJSON Schema
NameRequiredDescriptionDefault
toYes3-letter target currency code, e.g. "EUR"
fromYes3-letter source currency code, e.g. "USD"
amountYesAmount in the source currency

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the readOnly/openWorld annotations by disclosing the two-tier data pipeline (TrueFX live ticks for 10 majors, frankfurter.app ECB daily fallback for all other pairs, 30 total), and the failure behavior (available=false + error when no rate exists). This is exactly the operational context an agent needs for a live-data tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then source tiering and usage guidance, then return fields. Dense but every sentence carries information; the parenthetical 'genuinely live tick data' is slightly chatty but justified as differentiation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by enumerating the returned fields (converted, rate, source, asOf, available + error). Combined with source-tier and failure disclosure, nothing needed to invoke or interpret the call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds value the schema does not: it scopes which currency pairs are served and how each tier is sourced, which matters for choosing from/to codes. It does not add format syntax beyond the schema, so it stops short of 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (convert an amount between currencies) and immediately distinguishes itself from its non-live sibling workflow.run_forex_currency_converter by naming the live-rate mechanism. An agent can tell which of the two converter tools applies without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger ('what's $X worth in EUR?') and a useful exclusion ('where the rate itself isn't already known'). It does not explicitly route against the non-live sibling, but the live-rate framing provides clear context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_margin_levelForex Margin LevelA
Read-only
Inspect

Free margin and margin level % from account equity and used margin: equity/usedMargin*100, the same stop-out proximity metric every forex platform shows. Returns null (not Infinity) when usedMargin is 0, meaning no open position. Use when user asks "how close am I to a margin call?" or "what's my free margin?". Returns: freeMargin, marginLevelPct.

ParametersJSON Schema
NameRequiredDescriptionDefault
equityYesAccount equity (balance + floating P&L)
usedMarginYesMargin currently locked by open positions. 0 if none.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe, read-only, non-open-world calculation, so the bar is low. The description still adds real value by disclosing the usedMargin=0 edge case and the deliberate null-instead-of-Infinity return, which an agent cannot infer from annotations or schema alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the formula, then usage triggers, then return values, then the edge case – a sensible ordering. It is slightly dense with three separate concerns, but every clause carries information and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing the returned fields (freeMargin, marginLevelPct). Combined with the formula, edge-case note, and usage triggers, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are already documented, giving a baseline of 3. The description adds the relationship between them (equity divided by usedMargin times 100) and clarifies that usedMargin being 0 means no open position, which enriches semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the exact computation (equity/usedMargin*100) and names the resulting metrics (freeMargin, marginLevelPct), so the agent knows precisely what is produced. It also implicitly separates itself from the sibling run_forex_margin_required by framing itself as the stop-out proximity metric rather than a margin requirement calculation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger phrasings ("how close am I to a margin call?", "what's my free margin?") that map directly to invocation. It does not explicitly name alternatives such as run_forex_margin_required for exclusion, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_margin_requiredForex Margin RequiredA
Read-only
Inspect

Notional and margin required for a forex position, in the pair's own quote currency: no live FX rate needed. Never built for Phase 1 of this domain since a margin figure has no natural currency-neutral form the way pip value does. The live variant is workflow.run_forex_margin_required_live, which applies when the account currency differs from the pair's quote currency. Use when user asks "how much margin do I need for X lots of EUR/USD at 50:1?". Returns: notionalQuote, marginQuote, base, quote.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
priceYesEntry or current price
unitsYesPosition size in base-currency units
leverageYesLeverage multiplier, e.g. 50 for 50:1

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so the safety profile is covered. The description adds real behavioral context beyond that: it is a self-contained pure computation needing no live FX rate, and it enumerates the return fields (notionalQuote, marginQuote, base, quote). It stops short of stating units/precision conventions, hence a 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence and the alternative-tool sentence are well front-loaded and actionable. However, the sentence 'Never built for Phase 1 of this domain since a margin figure has no natural currency-neutral form the way pip value does' is internal design rationale that does not help an agent decide or invoke, weakening the signal-to-noise ratio.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description helpfully lists the returned fields and chains to the sibling variant. For a four-parameter read-only calculation this is nearly complete; only unit/tolerance conventions and the lots-vs-units question are left unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented, which sets the baseline at 3. The example ('X lots ... at 50:1') loosely maps units and leverage, but it also introduces 'lots' while the schema says 'units in base-currency units' — an unreconciled mismatch rather than added clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb/resource: computes notional and margin required for a forex position, and pins down the currency form ('pair's own quote currency: no live FX rate needed'). It clearly distinguishes itself from the sibling workflow.run_forex_margin_required_live, so an agent can pick correctly without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (the _live variant) and the selecting condition ('when the account currency differs from the pair's quote currency'), and even supplies a concrete trigger query ('how much margin do I need for X lots of EUR/USD at 50:1?'). Both when-to-use and when-to-use-the-other are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_margin_required_liveForex Margin Required (Live)A
Read-only
Inspect

Notional and margin required for a forex position, converted to a given account currency via a live FX rate (same TrueFX/frankfurter.app source as workflow.run_forex_pip_value_live). Use when user asks "how much margin do I need in my account currency?" and the account currency differs from the pair's own quote currency. Returns everything workflow.run_forex_margin_required does, plus accountCurrency, notionalAccount, marginAccount, fxRate, fxSource, available (false + error if no rate could be found).

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
priceYesEntry or current price
unitsYesPosition size in base-currency units
leverageYesLeverage multiplier, e.g. 50 for 50:1
accountCurrencyYes3-letter account currency code, e.g. "GBP"

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint). The description adds real value beyond them: it names the external FX source (TrueFX/frankfurter.app) and discloses the failure mode (available=false + error when no rate is found), which an agent needs to handle the open-world dependency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose, then usage, then return fields in three dense sentences; the output-field enumeration is long but earns its place given there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the added return fields (accountCurrency, notionalAccount, marginAccount, fxRate, fxSource, available) and the error condition, so an agent knows both input and output contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented; the description only restates accountCurrency's role and lists output fields. Baseline 3 applies since the schema carries the parameter burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (notional and margin required for a forex position) and the exact scope qualifier (converted to account currency via a live FX rate). It also explicitly distinguishes itself from workflow.run_forex_margin_required, which it extends.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger ("how much margin do I need in my account currency?") plus the selecting condition (account currency differs from the pair's quote currency), and names the non-live alternative for the other case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_pip_valueForex Pip ValueA
Read-only
Inspect

Value of 1 pip for a given forex pair and position size, in that pair's own quote currency (e.g. EUR/USD's pip value comes back in USD, USD/JPY's in JPY): no live FX rate needed, since a pair's pip value is naturally denominated in its own quote currency. Pip size is 0.0001 for non-JPY pairs, 0.01 for JPY-quoted pairs, a universal market convention verified against real broker documentation, not broker-specific. The live variant is workflow.run_forex_pip_value_live, which applies when the account currency differs from the pair's quote currency. Use when user asks "what's 1 pip worth on X lots of EUR/USD?". Returns: pipSize, pipValueQuote, base, quote.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
unitsYesPosition size in base-currency units (1 standard lot = 100,000, mini = 10,000, micro = 1,000, nano = 100)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it as a safe read-only operation. The description adds useful context beyond annotations: no live FX rate needed, pip size conventions (0.0001 vs 0.01), and the conditions when the live variant is required. However, it does not document error cases or limitations (e.g., unsupported pairs).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Information is front-loaded and reasonably efficient, but slightly verbose with parenthetical explanations and a meta-comment about verification. Every sentence contributes, though some could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a simple 2-parameter read-only calculator: covers purpose, the key alternative, and output fields. With no output schema, it lists the returned fields, which is helpful. Minor gap: no explicit handling of invalid pair formats or edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters. The description provides example output naming and the pip size convention, but doesn't add format details for the parameters beyond what the schema says. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: the value of 1 pip for a forex pair and position size, clearly denominated. Distinguishes itself from the live sibling and scopes the calculation to the pair's own quote currency.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative (workflow.run_forex_pip_value_live) and the condition that selects it (account currency differs from quote currency). Also gives a natural-language usage trigger.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_pip_value_liveForex Pip Value (Live)A
Read-only
Inspect

Value of 1 pip for a forex pair and position size, converted to a given account currency via a live FX rate: TrueFX for its 10 quoted majors (genuinely live tick data), frankfurter.app daily ECB reference rate as the fallback for every other currency (30 total). Use when user asks "what's 1 pip worth in my account currency?" and the account currency differs from the pair's own quote currency (if it matches, workflow.run_forex_pip_value alone is enough, no live fetch needed). Returns everything workflow.run_forex_pip_value does, plus accountCurrency, pipValueAccount, fxRate, fxSource (truefx/frankfurter/identity), available (false + error if no rate could be found for that currency).

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
unitsYesPosition size in base-currency units
accountCurrencyYes3-letter account currency code, e.g. "GBP"

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds real behavioral value beyond that: the two data sources (TrueFX live ticks for 10 majors, frankfurter.app ECB daily fallback for the other ~30 currencies) and the availability/failure contract (available=false plus error when no rate exists).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and the key routing rule before the data-source and return-value details. Dense but each clause carries information; the sourcing/return enumeration is slightly long but earns its place given there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description compensates by enumerating the return fields (accountCurrency, pipValueAccount, fxRate, fxSource, available/error). Combined with the routing rule and data-source disclosure, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema, giving a baseline of 3. The description reinforces how accountCurrency drives the FX conversion and that units is a base-currency position size, but adds no syntax or format detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: computes the value of 1 pip for a forex pair and position size, converted to an account currency via a live FX rate. It explicitly distinguishes itself from the sibling workflow.run_forex_pip_value and names the exact use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use when user asks what's 1 pip worth in my account currency?') and an explicit exclusion with the alternative: when the account currency matches the pair's quote currency, use workflow.run_forex_pip_value alone. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_pnlForex PnLA
Read-only
Inspect

Profit or loss for a closed or hypothetical forex trade, in pips and in the pair's own quote currency, long or short. Use when user asks "what did I make/lose on this trade?" or "what would X pips be worth on Y lots?". Returns: pips (signed, positive favors the position taken), pnlQuote.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
sideYes
unitsYesPosition size in base-currency units
exitPriceYes
entryPriceYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is low. The description usefully adds return semantics (pips signed, positive favors the position taken) and output fields (pnlQuote), but says nothing about quote-currency conversion behavior or handling of open vs closed trades.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the first defines the output and scope, the second gives usage triggers, then a terse returns clause. Front-loaded, zero waste, no repetition of schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description appropriately names the return fields (pips, pnlQuote) and their sign convention, and covers required inputs conceptually. It is complete enough to call correctly, but the missing price-field format details and open-vs-closed trade handling leave a modest gap for a 5-param, fully-required tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 40% – pair and units are documented in the schema, while side, entryPrice, and exitPrice are bare. The description partially compensates by referencing 'X pips', 'Y lots', and 'pair's own quote currency', but adds no format or unit guidance for the price fields, so it neither fully covers the gap nor is irrelevant.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computed resource (forex PnL in pips and quote currency) with an explicit scope (closed or hypothetical trade, long or short). Distinguishes itself from sibling run_forex_pip_value by combining pips and money PnL, but never names a sibling to differentiate, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use cues via example user phrasings ('what did I make/lose on this trade?', 'what would X pips be worth on Y lots?'). It does not, however, say when NOT to use it or route to alternatives like run_forex_pip_value or run_forex_risk_reward, which matters given the dense forex sibling cluster.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_position_size_liveForex Position Size (Live)A
Read-only
Inspect

Position size (units and standard lots) from a risk amount and stop distance, in a given account currency, via a live FX rate to convert pip value into that currency (same TrueFX/frankfurter.app source as the other forex _live tools). Never has a meaningful account-currency-agnostic form: sizing a position from a risk budget genuinely requires knowing what 1 pip is worth in the currency that budget is denominated in. Use when user asks "how many lots should I trade to risk $X on this setup?". Returns: stopDistancePips, units, lots, pipValuePerUnitAccount, fxRate, fxSource, available (false + error if entry equals stop, or if no rate could be found).

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
stopPriceYes
entryPriceYes
riskAmountYesAmount to risk, in accountCurrency
accountCurrencyYes3-letter account currency code, e.g. "USD"

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotations already covering read-only and open-world safety, the description adds meaningful behavioral context: it uses a live external FX source, converts pip value into the account currency, and returns available=false plus an error when entry equals stop or no rate is found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core computation before justifying the account-currency requirement, and the return fields are grouped at the end. The middle rationale sentence is slightly dense but earns its place by preventing misuse of the tool without a currency context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter live workflow with no output schema, the description supplies the essential missing context: the FX-rate dependency, required account currency, return field names, and failure conditions. Enough for an agent to invoke it correctly without further documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, so documentation must compensate. The description clarifies that riskAmount is expressed in accountCurrency, that stop distance is derived from entryPrice/stopPrice, and that a live rate converts pip value into that currency. It still does not define entry/stop directionality or pair-specific edge cases.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computation—position size in units and lots from risk amount and stop distance—and ties it to a live FX conversion. It distinguishes itself from generic position-sizing and other forex tools by naming the live-rate dependency and the account-currency requirement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete user-facing trigger: 'how many lots should I trade to risk $X on this setup?', and explains why accountCurrency is non-optional. It does not explicitly compare against the non-live forex position-size siblings or state when to prefer this live variant over them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_risk_rewardForex Risk RewardA
Read-only
Inspect

Risk and reward distance in pips from entry/stop/target, and the resulting ratio (reward/risk): a raw number, not a verdict. ratio is null when the stop sits exactly at entry (no risk distance), and also when validSetup is false (stop/target on the wrong side of entry for the given side, e.g. a long with its stop above entry) - check validSetup before trusting the ratio. Use when user asks "what's my risk/reward on this setup?". Returns: riskPips, rewardPips, validSetup, ratio.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
sideYes
stopPriceYes
entryPriceYes
targetPriceYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish this is a safe, closed-world read, and the description goes further by disclosing output semantics: ratio is a raw number, not a verdict, and is null in two distinct conditions (stop at entry, validSetup false). It does not state units beyond pips or formatting of the ratio.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core computation, then the null-ratio caveats, then the return list. Dense but every clause carries information; only the trailing "Returns:" enumeration is somewhat redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the work of naming the return fields (riskPips, rewardPips, validSetup, ratio) and explaining the null cases. Adequate for a 5-param calculator; a few numeric-format details are the only missing piece.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 20% (only pair is documented), so the description must carry the load. It usefully explains how stopPrice/targetPrice relate to entryPrice for a given side via the validSetup discussion, but gives no detail on price format, quote convention, or precision for the numeric fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a precise verb+resource: computes risk/reward distance in pips from entry, stop and target, plus the ratio. The forex/pips framing implicitly separates it from the generic workflow.run_risk_reward sibling, but that differentiation is never made explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete routing cue ("what's my risk/reward on this setup?") and a strong precondition: check validSetup before trusting the ratio, since ratio is null when the stop sits at entry or the stop/target are on the wrong side. It stops short of naming alternative siblings for related questions (e.g. pip value, position sizing).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_scenarioForex ScenarioA
Read-only
Inspect

PnL across a range of hypothetical price moves (in pips, signed by actual price direction, not pre-adjusted for side), for a single forex position size, long or short. Use when user asks "what if price moves X pips in either direction?". Returns: scenarios[] (deltaPips, exitPrice, pnlQuote).

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
sideYes
unitsYesPosition size in base-currency units
deltasPipsYesHypothetical price moves in pips, e.g. [-50, 0, 50]
entryPriceYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, non-destructive, closed-world behavior, so the safety profile needs no repetition. The description adds genuinely useful semantics the annotations cannot carry: the pip deltas are signed by actual price direction and are not pre-adjusted for side, plus the returned field shape. It stops short of describing edge cases or constraints on delta ranges.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: the core computation first, the trigger phrase second, the return shape last. Zero filler, and the sign-convention caveat is embedded where the agent will read it rather than buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description usefully names scenarios[] and its fields, and it pins down the sign convention that would otherwise cause wrong calls. It omits any mention of pip definition per pair or limits on the delta list, which are minor given the pip-value siblings exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 60%, so the description must compensate, and it does for the two riskiest parameters: it clarifies that deltasPips are raw signed moves and that a single side (long or short) applies per call. It does not address entryPrice or units beyond what the schema already states, leaving some coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computation (PnL across a range of hypothetical pip moves) for a specific resource (a single forex position, long or short), which an agent can distinguish from the many sibling workflow tools. It does not, however, explicitly contrast itself with the closest siblings such as run_forex_pnl or run_scenario_planning, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete trigger phrase ("what if price moves X pips in either direction?") that tells the agent when to reach for this tool. There is no explicit when-not guidance or named alternative (e.g., use run_forex_pnl for a single realized PnL), so routing among the dense set of forex siblings is only partially supported.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_swap_costForex Swap CostA
Read-only
Inspect

Total swap/rollover cost (or credit) for holding a forex position overnight, in the pair's own quote currency: no live FX rate needed. Swap rates are broker-set with no free live feed available, so swapPerLotPerNight is always a manual input (quoted per standard lot, matching how commission is quoted in workflow.run_forex_breakeven), not fetched. Negative = cost (you pay), positive = credit (you receive). The live variant is workflow.run_forex_swap_cost_live, which applies when the account currency differs from the pair's quote currency. Use when user asks "how much will holding this position overnight cost me?". Returns: lots, totalSwapQuote, base, quote.

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
unitsYesPosition size in base-currency units
nightsYesNumber of nights the position is held
swapPerLotPerNightYesSwap rate per standard lot (100,000 units) per night, in the pair's quote currency. Negative = cost, positive = credit. Broker-set: get this from the user's broker platform, there is no live source for it.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower; the description nonetheless adds real behavioral context: swap rates are broker-set with no free live feed, so swapPerLotPerNight is always a manual input quoted per standard lot, and the sign convention (negative=cost, positive=credit). It does not need to explain return format beyond the field list it already gives, but it stops short of noting anything about precision/rounding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then the manual-input caveat, sign convention, sibling disambiguation, trigger phrase, and return fields – well ordered and each sentence carries information. Slight redundancy with the schema on the negative/positive convention keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-required-parameter computation tool with no output schema, the description supplies the missing pieces: currency context for the result, the manual-input rationale for the only non-obvious parameter, the sign convention, the sibling routing rule, and the returned field list (lots, totalSwapQuote, base, quote). An agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3; the description earns above that by explaining why swapPerLotPerNight is manual (no live source), that it is quoted per standard lot and mirrors how commission is quoted in workflow.run_forex_breakeven, and by reinforcing the negative/positive sign semantics. This adds genuine meaning beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('total swap/rollover cost for holding a forex position overnight') with the unit of account (pair's own quote currency) and explicitly names the sibling it is not (workflow.run_forex_swap_cost_live). An agent can distinguish it from the many other forex workflow tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use when user asks "how much will holding this position overnight cost me?"') and a routing rule to the alternative ('The live variant ... applies when the account currency differs from the pair's quote currency'). Both when-to-use and when-to-use-the-other are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_forex_swap_cost_liveForex Swap Cost (Live)A
Read-only
Inspect

Total swap/rollover cost (or credit) for holding a forex position overnight, converted to a given account currency via a live FX rate (same TrueFX/frankfurter.app source as the other forex _live tools). swapPerLotPerNight is still always a manual input: swap rates are broker-set, with no free live feed available for them. Use when user asks "how much will holding this position overnight cost me in my account currency?" and it differs from the pair's own quote currency. Returns everything workflow.run_forex_swap_cost does, plus accountCurrency, totalSwapAccount, fxRate, fxSource, available (false + error if no rate could be found).

ParametersJSON Schema
NameRequiredDescriptionDefault
pairYesCurrency pair in BASE/QUOTE format, e.g. "EUR/USD"
unitsYesPosition size in base-currency units
nightsYesNumber of nights the position is held
accountCurrencyYes3-letter account currency code, e.g. "GBP"
swapPerLotPerNightYesSwap rate per standard lot per night, in the pair's quote currency. Negative = cost, positive = credit.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so the safety and network profile is covered. The description adds real value beyond that: it names the live FX source (same as other forex _live tools), discloses the failure mode (available=false + error if no rate found), and explains why swapPerLotPerNight must be supplied manually (broker-set, no free live feed).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the definition of the returned metric and the usage trigger, and the return-value sentence is compact. Slightly dense in the middle (FX source asides), but nearly every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description enumerates the extra return fields (accountCurrency, totalSwapAccount, fxRate, fxSource, available) and the error case, and explains what the non-live sibling already returns. Nothing an agent needs to call or interpret this tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented, including the sign convention for swapPerLotPerNight. The description only reinforces that swapPerLotPerNight is a manual, broker-set input—useful but not additive over the schema's own detail. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific output (total swap/rollover cost converted to account currency) and names the exact condition that separates it from its sibling workflow.run_forex_swap_cost (live FX conversion to a non-quote account currency). An agent can pick this tool apart from the non-live variant without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger with a quoted user question ('how much will holding this position overnight cost me in my account currency?') plus the qualifier that it applies when the account currency differs from the pair's quote currency. It also implicitly routes to workflow.run_forex_swap_cost when no conversion is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_funding_arbitrageFunding ArbitrageA
Read-only
Inspect

Calculate funding rate arbitrage profit: annualized yield, net profit, and breakeven days for a long/short basis trade across two exchanges: the bare numbers only, no plain-English verdict. For the same math plus a profitable/marginal/loss verdict, use workflow.run_carry_trade instead. Use when user asks "is this funding arb worth it?" or "how many days to break even on transfer fees?". Returns: netProfitUsdt, annualizedYieldPct, breakevenDays.

ParametersJSON Schema
NameRequiredDescriptionDefault
durationDaysYesHolding period in days
positionSizeYesPosition size in USDT
intervalHoursNoFunding interval: 8 (standard) or 1 (Hyperliquid)
transferFeePctNoOne-time transfer/setup fee as percentage, e.g. 0.1 for 0.1%
longFundingRateYesFunding rate on long side (% per interval, positive = you pay)
shortFundingRateYesFunding rate on short side (% per interval, positive = you receive)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=false, destructiveHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: this tool returns bare numbers with no interpretation, and it lists the three values produced. It does not discuss defaults for the optional intervalHours/transferFeePct or edge cases like negative rates, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the computation and its outputs, then routes to the sibling, then gives trigger phrasings — good ordering. The trailing 'Returns: netProfitUsdt, annualizedYieldPct, breakevenDays' partially repeats the opening sentence's output list, which is the only mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by naming the returned fields, and the input schema fully covers all six parameters. What is missing is guidance on optional-parameter defaults and behavior on unusual inputs, but for this calculation tool an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including the enum for intervalHours and the positive/negative sign convention for both funding rates, so the schema carries the parameter burden. The description adds no syntax, units, or defaults beyond what is already documented, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Calculate funding rate arbitrage profit: annualized yield, net profit, and breakeven days for a long/short basis trade across two exchanges') and explicitly names the sibling it is not (run_carry_trade). An agent can separate it from run_funding_breakeven, run_compound_funding, and run_cross_venue_arbitrage without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool and the exact condition that selects it ('For the same math plus a profitable/marginal/loss verdict, use workflow.run_carry_trade instead'), and gives concrete user phrasings that should route here. When-to-use, when-not-to-use, and the alternative are all explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_funding_breakevenFunding BreakevenA
Read-only
Inspect

Price move needed to cover funding cost + fees over a holding period. Use when user asks "how much does BTC need to move for me to profit after funding?" or "is funding killing my edge on this trade?". Returns: breakevenWithFunding, breakevenWithoutFunding, requiredMovePct.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
sizeYesPosition size in base currency
hold_hoursYesHold duration in hours
entry_priceYesEntry price
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. size is USD notional (contracts) for inverse; notional/funding_cost/fee_total/total_carry_cost come back denominated in the base coin.
fee_open_pctNoOpen fee rate (default 0.0002)
funding_rateYesFunding rate per 8h period (decimal, e.g. 0.0001)
fee_close_pctNoClose fee rate (default 0.0005)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, non-destructive, and closed-world, so the safety profile is covered. The description adds the return field names, which is useful given no output schema, but says nothing about unit conventions, sign of requiredMovePct, or how holding period interacts with discrete funding periods.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, purpose front-loaded, examples then return keys. Efficient, though the quoted example questions are slightly verbose relative to the rest.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter compute tool with no output schema, the description supplies return field names and triggering intent, and the schema documents the rest. It is nearly complete, with only units/sign conventions for the returned values left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 88% and contractType explicitly documents denomination behavior, so the schema carries the load. The description adds no parameter-level detail beyond 'holding period', so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific computation (price move needed to cover funding + fees over a holding period), which an agent can distinguish from generic 'breakeven' siblings. It does not explicitly differentiate from run_breakeven_planning, run_funding_cost, or run_compound_funding, which are the closest neighbors.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Two concrete example user phrasings are given, which map cleanly to invocation triggers. There are no explicit exclusions or named alternatives (e.g. use run_funding_cost instead when only the cost figure is wanted), so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_funding_costFunding CostA
Read-only
Inspect

Calculate the total funding cost (or income) for holding a perpetual futures position. Use when user asks "how much funding will I pay holding X days?" or "is funding eating my profit?". Returns: totalFundingUsdt (negative = you pay, positive = you receive), perIntervalUsdt.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysYesNumber of days to hold
sideYes
sizeBaseYesPosition size in base asset
entryPriceYesEntry price
fundingRateYesFunding rate per 8h period as fraction, e.g. 0.0001
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. For inverse, sizeBase is USD contracts and cost figures come out in the base coin.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, so the safe-computation profile is covered. The description adds real value beyond that by documenting the sign convention of the result (negative = pay, positive = receive), which is not derivable from annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the purpose front-loaded, followed by usage triggers and return semantics. No filler; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by naming the returned fields and their sign convention. Params are largely self-documented, so the main gap is minimal – an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 83%, with required params (side, sizeBase, entryPrice, fundingRate, days) and the contractType semantics already documented in the schema. The description adds no per-parameter detail beyond what is in the structured fields, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: calculating total funding cost/income for holding a perpetual futures position. The title and description align and the scope is unambiguous. It does not explicitly distinguish itself from close siblings like run_funding_breakeven or run_compound_funding, but the holding-period framing is distinctive enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger phrasing ("how much funding will I pay holding X days?", "is funding eating my profit?"), which tells an agent when to reach for this tool. It stops short of naming alternatives or conditions where another funding-related tool should be used instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_garchGARCH VolatilityA
Read-only
Inspect

GARCH(1,1) volatility model, fit by maximum likelihood on a return series: estimates omega/alpha/beta (the variance-persistence parameters) and forecasts next-period volatility. Use when user asks "what's my GARCH volatility forecast?" or "how persistent is volatility in this return series?". This is a backward-looking statistical fit, not a market prediction guarantee. Returns: mu, omega, alpha, beta, persistence (alpha+beta), unconditional_vol, forecast_vol, loglikelihood.

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsYesReturn series, one value per period, at least 50 values (GARCH needs real sample depth to identify persistence)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world), and the description adds a meaningful expectation caveat that the output is a backward-looking fit rather than a forecast guarantee. It also enumerates the returned fields, which no structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences ordered purpose, routing cues, caveat, returns; nothing is repeated from the title or schema, and the return list compensates for the absent output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating returned fields, and the single parameter is fully documented in-schema. It is nearly complete for a single-input statistical tool; only an explicit pointer to sibling volatility tools is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter's minItems constraint plus rationale (at least 50 values needed to identify persistence) is already documented in the schema. The description adds no syntax or preprocessing detail beyond what the schema states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific model (GARCH(1,1) fit by maximum likelihood), names the estimated quantities (omega/alpha/beta) and the output (next-period volatility forecast). This clearly separates it from siblings like run_implied_volatility or run_evt_tail_risk.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete user phrasings that should route here and sets a boundary that it is a statistical fit, not a prediction guarantee. It does not name alternative volatility tools (e.g. implied volatility, EVT tail risk) as an explicit when-not branch, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_hurst_exponentHurst ExponentA
Read-only
Inspect

Hurst exponent via rescaled-range (R/S) analysis: is this return/price series trending/persistent (H>0.5, a move tends to be followed by a move in the same direction), mean-reverting/anti-persistent (H<0.5), or consistent with a random walk (H~0.5)? Use when user asks "is this asset trending or mean-reverting?" or "does this series show long-range dependence?". Takes a return series, not raw price levels. Returns: hurst, interpretation (trending/mean_reverting/random_walk), window_sizes, rs_values.

ParametersJSON Schema
NameRequiredDescriptionDefault
valuesYesReturn series (not raw price levels)
min_windowNoSmallest R/S analysis window size (default 8)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare a safe read-only, closed-world, non-destructive operation, so the bar is lower. The description adds real behavioral value beyond that: the interpretation bands (H>0.5, H<0.5, H~0.5) and the critical input constraint that it takes a return series, not raw price levels, which directly affects correct invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core computation, then usage triggers, then the input caveat and return shape. Dense but almost every clause earns its place; the interpretation parentheticals are slightly verbose but useful for interpreting output since no output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the return fields (hurst, interpretation, window_sizes, rs_values) and the meaning of the interpretation values. Combined with the input constraint and usage triggers, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the input constraint 'Return series (not raw price levels)' is already present in the schema for values. The description restates this same constraint rather than adding syntax or guidance on min_window beyond what the schema default already provides, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource – computing the Hurst exponent via rescaled-range (R/S) analysis – and immediately distinguishes itself from siblings like run_garch and run_cointegration by naming the exact diagnostic it performs. The parenthetical interpretation thresholds give the agent a concrete sense of the output's meaning.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit trigger phrasing ('is this asset trending or mean-reverting?', 'does this series show long-range dependence?') that maps a user question directly onto the tool. It doesn't name alternative tools or when NOT to use it, but the triggering questions are specific enough to route correctly among the 60+ workflow siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_impermanent_lossImpermanent LossA
Read-only
Inspect

Impermanent loss for a liquidity-pool position: compares providing liquidity against simply holding the same tokens, at a manually-supplied entry and current price. Two modes: full_range (standard 50/50 constant-product pool, the textbook 2*sqrt(k)/(1+k)-1 closed form) or concentrated (a Uniswap-V3-style position confined to [lowerPrice, upperPrice] - IL is always worse than full_range for the same price move when the range is tight, and the position is fully single-asset, no longer earning fees, once price exits the range). impermanentLossPct is always <= 0 and is measured relative to the quote token (dimensionless, exact regardless of what the quote token is); the optional dollar figures additionally assume the quote token's own USD price stayed roughly stable (true for a stablecoin-quoted pool). Use when user asks "how much am I losing to impermanent loss?" or "is this LP position still worth it after fees?". Returns: impermanentLossPct, lpValueMultiplier, hodlValueMultiplier, inRange, lossUsd/netResultUsd (null unless depositValueUsd given).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoDefault full_range.
entryPriceYesBase token's price in quote-token terms when liquidity was deposited
lowerPriceNoRange lower bound - required for concentrated mode, must be below entryPrice
upperPriceNoRange upper bound - required for concentrated mode, must be above entryPrice
currentPriceYesBase token's current price in quote-token terms
feesEarnedUsdNoOptional trading fees earned so far in USD (default 0), folded into netResultUsd alongside lossUsd
depositValueUsdNoOptional: USD value deposited at entry, to also report dollar-denominated lpValueUsd/hodlValueUsd/lossUsd

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a safe, non-destructive, closed-world read, and the description adds substantial behavior beyond that: sign invariant (impermanentLossPct is always <= 0), normalization frame (relative to the quote token, dimensionless), the concentration caveat (IL worse than full_range when tight, position becomes fully single-asset and fee-less once price exits the range), and the stable-quote assumption behind the optional USD figures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core definition and mode split before the trigger phrasing and return list, and nearly every clause carries information. It is a single dense block with the textbook closed form noted parenthetically, which is arguably more formula detail than an agent needs to select the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema but a 7-parameter, mode-dependent computation, the description compensates by listing the returned fields (impermanentLossPct, lpValueMultiplier, hodlValueMultiplier, inRange, lossUsd/netResultUsd) and noting when they are null. Mode requirements and range constraints are covered between description and schema, leaving no material gap for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning the schema lacks: what 'mode' actually selects between, why the concentrated bounds change the economics, and what the dollar outputs assume. It does not restate the schema's own constraints verbatim, which is the right call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Impermanent loss for a liquidity-pool position: compares providing liquidity against simply holding the same tokens') and enumerates the two computation modes with their distinct formulas. The phrase 'manually-supplied entry and current price' implicitly separates it from the sibling workflow.run_impermanent_loss_live, so an agent can route between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete user-facing triggers ('how much am I losing to impermanent loss?', 'is this LP position still worth it after fees?'), which tells the agent when to select it. It does not state exclusions or name the live variant as the alternative for auto-fetched prices, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_impermanent_loss_liveImpermanent Loss (Live)A
Read-only
Inspect

Live variant of workflow.run_impermanent_loss: fetches each token's current live USD price (Solana or any of 5 EVM chains) and derives currentPrice as their ratio, instead of it being supplied manually. entryPrice stays a manual, historical input (a fact the caller must supply, not something a live quote should overwrite), same convention as every other live-capable tool in this product. Use when the caller has the pool's two token addresses but doesn't already know the current price ratio. Returns the same fields as workflow.run_impermanent_loss, plus currentPrice, baseSymbol, quoteSymbol, routable (false + error if either token's price can't be resolved).

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoDefault full_range.
baseMintYesAddress of the base token
baseChainNoChain of the base (volatile) token. Default solana.
quoteMintYesAddress of the quote token
entryPriceYesBase token's price in quote-token terms when liquidity was deposited (manual - a historical fact)
lowerPriceNoRange lower bound - required for concentrated mode
quoteChainNoChain of the quote token. Default solana.
upperPriceNoRange upper bound - required for concentrated mode
feesEarnedUsdNoOptional trading fees earned so far in USD (default 0)
depositValueUsdNoOptional: USD value deposited at entry

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, destructiveHint=false, and openWorldHint=true, covering the safety profile, so the description's added value is the failure surface: 'routable (false + error if either token's price can't be resolved)' and the invariant that entryPrice is never overwritten by a live quote. It stops short of describing latency, caching, or chain-specific resolution behavior, which matters for an open-world price fetch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core distinction (live price derivation vs. manual), followed by the usage condition and the return deltas. The parenthetical rationale about not overwriting entryPrice is slightly discursive but earns its place by preventing a plausible caller error; overall still efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by naming the returned deltas (currentPrice, baseSymbol, quoteSymbol, routable) and the error condition. Combined with 100% schema coverage and read-only/open-world annotations, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters including entryPrice as 'manual - a historical fact'. The description reinforces why entryPrice stays manual and clarifies that currentPrice is derived rather than an input, but adds no new syntax, units, or default guidance beyond what the schema/enums already provide. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource (fetches each token's current live USD price and derives currentPrice as their ratio) and explicitly frames itself as the live variant of workflow.run_impermanent_loss, which is itself in the sibling list. An agent can distinguish it from both the non-live variant and unrelated workflow tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit selection condition: 'Use when the caller has the pool's two token addresses but doesn't already know the current price ratio.' It also contrasts behavior with workflow.run_impermanent_loss (manual price supplied there, derived here), effectively naming the when-not case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_implied_volatilityImplied VolatilityA
Read-only
Inspect

Solves for the volatility that makes Black-Scholes reproduce an observed option price (Newton-Raphson with a bisection fallback for cases where vega is too flat to converge, e.g. deep ITM/OTM or very short-dated). Checks the price against its no-arbitrage bounds first and refuses to solve (converged: false + error) rather than return a garbage number when the price is impossible for the given spot/strike/rate. Use when user asks "what IV does this option price imply?" or gives a market price and wants the volatility, not the reverse. Returns: impliedVolatilityPct, iterations, method (newton-raphson/bisection), converged, priceAtSolution.

ParametersJSON Schema
NameRequiredDescriptionDefault
spotYesUnderlying spot price, USD
strikeYesStrike price, USD
optionTypeYes
daysToExpiryYesCalendar days until expiry (can be fractional)
targetPriceUsdYesThe observed option price, USD, to solve the implied volatility from
riskFreeRatePctNoRisk-free rate in percentage points. Default 0: standard crypto-options convention.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the read-only/non-destructive safety profile, and the description adds substantial behavioral context: the numerical method (Newton-Raphson with bisection fallback), the failure modes it handles (flat vega, deep ITM/OTM, short-dated), and the explicit refusal behavior with converged:false + error rather than a garbage result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then layers failure handling and usage routing in two compact follow-ups. Every clause carries information: algorithm choice, edge-case handling, and return shape.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description enumerates the return fields (impliedVolatilityPct, iterations, method, converged, priceAtSolution), and it covers failure semantics, so an agent has everything needed to call and interpret the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is already 83%, so the schema documents the inputs well. The description only references spot/strike/rate in passing for the no-arbitrage bounds and adds no syntax, units, or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: solving for the volatility that makes Black-Scholes reproduce an observed price. It clearly differentiates from the sibling run_black_scholes by inverting the direction of the computation ('not the reverse').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('what IV does this option price imply?' or a market price with volatility wanted) and an exclusion via 'not the reverse', which routes away from the forward-pricing sibling. It stops short of naming the alternative tool explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_kelly_frontierKelly Growth-Security FrontierA
Read-only
Inspect

Kelly growth-security frontier (MacLean, Ziemba & Blazenko 1992): for a strategy compounding at a fraction lambda of full Kelly, the probability wealth ever falls to a fraction alpha of its starting value is P = alpha^(2/lambda-1). Provide either lambda_fraction (to compute that probability) or max_probability (to solve for the largest lambda that keeps the ruin probability at or below it). Use when user asks "if I bet half-Kelly, what's my chance of ever losing half my bankroll?" or "what fraction of Kelly keeps my chance of a 50% drawdown under 5%?". Valid lambda range is (0, 2]; beyond 2 the probability is certain (1), not the raw formula value. Returns: probability, lambda_fraction (echoed or solved).

ParametersJSON Schema
NameRequiredDescriptionDefault
alphaYesFraction of starting capital, 0-1 exclusive (e.g. 0.5 = "ever falls to half my starting bankroll")
lambda_fractionNoFraction of full Kelly being bet (1 = full Kelly, 0.5 = half Kelly). Provide this OR max_probability, not both.
max_probabilityNoTarget ceiling on the ruin probability, 0-1 exclusive. Provide this to solve for the safe lambda_fraction instead of supplying it directly.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/non-destructive, so the safety profile is covered. The description adds genuine behavioral context the annotations don't: the closed-form formula, the boundary behavior that lambda > 2 yields probability 1 rather than the raw formula value, and what is returned. It stops short of discussing numerical edge cases like alpha approaching 0/1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the reference and formula, then modes, then examples, then return values. Dense but every sentence carries information; the two illustrative examples are the only mildly redundant element against the mode description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by explicitly listing the returns (probability, lambda_fraction echoed or solved) and disclosing boundary behavior. For a 3-parameter read-only calculator, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description still adds value by framing the mutual exclusivity of lambda_fraction and max_probability as a mode switch, and by restating the (0, 2] validity bound that constrains the input beyond the schema's raw type descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific computation (Kelly growth-security frontier probability) with an academic citation, and precisely states the two operating modes: compute probability from lambda_fraction, or solve lambda from max_probability. No sibling tool overlaps, so the agent can distinguish it purely from the description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit selection logic ('Provide either lambda_fraction ... or max_probability ...') plus two natural-language user-phrasing examples that anchor real invocation intents. It also states the valid lambda range (0, 2], which is a usage constraint rather than just a domain fact.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_liquidation_safetyLiquidation SafetyA
Read-only
Inspect

Calculate the liquidation price for an isolated-margin futures position. Use when user asks "where will I get liquidated?" or "how close is my liq price?". Returns: liquidationPrice, distancePct (how far from entry).

ParametersJSON Schema
NameRequiredDescriptionDefault
mmrNoMaintenance margin rate, default 0.005 (0.5%)
sideYes
leverageYesLeverage multiplier, e.g. 10 for 10x
entryPriceYesEntry price (positive)
contractTypeNolinear = USDT-margined (default), inverse = coin-margined (e.g. Deribit/Bybit/MEXC BTC-settled perps).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, closed-world), and the description adds useful return context (liquidationPrice, distancePct) in the absence of an output schema. It does not disclose assumptions such as margin-mode limits or default mmr behavior, so it adds value without being rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: purpose first, then usage triggers, then return fields. Every sentence carries distinct information with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned fields, and the input schema covers almost all parameters. A brief note on the isolated-margin assumption or default mmr handling would close the remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents mmr, leverage, entryPrice, and contractType including the linear/inverse distinction. The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Calculate the liquidation price') plus the scope constraint 'isolated-margin futures position', which implicitly separates it from siblings like run_max_leverage or run_pre_trade_check. It never names an alternative sibling explicitly, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete user-intent triggers ('where will I get liquidated?', 'how close is my liq price?'), which tells the agent when to reach for it. No when-not-to-use condition or named alternative is offered, so it is clear context rather than full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_market_cap_comparisonMarket Cap ComparisonA
Read-only
Inspect

Compares two tokens' live market caps (Solana or any of 5 EVM chains; the two tokens can be on different chains) and projects what an investment would be worth if the first token's market cap matched the second's. Narrative-agnostic ("if X reaches Y's market cap"): works for any token pair, not tied to one hype cycle or one chain. A snapshot ratio, not a forecast: assumes fixed supply on both sides. Use when user asks "what if this token reaches [other token]'s market cap?". Returns: multiplier, projectedValueUsd, projectedPriceUsd, profitUsd, comparable (false + error if either market cap can't be resolved).

ParametersJSON Schema
NameRequiredDescriptionDefault
tokenMintYesAddress of the token you hold or are evaluating (base58 for Solana, 0x... for EVM chains)
tokenChainNoChain of the token you hold. Default solana.
compareToMintYesAddress of the token whose market cap to compare against
investmentUsdYesInvestment amount in USD
compareToChainNoChain of the comparison token. Default solana. Can differ from tokenChain.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, non-destructive, open-world behavior. The description goes further by disclosing the fixed-supply assumption and the failure mode (comparable=false plus error when either market cap can't be resolved), which is genuinely useful operational context. It does not cover rate limits or latency, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then scope, then caveats, then return shape. Sentences are dense but each carries information; the parentheticals and quoted example slightly lengthen it without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description enumerates the return fields (multiplier, projectedValueUsd, projectedPriceUsd, profitUsd, comparable) and the error condition. Combined with full parameter coverage, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond it: chains may be Solana or one of 5 EVM chains and the two tokens can reside on different chains, which explains why tokenChain and compareToChain exist as separate parameters. That cross-chain nuance is not obvious from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('compares two tokens' live market caps ... projects what an investment would be worth'), plus the cross-chain scope. No sibling in the list performs market-cap comparison, and the description makes the capability unmistakable even without reading the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit trigger phrase ('Use when user asks "what if this token reaches [other token]'s market cap?"') and bounds the tool's applicability by declaring it narrative-agnostic and pair-agnostic. It also clarifies the model's assumption ('snapshot ratio, not a forecast'), which tells the agent when the output is and isn't appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_market_implied_oddsMarket Implied OddsA
Read-only
Inspect

Reads Kalshi's full live BTC or ETH year-end price ladder (a set of mutually-exclusive prediction markets covering the whole price range) and reports what the market itself implies: the median (50th-percentile) price bucket, the single most-likely (mode) bucket, and the probability of ending the year at or above any real bucket boundary. Deliberately does not compute an expected value or interpolate inside a bucket: the top/bottom buckets are open-ended, so any point estimate there would need an invented assumption; every number this tool returns traces back to one live, sourced price. Use when user asks "what does the market think BTC will be worth by year end?" or "what are the odds ETH ends the year above $X?". Returns: buckets[] (label, floor, cap, probabilityPct), medianBucketLabel, modeBucketLabel, vigPct, probabilityAtOrAbovePct + snappedThresholdUsd (only when thresholdUsd is supplied).

ParametersJSON Schema
NameRequiredDescriptionDefault
coinNoWhich coin's year-end ladder to read. Default BTC.
thresholdUsdNoOptional price threshold: returns the probability of ending the year at or above the nearest real bucket boundary at or below this value.

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover the safety profile (readOnly, openWorld, non-destructive); the description goes well beyond them by disclosing the deliberate methodological choices (no expected value, no intra-bucket interpolation, open-ended tail buckets) and that every returned figure traces to a live sourced price. This is exactly the kind of behavioral context annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and method rationale, and the trailing 'Returns:' list earns its place given there is no output schema. The prose sentence before it is long and slightly dense, but no sentence is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full return-value burden and does so explicitly (buckets[], medianBucketLabel, modeBucketLabel, vigPct, conditional probabilityAtOrAbovePct + snappedThresholdUsd). Combined with the when-to-use examples, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented, including the default-BTC behavior and the threshold snapping rule. The description largely restates the schema's threshold semantics and only marginally extends it by naming snappedThresholdUsd as an output, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (reads Kalshi's live BTC/ETH year-end price ladder) and immediately characterizes the data as a set of mutually-exclusive prediction markets, which differentiates it from the sibling prediction-market tooling (run_prediction_market_edge, run_odds_converter). An agent can identify the tool's function without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete triggering user utterances ('what does the market think BTC will be worth by year end?', 'what are the odds ETH ends the year above $X?'), which gives strong when-to-use signal. It does not explicitly name or exclude alternative tools, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_max_leverageMax LeverageA
Read-only
Inspect

Calculate the maximum safe leverage based on account size, max acceptable drawdown, and asset daily volatility. Use when user asks "what's the max leverage I should use on BTC?" or "how much leverage is safe given 3% daily volatility?". Returns: maxLeverage, marginAtRisk.

ParametersJSON Schema
NameRequiredDescriptionDefault
mmrNoMaintenance margin rate, default 0.005 (0.5%)
accountSizeYesTotal account size in USDT
volatilityPctYesExpected daily price volatility as percentage, e.g. 3 for 3%
maxDrawdownPctYesMaximum acceptable drawdown as percentage, e.g. 10 for 10%

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the return shape (maxLeverage, marginAtRisk), which is useful since no output schema exists, but discloses nothing about assumptions, units, or edge-case behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences: purpose, usage triggers, and return values, with the core purpose front-loaded. No filler, though the return-value sentence could arguably be trimmed for a read-only calculator.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deterministic four-parameter calculator with a fully described schema and no output schema, the description covers purpose, inputs, triggers, and result fields. Missing only minor detail like the role of the optional mmr parameter and unit conventions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters including the optional mmr default are documented in the schema. The description restates the three required inputs but adds no format, unit, or default guidance beyond what the schema already supplies, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Calculate the maximum safe leverage,' and names the three driving inputs (account size, max drawdown, daily volatility). It is clearly distinguishable from siblings like run_liquidation_safety or run_position_sizing, though it never explicitly names an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger phrasings ('what's the max leverage I should use on BTC?') that map user intent directly to the tool, which is strong positive guidance. It lacks any when-not condition or pointer to the sibling it should be preferred over, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_odds_converterOdds ConverterA
Read-only
Inspect

Converts a probability into decimal odds, American odds, and breakeven win rate: either from a manually supplied probability, or fetched live from Kalshi, Polymarket, ADI Predictstreet, Limitless, or Myriad (five independent crypto-price prediction market venues, all public keyless market data). When a Kalshi, Polymarket, Limitless, or Myriad source is supplied, also returns the vig (the exchange's built-in edge), computed from the market's own YES+NO prices, not estimated; ADI Predictstreet's crypto contracts currently have no live trading volume on any venue, so this returns available:false with an explanation rather than a fake price (use workflow.run_window_fair_value for a theoretical price on those instead). Use when user asks "what odds does a 35% probability work out to?" or "what's the vig on this Kalshi/Polymarket/Limitless/Myriad market?". Provide exactly one of probability/kalshiTicker/polymarketSlug/adiSymbol/limitlessSlug/myriadSlug. Returns: probability, decimalOdds, americanOdds, breakevenWinRatePct, vigPct (null unless a live two-sided source was used), source (manual/kalshi/polymarket/adi/limitless/myriad), identifier, label.

ParametersJSON Schema
NameRequiredDescriptionDefault
adiSymbolNoAn ADI Predictstreet market symbol (e.g. BTC1D-20260920T0000). Currently always returns available:false; these contracts have no live trading volume yet.
myriadSlugNoA Myriad Markets market slug, from the market URL (e.g. eth-at-three-digits-when-bitcoin-goes-below-50k), to fetch a live price from. On-chain (Abstract L2); not every market has a Yes/No outcome, some are multi-outcome ladders.
probabilityNoProbability as a decimal 0-1 (e.g. 0.35). Use this OR one of the live sources below, not both.
kalshiTickerNoA Kalshi market ticker (e.g. KXBTCY-27JAN0100-T149999.99) to fetch a live price from.
limitlessSlugNoA Limitless Exchange market slug, from the market URL (e.g. btc-up-or-down-5-min-1790249100), to fetch a live price from. On-chain (Base), mostly short-duration (5min/15min/daily) crypto up/down contracts.
polymarketSlugNoA Polymarket market slug, from the market URL (e.g. will-bitcoin-reach-100k-in-september-2026), to fetch a live price from.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, non-destructive, openWorld), so the description's remaining job is disclosing quirks. It does this well: vig is computed from the market's own YES+NO prices rather than estimated, ADI returns available:false with an explanation instead of a fake price, and all venues are public keyless data. Minor gap is no note on rate limits or latency for the live fetches.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and the return values, which is good, but the venue enumeration and the extended ADI caveat make it long and dense. The parenthetical definitions of the five venues could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter, zero-required, no-output-schema tool, it covers the essentials: which inputs are mutually exclusive, what each source returns, what vigPct null means, and which sibling to use when ADI has no volume. An agent could invoke it correctly from the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds value the schema does not: the exclusivity rule ('provide exactly one of probability/kalshiTicker/...') and the meaning of the null vigPct when no live two-sided source was used.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (converts) and a precise multi-part resource: a probability into decimal odds, American odds, and breakeven win rate, from either a manual input or five named live venues. It also names the sibling workflow.run_window_fair_value as the right tool for ADI contracts, giving an agent an unambiguous selection signal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrases ('what odds does a 35% probability work out to?') and routes ADI queries to workflow.run_window_fair_value, but does not define when NOT to use this tool versus workflow.run_market_implied_odds or workflow.run_prediction_market_edge, which look adjacent in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_open_analysisMarket Profile: Open AnalysisA
Read-only
Inspect

Market Profile open analysis: where and how price opened vs the prior session value area. Use for "how did BTC open today?" / "what does the open imply for the session?". Returns: open_location, open_type (OD/OTD/ORR/OAIR) with description/implication, confidence, key_levels (VAH/VAL/VPOC/IB), tails (session-high/low rejection tails and single prints), scenario_framing (bullish/bearish/neutral), invalidation level.

ParametersJSON Schema
NameRequiredDescriptionDefault
venueYesExchange to fetch candles from when candles[] not supplied
candlesNoOptional OHLCV for the session; omit to fetch from venue (reproducible + 0 COGS when supplied)
timeframeNoCandle timeframe (default 15m)
instrumentYesSymbol, e.g. BTCUSDT
prev_candlesNoOptional OHLCV for the previous session
session_dateYesSession date YYYY-MM-DD (UTC)
value_area_ruleNoValue-area fraction 0.5–0.9 (default 0.70)

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=true), so the bar is lower. The description adds genuinely useful context by enumerating the analytical outputs (open_location, open_type classification set, tails, scenario_framing, invalidation level), which the agent could not infer from schema or annotations. It does not mention cost/reproducibility of supplying candles, which the schema already states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose, then usage triggers, then a compact semicolon-delimited return inventory. The return list is dense but every field named is meaningful and there is no filler prose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return values and does so thoroughly (open classification, key levels, tails, scenario framing, invalidation). For a read-only analytical workflow with full schema coverage, the definition supplies everything an agent needs to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all seven parameters are documented in the schema (venue enum, timeframe enum, value_area_rule range, candles reproducibility note). The description adds almost no parameter-level meaning beyond that; baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Market Profile open analysis: where and how price opened vs the prior session value area.' The example queries ('how did BTC open today?') make the analysis target unambiguous. It does not, however, explicitly distinguish itself from adjacent siblings like workflow.run_session_structure, so a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The quoted user questions ('how did BTC open today?', 'what does the open imply for the session?') give concrete triggers for selecting this tool. There is no statement of when NOT to use it or which sibling to prefer for related session-structure questions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_options_payoffOptions PayoffA
Read-only
Inspect

Payoff, P&L, and breakeven price for a single-leg Deribit BTC/ETH option (long or short call/put) at a given scenario price at expiry. Deribit BTC/ETH options are coin-settled: premium, P&L, and the max profit/loss caps come back denominated in the base coin (BTC/ETH), not USD; a scenarioPnlUsd convenience field converts the coin P&L back to USD at the scenario price. Coin settlement means a long call's upside is capped (max profit = 1 − premium per unit, not unlimited) while a long put's upside is technically unbounded as price falls toward zero, the mirror image of a USD-settled option's payoff shape, not a bug. Use when user asks "what does my BTC call/put pay off at price X?" or "where's my breakeven on this option?". Returns: intrinsicPerUnitCoin, scenarioPnlCoin, scenarioPnlUsd, breakevenPrice, maxLossCoin/maxProfitCoin (null = unbounded), isProfitable.

ParametersJSON Schema
NameRequiredDescriptionDefault
strikeYesStrike price in USD
currencyNoUnderlying coin. Default BTC.
positionYes
quantityYesNumber of contracts (Deribit BTC/ETH options have contract_size 1.0, so this is coin-denominated size)
optionTypeYes
premiumCoinYesPremium paid/received per contract, in the base coin (matches Deribit's own quoted price, e.g. 0.02 BTC); must be < 1 for a call
scenarioPriceYesUnderlying price in USD to evaluate the payoff at

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, non-destructive, closed-world), so the description is free to add domain behavior — coin settlement, coin-denominated outputs, the USD convenience field, and the counterintuitive capped-call / unbounded-put payoff shape explicitly labeled 'not a bug'. That is exactly the context an agent needs to avoid mis-reading results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose and then layers the important settlement caveat; most sentences earn their place. It is dense and errs long, with the long-call/long-put asymmetry explained in one admittedly run-on sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the returned fields (intrinsicPerUnitCoin, scenarioPnlCoin, scenarioPnlUsd, breakevenPrice, maxLoss/maxProfitCoin, isProfitable) and the null=unbounded convention, which is everything an agent needs to interpret the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, so the schema documents most inputs itself; the description reinforces that premium/P&L are base-coin denominated and echoes the max-profit formula, but it does not clarify parameters the schema leaves thin (e.g. multi-unit scaling of quantity or the currency default beyond a passing mention). Adds marginal value over the schema, hence baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource: computes payoff, P&L, and breakeven for a single-leg Deribit BTC/ETH option at a scenario expiry price. The 'single-leg' scope cleanly separates it from run_spread_payoff and run_straddle_strangle, so an agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrases ('what does my BTC call/put pay off at price X?', 'where's my breakeven on this option?') which is strong when-to-use signal. It never names an alternative tool directly, relying on 'single-leg' to imply the exclusion, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_orderbook_impactOrderbook ImpactA
Read-only
Inspect

Order-book "walk the book" impact/capacity: given order-book levels (price, size in base-asset units) and EITHER a target notional or an impact budget in bps, computes VWAP and price impact for that size, or (via bisection) the largest notional that stays inside the impact budget. Use when user asks "what's my price impact if I trade $X" or "how much can I trade before impact exceeds Y bps". side="buy" walks the asks, side="sell" walks the bids; impact is measured from the book's own mid ((best_bid+best_ask)/2) and is always a positive "cost in bps" number regardless of side. Returns: mid_price, spread_bps, vwap, impact_bps - both null when EITHER the loaded book doesn't cover the requested notional (book_sufficient=false flags this specific case) OR the notional involved is ~0 (book_sufficient stays true then; happens for a near-zero notional_usd, or in capacity mode when even an infinitesimal trade already exceeds impact_budget_bps, in which case max_notional itself resolves to 0) - and in capacity mode, max_notional plus capacity_is_lower_bound (true if the ENTIRE supplied book was consumed within budget, meaning true market capacity may exceed what was supplied - this tool only sees the levels given to it).

ParametersJSON Schema
NameRequiredDescriptionDefault
asksYesAsk levels as [price, size_base] pairs, any order
bidsYesBid levels as [price, size_base] pairs, any order
sideYes"buy" walks the asks, "sell" walks the bids
notional_usdNoTarget notional to walk the book for. Provide this OR impact_budget_bps, not both.
impact_budget_bpsNoMax acceptable impact in bps; solves for the largest notional within it. Provide this OR notional_usd, not both.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Far exceeds what the annotations provide. Documents side semantics (buy walks asks, sell walks bids), that impact is always a positive cost number from the book's own mid, the edge cases where vwap/impact are null (book insufficient vs ~0 notional), and the capacity_is_lower_bound caveat when the whole supplied book is consumed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first clause and then qualifies behavior. Some sentences are dense and the return-value enumeration is long, but every statement earns its place and no filler is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value contract fully, including the null conditions, book_sufficient flag, and capacity_is_lower_bound semantics. An agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so the baseline is 3, but the description meaningfully enriches it: it explains side semantics, the either/or relationship between notional_usd and impact_budget_bps, and the return fields correlated with each parameter mode. It stops short of describing the [price, size_base] level format, which the schema handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('walk the book' impact/capacity) and precisely describes the two computational modes. The description would let an agent distinguish it from siblings like run_swap_price_impact without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrases ('what's my price impact if I trade $X', 'how much can I trade before impact exceeds Y bps') that map cleanly to the two modes. No explicit when-not or named alternative, but the usage context is very clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_pnl_planningPnL PlanningA
Read-only
Inspect

Calculate net PnL, ROE, fees and gross profit/loss for a futures trade. Use when user asks "what's my profit/loss on this trade?" Returns: grossPnl, fees, netPnl, netPnlUsdt, roe (%), maxLossBound (only non-null for an inverse/coin-margined short: the finite ceiling on net coin-denominated loss as price rises without limit, (size/entryPrice)×(1+feeOpenPct) - includes the opening fee, since it survives the exit-price-to-infinity limit while the closing fee vanishes - null for every other side/contractType combination, which either has no such bound or a trivial one).

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYesTrade direction
sizeYesPosition size: base asset qty for linear, USD contracts for inverse
exitPriceYesExit price (positive)
entryPriceYesEntry price (positive)
feeOpenPctNoOpening fee as fraction, e.g. 0.0002 = 0.02%
feeClosePctNoClosing fee as fraction
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. For inverse, pnl/fees are returned in the base coin, not USDT.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered. The description goes further by enumerating the returned fields and giving a genuine edge-case disclosure (maxLossBound is non-null only for inverse/coin-margined shorts, includes the opening fee, null otherwise). That is real behavioral context beyond structured data, though it is buried in a dense sentence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose sentence and return-field list are front-loaded and efficient, but the maxLossBound clause is a single ~70-word parenthetical with nested asides that is hard to parse and could be split or trimmed. Roughly half the description's weight sits in an over-stuffed sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly carries the return-value burden by naming grossPnl, fees, netPnl, netPnlUsdt, roe and maxLossBound. With annotations covering the safety profile and the schema covering all params, this is nearly complete; only the unwieldy maxLossBound phrasing weakens it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (side, size, entryPrice, exitPrice, feeOpenPct, feeClosePct, contractType) is already documented with units and semantics. The description adds only one parameter-related fact (inverse PnL/fees are in base coin) which the schema also states, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (calculate) and resource list (net PnL, ROE, fees, gross P/L) scoped to 'a futures trade', which distinguishes it from the sibling run_forex_pnl. An agent can identify the tool's function without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger phrase ('what's my profit/loss on this trade?') that maps user intent to this tool. It does not name when-not-to-use or point at alternatives such as run_forex_pnl or run_breakeven_planning, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_portfolio_riskPortfolio RiskA
Read-only
Inspect

Aggregates risk across multiple open positions in one call: total notional, total P&L, total margin in use, margin usage as a % of account balance (if given), and which single position sits closest to liquidation. Each position is computed through the same canonical PnL/liquidation math as the single-position tools, then rolled up. Linear (USDT-margined) positions sum into one USD total; inverse (coin-margined) positions are grouped by settlement coin instead, since a BTC-margined P&L cannot be summed with an ETH-margined one without a live conversion rate. Returns a verdict: healthy / watch / reduce / critical, driven by the closest liquidation distance and margin usage. Use when user asks "how exposed am I across all my positions?" or "which of my positions is closest to liquidation?". Position size follows the product-wide convention: base-asset quantity for linear, USD notional (contracts) for inverse.

ParametersJSON Schema
NameRequiredDescriptionDefault
positionsYes
account_balanceNoOptional account balance in USD, used to compute margin_usage_pct: linear margin plus every inverse position's own margin marked to market at its mark_price, all as one USD total

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnly/no-destructive/no-open-world, so the bar is lower, and the description adds real behavioral context: linear positions sum into one USD total while inverse positions are grouped by settlement coin because cross-coin P&L cannot be summed, plus the healthy/watch/reduce/critical verdict logic. It omits error/limit behavior (the 50-position cap lives only in the schema) and any latency or computation caveats.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the aggregation outputs before explaining the linear/inverse math distinction, and every sentence carries information. It is long, but the length is justified by the tool's complexity and the need to explain the settlement-coin grouping rationale. No filler or repetition of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the right thing by naming the returned fields and the verdict categories. It covers the core call path adequately for a 2-parameter tool; the remaining gaps are the inverse-position coin requirement and default behaviors, which are documented in the schema rather than the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so the description must carry weight; it restates the product-wide size convention (already in the size field description) and the account_balance→margin_usage_pct relationship (also already in the schema), adding only marginal clarification. It does not explain the required vs optional mix of the item objects, so the schema largely does the heavy lifting here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Aggregates risk across multiple open positions in one call') and enumerates exactly what is produced: total notional, total P&L, total margin, margin usage %, and closest-to-liquidation position. It also differentiates from the single-position tools by explaining the roll-up relationship, so an agent can place it among 70+ siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit trigger phrasings ('how exposed am I across all my positions?', 'which of my positions is closest to liquidation?') that map directly to user intent. It does not state when NOT to use it or name a concrete alternative (e.g. run_liquidation_safety for a single position), so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_portfolio_tearsheetPortfolio TearsheetA
Read-only
Inspect

Core risk/return tearsheet from a single return series: annualized return (compounded, not linear) and volatility, Sharpe (plain + Lo 2002-corrected + Pezier-White skew/kurtosis-adjusted), Sortino, max drawdown, Calmar ratio, the drawdown-ratio cluster (Ulcer Index/Martin ratio, Pain Index/Pain ratio, Burke ratio + modified), Omega-Sharpe ratio, Upside Potential Ratio, skewness, kurtosis, Probabilistic Sharpe Ratio, win rate, best/worst single-period return. Use when user asks for a full risk summary/report on a strategy or portfolio's returns, not just one metric. Returns all of the above in one call.

ParametersJSON Schema
NameRequiredDescriptionDefault
marNoMinimum acceptable return for the Sortino ratio, same periodicity as returns (default 0)
returnsYesReturn series, one value per period
benchmark_sharpeNoBenchmark Sharpe ratio for the PSR figure, same periodicity as returns (default 0)
periods_per_yearNoPeriods per year for annualization (default 365)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already assert readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that all figures are computed from one series and returned together in a single call, which is useful context, but it does not address assumptions constraining the input (e.g. that at least three periods are needed) or any performance/limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The lead phrase and the usage clause are front-loaded, and the long metric enumeration is compressed into one dense paragraph with no filler sentences. It is verbose but every listed metric does informational work, so it is appropriately sized rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description correctly carries the burden of describing return content by listing every metric returned. It leaves minor gaps (periodicity/units of the input series, the three-item minimum) that live only in the schema, but for a read-only analytics call it is largely self-sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so mar, benchmark_sharpe, periods_per_year and returns are all documented in the schema; baseline is 3. The description only says figures come "from a single return series" and adds no interpretation of the optional knobs (e.g. how mar feeds Sortino, or benchmark_sharpe feeding PSR) beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific artifact ("core risk/return tearsheet from a single return series") and enumerates exactly which metrics are produced, so an agent knows this is a comprehensive stats bundle rather than a single-metric calculator. It only implicitly differentiates from siblings like run_sharpe_stats or run_portfolio_risk via "not just one metric" rather than naming them, keeping it at a 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use when user asks for a full risk summary/report on a strategy or portfolio's returns, not just one metric" gives explicit when-to-use plus an exclusion criterion. It stops short of naming the alternative tools to reach for a single metric (e.g. run_sharpe_stats), so it is clear context without a full routing map.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_position_sizingPosition SizingA
Read-only
Inspect

Calculate the correct position size given a maximum risk in USDT and a stop-loss price. Use when user asks "how many coins should I buy?" or "size my position so I risk exactly $X". Returns: positionSize (base), positionUsdt, marginRequired.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
leverageNoLeverage, default 1
riskUsdtYesMaximum acceptable loss in USDT
stopLossYesStop-loss price
entryPriceYesEntry price
feeOpenPctNoOpening fee fraction, default 0.0002
feeClosePctNoClosing fee fraction, default 0.0005
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. For inverse, sizeQuote is USD contracts and margin is returned in the base coin.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered and the bar is lower. The description adds value annotations do not carry by naming the return shape (positionSize, positionUsdt, marginRequired), which is meaningful because there is no output schema. It omits edge-case behavior (e.g. zero risk distance, fee handling) but that is a minor gap given the annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core calculation front-loaded, then the usage triggers, then the returns. Nothing is wasted, though the return list could be trimmed or the trigger quotes consolidated without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a stateless calculation tool with 4 required params, an 88%-documented schema and no output schema, the description supplies the missing return contract and the invocation triggers. Edge-case and inverse-contract behavior are left to the schema, which is acceptable but leaves a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 88%, so the schema already documents entryPrice, stopLoss, riskUsdt, fees, leverage and contractType. The description only paraphrases riskUsdt and stopLoss, adding no syntax, units or constraint detail beyond what is already structured — baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (calculate) plus the resource (position size) and the two driving inputs (max risk in USDT, stop-loss price). An agent can identify the operation immediately, though it never names a sibling such as run_kelly_frontier or run_forex_position_size_live to disambiguate among the many sizing/risk tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete user-intent triggers ("how many coins should I buy?", "size my position so I risk exactly $X"), which is strong positive routing guidance. It stops short of stating when NOT to use it or which alternative to pick when the user is sizing by a different objective (e.g. Kelly or risk-parity).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_prediction_market_edgePrediction Market EdgeA
Read-only
Inspect

Compares your own probability estimate for an event against a prediction market's price (manual entry, or a live Kalshi ticker, Limitless slug, or Myriad slug) and sizes a bet using fractional Kelly criterion bet sizing (default: quarter-Kelly, a standard conservative haircut on full Kelly, stated explicitly as a convention). Returns zero recommended stake whenever your probability doesn't exceed the market's price: no edge, no bet. Use when user asks "does this bet have edge?" or "how much should I stake given my probability estimate vs the market's?". Provide exactly one of marketProbabilityPct/kalshiTicker/limitlessSlug/myriadSlug. Returns: edgePct, evPerDollarStaked, fullKellyFraction, cappedKellyFraction, recommendedStakeUsd, verdict (skip_this_one/think_twice/worth_the_risk/take_it).

ParametersJSON Schema
NameRequiredDescriptionDefault
myriadSlugNoA Myriad Markets market slug to fetch the market probability from live instead of supplying it manually. Not every market has a Yes/No outcome.
bankrollUsdYesBankroll available for this bet, in USD
kalshiTickerNoA Kalshi market ticker to fetch the market probability from live instead of supplying it manually.
limitlessSlugNoA Limitless Exchange market slug to fetch the market probability from live instead of supplying it manually.
kellyFractionCapNoFraction of full Kelly to actually stake, 0.01-1. Default 0.25 (quarter-Kelly).
yourProbabilityPctYesYour own probability estimate, 0.01-99.99
marketProbabilityPctNoThe market's probability (price), 0.01-99.99. Use this OR one of the live sources below, not both.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorld=true), so the description's job is to add context — which it does: it discloses the no-edge/no-bet rule (zero stake returned), states the Kelly convention explicitly, and notes live-source fetching. It does not clarify that no bet is actually placed, only a recommendation returned, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core comparison-and-sizing action, then usage triggers, then returns. Sentence two is somewhat heavy ("default: quarter-Kelly, a standard conservative haircut on full Kelly, stated explicitly as a convention") and duplicates the schema's default, so it is efficient but not lean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description compensates by enumerating every returned field (edgePct, evPerDollarStaked, fullKellyFraction, cappedKellyFraction, recommendedStakeUsd, verdict) plus the verdict enum values. An agent has everything needed to call and interpret the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real value beyond it: the exclusive-or rule across the four market-price sources and the quarter-Kelly default convention. Marginal value is genuine, though it repeats the schema's own default statement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource pair: compares the user's own probability estimate against a market price and sizes a stake via fractional Kelly. It is clearly separable from neighbours like run_kelly_frontier, run_position_sizing, and run_market_implied_odds because it is the only one framed around prediction-market edge and bet sizing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrasing ("does this bet have edge?", "how much should I stake…") and a hard input constraint ("Provide exactly one of marketProbabilityPct/kalshiTicker/limitlessSlug/myriadSlug"). It does not name an alternative sibling tool or state when NOT to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_pre_trade_checkPre Trade CheckA
Read-only
Inspect

Full pre-trade decision card: orchestrates position sizing, breakeven, liquidation, and funding cost in one call. Use when user describes a full trade setup and asks "should I take this trade?" or "run the numbers on this setup". Provide exchange+symbol to fetch live funding rate automatically. For an R:R-graded verdict on an entry, stop and target, see workflow.run_risk_reward. Returns: positionSize, breakeven, liquidationPrice, fundingCost, overnightBreakevenShift, verdict.

ParametersJSON Schema
NameRequiredDescriptionDefault
mmrNoMaintenance margin rate, default 0.005
sideYes
symbolNoPerpetual symbol, e.g. "BTCUSDT".
exchangeNoExchange code, e.g. "binance" or "bybit". Used to fetch live funding rate if funding_rate is omitted.
leverageYesLeverage multiplier
risk_pctYesRisk as % of balance, e.g. 1.0 = 1%
stop_lossYesStop-loss price (positive)
hold_hoursNoExpected hold time in hours for overnight shift calc. Default 8.
entry_priceYesEntry price (positive)
contractTypeNolinear = USDT-margined (default), inverse = coin-margined (e.g. Bybit BTCUSD). All returned figures (notional, margin, risk_amount, funding_cost_*) stay USD-denominated either way; recommended_size is USD notional (contracts) for inverse.
fee_open_pctNoOpening fee fraction, default 0.0002
funding_rateNoFunding rate per 8h as decimal, e.g. 0.0001. If omitted, fetched live from exchange.
fee_close_pctNoClosing fee fraction, default 0.0005
account_balanceYesTotal account balance in USDT

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false, and a closed world, so the safety profile is covered. The description adds real behavioral context beyond that: providing exchange+symbol will fetch a live funding rate automatically, and it enumerates the returned fields (positionSize, breakeven, liquidationPrice, fundingCost, overnightBreakevenShift, verdict) that no output schema documents. It stops short of covering failure modes (e.g. what happens if the live funding lookup fails) or default-dependent behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the tool's scope, then usage triggers, then the sibling pointer, then returns. Three dense sentences with almost no waste, though the trailing return-field list is close to a data dump and could be trimmed or moved to an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter orchestration tool with no output schema, the description conveys the orchestration scope, the live-fetch behavior, and the return shape, which is what an agent needs to select and invoke it. It is slightly thin on how the optional parameters (mmr, hold_hours, fee pcts) affect the verdict, but the schema covers those.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 93%, so the schema already documents nearly every parameter, including the funding-rate fallback and contractType semantics. The description's 'Provide exchange+symbol to fetch live funding rate automatically' largely restates what the exchange/symbol/funding_rate schema descriptions already say, adding little beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — a consolidated 'pre-trade decision card' that orchestrates position sizing, breakeven, liquidation, and funding cost in one call — and explicitly distinguishes itself from workflow.run_risk_reward, which gives an R:R-graded verdict instead. An agent can tell the two workflow tools apart without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the trigger condition in user language ('should I take this trade?', 'run the numbers on this setup') and points to the alternative workflow.run_risk_reward for entry/stop/target R:R grading. The when-to-use and when-to-use-something-else conditions are both explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_price_consensusPrice ConsensusA
Read-only
Inspect

Cross-exchange agreement check for one perpetual futures price: fetches the same asset USDT/USDC-margined perpetual from a reference set of 7 liquid exchanges (Binance, Bybit, OKX, Gate, KuCoin, Hyperliquid, MEXC; the chosen one is compared against the rest) and reports the median of the others, the spread across all, and how far the chosen exchange sits from that median. Says the exchanges disagree instead of guessing: status is agree (under 0.25%), warn (0.25% to 1%), diverge (1% or more) or insufficient (fewer than 3 other exchanges returned a comparable contract; inverse, dated and unlisted contracts are not compared, and that is not a verdict on the price). Use before acting on a live price: "is this Bybit BTC price in line with the market?" Symbol is the perp symbol as that exchange names it (e.g. BTCUSDT on bybit, ETH-USDT-SWAP on okx). Returns: status, summary, median, peerMedian, spreadPct, deviationPct, sources, quotes (price per exchange), thresholds_pct, as_of.

ParametersJSON Schema
NameRequiredDescriptionDefault
symbolYesThat exchange's own perpetual symbol, e.g. BTCUSDT (bybit/binance), ETH-USDT-SWAP (okx), BTCUSDC (hyperliquid)
exchangeYesExchange code of the price being checked, e.g. bybit, binance, okx, gate, kucoin, hyperliquid, mexc

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and destructiveHint=false, and the description adds substantial context beyond them: the reference exchange set, exact threshold bands (0.25%/1%), the meaning of the 'insufficient' status, and the exclusion of inverse, dated and unlisted contracts. This is exactly the kind of behavioral detail annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Long but dense and front-loaded: the core check is stated first, then the exchange set, then thresholds, then usage, then the return shape. Every sentence carries information, though the return-field enumeration is slightly list-heavy for a description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description enumerates the returned fields (status, summary, median, peerMedian, spreadPct, deviationPct, sources, quotes, thresholds_pct, as_of) and defines the status vocabulary, so an agent knows both inputs and outputs well enough to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning by stressing that the symbol must be the exchange-native perp symbol (with per-exchange examples such as BTCUSDT on bybit and ETH-USDT-SWAP on okx) and by explaining that the chosen exchange is compared against the rest. That prevents a common cross-exchange symbol mismatch.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('cross-exchange agreement check for one perpetual futures price') and immediately bounds the scope to 7 named exchanges with the chosen one compared against the rest. An agent can distinguish this from sibling workflows like run_cross_venue_arbitrage or run_pre_trade_check without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use before acting on a live price') plus a worked example question ('is this Bybit BTC price in line with the market?'). It also clarifies when NOT to treat the result as a verdict (insufficient status, non-comparable contract types), which is strong routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_risk_parityRisk Parity WeightsA
Read-only
Inspect

Risk-parity (equal or custom risk contribution) portfolio weights for N assets: given a covariance matrix (or N return series to compute one from), finds long-only weights where each asset contributes its target share of total portfolio risk. Use when user asks "what weights give each asset equal risk contribution?" or "how do I risk-parity-weight this portfolio?". A portfolio-construction calculation, not a buy/sell recommendation. Returns: weights, risk_contributions (should match risk_budgets exactly at convergence), portfolio_volatility.

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsNoOne return series per asset (2+ assets, all series the same length). Provide this OR covariance, not both.
covarianceNoDirect NxN covariance matrix, if not supplying returns[][] directly.
risk_budgetsNoTarget risk share per asset, one per asset (need not sum to 1, normalized internally). Default: equal (1/N each).

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint=true, destructiveHint=false, openWorldHint=false), so the bar is lower. The description adds genuine behavioral context beyond that: it specifies long-only output, an iterative convergence criterion, and the exact return fields including risk_contributions matching risk_budgets.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation and constraint, followed by usage triggers and the return contract. Dense but each sentence contributes; the parenthetical return-list is compact rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description carries the return-value burden and does so by enumerating weights, risk_contributions, and portfolio_volatility. Combined with full parameter coverage and a clear safety profile, an agent has what it needs; only edge-case behavior (non-convergence, asset-count limits) is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters including the returns-or-covariance exclusivity and the internal normalization of risk_budgets. The description restates the covariance/return-series duality but adds no syntax or format detail beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('finds long-only weights where each asset contributes its target share of total portfolio risk') and names the domain constraint (risk parity, equal or custom risk contribution). An agent can distinguish this from siblings like run_portfolio_risk or run_kelly_frontier without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit trigger phrasing ('what weights give each asset equal risk contribution?') and a clear exclusion ('A portfolio-construction calculation, not a buy/sell recommendation'). It stops short of naming a specific alternative sibling to route to, so it is clear context without full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_risk_rewardRisk RewardA
Read-only
Inspect

Full risk:reward analysis: takes a trade with entry, stop, AND target (all three). Calculates R:R ratio, position size, liquidation price, breakeven, and P&L at both stop and target. Returns a verdict: strong (3:1+) / good (2:1+) / marginal / poor, specifically graded on the R:R ratio. For a full setup check tied to a live exchange/symbol (including funding cost), see workflow.run_pre_trade_check; its verdict covers overall setup safety, not just R:R. Use when user asks "is this trade worth taking?" or "what's my risk reward on this setup?".

ParametersJSON Schema
NameRequiredDescriptionDefault
mmrNoMaintenance margin rate (default 0.005)
sideYes
leverageYesLeverage multiplier
risk_pctYesMax risk as % of account
stop_lossYesStop-loss price
entry_priceYesEntry price
take_profitYesTake-profit price
contractTypeNolinear = USDT-margined (default), inverse = coin-margined (e.g. Bybit BTCUSD). position_size/notional are USD notional (contracts) for inverse; pnl_at_stop/pnl_at_target come back denominated in the base coin.
fee_open_pctNoOpen fee rate (default 0.0002)
fee_close_pctNoClose fee rate (default 0.0005)
account_balanceYesAccount balance in USDT

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so safety is covered; the description adds value beyond that by disclosing the computed outputs and the verdict grading scheme (strong 3:1+ / good 2:1+ / marginal / poor). It does not cover edge cases like invalid stop/target geometry, but the behavioral context is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core operation, then inputs, then outputs, then the sibling disambiguation and trigger phrases. It is on the long side, but every sentence carries routing or output information, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must summarize returns; it does so by naming the verdict and computed metrics. For an 11-parameter calculation tool this is nearly complete, with only marginal detail (e.g., fee/MMR defaults) relying on the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 91%, so the schema already documents nearly all parameters. The description reinforces that entry, stop, AND target are all required, but adds no syntax or format meaning beyond that. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource with scope: 'Full risk:reward analysis: takes a trade with entry, stop, AND target (all three).' It enumerates the outputs (R:R ratio, position size, liquidation price, breakeven, P&L at stop/target) and names the sibling workflow.run_pre_trade_check, so an agent can distinguish it from alternatives without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the alternative ('see workflow.run_pre_trade_check') and the condition that selects it (live exchange/symbol including funding cost, overall setup safety rather than just R:R). It also gives concrete invocation triggers ('is this trade worth taking?', 'what's my risk reward on this setup?'). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_scale_outScale OutA
Read-only
Inspect

Scale-out planner: P&L, ROI, and cumulative P&L for each partial exit level. Use when user wants to take profit at multiple targets: "close 30% at $90k, 30% at $95k, 40% at $100k, what's my total P&L?". Returns: per-level pnl, weightedAvgExitPrice, totalRoi.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
exitsYes
total_sizeYesTotal position size in base currency
entry_priceYesEntry price
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. total_size is USD notional (contracts) for inverse, and per-level pnl comes back denominated in the base coin.
fee_open_pctNoOpen fee rate (default 0.0002)
fee_close_pctNoClose fee rate (default 0.0005)

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a safe, read-only, closed-world computation, so the safety burden is lifted. The description adds the return keys (per-level pnl, weightedAvgExitPrice, totalRoi), but there is no output schema and it discloses nothing about weighting assumptions, fee handling, or how inverse-contract denomination affects results beyond what the contractType schema field states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then return shape, in three compact sentences with an illustrative example that earns its space. No filler, though the example is somewhat verbose relative to the rest.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, listing the returned fields is the right compensation, and it does so. The main gap is that two parameters (side, exits[].price) are undocumented in the prose and no output schema exists to catch edge cases like minItems=2 or the inverse-denomination behavior that affects the returned pnl.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 71%, so most parameters are already documented in the schema, and the description's example implicitly conveys the exits array shape (price/pct pairs). It adds no syntax or semantics for side or exits[].price beyond the schema, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific tool type (scale-out planner) and states exactly what it computes: P&L, ROI and cumulative P&L per partial exit level. That is far more specific than the title "Scale Out", though it never distinguishes itself from adjacent siblings like run_exit_target or run_pnl_planning.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete trigger ("Use when user wants to take profit at multiple targets") reinforced by a worked example of the user utterance. There is no guidance on when NOT to use it or which sibling handles the single-target or planning case, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_scenario_planningScenario PlanningA
Read-only
Inspect

Run a scenario analysis: compute PnL for multiple price-change percentages at once. Use when user asks "show me my P&L if BTC moves -10%, -5%, +5%, +10%". Returns: array of { deltaPct, exitPrice, netPnl, roe }.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideYes
sizeYesPosition size in base asset
deltasPctYesList of price change percentages, e.g. [-10, -5, 0, 5, 10]. Must be at least -100 (exactly -100 only for linear contracts; inverse contracts need more than -100)
entryPriceYesEntry price
feeOpenPctNoOpening fee fraction
feeClosePctNoClosing fee fraction
contractTypeNolinear = USDT-margined (default), inverse = coin-margined. For inverse, size is USD contracts and pnl/fees come out in the base coin.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description earns credit for disclosing the return shape (array of { deltaPct, exitPrice, netPnl, roe }), which no output schema provides and which tells the agent exactly what comes back.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact clauses — what it does, when to use it, what it returns — front-loaded with no filler. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, description of the returned fields is the key missing piece and it is supplied. The conversion semantics for inverse contracts are left to the schema, and edge-case behavior (e.g. -100% on inverse) is undocumented, keeping it short of 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 86%, so the schema already carries the parameter meaning, including delta bounds and the inverse-vs-linear contract nuance. The description's example deltas merely restate what the schema shows, adding no new semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a scenario analysis: compute PnL for multiple price-change percentages at once'). The 'multiple percentages at once' framing separates it from single-point siblings like run_pnl_planning or run_breakeven_planning, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger phrase ('show me my P&L if BTC moves -10%, -5%, +5%, +10%'), which clearly signals when to reach for this tool. No explicit when-not or named alternative (e.g. vs run_pnl_planning) is offered, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_session_structureMarket Profile: Session StructureA
Read-only
Inspect

Market Profile day-type classifier: trend / balance / neutral_trend / normal / normal_var, from TPO, initial balance, range extension and value migration. Use for "is this a trend day or a balance day?". Returns: structure (the day-type label), description, bias, key_signals, key_levels (VAH/VAL/VPOC/IB/session high-low), tails (session-high/low rejection tails and single prints), scenario_framing, invalidation level.

ParametersJSON Schema
NameRequiredDescriptionDefault
venueYesExchange to fetch candles from when candles[] not supplied
candlesNoOptional OHLCV for the session; omit to fetch from venue (reproducible + 0 COGS when supplied)
timeframeNoCandle timeframe (default 15m)
instrumentYesSymbol, e.g. BTCUSDT
prev_candlesNoOptional OHLCV for the previous session
session_dateYesSession date YYYY-MM-DD (UTC)
value_area_ruleNoValue-area fraction 0.5–0.9 (default 0.70)

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so safety is covered. The description adds the classification inputs and a detailed return-field list, but says nothing about fetch latency, auth, or rate limits when candles are omitted and venue retrieval is required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph that is front-loaded with the purpose, then the method, then the use-case trigger, then the returns. It is longer than minimal but every clause carries information; minor tightening of the classification label list would help.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned fields (structure, bias, key_signals, key_levels, tails, scenario_framing, invalidation), which is exactly the compensation needed. It remains silent on error/fallback behavior when venue fetching fails, but is otherwise sufficient for a read-only classifier.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all seven parameters (including enums for venue and timeframe) are already documented, making 3 the baseline. The description's mention of TPO/initial balance/value migration hints at what drives the computation but adds no syntax or value guidance beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: a Market Profile day-type classifier producing trend/balance/neutral_trend/normal/normal_var. It names the exact feature set (TPO, initial balance, range extension, value migration) and the sibling namespace is large, yet this is clearly the only day-type classifier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete context trigger: 'Use for "is this a trend day or a balance day?"'. That maps the tool to a recognizable analytic question, but it names no exclusions or alternatives even though sibling tools like run_value_migration and run_breakout_acceptance overlap in the profile/auction space.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_sharpe_statsSharpe Ratio StatisticsA
Read-only
Inspect

Sharpe ratio with the Lo (2002) serial-correlation-aware annualization correction (the naive sqrt(periods_per_year) scaling overstates or understates the true annualized Sharpe when returns are autocorrelated), plus the Probabilistic Sharpe Ratio (Bailey & Lopez de Prado): the probability the true Sharpe exceeds a benchmark, adjusted for the sample's skewness/kurtosis and length, not just its point estimate. Use when user asks "what's my real annualized Sharpe, not the naive one?" or "how confident can I be this Sharpe ratio is actually good?". Returns: sharpe_period, sharpe_annualized_naive, sharpe_annualized_lo, autocorrelation_lag1, psr, skewness, kurtosis.

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsYesReturn series, one value per period
benchmark_sharpeNoBenchmark Sharpe ratio for the PSR test, same periodicity as returns (default 0)
periods_per_yearNoPeriods per year for annualization (default 365, crypto convention - trades every day)

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish that this is a read-only, non-destructive, closed-world computation. Beyond that, the description explains the method's purpose (correcting naive annualization for autocorrelation and adjusting PSR for skewness/kurtosis) and enumerates the returned fields, which is exactly the behavioral context needed since no output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's purpose, then usage, then returns. It is information-dense and every part is relevant, though the opening sentence is long and could be split for easier scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description lists the exact return fields (sharpe_period, sharpe_annualized_naive, sharpe_annualized_lo, autocorrelation_lag1, psr, skewness, kurtosis). Together with annotations covering safety and the schema covering inputs, an agent has everything needed to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description refers to periods_per_year in the formula and to a benchmark for PSR, but adds no syntax, format, or constraint details beyond what the schema provides, making the baseline of 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific computation (Sharpe ratio with Lo 2002 annualization correction and Probabilistic Sharpe Ratio), so the agent knows exactly what resource is produced. It does not explicitly differentiate from closely related siblings such as run_dsr (Deflated Sharpe Ratio), so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives two explicit user-question triggers: asking for the real annualized Sharpe and asking how confident one can be in the Sharpe. There is no statement of when NOT to use it or which sibling tool to prefer, so no exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_spread_payoffSpread PayoffA
Read-only
Inspect

Payoff, breakeven(s), and max profit/loss for a Deribit BTC/ETH vertical spread (2 legs, same option type, opposite direction, e.g. a bull call spread) or an iron condor/butterfly (4 legs: 2 calls + 2 puts) at a scenario price. Coin-settled, and the max profit/loss are genuinely NOT the flat, textbook USD-settled values: because each leg's own payoff is divided by the settlement price, (1) a debit vertical spread's peak payoff occurs exactly at its short strike, not "anywhere beyond it" - and its profit is a finite WINDOW that closes again at a high enough price, decaying back toward a full loss of the premium paid, and (2) an iron condor/butterfly's max loss is genuinely UNBOUNDED toward price->0 (maxLossCoin: null) if it has a put wing, unlike the "capped at wing width" USD-settled result - only the call side is actually bounded. Use when user asks about a bull/bear call/put spread, vertical spread, iron condor, or iron butterfly, e.g. "what's my max loss on this BTC call spread?" or "where do my iron condor breakevens sit?". Returns: structureType (vertical_spread/iron_condor/iron_butterfly, inferred from the legs given), netDebitCoin (negative = credit received), scenarioPayoffCoin, scenarioPnlCoin/Usd, breakevenPrices (0-3, ascending), maxLossCoin/maxProfitCoin (null = unbounded), isProfitable.

ParametersJSON Schema
NameRequiredDescriptionDefault
legsYesExactly 2 legs (same option type, opposite position - a vertical spread) or 4 legs (2 calls + 2 puts - an iron condor/butterfly). Each leg: { optionType: call|put, position: long|short, strike: number, premiumCoin: number }.
currencyNoUnderlying coin. Default BTC.
quantityYesNumber of spread/condor units
scenarioPriceYesUnderlying price in USD to evaluate the payoff at

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish the safe read-only profile, yet the description adds substantial behavioral context: coin-settlement semantics, why a debit vertical's profit is a finite window rather than unbounded, and why an iron condor's loss is genuinely unbounded toward price->0 (maxLossCoin: null). This is exactly the non-obvious domain behavior an agent cannot get from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the scoping sentence is front-loaded and the elaborate middle section is high-value domain caveat rather than filler. Some tightening is possible, but most sentences earn their place for a tool this non-intuitive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description enumerates the returned fields (structureType, netDebitCoin, scenarioPayoffCoin, scenarioPnlCoin/Usd, breakevenPrices, maxLossCoin/maxProfitCoin null semantics, isProfitable). For a complex multi-leg options tool, an agent has everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all four parameters including the leg shape and currency enum. The description largely mirrors the leg-count constraint and adds the netDebitCoin sign convention, but adds little parameter syntax beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computation (payoff, breakevens, max profit/loss) on a specific resource (Deribit BTC/ETH vertical spread or iron condor/butterfly), and pins the exact leg structure. An agent can distinguish it from run_options_payoff, run_straddle_strangle, and run_spread_reader without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it with trigger phrases ('Use when user asks about a bull/bear call/put spread, vertical spread, iron condor...') plus two concrete example queries. No explicit negative routing against sibling payoff tools, which keeps it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_spread_readerSpread ReaderA
Read-only
Inspect

Reads the same real-world bet's live price from 2-5 prediction-market venues at once (Kalshi, Polymarket, ADI Predictstreet, Limitless, Myriad) and reports the spread between the cheapest and most expensive. The caller supplies each venue's own identifier for what they've confirmed is the same underlying bet; this tool never auto-matches events across venues, only reads and compares prices for identifiers you provide. Use when user asks "is this bet priced differently on Kalshi vs Polymarket?" or "which venue has the best price on this?". Returns: quotes[] (venue, identifier, label, probabilityPct, available, error), availableCount, cheapestVenue, mostExpensiveVenue, spreadPct (percentage points, null if fewer than 2 quotes resolved).

ParametersJSON Schema
NameRequiredDescriptionDefault
quotesYesOne entry per venue you want to compare: 2 to 5 total. Each must genuinely be the same real-world bet; this tool does not verify that for you.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, openWorldHint). The description adds genuine behavioral context beyond them: per-quote availability/error handling, availableCount, and spreadPct being null when fewer than 2 quotes resolve. It also discloses the non-verification constraint and the live/external nature of the read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and venue list, then usage triggers, then returns. It is dense and slightly long, but since there is no output schema the returns clause earns its place. Minimal waste overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating the full return shape (quotes[], availableCount, cheapestVenue, mostExpensiveVenue, spreadPct) including the null edge case. Combined with the identifier contract and no-auto-match caveat, nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents the venue enum and identifier string. The description adds semantic responsibility beyond the schema by clarifying that each identifier must be the venue's own form of the same confirmed bet and that no verification occurs. Baseline 3 is raised by that added meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (reads live price), a precise resource (the same real-world bet across 2-5 named venues), and the output (spread between cheapest and most expensive). It explicitly distinguishes its scope from auto-matching siblings by stating it 'never auto-matches events across venues.' An agent can separate it from workflow.run_cross_venue_arbitrage or workflow.run_market_implied_odds without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete triggering user questions ('is this bet priced differently on Kalshi vs Polymarket?', 'which venue has the best price on this?'). It also states a clear boundary: the caller must supply confirmed identifiers, and the tool will not match events for you. When-to-use and when-not are both explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_straddle_strangleStraddle and StrangleA
Read-only
Inspect

Payoff, P&L, and both breakeven prices for a long or short straddle/strangle (a call + a put on the same Deribit BTC/ETH underlying, both legs the same direction) at a scenario price. A straddle is callStrike === putStrike; any callStrike > putStrike makes it a strangle, same formula either way. Coin-settled like workflow.run_options_payoff: a long position's max loss is the flat total premium paid (between the strikes, both legs worthless); max profit is technically unbounded, dominated by the put leg's payoff as price falls toward zero. Use when user asks about a straddle or strangle, e.g. "what does a BTC straddle pay off if price barely moves?" or "where are my breakevens on this strangle?". Returns: combinedIntrinsicCoin, scenarioPnlCoin, scenarioPnlUsd, upperBreakevenPrice, lowerBreakevenPrice, maxLossCoin/maxProfitCoin (null = unbounded), isStraddle, isProfitable.

ParametersJSON Schema
NameRequiredDescriptionDefault
currencyNoUnderlying coin. Default BTC.
positionYes
quantityYesNumber of straddle/strangle units (both legs sized equally)
putStrikeYesPut leg strike, USD. Must be <= callStrike.
callStrikeYesCall leg strike, USD. Equal to putStrike for a straddle, higher for a strangle.
scenarioPriceYesUnderlying price in USD to evaluate the payoff at
putPremiumCoinYesPut leg premium per contract, in the base coin
callPremiumCoinYesCall leg premium per contract, in the base coin (e.g. 0.02 BTC)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish a safe read-only, non-destructive profile, and the description adds substantial context beyond them: coin settlement, max loss equals flat total premium paid, max profit dominated by the put leg as price falls, and null-valued max fields meaning unbounded. It also enumerates the returned fields, which is valuable given there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded, opening with what is computed and leading into the straddle/strangle distinction. The example phrasings and return-field list are useful, though the single long paragraph is text-heavy and could be broken up; no sentence is truly wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by enumerating return fields (combinedIntrinsicCoin, scenarioPnlCoin/Usd, breakevens, max loss/profit, isStraddle, isProfitable) and explaining the payoff model. For an 8-parameter computation tool, nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high (88%), so the baseline is 3, but the description adds real meaning: quantity is 'both legs sized equally', putStrike must be <= callStrike, and the callStrike === putStrike rule defines a straddle vs strangle. It adds interpretation of premium semantics ('between the strikes, both legs worthless') beyond the field-level docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: computes payoff, P&L, and both breakevens for a straddle/strangle at a scenario price. It explicitly contrasts itself with sibling workflow.run_options_payoff and defines straddle vs strangle via strike relationship, so an agent can distinguish it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context with concrete user-phrasing examples ('what does a BTC straddle pay off if price barely moves?', 'where are my breakevens on this strangle?'). It references the closest sibling (run_options_payoff) as the coin-settlement analogue, but does not state explicit when-not/to-use-instead conditions versus other payoff tools like run_spread_payoff.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_swap_price_impactSwap Price ImpactA
Read-only
Inspect

Live price-impact quote for a Solana token swap: routed through Jupiter (the same aggregator real swaps use) across every pool it knows about, not a single-pool estimate. Use when user asks "how much slippage will I eat swapping X tokens?" or "what will I actually get if I sell N tokens?". Returns: outputAmount, priceImpactPct, effectivePrice, marketPriceUsd, liquidityUsd, routable (false + error if the size can't be routed at all).

ParametersJSON Schema
NameRequiredDescriptionDefault
mintYesSolana mint address of the token being sold (base58)
amountYesAmount of the token to swap, in human units (not raw base units)
outputAssetNoAsset to receive. Default USDC.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/openWorldHint/destructiveHint, and the description adds meaningful behavior beyond them: it discloses the routing mechanism (Jupiter aggregator, all known pools) and the failure mode ('routable (false + error if the size can't be routed at all)'). That failure disclosure is genuinely useful and not in the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core mechanic, then usage triggers, then the return list. The return enumeration is somewhat dense but earns its place since there is no output schema. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the return fields and the routable failure case, and annotations cover the safety profile. It is complete for calling the tool correctly, though it omits any note on response timing/refresh ('live') or rate limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented there, including the human-units caveat and the outputAsset default. The description adds nothing parameter-specific, so the baseline 3 is correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb+resource ('live price-impact quote for a Solana token swap') and immediately distinguishes its method from rivals: routed through Jupiter across every pool, 'not a single-pool estimate' — which separates it from siblings like run_orderbook_impact and run_impermanent_loss.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives two concrete user-question triggers ('how much slippage will I eat...', 'what will I actually get if I sell N tokens?'), making the use case clear. It does not name an alternative tool or state when NOT to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_token_risk_checkToken Risk CheckA
Read-only
Inspect

Token rug-pull MECHANISM check for a Solana token (mint address): can the deployer still mint supply, freeze wallets, pull liquidity, swap metadata, or has RugCheck flagged a known scam pattern (e.g. copycat token)? Fetches live facts from RugCheck (GoPlus as fallback) and returns a transparently-weighted composite score. Deliberately does NOT score holder concentration or "whale dump" impact: those are properties of any liquid market (a legit protocol's top holders are routinely treasury/vesting/exchange wallets), not rug signals; they are returned separately as informational market_context. Use when user asks "is this token a rug pull?" or "is [token] safe to buy?". This is a sourced, timestamped read of public facts, not a safety guarantee. Returns: score (0-100), verdict (clean/caution/high_risk/red_flags), verdict_summary, components breakdown, facts, market_context, sources.

ParametersJSON Schema
NameRequiredDescriptionDefault
mintYesSolana token mint address (base58)

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly/openWorld/destructive=false; the description adds substantial behavioral context beyond that: live facts sourced from RugCheck with GoPlus fallback, transparently-weighted composite score, timestamped snapshot, and an explicit caveat that it is 'not a safety guarantee'. It even discloses a deliberate scoring exclusion and why.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose and the failure modes, then the sourcing, then the exclusions, then the return shape. Dense but every clause carries information; the semicolon-chained 'does NOT score...' passage is long, costing a point on tightness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description takes on the burden and delivers it: it lists the returned fields (score, verdict, verdict_summary, components, facts, market_context, sources) and the verdict enum values. Given one simple required param, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter with 100% schema description coverage, so the schema fully documents 'mint' as a base58 Solana mint address. The description repeats but does not meaningfully extend the schema's parameter semantics; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Token rug-pull MECHANISM check for a Solana token (mint address)') and enumerates the exact failure modes it inspects (mint supply, freeze wallets, pull liquidity, swap metadata, scam patterns). An agent can distinguish it from the adjacent run_wallet_flag_check without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger conditions ('Use when user asks "is this token a rug pull?" or "is [token] safe to buy?"') and explicitly states what the tool does NOT cover (holder concentration / whale dump impact, deferred to market_context). It does not, however, name a sibling tool as an alternative for the excluded concerns.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_unsmoothingReturn UnsmoothingA
Read-only
Inspect

Return "unsmoothing" for infrequently-marked/illiquid or appraisal-based series: Getmansky-Lo-Makarov (2004) MA(2) smoothing index plus Blundell-Ward (1987) AR(1) volatility inflation, two complementary models answering "this return series looks smoother than it really is; what's the true volatility?". Use when user asks "how much is appraisal smoothing understating my real volatility?" or "what's my de-smoothed Sharpe ratio?". Returns: glm_theta (MA(2) weights), glm_smoothing_index (xi, 1=no smoothing, down to 1/3 for max MA(2) smoothing), glm_true_volatility_multiplier, glm_converged (false if the fit may be unreliable - treat that result with caution), bw_alpha (AR(1) coefficient = lag-1 autocorrelation, can be negative), bw_smoothing_detected (false when alpha<=0: no evidence of smoothing, bw_volatility_multiplier is then pinned to 1 with no correction applied rather than a misleading below-1 value), bw_volatility_multiplier, and each model's own true_stdev estimate.

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsYesThe observed (possibly smoothed) return series, at least 20 values

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover only the safety profile (readOnly, non-destructive, closed world), and the description adds rich behavioral detail: glm_converged=false means the fit may be unreliable and should be treated with caution, and bw_smoothing_detected=false pins the multiplier to 1 rather than emitting a misleading below-1 value. This edge-case disclosure is exactly the context annotations cannot provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a long description, but with no output schema present, the enumerated return-field list is load-bearing rather than padding. The purpose statement is front-loaded before the field glossary, and each sentence carries meaning, though the single dense paragraph could be more scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must explain return values, and it does so comprehensively, including convergence flags and the no-smoothing degenerate case. Nothing an agent needs in order to call this correctly or interpret its results appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (the 'returns' array documents at-least-20-values), so the schema already carries parameter meaning. The description only echoes this with 'observed (possibly smoothed) return series' and adds no format or preprocessing detail beyond the schema. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return unsmoothing') and immediately names the two models involved (Getmansky-Lo-Makarov 2004 MA(2) and Blundell-Ward 1987 AR(1)), plus the exact question it answers. An agent can distinguish it from siblings like run_garch or run_hurst_exponent without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit user-intent triggers ('how much is appraisal smoothing understating my real volatility?', 'what's my de-smoothed Sharpe ratio?') and scopes it to infrequently-marked/illiquid or appraisal-based series. It does not name a sibling alternative or a when-not-to-use condition, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_value_migrationMarket Profile: Value MigrationA
Read-only
Inspect

Market Profile value-area migration across sessions: is value migrating up, down, or overlapping (directional conviction vs balance)? Use for "is value moving higher day over day?". Returns: state, direction, migration_pct, key_levels (current vs. prior session VAH/VAL/VPOC), tails (session-high/low rejection tails and single prints), scenario_framing, invalidation level.

ParametersJSON Schema
NameRequiredDescriptionDefault
venueYesExchange to fetch candles from when candles[] not supplied
candlesNoOptional OHLCV for the session; omit to fetch from venue (reproducible + 0 COGS when supplied)
timeframeNoCandle timeframe (default 15m)
instrumentYesSymbol, e.g. BTCUSDT
prev_candlesNoOptional OHLCV for the previous session
session_dateYesSession date YYYY-MM-DD (UTC)
value_area_ruleNoValue-area fraction 0.5–0.9 (default 0.70)
lookback_sessionsNoSessions to compare, 1–5 (default 1)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so safety is covered. The description's value-add is enumerating output fields, but it says nothing about venue fetch behavior, latency, or reproducibility trade-offs (the '0 COGS when candles supplied' note lives only in the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the use case, then the return contract in one dense pass with essentially no filler. The return-field list is long but earns its place given there is no output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter analytical workflow with no output schema, the description supplies the return shape (state, direction, migration_pct, key_levels, tails, scenario_framing, invalidation) and the interpretive question, while annotations carry the safety profile. An agent has what it needs to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all eight parameters (venue, candles, timeframe, instrument, prev_candles, session_date, value_area_rule, lookback_sessions) are already documented. The description adds no parameter semantics beyond what the schema provides, which is the baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('value-area migration across sessions') and frames the analytical question (up/down/overlapping). It is clearly distinctive among siblings, but never names the nearest relative (e.g. workflow.run_session_structure), so an agent gets no explicit contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger ('Use for "is value moving higher day over day?"'), which is concrete enough to route the agent correctly. It offers no when-not condition and no named alternative, so usage is guided but not exclusionary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_var_cvarVaR and CVaRA
Read-only
Inspect

Parametric Value at Risk (VaR) and Conditional VaR / Expected Shortfall (CVaR), the variance-covariance method (assumes normally distributed returns), plus Modified VaR (Cornish-Fisher skew/kurtosis correction, Boudt/Peterson/Croux) when a return series is supplied. Supply either a return series or a mean/stdev pair directly, at a confidence level. Output is at the same periodicity as the input (no automatic annualization) - a daily return series gives a daily VaR/CVaR. Use when user asks "what's my VaR at 95%/99%?", "what's my expected shortfall on this position?", or "does my Sharpe/VaR estimate need a fat-tails correction?". Returns: var, cvar (both positive loss magnitudes; cvar >= var always), z, mean, stdev, skewness, excess_kurtosis (both null unless 3+ returns were supplied), var_modified (skew/kurtosis-adjusted VaR; null when skewness is, OR when this series' skew/kurtosis are too extreme for the Cornish-Fisher expansion to be a valid quantile - skewness/excess_kurtosis are still returned in that case).

ParametersJSON Schema
NameRequiredDescriptionDefault
meanNoMean return, if not supplying returns[] directly
stdevNoStandard deviation of returns, if not supplying returns[] directly
returnsNoReturn series (e.g. daily % returns as decimals). Provide this OR mean+stdev, not both.
confidenceNoConfidence level, 0-1 exclusive (default 0.95)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only, non-destructive, and non-open-world, but the description adds substantial behavioral detail: no automatic annualization, positivity of var/cvar, the cvar >= var invariant, and exactly when skewness, excess_kurtosis, and var_modified are null. This is rich disclosure beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is dense but front-loaded with the core purpose and method, then usage triggers, then output details. Given the absence of an output schema, the long Returns sentence is necessary and every clause carries useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a computation tool with four optional parameters, no output schema, and safety annotations already present, the description covers input modes, method variants, output fields, and edge-case null behavior. Nothing an agent needs to invoke it correctly appears to be missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter meanings are already fully documented. The description repeats the either/or input requirement and confidence level but adds little syntax or format detail beyond the schema, matching the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific computation and methods: parametric VaR/CVaR via variance-covariance, plus Modified VaR with Cornish-Fisher correction. This clearly separates it from broader risk tools like run_portfolio_risk or run_evt_tail_risk without needing to open the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage triggers such as 'what's my VaR at 95%/99%?' and 'what's my expected shortfall on this position?' and explains that either returns or mean/stdev are supplied. It does not explicitly state when not to use it or name an alternative sibling, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_wallet_flag_checkWallet Flag CheckA
Read-only
Inspect

Checks a wallet address (Solana or any of 5 EVM chains) against independent flag databases: GoPlus (malicious-address categories, all chains), Webacy (address analysis + sanctions check, all chains), and ScamSniffer (public phishing/drainer blacklist, EVM chains only), and returns each source's own facts separately, never merged into one invented score. Use when user asks "is this wallet address flagged?" or "is it safe to send to this address?". A clean result means "nothing found in these databases," not a certified-safe verdict. Returns: goplus (flags[], categoriesChecked), webacyGeneral (overallRisk, dprk/hack/ofacSanctioned, exchangeLabel), webacySanctions (status), scamSniffer (flagged; not applicable on Solana). Each source has an available flag: false + error if that source failed independently.

ParametersJSON Schema
NameRequiredDescriptionDefault
chainNoChain of the wallet address. Default solana.
walletAddressYesWallet address (base58 for Solana, 0x... for EVM chains)

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only/open-world safety profile, and the description adds substantial behavioral context beyond them: sources are never merged into one score, each source carries an `available` flag with independent error reporting, and ScamSniffer is inapplicable on Solana. This is precisely the kind of extra disclosure that earns credit over the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then use cases, caveat, and return shape. Dense but every line conveys actionable information (return fields stand in for the absent output schema). Slightly long, but nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing the return fields per source and the `available`/error semantics. Chain coverage and the Solana caveat are also addressed, leaving nothing an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds meaning to `chain` by enumerating 'Solana or any of 5 EVM chains' and clarifying that ScamSniffer is EVM-only. It adds modest value beyond the schema's enum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('checks a wallet address'), names the exact sources consulted, and scopes chain support ('Solana or any of 5 EVM chains'). It is clearly distinguishable from the adjacent token-oriented siblings like run_token_risk_check.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit trigger phrases ('is this wallet address flagged?', 'is it safe to send to this address?') and a caveat that a clean result is not a certified-safe verdict. It stops short of naming a sibling alternative or a when-not-to-use condition, so it falls just below the top mark.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

workflow.run_window_fair_valueWindow Fair ValueA
Read-only
Inspect

Theoretical fair value for a time-windowed crypto up/down contract (the shape ADI Predictstreet and Kalshi-style daily crypto markets use: pays out based on whether the settlement price finishes at/above or below a reference price pinned at window open, by a fixed close time): a cash-or-nothing digital option, priced with the standard N(d2) formula. Use this when there's no live market price to read (e.g. a venue's contract has real terms but zero trading volume) instead of a live-market odds tool. Volatility is a required manual input; there is no live implied-vol market on these contracts to pull it from. Use when user asks "what should this up/down contract be worth?" or "what's the fair probability BTC finishes above $X in N minutes?". Returns: d1, d2, probAbovePct, probBelowPct, fairPriceAboveCents, fairPriceBelowCents (cents convention, directly comparable to how these venues quote a contract).

ParametersJSON Schema
NameRequiredDescriptionDefault
currentPriceYesCurrent spot price of the coin, in USD.
volatilityPctYesAnnualized volatility, in percentage points (e.g. 50 for 50%). Required; no live source for this on these contracts.
minutesToCloseYesMinutes remaining until the window closes/settles.
referencePriceYesThe reference/pinned price the contract resolves against (the window's open price, or a stated strike).
riskFreeRatePctNoRisk-free rate in percentage points. Default 0: negligible for these short windows.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/openWorldHint/destructiveHint, so the safety profile is covered. The description adds meaningful behavioral context beyond that: volatility is a required manual input because there is no live implied-vol market to pull it from, and the payout is settled at a fixed close time. It does not discuss rate limits or failure modes, but for a pure calculation tool this is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose before the model/target-market detail, and every paragraph segment earns its place (instrument definition, routing, input requirement, return fields). It is dense and somewhat long, but not padded or repetitive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by enumerating the return fields (d1, d2, probAbovePct, probBelowPct, fairPriceAboveCents, fairPriceBelowCents) and the cents convention that makes them comparable to venue quotes. Combined with full schema coverage and annotations, nothing needed to call the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all five parameters. The description reinforces why volatilityPct is mandatory and confirms the cents convention, but adds little syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — theoretical fair value for a time-windowed crypto up/down contract — and pins down the exact instrument family (ADI Predictstreet / Kalshi-style daily crypto markets) and pricing model (cash-or-nothing digital, N(d2)). It explicitly distinguishes itself from the live-market odds sibling, so an agent can route to it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when-to-use condition ('no live market price to read, e.g. zero trading volume') and names the alternative it is preferred over ('instead of a live-market odds tool'). It also supplies example user phrasings, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Addedworkflow.run_price_consensus
  2. 1 tool update
    • Changedworkflow.run_scenario_planning1 field changed
      • changedInput schema / properties / deltasPct / description
        Previous value: -"List of price change percentages, e.g. [-10, -5, 0, 5, 10]"New value: +"List of price change percentages, e.g. [-10, -5, 0, 5, 10]. Must be at least -100 (exactly -100 only for linear contracts; inverse contracts need more than -100)"
  3. 1 tool update
    • Addedworkflow.run_cross_venue_arbitrage
  4. 2 tool updates
    • Addedworkflow.run_impermanent_loss
    • Addedworkflow.run_impermanent_loss_live
  5. 1 tool update
    • Addedworkflow.run_average_down
  6. 4 tool updates
    • Changedworkflow.run_breakout_acceptance1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.71)"New value: +"Value-area fraction 0.5–0.9 (default 0.70)"
    • Changedworkflow.run_open_analysis1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.71)"New value: +"Value-area fraction 0.5–0.9 (default 0.70)"
    • Changedworkflow.run_session_structure1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.71)"New value: +"Value-area fraction 0.5–0.9 (default 0.70)"
    • Changedworkflow.run_value_migration1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.71)"New value: +"Value-area fraction 0.5–0.9 (default 0.70)"
  7. 5 tool updates
    • Changedworkflow.run_breakout_acceptance1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.70)"New value: +"Value-area fraction 0.5–0.9 (default 0.71)"
    • Changedworkflow.run_open_analysis1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.70)"New value: +"Value-area fraction 0.5–0.9 (default 0.71)"
    • Changedworkflow.run_session_structure1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.70)"New value: +"Value-area fraction 0.5–0.9 (default 0.71)"
    • Addedworkflow.run_spread_payoff
    • Changedworkflow.run_value_migration1 field changed
      • changedInput schema / properties / value_area_rule / description
        Previous value: -"Value-area fraction 0.5–0.9 (default 0.70)"New value: +"Value-area fraction 0.5–0.9 (default 0.71)"
  8. 1 tool update
    • Addedworkflow.run_orderbook_impact
  9. 1 tool update
    • Addedworkflow.run_evt_tail_risk
  10. 1 tool update
    • Addedworkflow.run_unsmoothing

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    63 deterministic quant computation tools for autonomous financial agents. Options pricing, derivatives, risk metrics, portfolio optimization, statistics, crypto/DeFi, macro/FX, time value of money. 1,000 free calls/day, no signup required.
    74
    11
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    The verifiable risk engine for autonomous agents: deterministic, self-verifying financial calculations that an agent can delegate and prove. It covers liquidation and funding, position sizing and risk of ruin, options Greeks and margin, LP divergence, treasury concentration and depeg, execution quality checks, plus intelligence on options, DeFi, prediction markets, and transaction safety analysis.
    14 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Deterministic market-state engine for trading agents — zero LLM in the signal path. 8 tools: structural market state & phase, action gate (GO/WATCH/HOLD), entry/target/invalidation coordinates, bar-by-bar state timeline, composed view cards, and pre-trade intent validation. Every output traces to a bar-stamped ledger with a public daily self-scoring track record.
    3
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides real-time options analytics, pricing with Greeks, Monte Carlo simulations, volatility analysis, strategy backtesting, and risk metrics using actual market data from Yahoo Finance and Polygon.io.
    1
    -
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.