Skip to main content
Glama

crashtestyourstrategy

Challenge a strategy: find what breaks it (3-layer output)

challenge_strategy
Read-onlyIdempotent

Adversarial-evaluation primitive — the semantic integration layer of the platform. Given a strategy identifier, returns a 3-layer analysis: (1) outcome metrics in the worst regimes the strategy was evaluated against, (2) vulnerability profile in the 8-dimension strategy vulnerability ontology with severity classification, (3) descriptor attribution showing which regime descriptors most strongly couple to the strategy's failure. v1 supports only 'buy_and_hold' (the outcome matrix is built once per strategy); future versions will support arbitrary strategy specs once the parser-driven strategy backtest pipeline is wired in. Read ontology://strategy-vulnerabilities for the vulnerability vocabulary.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
strategy_idNoStrategy identifier; v1 supports only 'buy_and_hold'.buy_and_hold

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds meaningful context about the 3-layer output structure, the supported strategy limitation, and a reference to an ontology resource, enriching the behavioral profile without contradicting annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense with useful information and well-structured, starting with the core purpose and then elaborating with output layers and limitations. It could be slightly trimmed but remains efficient given the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and annotations covering safety, the description fully covers the tool's behavior, including its unique 3-layer output, supported strategies, and vocabulary reference. No significant gaps remain for an agent to invoke or interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema fully documents the single parameter with description and default, including the v1 limitation. The tool description repeats that limitation but adds no further parameter semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as an adversarial-evaluation primitive that returns a 3-layer analysis for a given strategy identifier. It enumerates the specific outputs (outcome metrics, vulnerability profile, descriptor attribution), distinguishing it from sibling tools like run_stress_test or portfolio_stress_test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for adversarial evaluation but does not explicitly state when to use this tool versus alternatives. It mentions the v1 limitation to 'buy_and_hold' as a constraint, but lacks direct comparisons or exclusions relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.1/5.0
Disambiguation4/5

Most tools target distinct resources/actions (regime introspection vs. portfolio stress vs. thesis management), but a few names like run_stress_test vs. portfolio_stress_test could cause confusion. Descriptions help clarify boundaries, but there is enough overlap to mark one point off.

Naming Consistency3/5

Names are mostly snake_case but mix verb_noun (get_dossier, run_stress_test) with noun phrases (factor_decomposition, market_regime_map). The verb style is inconsistent (get/list/run/describe/submit/challenge), though the pattern is readable. This falls between predictable and chaotic.

Tool Count4/5

16 tools is slightly above the typical 3-15 range, but the domain is broad (regime analysis, portfolio stress testing, strategy evaluation, feedback). Most tools are distinct and necessary; only a couple could be merged without loss of functionality.

Completeness4/5

The surface covers core workflows: discovering theses, stress-testing portfolios, analyzing regimes, evaluating strategy robustness, and collecting feedback. Minor gaps exist (e.g., no custom strategy builder, challenge_strategy only supports buy-and-hold), but these are explicitly noted as future work.