Skip to main content
Glama

Evaluate strategies on one market

evaluate_strategies
Read-onlyIdempotent

Run multiple trading strategies on one simulated market beside baseline agents, scoring each on return, P&L, cost in basis points, turnover and errors. Use as a first look before ranking across seeds.

Instructions

Run strategies on one simulated market, beside the baseline agents on the same market, and score each one: return, P&L, the cost of its own trading in basis points, turnover and errors. The right first look, but it is ONE seed, so use rank_strategies before believing an ordering. A strategy is data, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}, and validate_strategy checks one without running it. days 1 to 60 here (a few seconds), up to 252 through start_job; roster 2 to 120 names. Deterministic: the same arguments give the same scores.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
cashNoStarting cash for each entrant, in currency.
daysNoTrading days to run: 1 to 60 in a direct call, up to 252 (the certified horizon) through start_job.
seedNoSimulation seed, an integer from 0 to 2**64 - 1. The same seed and arguments give the same result.
universeNoA roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors.
strategiesYesStrategies to run, keyed by a name you choose. Each value is a strategy spec, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. Signal kinds: hold, random, momentum, mean_reversion, oracle, blend. At most 8. The baseline names (buy_and_hold, random, momentum, mean_reversion, oracle) are taken. Check a spec with validate_strategy before running it.
max_leverageNoCap on gross exposure as a multiple of net worth. null removes the cap, and the result then warns that trading size alone can win.
steps_per_dayNoDecision points per trading day, 1 to 22. Each entrant is asked for orders at each one. A step is 65 minutes, so 6 cover the trading session, and days times steps may be at most 360 in a direct call.
universe_seedNoSeed that generates the roster, separate from the simulation seed. Ignored when `universe` is given.
universe_sizeNoNames in a generated roster, 2 to 120. Ignored when `universe` is given.
universe_sectorsNoLowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so.
include_baselinesNoAdd the baseline agents (buy_and_hold, random, momentum, mean_reversion, oracle) to the same market. On by default, because a return means little without buy-and-hold's beside it.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, and the description adds substantial behavioral context on top: determinism ('the same arguments give the same scores'), runtime ('a few seconds'), horizon caps and job escalation, the 8-strategy cap, taken baseline names, and the warning that null max_leverage lets trading size alone win.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence carries the core action and outputs immediately, then alternatives and constraints follow in a logical order. A few facts (days 1-60, roster 2-120, the strategy example) duplicate the schema, which costs a little density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present the description need not explain return values, and it covers the remaining gaps an agent needs: scope (one seed), determinism, runtime, limits, escalation path, spec format, and when to prefer a sibling. Nothing required for a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents every parameter; the description's value-add is the concrete strategy-spec example, the enumeration of signal kinds, the note that baseline names are reserved, and the 2-120 roster bound. These mostly reinforce rather than extend schema text, so it sits just above the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (run/evaluate) and resource (strategies on one simulated market) plus the exact scoring outputs (return, P&L, cost in bps, turnover, errors). It explicitly says this is one market/one seed and contrasts itself with rank_strategies, so an agent can separate it from siblings without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('the right first look'), the reason not to trust the output alone ('ONE seed, so use rank_strategies before believing an ordering'), and the pre-flight alternative ('validate_strategy checks one without running it'). It also routes long horizons to start_job and bounds days to 1-60 for direct calls.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.