Rank strategies across many seeds
rank_strategiesScore and rank trading strategies across multiple simulation seeds after evaluation, comparing each against baselines with paired sign tests to avoid one-seed luck.
Instructions
Score strategies across MANY seeds, beside the baseline agents, and rank them with a paired sign test on each pair. Use it after evaluate_strategies, because one seed's ordering is often luck. Costs about one evaluate_strategies call per seed: 2 to 12 seeds (default six), days 1 to 60 here, up to 252 through start_job. Returns each entrant's record across the seeds (median P&L, seeds ahead of buy-and-hold) and each pair's sign test. Deterministic.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Trading days to run: 1 to 60 in a direct call, up to 252 (the certified horizon) through start_job. | |
| seeds | No | Simulation seeds, 2 to 12 of them. Every entrant trades the same market on each seed. Omit for [1, 2, 3, 4, 5, 6]. | |
| universe | No | A roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors. | |
| strategies | Yes | Strategies to run, keyed by a name you choose. Each value is a strategy spec, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. Signal kinds: hold, random, momentum, mean_reversion, oracle, blend. At most 8. The baseline names (buy_and_hold, random, momentum, mean_reversion, oracle) are taken. Check a spec with validate_strategy before running it. | |
| max_leverage | No | Cap on gross exposure as a multiple of net worth. null removes the cap, and the result then warns that trading size alone can win. | |
| steps_per_day | No | Decision points per trading day, 1 to 22. Each entrant is asked for orders at each one. A step is 65 minutes, so 6 cover the trading session, and days times steps may be at most 360 in a direct call. | |
| universe_seed | No | Seed that generates the roster, separate from the simulation seed. Ignored when `universe` is given. | |
| universe_size | No | Names in a generated roster, 2 to 120. Ignored when `universe` is given. | |
| universe_sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||