tradefloor
tradefloor
tradefloor is a market simulator you can run a strategy against. It has a Rust core and a Python API.
Give it a seed and a list of companies. It runs a market forward: prices, a limit order book, fills, and an economy that moves each day. Your orders match against the book's depth, so your trades move the price.
Real market data can't tell you what would have happened if you had traded
differently, or what caused a move. tradefloor can, because it computed every
price. You can fork a running market, change one thing in one branch (a rate
rise, a liquidity crisis, a different agent), and measure where the two
branches came apart. engine.truth() splits each move in the gap between a
price and the model's fair value into eleven factors, and
engine.explain(ticker, day) breaks down the move in the traded price, two
records no historical dataset carries.
Documentation is at https://docs.tradefloor.dev.
Install
pip install tradefloorThere are wheels for Linux, macOS and Windows on CPython 3.11+, and no
dependencies. The same engine is a Rust crate (cargo add tradefloor).
Optional extras add the MCP server (tradefloor[mcp]), Arrow output
(tradefloor[arrow]), the Gymnasium environment (tradefloor[rl]) and one
extra per agent framework.
The API may change before 1.0. Model changes ship as new presets, so a market with no agent orders in it replays exactly on its named preset in later releases. tradefloor was called pretium until 0.5.0. Versions up to 0.4.3 still install under that name, and results recorded with them still replay.
Related MCP server: Sablier MCP Server
A first run
import tradefloor as tf
universe = tf.Universe.random(40, seed=111)
spec = tf.StrategySpec.momentum(lookback_days=1.0, top_k=5)
scores = tf.evaluate({"mine": spec}, seed=7, universe=universe, days=10)
scores["mine"].return_pct # what it made
scores["mine"].impact_bps # what its own footprint cost
scores["mine"].strategy_fingerprint # sha256, cite this
scores["mine"].errors # each step that raised or was refused
scores["mine"].sharpe # annualised, from the daily closes
scores["mine"].time_in_market # share of steps holding a positionThat result comes from one random market, so it says as much about the seed
as about the strategy. tf.rank runs many seeds and compares strategies with
a paired sign test. Add tf.baselines.reference_agents() to the entrants and
tf.versus_buy_and_hold(scores) reads each score against buy-and-hold on
the same market.
A Python agent is any object with act(obs) that returns orders: a number
of shares for a market order, tf.Limit(quantity, price) or tf.Cancel().
It sees a read-only view of the market and its own portfolio, and
obs.history holds a daily bar per name. A bar's close is the day's last
print. On pt-v20 the close then re-marks every name, so the next day starts
15 bp away at the median on a 20-name roster.
docs/AGENTS.md
covers what the view holds, how trades are charged, the framework adapters
(OpenAI Agents SDK, PydanticAI, LangGraph, FinRobot) and how scoring works.
The demo
The examples are in this repository and not in the package, so clone it first:
git clone https://github.com/simoncoombes/tradefloor
cd tradefloor
python examples/rate-shock/counterfactual.pyIt runs an agent in a controlled market, checkpoints the world and forks it, raises rates by 200 bp in one branch, and compares what the same agent does next. It prints nine checks that the two branches started identical, the step at which the agent's behavior changed, and the two branches side by side. The run takes under five seconds of CPU and needs no keys and no network. The walkthrough is Your first counterfactual experiment.
Contents
| why each price moved: eleven factors that sum to the mispricing's move, to 1e-16 |
| how each trade price came about: the shock, and the order book depth that absorbed it |
counterfactual TCA | your trading cost, from the same seed run with your orders and without them |
| many seeds, paired sign tests |
| what a reader needs to replay a run, checked by |
| fork a running experiment, change one variable, and measure where the two came apart |
scenarios | seven packaged shocks, and a file format for your own |
MCP server | thirteen read-only tools for a coding agent, scenarios included |
more | a Gymnasium environment, Arrow output, checkpoints, SEC EDGAR data, simulated rate indices, a browser build |
Drive it from an agent
pip install "tradefloor[mcp]"
claude mcp add tradefloor -- tradefloor-mcptradefloor-mcp speaks MCP over stdio, and tradefloor mcp starts the same
server. Strategies, universes and scenarios are data, so a tool argument
cannot reach code. Each result carries its own caveats. See
the MCP page.
Scenarios
engine = tf.Engine(seed=42, universe=universe)
engine.run_days(20) # a shared history first
scenario = tf.Scenario.load("liquidity_crisis") # ships with the package
control, stress = tf.branch(engine, 2)
for day in range(80):
scenario.apply(stress, day)
... # run both branchesA scenario is a file of changes to the market and the assumptions behind them. Each change targets a field the engine reads, and the file keeps the shock apart from the knock-on effects you assume follow it:
tradefloor scenario show oil_price_spike
Exogenous shocks
----------------------------------------------------------
day 50+ commodity.oil x1.4
Assumed transmission
----------------------------------------------------------
day 55..74 ramp macro.inflation +1.50pp
day 55+ macro.corporate_yield +0.50ppat in a scenario counts days from the first day it is applied, so on a
branch it counts from the branch. Most packaged files first fire on day 50.
To fire one on the first day after a fork, use scenario.starting_at(0), or
world.apply(scenario, at=0) on a World. The gaps between its events stay
the same, and the run's record keeps the packaged file's fingerprint and the
days each event fired.
tradefloor does not predict what a war, an election, an oil shock or a
recession will do to markets. You state the assumptions and it measures how an
agent behaves under them. tradefloor scenario list names the seven packaged
scenarios, and tradefloor scenario targets lists every field a scenario can
change.
Reproducibility
The same seed gives the same market on every platform. tradefloor ships its
own exp, log, pow, sin and cos, so the system's math library cannot
change a result, and each release runs a fixed simulation on five platforms
and stops if any result differs.
A shipped preset never changes, so a market with no agent orders in it replays
exactly on its named preset in every later release. Each release checks that
with a digest per preset. A run with agent orders in it replays exactly on the
same release. Across releases the promise is narrower. 0.8.5 changed how an
agent's fills reach the market, on every preset, so a traded run recorded
before 0.8.5 matches up to its first trade and differs after it. The default
preset is pt-v20, and any earlier one can be named:
eng = tf.Engine(seed=42, universe=u, model="pt-v10")To let a reader rerun a result, publish its RunManifest. It records the
version, preset, seed, universe, macro state and scenario, and reproduce()
stops on a mismatch. A manifest checks the market and carries no score. Its
result block holds the market's digest, the number of days and
draws_consumed. tf.evaluate and tf.rank write no manifest, so a
published score has to be rerun to be checked.
docs/REPRODUCIBILITY.md
has the full contract, including what a saved engine state promises when it
is restored, and
docs/SUPPORT.md
says which release to pin for a long study.
Realism
tradefloor checks its market against real ones with three named sets of
statistics, listed in
docs/STATISTICS.md.
On the default preset, pt-v20, all 19 statistics of the one-year table
(volatility, fat tails, how much stocks move together, how far the VIX jumps
after a fall) are inside the range real markets show over a year. All 14
graded statistics of the two-year panel are inside their two-year ranges.
The long-run criteria are 40 rows over 21 years for pt-v20, covering crash
depth, how long fear lasts, bear markets per decade, the 2008 and 2020
replays, the rate indices and the cost of size in the book.
pt-v20 meets all 40.
Read those claims narrowly:
The 19 of 19 is a verdict on figures pooled over 30 seeds. One seed's year often misses some of its 14 shape statistics. On seeds 101 to 116, all 14 were in range on 5 of the 16, and one seed had 8 of 14. If you run one market per condition, read
tf.envelope.intervals()for each statistic's spread across seeds.A shape statistic's range is the median of 35 real one-year windows plus or minus 2.1 trimmed standard deviations, so passing one is weak evidence. Volatility clustering is one case.
abs_return_acf1reads 0.028, below every real 2015 to 2025 window (the lowest is 0.039), and it passes because its range reaches lower than those windows do.The one-year table helped choose most of pt-v20's coefficients, so the held-out checks are the fresh seeds and the fresh set of companies the panel is repeated on.
One year is the certified horizon. Two years is graded on the two-year panel, and longer runs only by the long-run criteria. Every run on a roster opens at nearly the same VIX (17.66 on the certified roster), so the one-year figures describe years that start calm.
A driven scenario moves prices at a quarter to a half of the real size, in the right direction. Use a scenario to detect a response, and do not read its size as a forecast.
Volatility memory is weaker than real at every lag, about a quarter of real at lag 1. Nothing below the 65-minute step is calibrated.
An order sliced over a day costs far less than published studies find: 0.04 of a daily standard deviation for 10% of a day's volume in 36 slices, against 0.15 to 0.3. A schedule optimiser will overstate the value of trading slowly.
Your fills pay for the book depth they take, but that temporary impact barely reaches the printed prices. The lasting part is linear and fades, and no other trader adapts to you, so no liquidity spiral or predatory trading can arise.
tf.envelope.check(horizon_days=...) refuses a question that falls outside
a measured limit.
docs/REALISM.md
has every number behind these claims and the full table of limits.
Before you publish a result
An agent scored on naming the factor behind each day's move gets an
explanation_accuracy. On pt-v20 a constant answer scores 0.95 to 1.0, so quoteexplanation_edge, the accuracy minus that baseline, and never the accuracy alone.Agents in one
tf.evaluateortf.rankcall run one after another in one Python process, on one seed. An earlier agent can leave the price path in a class variable for a later one. An agent written to cheat can read the seed from the harness's frames throughsys._getframeand run a copy of the market ahead. Nothing flags either. The read-only market view guards only against accidents, so run each agent you did not write in its own process, through the MCP server.Every
evaluateandrankrun starts at day 0, so a rule that needs 20 days of prices sits out the first 20 while buy-and-hold is invested. Passhistory_days=20to run the market 20 days first with nobody trading.In a
Worldwith several agents, orders placed at the same step execute in label order, alphabetical, for the whole run. Rotate the labels across runs when you compare different agents in one market.There are no commissions, no borrow fee on a short and no stop orders. A stop you check at each step fills a median 26.5 bp past its level at six steps a day. Uninvested cash earns nothing unless you pass
cash_interest=True, and a negative cash balance pays the policy rate, which is below a broker's margin rate.
docs/AGENTS.md has the measurements behind each of these.
Examples
The twelve numbered examples/ are in reading order, and the test suite runs them:
Start here: one company, one year, two crises, one chart | |
Universe, engine, order book, determinism | |
Specs, baselines, ranking across seeds | |
The eleven factors that sum to the mispricing's move | |
The realism panel and the limits | |
The Gymnasium environment, and what size costs | |
TCA and the counterfactual run | |
A whole study in one file. It takes about forty seconds of CPU and needs | |
An LLM agent trading the market through the harness | |
A real 2020-21 macro path, and which fields transmit. Pinned to | |
Fork a market, raise the rate in one branch, and compare the futures | |
A scenario file applied to one branch of a fork, and what it cost |
The rate-shock/
study is the demo above, and
integrations/
runs the same kind of experiment through each agent framework, offline and
without an API key.
Documentation
https://docs.tradefloor.dev covers install, the API, the guides and how the model is measured. These pages in this repository have the detail behind the sections above:
docs/MODEL.md: the model as equations, with every coefficient's value on the default preset and where it came from
docs/STATISTICS.md: the named sets of realism statistics
docs/REALISM.md: the realism results and every measured limit
docs/REPRODUCIBILITY.md: what replays exactly, across platforms, releases and restored state
docs/AGENTS.md: writing, scoring and comparing agents
docs/SUPPORT.md: which release lines get fixes, and for how long
Contributing and support
CONTRIBUTING.md explains how to build and test the project. Its main rule is that any change to the simulated trajectory is a breaking change, however small, so a model change ships as a new preset. RELEASING.md is the release checklist.
Report a vulnerability through GitHub's security advisory form, not a public issue. SECURITY.md says what is in scope. Bugs and questions go to GitHub issues.
0.8.5 and the 0.8 patches after it are the long-term support line, with bug and security fixes for 24 months. docs/SUPPORT.md says what a support line promises and which release to pin for a long study.
Citing tradefloor
Cite the version you ran and name the preset. The same version can run several presets, and results depend on the preset.
@software{tradefloor,
author = {Coombes, Simon},
title = {tradefloor: a deterministic market simulator with a limit order book},
version = {0.9.0},
year = {2026},
url = {https://github.com/simoncoombes/tradefloor},
note = {Model preset pt-v20}
}CITATION.cff carries the same details, and GitHub's "Cite this repository" button reads it.
In the text, say which model you used, for example: "tradefloor 0.9.0, preset pt-v20, specified in its docs/MODEL.md". docs/REPRODUCIBILITY.md says how to publish a result so a reader can rerun it, and how to show a score was not tuned to its seeds.
License
tradefloor is licensed under MIT OR Apache-2.0, at your option. See
LICENSE-MIT
and
LICENSE-APACHE.
GitHub's sidebar reads Apache-2.0 because its license detection picks one
file and stops. The grant that applies is the dual one, stated in
pyproject.toml, rust/Cargo.toml and this section.
Available Tools
13 toolsbuild_scenarioBuild and preview a custom scenarioARead-onlyIdempotent
Author a custom scenario and see what it resolves to before running it. Give a macro PATH as hold, ramp and step instructions in steps, or explicit INTERVENTIONS as shocks and assumed transmission. Returns the resolved document, its fingerprint and any warnings; pass that document to run_stress_scenario as scenario. Days count from 0, so an event at day 50 needs a run of at least 51 days. Runs no market.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | The run length you intend, 1 to 60 (up to 252 through start_job). A path's table covers these days, and a scenario whose events all fall after them is refused. | |
| label | No | A name for the scenario, carried into results. | |
| steps | No | A macro PATH, pinned for the whole run. Each step is {"kind": "hold", "fields": {"vix": 30.0}}, {"kind": "ramp", "field": F, "start": X, "end": Y, "over": N, "begin": D} or {"kind": "step", "field": F, "before": X, "after": Y, "at": D}. Fields: vix, federal_funds_rate, corporate_bond_yield, inflation_rate, qe_pe_boost, qe_assets_ratio, fear_greed_index, gdp_growth, unemployment_rate, tariff_rate, oil_price, cycle, epicentre, treasury_yield_2y, treasury_yield_10y. | |
| shocks | No | INTERVENTIONS the scenario asserts happened, each {"target": T, "operation": "set|add|multiply", "value": V, "at": D, "duration": N, "shape": "impulse|hold|ramp|permanent"}. list_scenarios gives every target and what it was measured to move. | |
| transmission | No | Interventions the scenario ASSUMES followed, in the same form as shocks. The simulator treats them alike; the split records which effects are assumptions. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and 'Runs no market' reinforces that this is a pure build step. It adds real behavioral context beyond annotations: the returned artifacts (resolved document, fingerprint, warnings) and the day-counting rule that can cause a scenario to be refused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose then the two input modes then the day-count caveat and the no-market disclaimer. Dense but every clause carries information; slightly packed for a single paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed, yet the description still notes what comes back. Combined with the day-count caveat and sibling routing, an agent has what it needs to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema: the day-0 offset rule ('an event at day 50 needs at least 51 days') and the purpose of the steps/shocks/transmission split. This compensates usefully for what the schema alone conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Author a custom scenario') plus the distinctive capability of previewing the resolution before running it. It explicitly distinguishes itself from run_stress_scenario by naming it as the downstream consumer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two authoring modes (steps as a macro PATH, or shocks/transmission as INTERVENTIONS) and that the returned document is passed to run_stress_scenario. It does not explicitly say when to prefer steps over shocks, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_universeBuild and preview a rosterARead-onlyIdempotent
Build a roster of companies and preview it, either generated from a size and seed (optionally concentrated on chosen sectors) or from explicit instruments you supply. Use it when the default random roster will not do, for example to test one sector or your own companies. Returns a universe document that every run tool takes as universe, with the roster's fingerprint and any envelope warning. Runs no market.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed that generates the roster, 0 to 2**64 - 1. | |
| size | No | Names in a generated roster, 2 to 120. | |
| limit | No | How many instruments the preview lists. | |
| sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. | |
| instruments | No | Explicit rows, 2 to 120, in roster order. Each needs ticker, sector (a lowercase id), initial_price and shares_outstanding. Give eps (earnings per share, valued at a sector P/E) or, for a loss-making company, book_value_per_share: a row with neither is refused, because the model would value it at one cent. Optional: revenue_growth, avg_volume, beta, short_interest (a share count, not a fraction). When given, size, seed and sectors are ignored. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), and the description adds meaningful context beyond them: 'Runs no market,' the returned `universe` document contract that every run tool consumes, the roster fingerprint, the envelope warning, and the fact that sector concentration is a 'named envelope gap' reported in the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and construction modes, then the return contract and the 'no market' caveat. No filler sentences; every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description correctly avoids detailing return values while still pointing to the `universe` key and the fingerprint/warning fields. Usage, behavior, and the downstream contract are all covered for a 5-parameter builder tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and each parameter is richly documented in the schema itself, including the mutual-exclusion rule ('When given, size, seed and sectors are ignored'). The description only restates the same two-mode grouping at a high level, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Build a roster of companies and preview it') and immediately names the two construction modes: generated from size/seed (optionally sector-concentrated) or from explicit instruments. This distinguishes it from the neighboring build_scenario / run_stress_scenario tools, which consume a universe rather than construct one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: 'Use it when the default random roster will not do, for example to test one sector or your own companies.' It also clarifies the tool runs no market, so an agent knows not to expect simulation output. It stops short of naming a sibling alternative, so it earns 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_envelopeCheck a question against the realism envelopeARead-onlyIdempotent
Check whether a question falls inside the range the simulator's realism was measured for, BEFORE running it. Use it whenever a conclusion leans on a horizon longer than a year, on particular statistics, on a sector-concentrated roster or on the size of a scenario's effect. Returns ok or a refusal that names the measurement behind it. Runs no market, so it answers at once.
| Name | Required | Description | Default |
|---|---|---|---|
| statistics | No | Panel statistics your conclusion leans on. Known: abs_return_acf1, abs_return_acf20, abs_return_acf5, annualised_vol_pct, corr_asymmetry, corr_asymmetry_lagged, corr_persistence_acf1, cross_sectional_corr, excess_kurtosis, fear_gauge_dn1, fear_gauge_dn3, index_drift_pct, index_tail_dn3_pct, leverage_effect, return_acf1, sector_excess_corr, volume_abs_return_corr, volume_change_acf1. An unknown name is refused. | |
| horizon_days | Yes | The run length you are asking about, in trading days. The certified horizon is 252. | |
| macro_regime | No | true if the conclusion depends on the economy reaching a particular state, such as high inflation, stagflation or a policy crisis. run_stress_scenario sets this itself when a scenario drives inflation, growth or the cycle. | |
| scenario_magnitude | No | true if the conclusion depends on the SIZE of a scenario's effect rather than its direction. | |
| sector_concentrated | No | true if the roster is sector-concentrated, or the name of a measured mix: all_technology, defensive, sp500_like, tech_heavy. A mix is certified only on the preset it was measured on. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/no open world, so safety is covered. The description adds non-structured facts: it returns ok or a refusal naming the measurement behind it, and that it runs no market so it answers immediately — latency and outcome shape the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the action and the temporal constraint (BEFORE running it), then the trigger conditions, then the return/latency payoff. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be spelled out, yet the description still summarizes the outcome. Purpose, timing, triggers and latency are all covered, leaving nothing an agent needs before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description nonetheless ties parameters to decision conditions — horizon longer than a year, particular statistics, sector-concentrated roster, effect size — which adds interpretive meaning over the field-by-field schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action (check) against a specific resource (the simulator's realism envelope) and states the scope: whether a question falls inside the range realism was measured for. This is clearly distinguishable from siblings like validate_strategy or describe_simulator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering conditions — horizon longer than a year, particular statistics, sector-concentrated roster, scenario effect size — and the sequencing rule (BEFORE running it). It does not name an alternative tool or state when-not to use it, so it falls short of a full routing instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_jobCheck a background jobARead-onlyIdempotent
Check a background job started by start_job. Returns its status and, once it has finished, the full result in the same form the direct tool returns. Omit job_id to list every job this server still holds. Changes nothing, so it is safe to poll.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | No | An id from start_job, such as "job-1". Omit to list every job this server process holds. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so 'changes nothing' largely restates structured data; however it converts that into actionable polling advice ('safe to poll') and describes the two-phase return (status first, full result once finished), which is real behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler: origin and return behavior first, listing mode next, safety last. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema, complete parameter docs, and read-only annotations, the description only needs to cover origin, modes, and polling safety – and it does all three. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema itself already documents the optional job_id and the omit-to-list behavior, so the description's restatement adds little. Baseline 3 applies since the schema carries the parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('check a background job'), plus the origin relationship to start_job and both operating modes (status by job_id, list-all when omitted). An agent can distinguish this from start_job and the direct tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the triggering context clearly: it is for jobs created by start_job, and omitting job_id switches to the listing mode. It implies the alternative is calling the direct tool, but never says when to poll versus when to call the direct tool, so no explicit exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_simulatorDescribe the simulatorARead-onlyIdempotent
Describe what this simulator is, what its realism checks certify, what it cannot do, the caps on every tool and how long a run takes. Call it first in a session, before any other tool. It takes no arguments, runs no market and returns the same text on every call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered; the description goes further by disclosing that it "takes no arguments, runs no market and returns the same text on every call," which is genuine behavioral context (deterministic output, no side effects on simulator state). It stops short of saying anything about the size/format of the returned text, which is acceptable given the output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The scope enumeration is front-loaded in the first sentence, followed by the single usage instruction and the no-arguments/determinism caveat. Every clause carries information; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument documentation tool with an output schema present, the description covers scope, invocation order, and determinism, which is everything needed to select and call it. Return-value structure is correctly left to the output schema rather than duplicated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4; the description's statement that it "takes no arguments" is consistent with the empty schema and removes any doubt about needing input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (describe) and an enumerated scope: what the simulator is, what its realism checks certify, what it cannot do, tool caps, and run duration. This is unmistakably distinct from every sibling (validate_strategy, run_stress_scenario, etc.), which are action tools rather than meta-documentation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Call it first in a session, before any other tool" is an explicit, unambiguous ordering instruction that tells the agent exactly when this tool applies. No alternative needs to be named because it is a prerequisite orientation call rather than a competing workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_strategiesEvaluate strategies on one marketARead-onlyIdempotent
Run strategies on one simulated market, beside the baseline agents on the same market, and score each one: return, P&L, the cost of its own trading in basis points, turnover and errors. The right first look, but it is ONE seed, so use rank_strategies before believing an ordering. A strategy is data, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}, and validate_strategy checks one without running it. days 1 to 60 here (a few seconds), up to 252 through start_job; roster 2 to 120 names. Deterministic: the same arguments give the same scores.
| Name | Required | Description | Default |
|---|---|---|---|
| cash | No | Starting cash for each entrant, in currency. | |
| days | No | Trading days to run: 1 to 60 in a direct call, up to 252 (the certified horizon) through start_job. | |
| seed | No | Simulation seed, an integer from 0 to 2**64 - 1. The same seed and arguments give the same result. | |
| universe | No | A roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors. | |
| strategies | Yes | Strategies to run, keyed by a name you choose. Each value is a strategy spec, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. Signal kinds: hold, random, momentum, mean_reversion, oracle, blend. At most 8. The baseline names (buy_and_hold, random, momentum, mean_reversion, oracle) are taken. Check a spec with validate_strategy before running it. | |
| max_leverage | No | Cap on gross exposure as a multiple of net worth. null removes the cap, and the result then warns that trading size alone can win. | |
| steps_per_day | No | Decision points per trading day, 1 to 22. Each entrant is asked for orders at each one. A step is 65 minutes, so 6 cover the trading session, and days times steps may be at most 360 in a direct call. | |
| universe_seed | No | Seed that generates the roster, separate from the simulation seed. Ignored when `universe` is given. | |
| universe_size | No | Names in a generated roster, 2 to 120. Ignored when `universe` is given. | |
| universe_sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. | |
| include_baselines | No | Add the baseline agents (buy_and_hold, random, momentum, mean_reversion, oracle) to the same market. On by default, because a return means little without buy-and-hold's beside it. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, and the description adds substantial behavioral context on top: determinism ('the same arguments give the same scores'), runtime ('a few seconds'), horizon caps and job escalation, the 8-strategy cap, taken baseline names, and the warning that null max_leverage lets trading size alone win.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence carries the core action and outputs immediately, then alternatives and constraints follow in a logical order. A few facts (days 1-60, roster 2-120, the strategy example) duplicate the schema, which costs a little density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description need not explain return values, and it covers the remaining gaps an agent needs: scope (one seed), determinism, runtime, limits, escalation path, spec format, and when to prefer a sibling. Nothing required for a correct call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents every parameter; the description's value-add is the concrete strategy-spec example, the enumeration of signal kinds, the note that baseline names are reserved, and the 2-120 roster bound. These mostly reinforce rather than extend schema text, so it sits just above the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (run/evaluate) and resource (strategies on one simulated market) plus the exact scoring outputs (return, P&L, cost in bps, turnover, errors). It explicitly says this is one market/one seed and contrasts itself with rank_strategies, so an agent can separate it from siblings without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('the right first look'), the reason not to trust the output alone ('ONE seed, so use rank_strategies before believing an ordering'), and the pre-flight alternative ('validate_strategy checks one without running it'). It also routes long horizons to start_job and bounds days to 1-60 for direct calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explainTrace a move to its random drawsARead-onlyIdempotent
Trace one name's price move on one day down to the random draws that caused it, as a tree: the day's log move at the top, then each factor, then the draw addresses beneath them. Every node can be replayed, and every number is measured by running the day again. Use explain_price_move to see which factors moved prices across the roster; use this for one name when you need to know which draws moved those factors. depth sets how much of the tree the render text shows. Builds and runs its own market for up to 60 days; read-only and deterministic.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | The trading day to explain, 1 to 60. | |
| seed | No | Simulation seed, an integer from 0 to 2**64 - 1. The same seed and arguments give the same result. | |
| depth | No | How many levels of the tree the `render` text shows, 0 to 4. The `tree` field is always whole. | |
| ticker | No | One ticker from the roster. Omit to take the largest moves (explain_price_move) or the first name (explain). | |
| universe | No | A roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors. | |
| universe_seed | No | Seed that generates the roster, separate from the simulation seed. Ignored when `universe` is given. | |
| universe_size | No | Names in a generated roster, 2 to 120. Ignored when `universe` is given. | |
| universe_sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true, idempotent=true, closed-world, so the safety profile is covered; the description nonetheless adds real behavioral context beyond them: it builds and runs its own market capped at 60 days, is deterministic, and every node can be replayed with numbers measured by re-running the day. The 'read-only' restatement is redundant with annotations, which keeps it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads what the tool produces before the routing guidance, and every sentence carries information (output shape, sibling boundary, depth scope, build cost, determinism). Slightly dense in the middle clause, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value documentation is unnecessary, and the description covers the remaining agent-facing concerns: cost of building its own market, the 60-day horizon, determinism, and the tree/render distinction. Only the seed/universe interaction is left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameters are already documented, including ticker's omit-behavior and universe's replacement semantics. The description adds the meaning of depth in terms of the `render` text specifically ('sets how much of the tree the render text shows'), but does not meaningfully elaborate the seed/universe parameters. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Trace one name's price move on one day down to the random draws'), and describes the shape of the answer (a tree: log move, factors, draw addresses). It explicitly distinguishes itself from the sibling explain_price_move, so the agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the alternative and the exact condition that selects it: 'Use explain_price_move to see which factors moved prices across the roster; use this for one name when you need to know which draws moved those factors.' Both the when and the not-when are present, plus the granularity difference (roster vs one name).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_price_moveExplain a price move by factorARead-onlyIdempotent
Break one day's move for each name into the 11 factor contributions that sum to the day's change in the mispricing, the log gap between the model price and fair value. Use it to ask which factors moved prices; use explain to trace one name's move down to the random draws behind it. They are the simulator's own bookkeeping, and they are not the whole price move. On the default preset most of the day's news and noise moves fair value, fair_value_shift takes that part out of the mispricing, and the fair-value move itself is not split up. Without a ticker it returns the top_n largest moves. Builds and runs its own market for up to 60 days; read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | The trading day to explain, 1 to 60. | |
| seed | No | Simulation seed, an integer from 0 to 2**64 - 1. The same seed and arguments give the same result. | |
| top_n | No | How many instruments to return, largest moves first, when no ticker is given. 1 or more. | |
| ticker | No | One ticker from the roster. Omit to take the largest moves (explain_price_move) or the first name (explain). | |
| universe | No | A roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors. | |
| universe_seed | No | Seed that generates the roster, separate from the simulation seed. Ignored when `universe` is given. | |
| universe_size | No | Names in a generated roster, 2 to 120. Ignored when `universe` is given. | |
| universe_sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world. The description adds genuinely useful context beyond them: the factors are 'the simulator's own bookkeeping,' they are 'not the whole price move,' and on the default preset fair_value_shift removes news/noise from the mispricing. It also notes the 60-day market build. Return format is left to the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first sentence, then layers caveats and sibling routing. Dense and mostly earns its length, though the fair-value/mispricing explanation is slightly winding and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter factor-decomposition tool with an output schema, the description covers the interpretation caveats, sibling routing, no-ticker behavior, and read-only/self-contained nature. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the ticker/top_n interplay ('Without a ticker it returns the top_n largest moves') and the fair-value decomposition caveat that shapes interpretation of the returned factors.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource: break one day's move into 11 factor contributions that sum to the change in the mispricing. It also names the sibling it is not (explain) and clarifies the scope (mispricing, log gap between model price and fair value). An agent can distinguish it from explain without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes: 'use it to ask which factors moved prices; use explain to trace one name's move down to the random draws behind it.' It also documents the no-ticker behavior (returns top_n largest moves) and the fair_value_shift caveat, giving clear selection conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_scenariosList scenarios and intervention targetsARead-onlyIdempotent
List the shipped stress scenarios, the scenario constructors and every intervention target, with what each target was measured to reach. Read it before build_scenario or run_stress_scenario: each shipped scenario's first_event_day sets the shortest useful run, and four targets have effects too small to see over a hundred days. Takes no arguments and runs no market.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, and the description adds real domain context beyond them: it runs no market (no cost/side effects) and warns that four targets have effects too small to observe over a hundred days. These are non-obvious traits an agent would not infer from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what it lists, then usage guidance, then the behavioral caveats. Three dense sentences with no filler; each clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return format needn't be restated, and the description still summarizes the return content (targets and what each was measured to reach). Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool takes zero parameters and schema coverage is 100%, so baseline is 4. The description's 'Takes no arguments' simply confirms the empty schema rather than adding parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and three concrete resources (shipped stress scenarios, scenario constructors, intervention targets) plus what each target was measured to reach. This clearly distinguishes it from siblings build_scenario and run_stress_scenario, which the description names directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to read before build_scenario or run_stress_scenario and gives the reason (first_event_day sets the shortest useful run; four targets have imperceptible effects). Names the alternatives and the condition that selects this tool over them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rank_strategiesRank strategies across many seedsARead-onlyIdempotent
Score strategies across MANY seeds, beside the baseline agents, and rank them with a paired sign test on each pair. Use it after evaluate_strategies, because one seed's ordering is often luck. Costs about one evaluate_strategies call per seed: 2 to 12 seeds (default six), days 1 to 60 here, up to 252 through start_job. Returns each entrant's record across the seeds (median P&L, seeds ahead of buy-and-hold) and each pair's sign test. Deterministic.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Trading days to run: 1 to 60 in a direct call, up to 252 (the certified horizon) through start_job. | |
| seeds | No | Simulation seeds, 2 to 12 of them. Every entrant trades the same market on each seed. Omit for [1, 2, 3, 4, 5, 6]. | |
| universe | No | A roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors. | |
| strategies | Yes | Strategies to run, keyed by a name you choose. Each value is a strategy spec, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. Signal kinds: hold, random, momentum, mean_reversion, oracle, blend. At most 8. The baseline names (buy_and_hold, random, momentum, mean_reversion, oracle) are taken. Check a spec with validate_strategy before running it. | |
| max_leverage | No | Cap on gross exposure as a multiple of net worth. null removes the cap, and the result then warns that trading size alone can win. | |
| steps_per_day | No | Decision points per trading day, 1 to 22. Each entrant is asked for orders at each one. A step is 65 minutes, so 6 cover the trading session, and days times steps may be at most 360 in a direct call. | |
| universe_seed | No | Seed that generates the roster, separate from the simulation seed. Ignored when `universe` is given. | |
| universe_size | No | Names in a generated roster, 2 to 120. Ignored when `universe` is given. | |
| universe_sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower; the description adds genuine extra context: cost ('about one evaluate_strategies call per seed'), the 2-12 seed range, the days cap, and 'Deterministic.' It stops short of describing what happens on a failed entrant or partial-run behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action, then the rationale, then the cost and limits. Dense but every clause carries information; the run-on construction of the limits sentence is the only mild inefficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with a full output schema and complete parameter docs, the description supplies the missing operational context (cost per seed, seed/day ceilings, determinism, sequencing after evaluate_strategies). Nothing an agent needs to invoke it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter, and the description mostly restates those bounds (seeds default six, days 1-60/252, at most 8 strategies). Baseline 3 is appropriate; it adds little beyond what the schema carries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Score strategies across MANY seeds ... rank them with a paired sign test') and explicitly positions itself against the baseline agents and the sibling evaluate_strategies. An agent can distinguish it from evaluate_strategies and start_job without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('Use it after evaluate_strategies, because one seed's ordering is often luck') and names the escalation path for longer horizons ('up to 252 through start_job'). This is when-to-use plus alternatives, not just implied context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_stress_scenarioRun strategies through a stress scenarioARead-onlyIdempotent
Run strategies through a macro stress scenario, always beside the same market unshocked, and compare each strategy across the two. scenario is a shipped document by name (list_scenarios gives each one's first_event_day, and the run must be longer than that), a constructor by name (vix_shock, rate_ramp, timed with peak_day), or a document from build_scenario. A scenario whose events all fall after the run is refused. Pass fork_day to run both markets together first and start the scenario on that day, as a fork of one shared history. Use the result to detect a response, not to forecast its size. days 1 to 60 here, up to 252 through start_job. Deterministic.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Trading days to run: 1 to 60 in a direct call, up to 252 (the certified horizon) through start_job. | |
| seed | No | Simulation seed, an integer from 0 to 2**64 - 1. The same seed and arguments give the same result. | |
| fork_day | No | Run both markets together for this many days, then split them and start the scenario on the first day after the split: day `fork_day` of the run. The scenario's events keep their spacing, and every strategy trades the shared days identically in both arms. 0 to days - 1. Only for a scenario made of interventions (a shipped document, or build_scenario with `shocks`); a macro path or a constructor pins the macro from day 0, so it has no fork point. | |
| peak_day | No | Constructors only. vix_shock: the day the VIX jumps to its peak (default 10). rate_ramp: the day the policy rate and corporate yield reach their end level (default 30). A vix_shock peak must fall inside the run. | |
| scenario | Yes | A shipped document by name (list_scenarios lists them with the day each one's first event falls on, and the run must be longer than that), a constructor by name (vix_shock, rate_ramp, timed with peak_day), or a scenario document from build_scenario. | |
| universe | No | A roster document, usually the `universe` field of a build_universe result. Either {"size": n, "seed": s, "sectors": [...]} or {"instruments": [...]}. When given it replaces universe_size, universe_seed and universe_sectors. | |
| strategies | No | Optional strategies to run beside the baselines, in the evaluate_strategies form. The baseline names are taken. | |
| universe_seed | No | Seed that generates the roster, separate from the simulation seed. Ignored when `universe` is given. | |
| universe_size | No | Names in a generated roster, 2 to 120. Ignored when `universe` is given. | |
| universe_sectors | No | Lowercase sector ids to concentrate a generated roster on, for example ["technology", "energy"]. The ids: technology, financial_services, healthcare, energy, consumer_discretionary, consumer_staples, industrials, materials, real_estate, utilities, telecommunications, transportation. A concentrated roster is a named envelope gap, and the result says so. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds real behavioral context beyond them: determinism, the hard refusal rule for out-of-window scenarios, the fork mechanic where both arms trade shared days identically, and the note that a concentrated roster is a named envelope gap surfaced in the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Information-dense and front-loaded with purpose and the two-arm comparison, but it restates material already in the schema (the scenario-source enumeration and the 1-to-60/252 day limits appear in both places), which costs a little efficiency without hurting comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 10 parameters, an output schema, and full schema descriptions, the description supplies exactly the missing operational layer: valid scenario sources, refusal conditions, fork applicability, horizon routing to start_job, and determinism. Nothing an agent needs to call this correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description still adds meaning, clarifying that fork_day only applies to intervention-based scenarios (a macro path or constructor pins the macro from day 0), that peak_day is constructor-only, and that days is capped at 60 here versus 252 via start_job.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run strategies through a macro stress scenario') and immediately pins the distinctive design: every run sits beside an unshocked twin market and the two are compared. This separates it from sibling evaluation tools like evaluate_strategies and rank_strategies, which lack the shock/control contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates the three accepted scenario sources (shipped document via list_scenarios, constructor vix_shock/rate_ramp, or build_scenario output), states the refusal condition when all events fall after the run, explains when to pass fork_day, and routes the 252-day horizon to start_job. It even sets the interpretive frame: 'detect a response, not to forecast its size.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_jobStart a background simulationA
Start a long run of evaluate_strategies, rank_strategies or run_stress_scenario in the background and get a job id back immediately. This is the ONLY way to run to the certified 252-day horizon; a direct call is capped at 60 days so it can answer inside a conversation. The arguments are checked before the job starts, and the response estimates its run time. At most 2 jobs run at once and the last 32 are kept, in this server's memory only. Poll with check_job.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | The tool to run in the background. | |
| arguments | No | That tool's arguments, as a direct call takes them. days may go to 252. An unknown argument or a wrong type is refused before the job starts. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety profile (non-read-only, non-destructive, non-idempotent, closed-world). The description adds substantial behavior the annotations cannot convey: immediate job-id return, pre-flight argument validation, a run-time estimate in the response, a concurrency cap of 2 simultaneous jobs, retention of only the last 32, and that state lives in server memory only. This is exactly the added value the dimension rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the primary action and the job-id return, then constraints, then the polling pointer. No filler and nothing repeated from the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but the description still flags the job id and run-time estimate. For a long-running, resource-limited, in-memory job launcher, the concurrency limit, retention count, and polling instruction cover everything an agent needs to call and follow up correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description still adds meaning: 'days may go to 252' clarifies the range limit that the schema does not encode, and it confirms validation semantics ('an unknown argument or a wrong type is refused before the job starts') for the pass-through arguments object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (start) and resource (a background job running one of three named tools), and enumerates exactly which tools it wraps. An agent can distinguish it from evaluate_strategies/rank_strategies/run_stress_scenario directly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('the ONLY way to run to the certified 252-day horizon') and the reason the direct call is insufficient (capped at 60 days so it can answer inside a conversation). It also names the follow-up tool ('Poll with check_job'), leaving no routing inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_strategyValidate a strategy specARead-onlyIdempotent
Parse and fingerprint one strategy spec WITHOUT running it. Use it to iterate on a spec cheaply before evaluate_strategies or rank_strategies: a grammar error comes back naming the field that was wrong, and a valid spec comes back normalised with its fingerprint. A spec looks like {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. Runs no market.
| Name | Required | Description | Default |
|---|---|---|---|
| spec | Yes | One strategy spec, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. `spec_version` may be left out and is then set to the current version. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description goes further by disclosing failure behavior (grammar errors name the offending field) and success behavior (normalised spec plus fingerprint), and it asserts no market is run. Return format is partly carried by the output schema, so 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the 'without running it' constraint in the first clause, then handles alternatives and behavior in two efficient sentences. The inline example is duplicated from the schema, a minor redundancy, but each sentence otherwise earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter validation tool with an output schema and rich annotations, the description covers purpose, alternatives, failure and success behavior, and the no-side-effect guarantee. Nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is fully documented, including the optional spec_version default. The description repeats the example spec that the schema already supplies, adding no format or syntax meaning beyond it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (parse and fingerprint) and resource (one strategy spec), plus the crucial scope qualifier 'WITHOUT running it.' It explicitly differentiates from the sibling tools evaluate_strategies and rank_strategies, so an agent can select it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use: 'iterate on a spec cheaply before evaluate_strategies or rank_strategies.' It names the alternative tools and the condition (cheap iteration) that selects this one, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.1- Changed
run_stress_scenario1 field changed- added
Input schema / properties / fork_dayAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Run both markets together for this many days, then split them and start the scenario on the first day after the split: day `fork_day` of the run. The scenario's events keep their spacing, and every strategy trades the shared days identically in both arms. 0 to days - 1. Only for a scenario made of interventions (a shipped document, or build_scenario with `shocks`); a macro path or a constructor pins the macro from day 0, so it has no fork point.", + "title": "Fork Day" +}
13 tool updates
v0.1.0- First observed
build_scenario - First observed
build_universe - First observed
check_envelope - First observed
check_job - First observed
describe_simulator - First observed
evaluate_strategies - First observed
explain - First observed
explain_price_move - First observed
list_scenarios - First observed
rank_strategies - First observed
run_stress_scenario - First observed
start_job - First observed
validate_strategy
TDQS
Scored across 13 tools
Each tool targets a distinct phase: info, validation, single-seed evaluation, multi-seed ranking, scenario authoring/execution, explainability, universe building, and job management. Overlaps like evaluate_strategies/rank_strategies/start_job are explicitly clarified by scope (one seed vs many seeds vs long-running background).
All names are snake_case and most follow a verb_noun pattern (validate_strategy, run_stress_scenario, start_job). The only deviation is the bare verb 'explain', which slightly breaks the noun-phrase pattern.
13 tools fit the simulator's breadth: info, validation, evaluation, ranking, scenarios, explainability, universe, and jobs. No tool appears redundant, and the count is well-scoped for a complex domain.
Core lifecycle is covered: preflight, spec validation, evaluation/ranking, scenario building/running, explainability, custom universes, and long-run job management. Minor gaps remain, such as no job cancellation/deletion and no explicit strategy-kind enumeration, but agents can work around them.
Maintenance
Related MCP Connectors
Benchmark for AI trading agents: historic market scenarios, public leaderboard.
Research-only MCP server: your AI as a quant research desk. 90 tools, no trades, no brokers.
MCP server for OpenMM — exposes market data, account, trading, and strategy tools to AI agents
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceMCP server that provides AI agents with financial tools including real-time quotes, backtesting, technical analysis, and multi-exchange data via a simple CLI interface.1MIT

Sablier MCP Serverofficial
AlicenseAqualityDmaintenanceAn MCP server that lets AI assistants analyze portfolios, stress-test scenarios, generate synthetic market paths, and scan SEC filings — in under 2 minutes.833MIT- AlicenseAqualityCmaintenanceAn MCP server exposing a registry of paper-backed quantitative trading methods plus a deterministic, no-LLM decision helper for reproducible trading research.1315 PyPIMIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that lets you talk to your AI trading assistant in plain English to research stocks, generate trade recommendations, manage a portfolio, and execute trades through natural language.MIT