Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
describe_simulatorA

Describe what this simulator is, what its realism checks certify, what it cannot do, the caps on every tool and how long a run takes. Call it first in a session, before any other tool. It takes no arguments, runs no market and returns the same text on every call.

check_envelopeA

Check whether a question falls inside the range the simulator's realism was measured for, BEFORE running it. Use it whenever a conclusion leans on a horizon longer than a year, on particular statistics, on a sector-concentrated roster or on the size of a scenario's effect. Returns ok or a refusal that names the measurement behind it. Runs no market, so it answers at once.

validate_strategyA

Parse and fingerprint one strategy spec WITHOUT running it. Use it to iterate on a spec cheaply before evaluate_strategies or rank_strategies: a grammar error comes back naming the field that was wrong, and a valid spec comes back normalised with its fingerprint. A spec looks like {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}. Runs no market.

evaluate_strategiesA

Run strategies on one simulated market, beside the baseline agents on the same market, and score each one: return, P&L, the cost of its own trading in basis points, turnover and errors. The right first look, but it is ONE seed, so use rank_strategies before believing an ordering. A strategy is data, for example {"signal": {"kind": "momentum", "lookback_days": 1.0}, "portfolio": {"top_k": 5, "gross": 1.0}}, and validate_strategy checks one without running it. days 1 to 60 here (a few seconds), up to 252 through start_job; roster 2 to 120 names. Deterministic: the same arguments give the same scores.

rank_strategiesA

Score strategies across MANY seeds, beside the baseline agents, and rank them with a paired sign test on each pair. Use it after evaluate_strategies, because one seed's ordering is often luck. Costs about one evaluate_strategies call per seed: 2 to 12 seeds (default six), days 1 to 60 here, up to 252 through start_job. Returns each entrant's record across the seeds (median P&L, seeds ahead of buy-and-hold) and each pair's sign test. Deterministic.

list_scenariosA

List the shipped stress scenarios, the scenario constructors and every intervention target, with what each target was measured to reach. Read it before build_scenario or run_stress_scenario: each shipped scenario's first_event_day sets the shortest useful run, and four targets have effects too small to see over a hundred days. Takes no arguments and runs no market.

build_scenarioA

Author a custom scenario and see what it resolves to before running it. Give a macro PATH as hold, ramp and step instructions in steps, or explicit INTERVENTIONS as shocks and assumed transmission. Returns the resolved document, its fingerprint and any warnings; pass that document to run_stress_scenario as scenario. Days count from 0, so an event at day 50 needs a run of at least 51 days. Runs no market.

run_stress_scenarioA

Run strategies through a macro stress scenario, always beside the same market unshocked, and compare each strategy across the two. scenario is a shipped document by name (list_scenarios gives each one's first_event_day, and the run must be longer than that), a constructor by name (vix_shock, rate_ramp, timed with peak_day), or a document from build_scenario. A scenario whose events all fall after the run is refused. Pass fork_day to run both markets together first and start the scenario on that day, as a fork of one shared history. Use the result to detect a response, not to forecast its size. days 1 to 60 here, up to 252 through start_job. Deterministic.

explain_price_moveA

Break one day's move for each name into the 11 factor contributions that sum to the day's change in the mispricing, the log gap between the model price and fair value. Use it to ask which factors moved prices; use explain to trace one name's move down to the random draws behind it. They are the simulator's own bookkeeping, and they are not the whole price move. On the default preset most of the day's news and noise moves fair value, fair_value_shift takes that part out of the mispricing, and the fair-value move itself is not split up. Without a ticker it returns the top_n largest moves. Builds and runs its own market for up to 60 days; read-only.

explainA

Trace one name's price move on one day down to the random draws that caused it, as a tree: the day's log move at the top, then each factor, then the draw addresses beneath them. Every node can be replayed, and every number is measured by running the day again. Use explain_price_move to see which factors moved prices across the roster; use this for one name when you need to know which draws moved those factors. depth sets how much of the tree the render text shows. Builds and runs its own market for up to 60 days; read-only and deterministic.

build_universeA

Build a roster of companies and preview it, either generated from a size and seed (optionally concentrated on chosen sectors) or from explicit instruments you supply. Use it when the default random roster will not do, for example to test one sector or your own companies. Returns a universe document that every run tool takes as universe, with the roster's fingerprint and any envelope warning. Runs no market.

start_jobA

Start a long run of evaluate_strategies, rank_strategies or run_stress_scenario in the background and get a job id back immediately. This is the ONLY way to run to the certified 252-day horizon; a direct call is capped at 60 days so it can answer inside a conversation. The arguments are checked before the job starts, and the response estimates its run time. At most 2 jobs run at once and the last 32 are kept, in this server's memory only. Poll with check_job.

check_jobA

Check a background job started by start_job. Returns its status and, once it has finished, the full result in the same form the direct tool returns. Omit job_id to list every job this server still holds. Changes nothing, so it is safe to poll.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4.5/5.0

Scored across 13 tools

Disambiguation5/5

Each tool targets a distinct phase: info, validation, single-seed evaluation, multi-seed ranking, scenario authoring/execution, explainability, universe building, and job management. Overlaps like evaluate_strategies/rank_strategies/start_job are explicitly clarified by scope (one seed vs many seeds vs long-running background).

Naming Consistency4/5

All names are snake_case and most follow a verb_noun pattern (validate_strategy, run_stress_scenario, start_job). The only deviation is the bare verb 'explain', which slightly breaks the noun-phrase pattern.

Tool Count5/5

13 tools fit the simulator's breadth: info, validation, evaluation, ranking, scenarios, explainability, universe, and jobs. No tool appears redundant, and the count is well-scoped for a complex domain.

Completeness4/5

Core lifecycle is covered: preflight, spec validation, evaluation/ranking, scenario building/running, explainability, custom universes, and long-run job management. Minor gaps remain, such as no job cancellation/deletion and no explicit strategy-kind enumeration, but agents can work around them.

Maintenance

ActivityActive
ResponsivenessResponsive