tastytrade-research-mcp
Research-only MCP server for tastytrade historical options research — it supplies evidence (backtests, candles, package pricing, SPX reconstruction, fill verification) and never places, replaces, or cancels brokerage orders.
Backtester jobs: list available symbols/date ranges, list backtests, create/poll/cancel research backtests, and fetch execution logs.
Trade simulation: simulate one exact historical trade via
POST /simulate-tradeand get its historical path/results.Package pricing: price 2–4 leg debit/credit verticals, iron condors, and double diagonals with exact decimal arithmetic and explicit native-package vs. synthetic-natural vs. midpoint provenance.
SPX spread workflows: normalize legs, run exact-leg historical simulations, and submit aggregate SPX Backtester jobs via relative selectors (double diagonals stay exact-simulation only; 0DTE requires explicit opt-in).
Historical candles: pull normalized DXLink OHLCV candles for an exact UTC/session window with source timestamps, gap warnings,
resampled: false, and bounded receive/buffer/output/timeout limits.Fill verification: check whether a frozen paper limit was touched over
[submitted_at, valid_until], distinguishingLIMIT_TOUCHfromCONSERVATIVE_CROSSand never rewriting the live paper event.Beyond this schema (per README): live SPX/SPXW Gamma+OI snapshots, versioned signed-GEX heuristics, historical SPXW candidate discovery/universes, historical package checkpoint/path/horizon reconstruction, execution-evidence normalization and simulation, and replay reporting.
Research-only guarantees: results are valuation/regression evidence with explicit provenance, freshness, and fail-closed warnings.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@tastytrade-research-mcpWhat backtest coverage is available for SPY?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
tastytrade-research-mcp
Research-only Model Context Protocol (MCP) server for tastytrade historical options research, bounded live SPX/SPXW option evidence, package-price evidence, and strategy regression testing.
The project is intentionally separate from the official tastytrade/tastytrade-mcp:
official tastytrade MCP: general live market data, account workflows, and broker dry-run validation
this project: historical candles, Backtester access, bounded direct SPX/SPXW Gamma+OI snapshots, package-pricing research, and post-session fill verification
never exposed here: brokerage order placement, replacement, or cancellation
MCP tools
Provider-native Backtester tools
Tool | Upstream endpoint | Purpose |
|
| Discover symbols and historical coverage |
|
| List submitted research jobs |
|
| Submit a provider-native historical strategy |
|
| Poll status and retrieve results |
|
| Retrieve trials and execution logs |
|
| Cancel only a Backtester job |
|
| Simulate one exact historical trade |
Research and regression tools
Tool | Purpose |
| Auto-chunk and merge exact SPX/SPXW DXLink Quote, Greeks, and |
| Apply an explicit versioned research signing hypothesis to one unified snapshot, with separate current-Gamma signed GEX and optional bounded spot-repriced heuristic gamma flip |
| Price verticals, iron condors, and double diagonals with explicit native/synthetic provenance |
| Reconstruct timestamp-safe historical SPXW candidates under a versioned resolution profile, with exact-timestamp Backtester fallback |
| Run the same candidate discovery over an explicit trading calendar with bounded concurrency, deadlines, immutable-cache diagnostics, and resumable partial progress |
| Return a bounded multi-strike, multi-expiration SPXW universe and optional versioned RESEARCH_ONLY DD selected-leg/matched-delta IV handoff |
| Reconstruct an exact-leg SPX package reference at an RFC3339 or IANA-local checkpoint with explicit age and skew controls |
| Return profile-aligned completed-candle package points and explicit gaps without interpolation |
| Reconstruct a frozen exact-leg candidate inventory at caller-supplied ENTRY/+3/+5 trading sessions, with an opt-in research-only model fallback for candle-missing legs |
| Normalize immutable exact-leg quote snapshots/windows into signed, provider-neutral simulated-execution inputs without claiming a fill |
| Apply one caller-frozen, hashed execution profile to immutable exact-leg evidence and return a separate simulated fill/P&L result |
| Aggregate frozen baseline/research decisions and immutable execution simulations into deterministic +3/+5-day metrics without changing grading or paper state |
| Deterministically normalize SPX legs without calling an upstream service |
| Run exact-leg SPX historical simulation and normalize its result |
| Submit supported SPX structures through relative Backtester selectors |
| Check a frozen paper limit against a caller-supplied historical package path or a forward Backtester path |
| Retrieve normalized DXLink OHLCV candles without resampling |
Use MCP tools/list for the complete JSON input schemas.
The direct live snapshot, bounded DXLink batch provenance, event/cohort
freshness contract, OI-based unsigned Gamma concentration methodology,
semantic limits, regression handoff, and opt-in live gate are documented in
docs/live-option-snapshot.md.
The separate Level 3 signing hypothesis, model/result identities,
current-Gamma versus spot-repriced evidence, bounded heuristic gamma-flip
search, and research-only live gate are documented in
docs/heuristic-signed-gex.md.
The shared profile contract and the optional seven-date live capability gate
are documented in
docs/resolution-profiles.md.
The separate Double Diagonal IV measurement contract, legacy-preserving
migration, and reproducible grading/regression handoff are documented in
docs/dd-iv-measurements.md.
Private immutable source-cache configuration, exact offline replay, and
migration guidance are documented in
docs/evidence-cache.md.
The external historical-source capability review, explicit no-spend decision,
and approval gates are documented in
docs/historical-provider-decision.md.
The quote-backed exact-leg handoff, signed inventory convention, and
valuation/simulation/broker separation are documented in
docs/historical-execution-evidence.md.
Caller-frozen execution profiles, fill statuses, exact P&L arithmetic, and
fee semantics are documented in
docs/historical-execution-model.md.
The replay acceptance matrix, denominator rules, one-week smoke manifest, and
runner guidance are documented in
docs/historical-replay.md.
Related MCP server: Nubra MCP Server
Private historical evidence cache
The historical candle, SPX candidate/universe, and exact-package tools accept
an opt-in evidence_cache policy. The server exposes 24 tools, including the
local-only historical quote normalizer, deterministic execution-model
adapter, and historical replay report builder.
READ_WRITE stores sanitized source observations and normalized results as
separate content-addressed objects; REFRESH creates a new immutable revision
and diff; CACHE_ONLY replays only caller-supplied exact manifest IDs and
never contacts the provider.
The filesystem backend is disabled until
TASTYTRADE_EVIDENCE_CACHE_DIR points to a private volume. It uses atomic
writes, read-back checksum verification, bounded provider concurrency,
request deduplication, short-lived retryable-failure indexes, and a hard disk
quota without evicting immutable evidence. Cache identity includes exact
symbols, range/as-of, provider/dataset/license scope, aggregation/session/
alignment/price type, the full resolution profile, resource policy, and
normalization/model/source revisions. Retrieval time is frozen in each
immutable manifest and never replaces bar availability time.
Run npm run report:evidence-cache for a deterministic synthetic cache
hit/miss, byte-count, normalized-hash, and provider-call-reduction report.
Approval-gated bounded history
The repository contains a provider-neutral bounded-history contract for a future approved source. It is intentionally dormant: no vendor transport, credential, paid fallback, environment configuration, or MCP tool is enabled. Construction requires an exact approval scope, confirmed field and expired-symbol entitlements, a provider-specific native-field allowlist, a fixed HTTPS host, separate credentials, and explicit cache storage/reuse permission. Synthetic tests validate the contract but do not establish external SPX/SPXW coverage.
Execution evidence contract
Every normalized research result uses execution-evidence contract 1.0.0.
The contract distinguishes:
NATIVE_PACKAGESYNTHETIC_NATURALSYNTHETIC_MID_REFERENCEHISTORICAL_OPTION_PACKAGE_REFERENCEHISTORICAL_PATHBACKTESTER_SIMULATIONBROKER_DRY_RUN
The machine-readable schema is
docs/execution-evidence.schema.json.
Migration guidance for existing paper-simulation logs is in
docs/execution-evidence-migration.md.
BROKER_DRY_RUN validates a caller-supplied order; it never implies
fillability. Synthetic midpoint evidence is valuation-only.
Package pricing
Package arithmetic uses an in-repository exact decimal implementation rather than binary floating-point arithmetic.
NATIVE_PACKAGEis selected only when the caller supplies a valid, non-crossed upstream package market.SYNTHETIC_NATURALbuys each leg at its ask and sells each leg at its bid.SYNTHETIC_MID_REFERENCEuses leg midpoints and always reportsguaranteed_executable: false.Native package bid/ask fields are never populated with synthetic arithmetic.
Every leg retains its exact symbol, action, quantity, expiration, timestamp, and provider source.
as_of, freshness, and temporal alignment are computed from the decision-critical observations.Missing, crossed, stale, future-dated, and materially misaligned quotes are returned as explicit warnings. Unusable evidence is never success-shaped.
Supported families are DEBIT_VERTICAL, CREDIT_VERTICAL, IRON_CONDOR, and
multi-expiration DOUBLE_DIAGONAL.
SPX spread adapter
The high-level adapter preserves each exact provider/OCC option symbol in
/simulate-trade requests and returns stable 1.0.0 result envelopes. Each
result repeats the immutable normalized legs and carries a deterministic
SHA-256 request_id, so persisted evidence remains attributable even though
the provider response itself contains only prices and timestamps.
Aggregate /backtests requests use relative selectors (delta,
percentageOTM, currentPriceOffset, or premium) and therefore cannot
preserve an exact historical strike or expiration. The adapter exposes that
limitation in its capability flags.
Double diagonals remain exact-simulation only because the aggregate
Backtester cannot faithfully preserve their multi-expiration identity. 0DTE
legs are rejected unless the caller explicitly sets allow_0dte: true.
This MCP supplies evidence only. It does not assign grades or make production entry decisions.
Historical SPX candidate discovery
tastytrade_discover_historical_spx_candidates is the
REGRESSION_RESEARCH bridge between timestamp-safe selector evidence and
exact contract identity.
tastytrade_discover_historical_spx_candidates_range applies that same
single-checkpoint contract to an explicit caller-supplied trading calendar.
It preserves chronological output order while using a bounded worker pool,
isolates checkpoint failures, retries only transient provider failures, and
returns an opaque continuation cursor for unresolved or deferred sessions.
Continuation scheduling completes a first pass over never-attempted sessions
before retrying failures. Retry-eligible checkpoints run before checkpoints
whose provider cooldown is still active, and unresolved retries rotate to the
queue tail so rate-limited dates cannot block later checkpoints.
All Backtester endpoints in one server process share a request-start gate.
It spaces starts by at least 100 ms by default, honors Retry-After before
X-RateLimit-Reset, and falls back to bounded jittered 10/30/90-second
cooldowns. Override only the pacing interval with
TASTYTRADE_BACKTESTER_MIN_REQUEST_INTERVAL_MS; continuation retry timestamps
remain provider- or scheduler-derived.
Per-checkpoint stage timings identify cache lookup, provider bootstrap,
contract-universe construction, candle reconstruction, selector evaluation,
and the stage interrupted by a deadline. Operational settings and diagnostic
callbacks do not alter the logical batch request ID or candle-cache
fingerprints.
The official option-chain and REST quote endpoints do not document an
historical as_of parameter. Backtester logs currently expose exact selected
strike, expiration, side, and a provider-internal symbol, but those log fields
are undocumented. The documented Backtester EntryConditions also has no
time-of-day field, and a live request containing undocumented
entryTime: "14:30:00Z" was silently ignored: the SPX trial still opened at
19:45:00Z.
For DELTA and PERCENTAGE_OTM, the adapter reconstructs a bounded SPXW
universe from generated OCC/streamer symbols and completed DXLink candles.
Candidate discovery uses the provider-supported maximum of 100 option symbols
per candle batch; the standalone bounded-universe tool retains its existing
20-symbol batch policy.
Omitting resolution_profile preserves the 5-minute cohort; strict native-hour RTH
research requires explicit HOURLY_VALUATION_RESEARCH. When DXLink labels SPXW
hour bars on its provider clock instead of the 09:30 ET session anchor, callers
must opt into the separate HOURLY_PROVIDER_ALIGNED_RESEARCH cohort; those bars
are never relabeled as session-aligned. It enforces
available_at <= as_of, profile age/skew limits, aligned call/put parity for
the forward, and contract-candle IV for Black-76-style delta. Selection is
deterministic by selector error, observation age, DTE distance, and strike.
Generated contract identity becomes eligible only when DXLink returns historical evidence for that exact symbol. The result preserves checkpoint selection time, observation age, price, contract IV, reconstructed delta, volume/OI when present, OCC identity, and field-level provenance.
Exact-timestamp Backtester selection remains the fallback for unsupported
selectors or insufficient reconstruction evidence. It never sends or trusts
entryTime, and it accepts only a trial and opening order exactly at
as_of. Stale and future trials remain fail-closed.
The capability returns HISTORICAL_SELECTOR_CANDIDATE_SET, never a
full-chain snapshot. Historical bid/ask, ATM IV surface, skew, and term
structure remain unavailable. Exact symbols are /simulate-trade compatible,
but the capability separately reports whether an exact simulation snapshot
exists at the arbitrary checkpoint.
The live 2026-08-25 07:30 PT smoke test returns four timestamp-safe CALL/PUT Delta-20 and 1%-OTM candidates without creating Backtester jobs.
The full provider findings, contract, anti-lookahead rules, and limitations
are documented in
docs/historical-spx-candidates.md.
Historical SPX candidate universe
tastytrade_get_historical_spx_candidate_universe expands checkpoint-safe
evidence from selector winners into a caller-bounded strike grid. By default
it covers the nearest min/mid/max DTE SPXW expirations, both option sides, and
a 25-point strike step.
Generated OCC identity is returned only after DXLink supplies a completed historical candle for that exact symbol. Contracts without timestamp-safe evidence are omitted and counted as coverage gaps. The endpoint preserves price, reconstructed delta, contract IV, OI, volume, observation age, and field-level provenance when available; it never selects the final spread. Delta reconstruction prefers a timestamp-aligned put/call parity forward. When parity is unavailable but historical IV exists, the universe endpoint uses the completed checkpoint SPX price as an explicitly warned zero-carry forward approximation. Missing IV remains null, and this fallback does not change the stricter selector-discovery path.
The default freshness maximum is 60 minutes. A caller may explicitly extend
it to 24 hours; older pre-checkpoint observations are then retained only as
STALE/LOW evidence. Coverage gaps, provider batch errors, and per-field
availability counts remain machine-readable, and aggregate field capability
flags are true only when every returned contract supports the field.
Both historical SPX tools accept either RFC3339 as_of or an unambiguous
IANA local_checkpoint. They preserve bar_start, bar_end,
available_at, and retrieved_at, and return the normalized resolution
profile with deterministic requested/effective cohort IDs. An opaque
versioned candidate_construction_profile is returned unchanged; this MCP
does not duplicate grading, DD-bucket, or final-leg-selection rules.
An optional dd_iv_measurement request freezes an exact four-leg Double
Diagonal candidate and returns distinct SELECTED_LEG_IV_DIFFERENCE and
MATCHED_DELTA / MATCHED_FORWARD_MONEYNESS cohorts. The handoff is always
RESEARCH_ONLY, preserves provider/derived models, timestamp lineage, and
optional immutable evidence-manifest IDs, and explicitly does not replace
legacy production term_structure.
See
docs/historical-spx-universe.md
for the input contract, reconstruction rules, and live checkpoint findings.
Historical exact-leg package reconstruction
tastytrade_get_historical_option_package_at_checkpoint selects only
completed option candles available by the requested checkpoint and preserves
each leg's exact OCC identity, derived streamer symbol, historical close, IV,
bar-start timestamp, availability timestamp, age, and provenance. Explicit
age and temporal-skew limits fail closed. Candle-derived values are always
VALUATION_ONLY; historical bid/ask is not fabricated.
The checkpoint and path results include deterministic request IDs plus requested/native/effective aggregation, session, alignment, age/skew, fallback, and cohort metadata. The default checkpoint behavior remains the existing 5-minute compatibility path. Native-hour New York RTH valuation is explicit opt-in and does not silently fall into the 5-minute cohort. Checkpoint legs also return a structured reconstruction status and failure reason. Their provenance carries the requested lifecycle, resolution profile, provider failure reasons, source revision, and exact cache manifest/content IDs when available.
tastytrade_get_historical_option_package_horizons accepts a bounded frozen
candidate inventory and a caller-supplied ordered trading-session calendar.
It resolves ENTRY, OUTCOME_3_TRADING_DAYS, and
OUTCOME_5_TRADING_DAYS by session index rather than assuming weekdays are
trading days. Every horizon reuses the same OCC symbols, actions, quantities,
roles, strategy family, and resolution profile. The result keeps candle
references labeled VALUATION_ONLY / CANDLE_REFERENCE and reports complete
package counts, missing legs by role, missing-reason counts, and coverage by
strategy, expiration, entry DTE, and resolution profile. Exact 21/35-DTE
Double Diagonals are supported without changing grading or leg selection.
An optional valuation_fallback can supply checkpoint-frozen SPX levels,
per-exact-leg IV or surface inputs, rates, dividends, immutable source IDs,
and an IV-shift uncertainty assumption. The strict candle status, package,
legs, and failure taxonomy remain unchanged. A separate valuation object
classifies each result as EXACT_PACKAGE_REFERENCE,
MIXED_OBSERVED_MODELED, or MODEL_SURFACE; observed leg values are copied
unchanged and only unavailable exact legs are modeled. Cache/provider errors,
stale observations, and alignment failures remain non-modelable.
Any package containing a modeled leg is MODEL_REFERENCE /
VALUATION_ONLY, carries guaranteed_executable: false, and is explicitly
not bid/ask, NBBO, midpoint, touch, or fill evidence. Its
execution_evidence_input can be passed to
tastytrade_normalize_historical_execution_evidence and then evaluated only
through a separately frozen REFERENCE_COST execution profile. Coverage
keeps strict package counts separate from valued exact, mixed, and fully
modeled cohorts.
tastytrade_get_historical_option_package_path emits a point only when every
leg has an exact timestamp-aligned completed bar. It never interpolates or
forward-fills missing legs. If a bounded old 1-minute DXLink replay exhausts a
configured local resource budget or the provider returns a clipped snapshot,
the result records the failed attempt and explicitly selects the finest
retrievable coarser resolution. A local budget result is not described as a
provider limit.
The returned fill_verification_path can be passed directly to
tastytrade_verify_historical_fill. Full rules and the 2026-08-27 07:30 PT
acceptance finding are documented in
docs/historical-option-package.md.
Historical fill verification
tastytrade_verify_historical_fill:
evaluates only
[submitted_at, valid_until];accepts either caller-supplied historical package points or the existing exact-leg Backtester mode;
supports both
ENTRYandEXIT;preserves
paper_order_id,checkpoint_id, andposition_idreferences;returns legacy
statusvaluesTOUCHED,NOT_TOUCHED, orNOT_VERIFIABLE, plusassessment_status: NOT_ASSESSABLEfor an insufficient path;distinguishes
LIMIT_TOUCHfromCONSERVATIVE_CROSS;reports an exact observed touch timestamp when defensible, otherwise a bounded interval for a sparse path;
records disagreement with a live paper assumption without mutating the original paper event.
All verification is marked POST_SESSION_REGRESSION and includes
HISTORICAL_EVIDENCE_ONLY_DO_NOT_REWRITE_LIVE_EVENT.
DXLink historical candles
tastytrade_get_historical_candles obtains an API quote token from
GET /api-quote-tokens, connects only to a wss:// host under
dxfeed.com, completes the DXLink handshake, and waits for the indexed-event
snapshot boundary.
The normalized output includes:
deterministic request and cohort identity;
dxlink_authrecovery provenance and a sanitizedprovider_errorwhen authentication required recovery;requested and actual UTC ranges;
source_time,bar_start,bar_end,available_at, andretrieved_atper bar;explicit instrument type and interval;
ALL, USREGULAR, or caller-definedCUSTOMsession filtering with an IANA timezone;OHLC, volume, VWAP, bid/ask volume, implied volatility, and open interest when supplied;
missing-bar, empty-result, and snapshot-truncation warnings;
resampled: false.
REGULAR requests use provider-native a=s,tho=true candle attributes, so
bars are aligned to and built only from the instrument's regular trading
session. CUSTOM windows filter complete provider bars by their source
timestamp; they are never reaggregated and include an explicit warning about
that limitation.
One-unit periods use DXLink's native normalized spelling (m, h, d, or
w). In particular, public interval 1h subscribes to native HOUR {=h};
60m remains {=60m} and is not treated as equivalent. Request/response
matching accepts only provider-defined canonical aliases such as omitted
defaults and attribute ordering differences.
The server sends fromTime as epoch milliseconds, matching the current
production DXLink service. The published AsyncAPI description currently says
seconds, but seconds cause the service to replay the full available history.
Production currently appears to ignore toTime. The client reports a
continuous-calendar replay estimate for planning, but that estimate is
explicitly advisory: it does not account for sessions, closures, sparse
options, or expiration, and never blocks a request before connecting.
Streaming state is indexed and deduplicated per symbol, retains only rows in the requested range and session, serializes message processing, and does not infer completion from timestamp order. Completion still requires DXLink snapshot protocol evidence. Each result includes:
status,snapshot_complete,snapshot_truncated, andprovider_snapshot_complete;machine-readable
failure_reasons;per-symbol and aggregate request counters for received events, valid events, unique observations, retained rows/bytes, and returned rows;
configured limits and the advisory calendar-slot estimates.
snapshot_complete means a provider END marker was observed without local or
provider truncation. It does not assert that every interval traded;
status and failure_reasons separately report missing or uncovered evidence.
The independent client controls are:
Input | Scope | Default | Maximum |
| retained/returned rows per symbol | 10,000 | 250,000 |
| aggregate Candle protocol rows per request | 10,000 | 1,000,000 |
| aggregate queued wire data plus accounted retained state | 16 MiB | 128 MiB |
| complete DXLink snapshot lifecycle | 15,000 ms | 60,000 ms |
Received-event counts include marker, remove, invalid, duplicate, and unmatched Candle rows. Unique observations count distinct in-window indexed events admitted to bounded state; retained rows reflect removals, and returned rows reflect the final normalized output. Retained-byte accounting uses serialized field sizes plus conservative per-object/index overhead rather than claiming exact V8 heap usage.
max_candles remains as a deprecated compatibility shorthand, capped at
20,000, that applies the same value to per-symbol output and aggregate receive
budgets. It cannot be combined with either explicit field. timeout_ms is a
deprecated alias for deadline_ms and cannot be combined with it.
LOCAL_RECEIVE_BUDGET_EXCEEDED, LOCAL_BUFFER_BUDGET_EXCEEDED, and
LOCAL_OUTPUT_BUDGET_EXCEEDED are client facts. PROVIDER_SNAPSHOT_SNIPPED
is provider protocol evidence. REQUESTED_WINDOW_NOT_COVERED,
SNAPSHOT_TIMEOUT, and MISSING_CONTRACT_EVIDENCE describe observed
availability; none is automatically classified as an entitlement failure.
Completed symbols remain usable when another batch symbol fails locally or is
snipped.
Canonical identity, snapshot transaction rules, the sanitized timeout root
cause, and live native-hour findings are documented in
docs/historical-candles.md.
Rate limits and retries
API quote tokens use the provider's
expires-attimestamp with a one-minute safety margin. The previous 23-hour bound remains only as a fallback when the provider omits expiry metadata.Historical candles and live option snapshots share one quote-token lifecycle in the production service. A post-
AUTHUNAUTHORIZEDinvalidates both the cached quote token and cached OAuth access token, reacquires credentials, rebuilds the WebSocket connection, and retries exactly once.Successful recovery returns
dxlink_auth.status = REFRESHEDwithretry_count = 1. A second rejection returns explicitAUTH_FAILEDdiagnostics; a timeout beforeAUTH_STATE/AUTHORIZEDisNOT_CONFIRMED. Tokens are never included.The quote-token REST request retries network errors,
429, and5xxup to three attempts with capped exponential backoff and jitter.A candle snapshot has a configurable
deadline_ms(maximum 60 seconds).Non-authentication WebSocket failures are returned explicitly and are not silently retried or merged.
Received events, returned rows, and memory are bounded independently as documented above.
DXLink permits at most 5 concurrent sessions and 100 Candle subscriptions per session. Callers should batch work rather than fan out unbounded calls.
Security model
OAuth client credentials and refresh tokens are sent only to
api.tastyworks.com. The short-lived OAuth token is sent to the fixed
Backtester host and to the tastytrade quote-token endpoint. The resulting
quote token is sent only over wss:// to a host under dxfeed.com.
Do not commit credentials. If using Node's --env-file, unset inherited
TASTYTRADE_* variables first because inherited values override the file.
All API timestamps must be RFC3339 values containing Z or an explicit UTC
offset; timezone-less timestamps are rejected.
The remote HTTP entrypoint supports two authentication modes:
api-keyfor local development and emergency rollback, using a constant-time comparison againstMCP_API_KEY;oauthfor production, validating Entra JWT signature, issuer, audience, expiration, andmcp.readscope through cached JWKS.
OAuth mode publishes RFC 9728 protected-resource metadata at both supported
well-known paths and includes that URL and the required scope in every 401
challenge. MCP is accepted only at POST /mcp, request bodies are limited to
1 MiB, and GET /healthz remains unauthenticated with health metadata only.
Requirements and setup
Node.js 22+
tastytrade OAuth API grant:
TASTYTRADE_CLIENT_IDTASTYTRADE_CLIENT_SECRETTASTYTRADE_REFRESH_TOKEN
A fully onboarded tastytrade customer for DXLink quote tokens
npm ci
npm run build
export TASTYTRADE_CLIENT_ID="..."
export TASTYTRADE_CLIENT_SECRET="..."
export TASTYTRADE_REFRESH_TOKEN="..."
npm startFor remote Streamable HTTP:
export MCP_API_KEY="$(openssl rand -hex 32)"
export MCP_HTTP_HOST=127.0.0.1
export MCP_HTTP_PORT=8000
npm run start:httpConnect to http://127.0.0.1:8000/mcp with
Authorization: Bearer <MCP_API_KEY>. Azure Container Apps deployment and
Key Vault guidance are documented in
docs/azure-deployment.md.
Production OAuth configuration additionally requires MCP_PUBLIC_URL,
OAUTH_ISSUER, OAUTH_JWKS_URL, OAUTH_AUDIENCE,
OAUTH_REQUIRED_SCOPE, and OAUTH_TOKEN_SCOPE. ChatGPT discovers the Entra
authorization server from
/.well-known/oauth-protected-resource, then sends its access token in the
standard Authorization header.
For an MCP client:
{
"mcpServers": {
"tastytrade-research": {
"command": "node",
"args": ["/absolute/path/to/tastytrade-research-mcp/dist/index.js"],
"env": {
"TASTYTRADE_CLIENT_ID": "...",
"TASTYTRADE_CLIENT_SECRET": "...",
"TASTYTRADE_REFRESH_TOKEN": "..."
}
}
}
}Development
npm ci
npm run typecheck
npm testThe regression suite covers clean and degraded package markets, all supported spread families, regular and custom candle sessions, both fill models, ambiguous paths, Backtester normalization, and MCP tool dispatch.
Upstream documentation
Backtesting guide: https://developer.tastytrade.com/docs/guides/backtesting/
Backtesting API: https://developer.tastytrade.com/reference/backtesting/
Streaming guide: https://developer.tastytrade.com/docs/guides/stream-market-data/
Streaming concepts: https://developer.tastytrade.com/docs/concepts/streaming/
Rate limits: https://developer.tastytrade.com/docs/guides/rate-limits-and-backoff/
Disclaimer
This is an independent research project and is not an official tastytrade product. Historical simulations can differ materially from live execution because of fills, liquidity, spreads, slippage, data availability, and model assumptions.
MIT License.
Available Tools
13 toolstastytrade_cancel_backtestA
Cancel a running research backtest. This affects only the Backtester job and never a brokerage order.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Backtest ID returned by tastytrade. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The description clarifies the scope (only backtest jobs, not orders), which is critical behavioral context. However, it does not mention side effects such as whether the cancellation is irreversible or if resources are freed, but given the tool is a cancellation, the key transparency is the scope distinction, which is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is concise and front-loaded with the core purpose ('Cancel a running research backtest'), followed by an important clarification. There is zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple cancel operation with one well-documented parameter, the description effectively covers the core usage. The only minor gap is not specifying what happens after cancellation (e.g., no return value description), but given the lack of output schema and the simplicity, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage for the single parameter, describing it as 'Backtest ID returned by tastytrade.' The description does not add extra semantics beyond the schema, but since the schema is complete, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (cancel), the resource (a running research backtest), and explicitly distinguishes it from brokerage orders. This makes it unambiguous and differentiates it from sibling tools like tastytrade_list_backtests or tastytrade_create_backtest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool (to cancel a running backtest) and what it does NOT affect (never a brokerage order). This is a clear 'when/not' guideline, which is particularly important given the presence of tastytrade_simulate_trade and other order-related tools in the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_create_backtestA
Create a tastytrade historical options backtest. This is research-only and does not place brokerage orders.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Raw request body accepted by tastytrade POST /backtests. Kept provider-native so new upstream fields remain usable. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the disclosure burden and does meaningfully reveal that no brokerage orders are placed, which is a non-obvious and important behavioral trait. Still, it does not mention whether the backtest is executed asynchronously, how long it may take, or what the response represents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. It front-loads the action and resource, then immediately states the key caveat, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a thin API wrapper, especially with sibling names like tastytrade_get_backtest and tastytrade_cancel_backtest hinting at the lifecycle. However, it does not explain what the tool returns or that the backtest may run asynchronously, which would help an agent know how to poll or retrieve results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter and the schema already describes it thoroughly as the raw POST /backtests request body. The tool description itself adds no parameter-level guidance, so it stays at the baseline for full schema coverage. The provider-native note in the schema is helpful but is already part of the structured input definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create') and the resource ('a tastytrade historical options backtest'), so the core purpose is unambiguous. It also adds a research-only qualifier that helps separate it from trading/order tools, though it does not explicitly differentiate from the sibling tastytrade_create_spx_spread_backtest.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'This is research-only and does not place brokerage orders' gives a useful boundary for when not to use it, implying it is for study rather than execution. However, it provides no explicit guidance about when to choose this tool over alternatives like tastytrade_create_spx_spread_backtest or tastytrade_simulate_trade.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_create_spx_spread_backtestA
Create an aggregate SPX spread Backtester job when the structure can be represented faithfully enough by relative leg selectors. Double diagonals remain exact-simulation only.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden, and it does meaningful work by revealing that this is an aggregate/approximate backtest rather than exact simulation, and that double diagonals are excluded. It does not fully describe the job lifecycle or return behavior, but it discloses the most important behavioral limitation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The primary use condition is front-loaded, and the exclusion is stated immediately after. Every sentence contributes selection or invocation guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is strong for tool selection but weaker for invocation and post-creation expectations. There is no output schema, no annotations, and a complex nested request object, so the description should say more about what happens after the job is created and how results are retrieved. The double-diagonal guardrail is helpful but not sufficient for full contextual completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The phrase 'relative leg selectors' adds meaning to the backtest_selector concept inside the request, and the aggregate-vs-exact distinction helps interpret the family field. However, it does not explain required fields like intended_price, price_effect, entry_at, or exit_at, leaving a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Create'), a specific resource ('aggregate SPX spread Backtester job'), and the key mechanism ('relative leg selectors'). It also distinguishes this tool from exact-simulation tools by declaring that double diagonals are not supported here.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-to-use condition: use this when the structure is faithfully representable by relative leg selectors. It also gives an explicit when-not: double diagonals are exact-simulation only. However, it does not name the exact-simulation sibling tool directly, and 'faithfully enough' is somewhat subjective.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_get_backtestA
Get a tastytrade backtest by ID, including its current status and results when completed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Backtest ID returned by tastytrade. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It discloses that the response includes current status and results only when completed, which is meaningful. However, it does not mention error behavior, whether the operation is read-only, or what happens for an unknown ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It immediately identifies the action, the resource, and the key output characteristics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter getter with no output schema, the description provides a useful summary of what the tool returns. It is slightly sparse on edge-case behavior, but the core calling context is sufficiently covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with the 'id' parameter already described as 'Backtest ID returned by tastytrade.' The tool description adds no additional parameter meaning, so the schema does the heavy lifting and the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get'), a resource ('a tastytrade backtest'), and the identifier-based access method ('by ID'). It also clarifies what the tool returns ('current status and results when completed'), distinguishing it from list and create siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool is used when you already have a backtest ID and want its status or results. It does not explicitly name alternatives or provide exclusion criteria, but the 'by ID' framing gives enough context to differentiate it from listing or creating backtests.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_get_backtest_available_datesA
List tastytrade Backtester symbols and their available historical date ranges. Use this before regression tests to verify SPY, XSP, SPX, or other symbol coverage.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. The word 'List' implies a read-only operation, and the mention of date ranges adds useful context, but the description does not disclose output structure, pagination, or any other behavioral details beyond the core action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core action and subject are front-loaded, and the usage guidance follows naturally.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description is complete: it states the resource, the information returned, and the intended workflow context. An agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parametersampions and schema coverage is 100%, so there is nothing for the description to clarify about inputs. The baseline for zero-parameter tools is 4, and the description adequately describes what the tool returns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists tastytrade Backtester symbols and their available historical date ranges, which is a distinct resource from sibling tools like tastytrade_list_backtests. The verb 'List' is specific and the subject is concrete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells agents to use this before regression tests to verify symbol coverage, giving a clear intended context. It does not explicitly name alternatives or when not to use it, but for a zero-parameter listing tool the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_get_backtest_logsC
Get tastytrade Backtester execution logs for a backtest ID.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Backtest ID returned by tastytrade. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Get' implies a read operation, but the description does not mention absence of side effects, output format, pagination, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. The core action and resource are front-loaded, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter retrieval tool, the description is minimally viable: it names the input and the expected resource. However, without annotations or an output schema, it leaves usage context and behavioral details undisclosed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: the single 'id' parameter is already described as the backtest ID returned by tastytrade. The description adds no extra meaning beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('tastytrade Backtester execution logs') for a given 'backtest ID'. This is clear and accurate, though it does not explicitly differentiate itself from the sibling tastytrade_get_backtest tool beyond the word 'logs'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as tastytrade_get_backtest or tastytrade_list_backtests. There are no prerequisites, exclusions, or context about whether logs are available before/after a backtest completes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_get_historical_candlesA
Retrieve normalized historical OHLCV candles from tastytrade DXLink for an exact UTC and session window, with source timestamps and explicit gap warnings. No resampling is performed.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses normalization, source timestamps, explicit gap warnings, and absence of resampling, which are meaningful behavioral traits. It does not cover auth, rate limits, or error behavior, but the disclosed traits are substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly packed sentence with no filler. It front-loads the core purpose and then adds only high-value caveats: source timestamps, gap warnings, and no resampling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is viable but not fully complete for a tool with a nested request schema, many parameters, and no output schema. It mentions return-related characteristics like source timestamps and gap warnings, but omits output structure, pagination, error behavior, and session timezone semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter semantics, but it only hints at 'exact UTC' and 'session window.' It does not explain interval units, instrument_type values, timeout_ms, max_candles, or session object details, leaving much of the parameter meaning to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Retrieve normalized historical OHLCV candles from tastytrade DXLink.' It further distinguishes this tool from siblings by emphasizing exact UTC/session windows, source timestamps, gap warnings, and no resampling, making its role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use this tool: when an exact UTC and session window is needed, with normalized candles, source timestamps, and gap warnings. It does not explicitly name alternatives or exclusion criteria, but the context is strong enough to guide selection among the backtest/simulation siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_list_backtestsA
List backtests submitted for the authenticated tastytrade API grant.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must carry behavioral context. It indicates authentication scope and the 'submitted' filter, and 'List' implies a read-only operation, but it does not disclose pagination, ordering, response shape, or whether in-progress backtests are included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It conveys the action, resource, and scope economically, which is appropriate for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple parameterless list operation, the description gives enough to know what the tool returns conceptually. It lacks explicit return-field or pagination details, but with no output schema, that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is fully empty, so there is nothing for the description to explain. The baseline of 4 applies because parameter semantics are simply not applicable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('backtests'), and adds a scope qualifier ('submitted for the authenticated tastytrade API grant'). This makes it easy to distinguish from sibling tools such as get_backtest, create_backtest, and get_backtest_logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this to see all submitted backtests for the grant. However, it does not explicitly state when not to use it or point to a sibling alternative, so an agent has to infer the boundary from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_prepare_spx_spreadB
Normalize an SPX debit vertical, credit vertical, iron condor, or double diagonal into deterministic Backtester and exact-leg simulation requests without submitting it.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description is the sole behavioral source. It reveals that the tool does not submit the order and that it normalizes input for both backtester and exact-leg simulation, which is useful. However, it does not disclose whether it validates inputs, checks symbol existence, or what happens with invalid spread configurations. For a non-annotated tool, this leaves key behavioral aspects implied rather than stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the types of spreads and the core output (backtester and exact-leg simulation requests) and ends with the critical non-submitting behavior. It is concise and accomplishes its purpose in one breath, though it could have added a brief list of outputs for even faster scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (deeply nested schema over four spread families, multiple simulation paths, and no output schema), the description is minimal. It does not mention the distinction between backtester and exact-leg simulation, nor the validation or normalization steps, nor what the return value looks like. Siblings like tastytrade_simulate_spx_spread and tastytrade_create_spx_spread_backtest exist, so an agent needs to know what this returns to use it as a prerequisite. The description is not sufficient on its own.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but it is an extremely detailed schema with nested properties, enums for family/action/price_effect, and even a backtest_selector object. The description only says 'normalize' and lists the spread types, which adds minimal value beyond what the schema already shows. It doesn't explain the distinction between the backtest and exact-leg outputs, nor does it clarify the interplay of fields like backtest_selector and days_until_expiration. The schema carries most of the semantic burden, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool normalizes SPX spread variants (debit/credit vertical, iron condor, double diagonal) into two specific kinds of requests (Backtester and exact-leg simulation), and explicitly notes it does not submit the trade. This distinguishes it from siblings like tastytrade_simulate_spx_spread (which submits) and tastytrade_create_spx_spread_backtest (which creates a backtest).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'normalize... without submitting it' but does not explicitly state when to use it versus alternatives like tastytrade_simulate_spx_spread (which presumably submits) or tastytrade_create_spx_spread_backtest (which creates a backtest). An agent must infer that this is a preprocessing step, but there is no explicit guidance on when to call this first versus directly calling the simulation/backtest creation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_price_option_packageC
Price a 2-4 leg option package with explicit native-package, synthetic-natural, and midpoint-reference provenance using exact decimal arithmetic.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does mention 'exact decimal arithmetic' and 'provenance', which are useful computational details, but it does not disclose whether the tool performs a read-only calculation, requires live quotes, has side effects, or how it handles missing or stale data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. The main action and scope appear immediately, though the dense provenance phrase adds complexity without much clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a complex nested schema, no annotations, and no output schema, yet the description only covers leg count and arithmetic style. It omits the request wrapper, required fields, family semantics, and the meaning of the provenance modes, leaving the agent reliant on schema inspection for critical context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only hints at '2-4 leg' and 'native-package'. It does not explain the required request object, the family enum, the structure of legs, or the meaning of fields like references, max_quote_age_ms, and max_temporal_skew_ms.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Price'), the resource ('option package'), and the scope ('2-4 leg'), which distinguishes it from sibling backtest and simulation tools. However, the phrase 'explicit native-package, synthetic-natural, and midpoint-reference provenance' is jargon-heavy and may obscure the core purpose for an agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for pricing option packages, but it gives no explicit guidance on when to choose it over alternatives like tastytrade_prepare_spx_spread or tastytrade_simulate_trade. There are no when-not-to-use conditions or references to related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_simulate_spx_spreadC
Run exact-leg historical simulation for a normalized SPX defined-risk spread and return a stable regression result contract.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It states it is a simulation and returns a contract, but does not disclose whether it is read-only, any side effects, authentication needs, rate limits, or what happens to data. The term 'stable' is vague and does not explain behavioral traits beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It front-loads the action and output. However, it is extremely short given the complexity of the tool, but conciseness is high because every word is purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complex nested 'request' parameter with many required fields, and no annotations or output schema, the description is completely inadequate. An agent cannot know what to put in the request, what the contract contains, or any constraints. It is missing almost all contextual information needed to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It mentions none of the parameters, including the required 'request' object and its fields like 'family', 'legs', 'entry_at', etc. The description adds no semantic meaning to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Run' and a precise object: 'exact-leg historical simulation for a normalized SPX defined-risk spread'. Also specifies the return: 'stable regression result contract'. This is distinct from sibling tools like tastytrade_simulate_trade and tastytrade_create_spx_spread_backtest by mentioning 'exact-leg' and 'normalized'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. Does not mention scenarios, prerequisites, or exclusions. Sibling tools like tastytrade_simulate_trade and tastytrade_create_spx_spread_backtest are not referenced, leaving the agent to infer selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_simulate_tradeA
Simulate a single historical option trade with tastytrade Backtester and return its historical path/results. Research-only; no brokerage order is placed.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes | Raw request body accepted by tastytrade POST /simulate-trade. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It explicitly discloses that no brokerage order is placed and that this is research-only, which is a critical safety-related behavior. It also mentions the return of historical path/results, though it does not cover error behavior, authentication requirements, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core action and output are front-loaded, and the research-only safety note is placed clearly. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no annotations, no output schema, and a single opaque 'request' parameter that is essentially a raw API body. The description provides the high-level purpose and safety profile but does not explain how to construct the request, what fields are required, or what the returned path/results will look like. An agent would likely need external API documentation to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minimal meaning beyond the schema, only indicating the trade is historical and single. The 'request' parameter is still opaque ('Raw request body accepted by tastytrade POST /simulate-trade'), and the description does not compensate with additional parameter-level details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Simulate'), a specific resource ('a single historical option trade with tastytrade Backtester'), and the expected output ('historical path/results'). It also clearly differentiates from sibling tools by emphasizing 'single' trade and 'Research-only', which distinguishes it from broader backtest creation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context with 'Research-only; no brokerage order is placed', but it does not explicitly state when to use this tool versus alternatives like tastytrade_create_backtest or tastytrade_simulate_spx_spread. The guidance is mostly implicit and relies on the agent inferring the difference from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tastytrade_verify_historical_fillB
Verify whether a frozen paper-order limit was touched during a forward historical interval. Returns separate post-session evidence and never rewrites the live paper event.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden of behavioral disclosure. It does disclose a key non-mutating behavior: 'never rewrites the live paper event,' which is important for an agent to understand side effects. It also mentions that it returns 'separate post-session evidence.' However, it does not describe other behavioral traits such as rate limits, authentication requirements, or failure modes. The description is partially transparent but leaves gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the primary purpose. The two sentences have no redundancy and effectively convey the core function and a key behavioral guarantee. It is appropriately sized for a tool of moderate complexity, though it could include a bit more guidance without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a single request object with many nested required fields, no output schema), the description is far from complete. It does not explain what 'frozen paper-order limit' means, what 'forward historical interval' implies, or how to structure the request. It also does not describe the output format or any error conditions. An agent attempting to call this tool correctly would be missing essential context, especially without annotations or an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides zero information about the 'request' parameter or its nested fields. The schema is complex with many required properties, but the description does not explain what any of them mean, how to construct a valid request, or which combinations are valid. With schema description coverage at 0%, the description must compensate, but it fails to do so. The agent would have to rely entirely on the schema, which has minimal field descriptions (only 'provider_symbol' has a description).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('verify') and a precise resource ('frozen paper-order limit') within a defined context ('forward historical interval'). It also clarifies the output ('separate post-session evidence') and the non-mutating behavior ('never rewrites the live paper event'). This clearly differentiates it from sibling tools like backtest creation or simulation, which are about generating or running tests rather than verifying an existing frozen order.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool versus alternatives. It does not mention conditions that would select this tool over others (e.g., when you need to check if a limit was touched in a past session) nor does it name any sibling tools. The verb 'verify' and the mention of 'frozen paper-order limit' imply a verification use case, but without explicit context or exclusions, the agent is left to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v0.2.0- First observed
tastytrade_cancel_backtest - First observed
tastytrade_create_backtest - First observed
tastytrade_create_spx_spread_backtest - First observed
tastytrade_get_backtest - First observed
tastytrade_get_backtest_available_dates - First observed
tastytrade_get_backtest_logs - First observed
tastytrade_get_historical_candles - First observed
tastytrade_list_backtests - First observed
tastytrade_prepare_spx_spread - First observed
tastytrade_price_option_package - First observed
tastytrade_simulate_spx_spread - First observed
tastytrade_simulate_trade - First observed
tastytrade_verify_historical_fill
TDQS
Scored across 13 tools
Most tools are clearly distinct by resource and action, but the backtest/simulation cluster (create_backtest, simulate_trade, simulate_spx_spread, create_spx_spread_backtest) risks confusion despite clarifying descriptions. The price/prepare/simulate spread pipeline is well-separated with explicit roles.
All tools follow a consistent tastytrade_ prefix with snake_case verb_noun naming (list_backtests, create_backtest, get_backtest_logs, cancel_backtest). Verbs and objects are uniform and predictable across the entire set.
13 tools is well within the ideal 3-15 range for a research-focused server. Each tool contributes to distinct workflows such as backtest management, option pricing, spread simulation, and historical data access without redundancy.
The core research workflow is well covered: backtest lifecycle (list, create, get, cancel, logs), historical simulation, option pricing, spread normalization, and candle retrieval. Minor gaps exist such as lack of a delete-backtest or batch operation, but no critical dead ends are apparent.
Maintenance
Related MCP Connectors
Backtest trading ideas and inspect strategies, trades and results. OAuth; invite-only beta.
Run crypto trading strategy backtests through EmidLabs's Backtesting API.
Honest backtests and paper trading for US options and equities, driven by your agent.
Unified financial infrastructure connecting AI agents directly to trade live/demo brokerage accounts, Web3 non-custodial wallets, real-time market data across equities, ETFs, crypto, forex, options, DeFi swaps, and prediction markets, institutional research feeds, and algorithmic strategy backtesters.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceexposes a remote MCP endpoint so agents can: run strategy backtests by symbol/timeframe/date range, pass strategy inputs programmatically, receive structured backtest results (trades, win rate, profit, drawdown), keep long-running runs observable via progress notifications, support Binance Futures tickers only, enforce a maximum of 1440 candles per backtest, apply a rate limit of 3 backtests per6-
- AlicenseNot gradedqualityDmaintenanceEnables interaction with the Nubra trading platform for authentication, instrument lookup, quotes, historical data, options analytics, portfolio management, report generation, screening, backtesting, and order placement via UAT environment.MIT
- AlicenseNot gradedqualityDmaintenanceEnables backtesting of limit-order strategies on Polymarket's BTC 5-minute markets using historical data, with tools to browse markets, get price series, and run simulations.MIT
- AlicenseNot gradedqualityBmaintenanceEnables interactive access to TradeSearcher strategies and backtests via CLI and MCP, allowing agents to search, backtest, and compare trading strategies.MIT