Skip to main content
Glama
hplin

tastytrade-research-mcp

by hplin

tastytrade-research-mcp

Research-only Model Context Protocol (MCP) server for tastytrade historical options research, bounded live SPX/SPXW option evidence, package-price evidence, and strategy regression testing.

The project is intentionally separate from the official tastytrade/tastytrade-mcp:

  • official tastytrade MCP: general live market data, account workflows, and broker dry-run validation

  • this project: historical candles, Backtester access, bounded direct SPX/SPXW Gamma+OI snapshots, package-pricing research, and post-session fill verification

  • never exposed here: brokerage order placement, replacement, or cancellation

MCP tools

Provider-native Backtester tools

Tool

Upstream endpoint

Purpose

tastytrade_get_backtest_available_dates

GET /available-dates

Discover symbols and historical coverage

tastytrade_list_backtests

GET /backtests

List submitted research jobs

tastytrade_create_backtest

POST /backtests

Submit a provider-native historical strategy

tastytrade_get_backtest

GET /backtests/{id}

Poll status and retrieve results

tastytrade_get_backtest_logs

GET /backtests/{id}/logs

Retrieve trials and execution logs

tastytrade_cancel_backtest

POST /backtests/{id}/cancel

Cancel only a Backtester job

tastytrade_simulate_trade

POST /simulate-trade

Simulate one exact historical trade

Research and regression tools

Tool

Purpose

tastytrade_get_live_option_snapshot

Auto-chunk and merge exact SPX/SPXW DXLink Quote, Greeks, and Summary.openInterest cohorts, then calculate a research-only unsigned Gamma concentration proxy

tastytrade_compute_heuristic_signed_gex

Apply an explicit versioned research signing hypothesis to one unified snapshot, with separate current-Gamma signed GEX and optional bounded spot-repriced heuristic gamma flip

tastytrade_price_option_package

Price verticals, iron condors, and double diagonals with explicit native/synthetic provenance

tastytrade_discover_historical_spx_candidates

Reconstruct timestamp-safe historical SPXW candidates under a versioned resolution profile, with exact-timestamp Backtester fallback

tastytrade_discover_historical_spx_candidates_range

Run the same candidate discovery over an explicit trading calendar with bounded concurrency, deadlines, immutable-cache diagnostics, and resumable partial progress

tastytrade_get_historical_spx_candidate_universe

Return a bounded multi-strike, multi-expiration SPXW universe and optional versioned RESEARCH_ONLY DD selected-leg/matched-delta IV handoff

tastytrade_get_historical_option_package_at_checkpoint

Reconstruct an exact-leg SPX package reference at an RFC3339 or IANA-local checkpoint with explicit age and skew controls

tastytrade_get_historical_option_package_path

Return profile-aligned completed-candle package points and explicit gaps without interpolation

tastytrade_get_historical_option_package_horizons

Reconstruct a frozen exact-leg candidate inventory at caller-supplied ENTRY/+3/+5 trading sessions, with an opt-in research-only model fallback for candle-missing legs

tastytrade_normalize_historical_execution_evidence

Normalize immutable exact-leg quote snapshots/windows into signed, provider-neutral simulated-execution inputs without claiming a fill

tastytrade_simulate_historical_execution

Apply one caller-frozen, hashed execution profile to immutable exact-leg evidence and return a separate simulated fill/P&L result

tastytrade_build_historical_replay_report

Aggregate frozen baseline/research decisions and immutable execution simulations into deterministic +3/+5-day metrics without changing grading or paper state

tastytrade_prepare_spx_spread

Deterministically normalize SPX legs without calling an upstream service

tastytrade_simulate_spx_spread

Run exact-leg SPX historical simulation and normalize its result

tastytrade_create_spx_spread_backtest

Submit supported SPX structures through relative Backtester selectors

tastytrade_verify_historical_fill

Check a frozen paper limit against a caller-supplied historical package path or a forward Backtester path

tastytrade_get_historical_candles

Retrieve normalized DXLink OHLCV candles without resampling

Use MCP tools/list for the complete JSON input schemas. The direct live snapshot, bounded DXLink batch provenance, event/cohort freshness contract, OI-based unsigned Gamma concentration methodology, semantic limits, regression handoff, and opt-in live gate are documented in docs/live-option-snapshot.md. The separate Level 3 signing hypothesis, model/result identities, current-Gamma versus spot-repriced evidence, bounded heuristic gamma-flip search, and research-only live gate are documented in docs/heuristic-signed-gex.md. The shared profile contract and the optional seven-date live capability gate are documented in docs/resolution-profiles.md. The separate Double Diagonal IV measurement contract, legacy-preserving migration, and reproducible grading/regression handoff are documented in docs/dd-iv-measurements.md. Private immutable source-cache configuration, exact offline replay, and migration guidance are documented in docs/evidence-cache.md. The external historical-source capability review, explicit no-spend decision, and approval gates are documented in docs/historical-provider-decision.md. The quote-backed exact-leg handoff, signed inventory convention, and valuation/simulation/broker separation are documented in docs/historical-execution-evidence.md. Caller-frozen execution profiles, fill statuses, exact P&L arithmetic, and fee semantics are documented in docs/historical-execution-model.md. The replay acceptance matrix, denominator rules, one-week smoke manifest, and runner guidance are documented in docs/historical-replay.md.

Related MCP server: Nubra MCP Server

Private historical evidence cache

The historical candle, SPX candidate/universe, and exact-package tools accept an opt-in evidence_cache policy. The server exposes 24 tools, including the local-only historical quote normalizer, deterministic execution-model adapter, and historical replay report builder. READ_WRITE stores sanitized source observations and normalized results as separate content-addressed objects; REFRESH creates a new immutable revision and diff; CACHE_ONLY replays only caller-supplied exact manifest IDs and never contacts the provider.

The filesystem backend is disabled until TASTYTRADE_EVIDENCE_CACHE_DIR points to a private volume. It uses atomic writes, read-back checksum verification, bounded provider concurrency, request deduplication, short-lived retryable-failure indexes, and a hard disk quota without evicting immutable evidence. Cache identity includes exact symbols, range/as-of, provider/dataset/license scope, aggregation/session/ alignment/price type, the full resolution profile, resource policy, and normalization/model/source revisions. Retrieval time is frozen in each immutable manifest and never replaces bar availability time. Run npm run report:evidence-cache for a deterministic synthetic cache hit/miss, byte-count, normalized-hash, and provider-call-reduction report.

Approval-gated bounded history

The repository contains a provider-neutral bounded-history contract for a future approved source. It is intentionally dormant: no vendor transport, credential, paid fallback, environment configuration, or MCP tool is enabled. Construction requires an exact approval scope, confirmed field and expired-symbol entitlements, a provider-specific native-field allowlist, a fixed HTTPS host, separate credentials, and explicit cache storage/reuse permission. Synthetic tests validate the contract but do not establish external SPX/SPXW coverage.

Execution evidence contract

Every normalized research result uses execution-evidence contract 1.0.0. The contract distinguishes:

  • NATIVE_PACKAGE

  • SYNTHETIC_NATURAL

  • SYNTHETIC_MID_REFERENCE

  • HISTORICAL_OPTION_PACKAGE_REFERENCE

  • HISTORICAL_PATH

  • BACKTESTER_SIMULATION

  • BROKER_DRY_RUN

The machine-readable schema is docs/execution-evidence.schema.json. Migration guidance for existing paper-simulation logs is in docs/execution-evidence-migration.md.

BROKER_DRY_RUN validates a caller-supplied order; it never implies fillability. Synthetic midpoint evidence is valuation-only.

Package pricing

Package arithmetic uses an in-repository exact decimal implementation rather than binary floating-point arithmetic.

  • NATIVE_PACKAGE is selected only when the caller supplies a valid, non-crossed upstream package market.

  • SYNTHETIC_NATURAL buys each leg at its ask and sells each leg at its bid.

  • SYNTHETIC_MID_REFERENCE uses leg midpoints and always reports guaranteed_executable: false.

  • Native package bid/ask fields are never populated with synthetic arithmetic.

  • Every leg retains its exact symbol, action, quantity, expiration, timestamp, and provider source.

  • as_of, freshness, and temporal alignment are computed from the decision-critical observations.

  • Missing, crossed, stale, future-dated, and materially misaligned quotes are returned as explicit warnings. Unusable evidence is never success-shaped.

Supported families are DEBIT_VERTICAL, CREDIT_VERTICAL, IRON_CONDOR, and multi-expiration DOUBLE_DIAGONAL.

SPX spread adapter

The high-level adapter preserves each exact provider/OCC option symbol in /simulate-trade requests and returns stable 1.0.0 result envelopes. Each result repeats the immutable normalized legs and carries a deterministic SHA-256 request_id, so persisted evidence remains attributable even though the provider response itself contains only prices and timestamps.

Aggregate /backtests requests use relative selectors (delta, percentageOTM, currentPriceOffset, or premium) and therefore cannot preserve an exact historical strike or expiration. The adapter exposes that limitation in its capability flags.

Double diagonals remain exact-simulation only because the aggregate Backtester cannot faithfully preserve their multi-expiration identity. 0DTE legs are rejected unless the caller explicitly sets allow_0dte: true.

This MCP supplies evidence only. It does not assign grades or make production entry decisions.

Historical SPX candidate discovery

tastytrade_discover_historical_spx_candidates is the REGRESSION_RESEARCH bridge between timestamp-safe selector evidence and exact contract identity. tastytrade_discover_historical_spx_candidates_range applies that same single-checkpoint contract to an explicit caller-supplied trading calendar. It preserves chronological output order while using a bounded worker pool, isolates checkpoint failures, retries only transient provider failures, and returns an opaque continuation cursor for unresolved or deferred sessions. Continuation scheduling completes a first pass over never-attempted sessions before retrying failures. Retry-eligible checkpoints run before checkpoints whose provider cooldown is still active, and unresolved retries rotate to the queue tail so rate-limited dates cannot block later checkpoints. All Backtester endpoints in one server process share a request-start gate. It spaces starts by at least 100 ms by default, honors Retry-After before X-RateLimit-Reset, and falls back to bounded jittered 10/30/90-second cooldowns. Override only the pacing interval with TASTYTRADE_BACKTESTER_MIN_REQUEST_INTERVAL_MS; continuation retry timestamps remain provider- or scheduler-derived. Per-checkpoint stage timings identify cache lookup, provider bootstrap, contract-universe construction, candle reconstruction, selector evaluation, and the stage interrupted by a deadline. Operational settings and diagnostic callbacks do not alter the logical batch request ID or candle-cache fingerprints.

The official option-chain and REST quote endpoints do not document an historical as_of parameter. Backtester logs currently expose exact selected strike, expiration, side, and a provider-internal symbol, but those log fields are undocumented. The documented Backtester EntryConditions also has no time-of-day field, and a live request containing undocumented entryTime: "14:30:00Z" was silently ignored: the SPX trial still opened at 19:45:00Z.

For DELTA and PERCENTAGE_OTM, the adapter reconstructs a bounded SPXW universe from generated OCC/streamer symbols and completed DXLink candles. Candidate discovery uses the provider-supported maximum of 100 option symbols per candle batch; the standalone bounded-universe tool retains its existing 20-symbol batch policy. Omitting resolution_profile preserves the 5-minute cohort; strict native-hour RTH research requires explicit HOURLY_VALUATION_RESEARCH. When DXLink labels SPXW hour bars on its provider clock instead of the 09:30 ET session anchor, callers must opt into the separate HOURLY_PROVIDER_ALIGNED_RESEARCH cohort; those bars are never relabeled as session-aligned. It enforces available_at <= as_of, profile age/skew limits, aligned call/put parity for the forward, and contract-candle IV for Black-76-style delta. Selection is deterministic by selector error, observation age, DTE distance, and strike.

Generated contract identity becomes eligible only when DXLink returns historical evidence for that exact symbol. The result preserves checkpoint selection time, observation age, price, contract IV, reconstructed delta, volume/OI when present, OCC identity, and field-level provenance.

Exact-timestamp Backtester selection remains the fallback for unsupported selectors or insufficient reconstruction evidence. It never sends or trusts entryTime, and it accepts only a trial and opening order exactly at as_of. Stale and future trials remain fail-closed.

The capability returns HISTORICAL_SELECTOR_CANDIDATE_SET, never a full-chain snapshot. Historical bid/ask, ATM IV surface, skew, and term structure remain unavailable. Exact symbols are /simulate-trade compatible, but the capability separately reports whether an exact simulation snapshot exists at the arbitrary checkpoint.

The live 2026-08-25 07:30 PT smoke test returns four timestamp-safe CALL/PUT Delta-20 and 1%-OTM candidates without creating Backtester jobs.

The full provider findings, contract, anti-lookahead rules, and limitations are documented in docs/historical-spx-candidates.md.

Historical SPX candidate universe

tastytrade_get_historical_spx_candidate_universe expands checkpoint-safe evidence from selector winners into a caller-bounded strike grid. By default it covers the nearest min/mid/max DTE SPXW expirations, both option sides, and a 25-point strike step.

Generated OCC identity is returned only after DXLink supplies a completed historical candle for that exact symbol. Contracts without timestamp-safe evidence are omitted and counted as coverage gaps. The endpoint preserves price, reconstructed delta, contract IV, OI, volume, observation age, and field-level provenance when available; it never selects the final spread. Delta reconstruction prefers a timestamp-aligned put/call parity forward. When parity is unavailable but historical IV exists, the universe endpoint uses the completed checkpoint SPX price as an explicitly warned zero-carry forward approximation. Missing IV remains null, and this fallback does not change the stricter selector-discovery path.

The default freshness maximum is 60 minutes. A caller may explicitly extend it to 24 hours; older pre-checkpoint observations are then retained only as STALE/LOW evidence. Coverage gaps, provider batch errors, and per-field availability counts remain machine-readable, and aggregate field capability flags are true only when every returned contract supports the field.

Both historical SPX tools accept either RFC3339 as_of or an unambiguous IANA local_checkpoint. They preserve bar_start, bar_end, available_at, and retrieved_at, and return the normalized resolution profile with deterministic requested/effective cohort IDs. An opaque versioned candidate_construction_profile is returned unchanged; this MCP does not duplicate grading, DD-bucket, or final-leg-selection rules.

An optional dd_iv_measurement request freezes an exact four-leg Double Diagonal candidate and returns distinct SELECTED_LEG_IV_DIFFERENCE and MATCHED_DELTA / MATCHED_FORWARD_MONEYNESS cohorts. The handoff is always RESEARCH_ONLY, preserves provider/derived models, timestamp lineage, and optional immutable evidence-manifest IDs, and explicitly does not replace legacy production term_structure.

See docs/historical-spx-universe.md for the input contract, reconstruction rules, and live checkpoint findings.

Historical exact-leg package reconstruction

tastytrade_get_historical_option_package_at_checkpoint selects only completed option candles available by the requested checkpoint and preserves each leg's exact OCC identity, derived streamer symbol, historical close, IV, bar-start timestamp, availability timestamp, age, and provenance. Explicit age and temporal-skew limits fail closed. Candle-derived values are always VALUATION_ONLY; historical bid/ask is not fabricated.

The checkpoint and path results include deterministic request IDs plus requested/native/effective aggregation, session, alignment, age/skew, fallback, and cohort metadata. The default checkpoint behavior remains the existing 5-minute compatibility path. Native-hour New York RTH valuation is explicit opt-in and does not silently fall into the 5-minute cohort. Checkpoint legs also return a structured reconstruction status and failure reason. Their provenance carries the requested lifecycle, resolution profile, provider failure reasons, source revision, and exact cache manifest/content IDs when available.

tastytrade_get_historical_option_package_horizons accepts a bounded frozen candidate inventory and a caller-supplied ordered trading-session calendar. It resolves ENTRY, OUTCOME_3_TRADING_DAYS, and OUTCOME_5_TRADING_DAYS by session index rather than assuming weekdays are trading days. Every horizon reuses the same OCC symbols, actions, quantities, roles, strategy family, and resolution profile. The result keeps candle references labeled VALUATION_ONLY / CANDLE_REFERENCE and reports complete package counts, missing legs by role, missing-reason counts, and coverage by strategy, expiration, entry DTE, and resolution profile. Exact 21/35-DTE Double Diagonals are supported without changing grading or leg selection.

An optional valuation_fallback can supply checkpoint-frozen SPX levels, per-exact-leg IV or surface inputs, rates, dividends, immutable source IDs, and an IV-shift uncertainty assumption. The strict candle status, package, legs, and failure taxonomy remain unchanged. A separate valuation object classifies each result as EXACT_PACKAGE_REFERENCE, MIXED_OBSERVED_MODELED, or MODEL_SURFACE; observed leg values are copied unchanged and only unavailable exact legs are modeled. Cache/provider errors, stale observations, and alignment failures remain non-modelable.

Any package containing a modeled leg is MODEL_REFERENCE / VALUATION_ONLY, carries guaranteed_executable: false, and is explicitly not bid/ask, NBBO, midpoint, touch, or fill evidence. Its execution_evidence_input can be passed to tastytrade_normalize_historical_execution_evidence and then evaluated only through a separately frozen REFERENCE_COST execution profile. Coverage keeps strict package counts separate from valued exact, mixed, and fully modeled cohorts.

tastytrade_get_historical_option_package_path emits a point only when every leg has an exact timestamp-aligned completed bar. It never interpolates or forward-fills missing legs. If a bounded old 1-minute DXLink replay exhausts a configured local resource budget or the provider returns a clipped snapshot, the result records the failed attempt and explicitly selects the finest retrievable coarser resolution. A local budget result is not described as a provider limit.

The returned fill_verification_path can be passed directly to tastytrade_verify_historical_fill. Full rules and the 2026-08-27 07:30 PT acceptance finding are documented in docs/historical-option-package.md.

Historical fill verification

tastytrade_verify_historical_fill:

  • evaluates only [submitted_at, valid_until];

  • accepts either caller-supplied historical package points or the existing exact-leg Backtester mode;

  • supports both ENTRY and EXIT;

  • preserves paper_order_id, checkpoint_id, and position_id references;

  • returns legacy status values TOUCHED, NOT_TOUCHED, or NOT_VERIFIABLE, plus assessment_status: NOT_ASSESSABLE for an insufficient path;

  • distinguishes LIMIT_TOUCH from CONSERVATIVE_CROSS;

  • reports an exact observed touch timestamp when defensible, otherwise a bounded interval for a sparse path;

  • records disagreement with a live paper assumption without mutating the original paper event.

All verification is marked POST_SESSION_REGRESSION and includes HISTORICAL_EVIDENCE_ONLY_DO_NOT_REWRITE_LIVE_EVENT.

tastytrade_get_historical_candles obtains an API quote token from GET /api-quote-tokens, connects only to a wss:// host under dxfeed.com, completes the DXLink handshake, and waits for the indexed-event snapshot boundary.

The normalized output includes:

  • deterministic request and cohort identity;

  • dxlink_auth recovery provenance and a sanitized provider_error when authentication required recovery;

  • requested and actual UTC ranges;

  • source_time, bar_start, bar_end, available_at, and retrieved_at per bar;

  • explicit instrument type and interval;

  • ALL, US REGULAR, or caller-defined CUSTOM session filtering with an IANA timezone;

  • OHLC, volume, VWAP, bid/ask volume, implied volatility, and open interest when supplied;

  • missing-bar, empty-result, and snapshot-truncation warnings;

  • resampled: false.

REGULAR requests use provider-native a=s,tho=true candle attributes, so bars are aligned to and built only from the instrument's regular trading session. CUSTOM windows filter complete provider bars by their source timestamp; they are never reaggregated and include an explicit warning about that limitation.

One-unit periods use DXLink's native normalized spelling (m, h, d, or w). In particular, public interval 1h subscribes to native HOUR {=h}; 60m remains {=60m} and is not treated as equivalent. Request/response matching accepts only provider-defined canonical aliases such as omitted defaults and attribute ordering differences.

The server sends fromTime as epoch milliseconds, matching the current production DXLink service. The published AsyncAPI description currently says seconds, but seconds cause the service to replay the full available history. Production currently appears to ignore toTime. The client reports a continuous-calendar replay estimate for planning, but that estimate is explicitly advisory: it does not account for sessions, closures, sparse options, or expiration, and never blocks a request before connecting.

Streaming state is indexed and deduplicated per symbol, retains only rows in the requested range and session, serializes message processing, and does not infer completion from timestamp order. Completion still requires DXLink snapshot protocol evidence. Each result includes:

  • status, snapshot_complete, snapshot_truncated, and provider_snapshot_complete;

  • machine-readable failure_reasons;

  • per-symbol and aggregate request counters for received events, valid events, unique observations, retained rows/bytes, and returned rows;

  • configured limits and the advisory calendar-slot estimates.

snapshot_complete means a provider END marker was observed without local or provider truncation. It does not assert that every interval traded; status and failure_reasons separately report missing or uncovered evidence.

The independent client controls are:

Input

Scope

Default

Maximum

max_output_candles

retained/returned rows per symbol

10,000

250,000

max_received_events

aggregate Candle protocol rows per request

10,000

1,000,000

max_buffer_bytes

aggregate queued wire data plus accounted retained state

16 MiB

128 MiB

deadline_ms

complete DXLink snapshot lifecycle

15,000 ms

60,000 ms

Received-event counts include marker, remove, invalid, duplicate, and unmatched Candle rows. Unique observations count distinct in-window indexed events admitted to bounded state; retained rows reflect removals, and returned rows reflect the final normalized output. Retained-byte accounting uses serialized field sizes plus conservative per-object/index overhead rather than claiming exact V8 heap usage.

max_candles remains as a deprecated compatibility shorthand, capped at 20,000, that applies the same value to per-symbol output and aggregate receive budgets. It cannot be combined with either explicit field. timeout_ms is a deprecated alias for deadline_ms and cannot be combined with it.

LOCAL_RECEIVE_BUDGET_EXCEEDED, LOCAL_BUFFER_BUDGET_EXCEEDED, and LOCAL_OUTPUT_BUDGET_EXCEEDED are client facts. PROVIDER_SNAPSHOT_SNIPPED is provider protocol evidence. REQUESTED_WINDOW_NOT_COVERED, SNAPSHOT_TIMEOUT, and MISSING_CONTRACT_EVIDENCE describe observed availability; none is automatically classified as an entitlement failure. Completed symbols remain usable when another batch symbol fails locally or is snipped.

Canonical identity, snapshot transaction rules, the sanitized timeout root cause, and live native-hour findings are documented in docs/historical-candles.md.

Rate limits and retries

  • API quote tokens use the provider's expires-at timestamp with a one-minute safety margin. The previous 23-hour bound remains only as a fallback when the provider omits expiry metadata.

  • Historical candles and live option snapshots share one quote-token lifecycle in the production service. A post-AUTH UNAUTHORIZED invalidates both the cached quote token and cached OAuth access token, reacquires credentials, rebuilds the WebSocket connection, and retries exactly once.

  • Successful recovery returns dxlink_auth.status = REFRESHED with retry_count = 1. A second rejection returns explicit AUTH_FAILED diagnostics; a timeout before AUTH_STATE/AUTHORIZED is NOT_CONFIRMED. Tokens are never included.

  • The quote-token REST request retries network errors, 429, and 5xx up to three attempts with capped exponential backoff and jitter.

  • A candle snapshot has a configurable deadline_ms (maximum 60 seconds).

  • Non-authentication WebSocket failures are returned explicitly and are not silently retried or merged.

  • Received events, returned rows, and memory are bounded independently as documented above.

  • DXLink permits at most 5 concurrent sessions and 100 Candle subscriptions per session. Callers should batch work rather than fan out unbounded calls.

Security model

OAuth client credentials and refresh tokens are sent only to api.tastyworks.com. The short-lived OAuth token is sent to the fixed Backtester host and to the tastytrade quote-token endpoint. The resulting quote token is sent only over wss:// to a host under dxfeed.com.

Do not commit credentials. If using Node's --env-file, unset inherited TASTYTRADE_* variables first because inherited values override the file. All API timestamps must be RFC3339 values containing Z or an explicit UTC offset; timezone-less timestamps are rejected.

The remote HTTP entrypoint supports two authentication modes:

  • api-key for local development and emergency rollback, using a constant-time comparison against MCP_API_KEY;

  • oauth for production, validating Entra JWT signature, issuer, audience, expiration, and mcp.read scope through cached JWKS.

OAuth mode publishes RFC 9728 protected-resource metadata at both supported well-known paths and includes that URL and the required scope in every 401 challenge. MCP is accepted only at POST /mcp, request bodies are limited to 1 MiB, and GET /healthz remains unauthenticated with health metadata only.

Requirements and setup

  • Node.js 22+

  • tastytrade OAuth API grant:

    • TASTYTRADE_CLIENT_ID

    • TASTYTRADE_CLIENT_SECRET

    • TASTYTRADE_REFRESH_TOKEN

  • A fully onboarded tastytrade customer for DXLink quote tokens

npm ci
npm run build

export TASTYTRADE_CLIENT_ID="..."
export TASTYTRADE_CLIENT_SECRET="..."
export TASTYTRADE_REFRESH_TOKEN="..."

npm start

For remote Streamable HTTP:

export MCP_API_KEY="$(openssl rand -hex 32)"
export MCP_HTTP_HOST=127.0.0.1
export MCP_HTTP_PORT=8000

npm run start:http

Connect to http://127.0.0.1:8000/mcp with Authorization: Bearer <MCP_API_KEY>. Azure Container Apps deployment and Key Vault guidance are documented in docs/azure-deployment.md.

Production OAuth configuration additionally requires MCP_PUBLIC_URL, OAUTH_ISSUER, OAUTH_JWKS_URL, OAUTH_AUDIENCE, OAUTH_REQUIRED_SCOPE, and OAUTH_TOKEN_SCOPE. ChatGPT discovers the Entra authorization server from /.well-known/oauth-protected-resource, then sends its access token in the standard Authorization header.

For an MCP client:

{
  "mcpServers": {
    "tastytrade-research": {
      "command": "node",
      "args": ["/absolute/path/to/tastytrade-research-mcp/dist/index.js"],
      "env": {
        "TASTYTRADE_CLIENT_ID": "...",
        "TASTYTRADE_CLIENT_SECRET": "...",
        "TASTYTRADE_REFRESH_TOKEN": "..."
      }
    }
  }
}

Development

npm ci
npm run typecheck
npm test

The regression suite covers clean and degraded package markets, all supported spread families, regular and custom candle sessions, both fill models, ambiguous paths, Backtester normalization, and MCP tool dispatch.

Upstream documentation

Disclaimer

This is an independent research project and is not an official tastytrade product. Historical simulations can differ materially from live execution because of fills, liquidity, spreads, slippage, data availability, and model assumptions.

MIT License.

Available Tools

13 tools
tastytrade_cancel_backtestA

Cancel a running research backtest. This affects only the Backtester job and never a brokerage order.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesBacktest ID returned by tastytrade.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The description clarifies the scope (only backtest jobs, not orders), which is critical behavioral context. However, it does not mention side effects such as whether the cancellation is irreversible or if resources are freed, but given the tool is a cancellation, the key transparency is the scope distinction, which is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that is concise and front-loaded with the core purpose ('Cancel a running research backtest'), followed by an important clarification. There is zero waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple cancel operation with one well-documented parameter, the description effectively covers the core usage. The only minor gap is not specifying what happens after cancellation (e.g., no return value description), but given the lack of output schema and the simplicity, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage for the single parameter, describing it as 'Backtest ID returned by tastytrade.' The description does not add extra semantics beyond the schema, but since the schema is complete, the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (cancel), the resource (a running research backtest), and explicitly distinguishes it from brokerage orders. This makes it unambiguous and differentiates it from sibling tools like tastytrade_list_backtests or tastytrade_create_backtest.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool (to cancel a running backtest) and what it does NOT affect (never a brokerage order). This is a clear 'when/not' guideline, which is particularly important given the presence of tastytrade_simulate_trade and other order-related tools in the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_create_backtestA

Create a tastytrade historical options backtest. This is research-only and does not place brokerage orders.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesRaw request body accepted by tastytrade POST /backtests. Kept provider-native so new upstream fields remain usable.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the disclosure burden and does meaningfully reveal that no brokerage orders are placed, which is a non-obvious and important behavioral trait. Still, it does not mention whether the backtest is executed asynchronously, how long it may take, or what the response represents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. It front-loads the action and resource, then immediately states the key caveat, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a thin API wrapper, especially with sibling names like tastytrade_get_backtest and tastytrade_cancel_backtest hinting at the lifecycle. However, it does not explain what the tool returns or that the backtest may run asynchronously, which would help an agent know how to poll or retrieve results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is only one parameter and the schema already describes it thoroughly as the raw POST /backtests request body. The tool description itself adds no parameter-level guidance, so it stays at the baseline for full schema coverage. The provider-native note in the schema is helpful but is already part of the structured input definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Create') and the resource ('a tastytrade historical options backtest'), so the core purpose is unambiguous. It also adds a research-only qualifier that helps separate it from trading/order tools, though it does not explicitly differentiate from the sibling tastytrade_create_spx_spread_backtest.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'This is research-only and does not place brokerage orders' gives a useful boundary for when not to use it, implying it is for study rather than execution. However, it provides no explicit guidance about when to choose this tool over alternatives like tastytrade_create_spx_spread_backtest or tastytrade_simulate_trade.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_create_spx_spread_backtestA

Create an aggregate SPX spread Backtester job when the structure can be represented faithfully enough by relative leg selectors. Double diagonals remain exact-simulation only.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden, and it does meaningful work by revealing that this is an aggregate/approximate backtest rather than exact simulation, and that double diagonals are excluded. It does not fully describe the job lifecycle or return behavior, but it discloses the most important behavioral limitation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The primary use condition is front-loaded, and the exclusion is stated immediately after. Every sentence contributes selection or invocation guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is strong for tool selection but weaker for invocation and post-creation expectations. There is no output schema, no annotations, and a complex nested request object, so the description should say more about what happens after the job is created and how results are retrieved. The double-diagonal guardrail is helpful but not sufficient for full contextual completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The phrase 'relative leg selectors' adds meaning to the backtest_selector concept inside the request, and the aggregate-vs-exact distinction helps interpret the family field. However, it does not explain required fields like intended_price, price_effect, entry_at, or exit_at, leaving a meaningful gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Create'), a specific resource ('aggregate SPX spread Backtester job'), and the key mechanism ('relative leg selectors'). It also distinguishes this tool from exact-simulation tools by declaring that double diagonals are not supported here.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear when-to-use condition: use this when the structure is faithfully representable by relative leg selectors. It also gives an explicit when-not: double diagonals are exact-simulation only. However, it does not name the exact-simulation sibling tool directly, and 'faithfully enough' is somewhat subjective.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_get_backtestA

Get a tastytrade backtest by ID, including its current status and results when completed.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesBacktest ID returned by tastytrade.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It discloses that the response includes current status and results only when completed, which is meaningful. However, it does not mention error behavior, whether the operation is read-only, or what happens for an unknown ID.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It immediately identifies the action, the resource, and the key output characteristics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter getter with no output schema, the description provides a useful summary of what the tool returns. It is slightly sparse on edge-case behavior, but the core calling context is sufficiently covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the 'id' parameter already described as 'Backtest ID returned by tastytrade.' The tool description adds no additional parameter meaning, so the schema does the heavy lifting and the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a resource ('a tastytrade backtest'), and the identifier-based access method ('by ID'). It also clarifies what the tool returns ('current status and results when completed'), distinguishing it from list and create siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies the tool is used when you already have a backtest ID and want its status or results. It does not explicitly name alternatives or provide exclusion criteria, but the 'by ID' framing gives enough context to differentiate it from listing or creating backtests.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_get_backtest_available_datesA

List tastytrade Backtester symbols and their available historical date ranges. Use this before regression tests to verify SPY, XSP, SPX, or other symbol coverage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. The word 'List' implies a read-only operation, and the mention of date ranges adds useful context, but the description does not disclose output structure, pagination, or any other behavioral details beyond the core action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core action and subject are front-loaded, and the usage guidance follows naturally.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema tool, the description is complete: it states the resource, the information returned, and the intended workflow context. An agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parametersampions and schema coverage is 100%, so there is nothing for the description to clarify about inputs. The baseline for zero-parameter tools is 4, and the description adequately describes what the tool returns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists tastytrade Backtester symbols and their available historical date ranges, which is a distinct resource from sibling tools like tastytrade_list_backtests. The verb 'List' is specific and the subject is concrete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells agents to use this before regression tests to verify symbol coverage, giving a clear intended context. It does not explicitly name alternatives or when not to use it, but for a zero-parameter listing tool the guidance is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_get_backtest_logsC

Get tastytrade Backtester execution logs for a backtest ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesBacktest ID returned by tastytrade.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Get' implies a read operation, but the description does not mention absence of side effects, output format, pagination, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words. The core action and resource are front-loaded, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter retrieval tool, the description is minimally viable: it names the input and the expected resource. However, without annotations or an output schema, it leaves usage context and behavioral details undisclosed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: the single 'id' parameter is already described as the backtest ID returned by tastytrade. The description adds no extra meaning beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('tastytrade Backtester execution logs') for a given 'backtest ID'. This is clear and accurate, though it does not explicitly differentiate itself from the sibling tastytrade_get_backtest tool beyond the word 'logs'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as tastytrade_get_backtest or tastytrade_list_backtests. There are no prerequisites, exclusions, or context about whether logs are available before/after a backtest completes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_get_historical_candlesA

Retrieve normalized historical OHLCV candles from tastytrade DXLink for an exact UTC and session window, with source timestamps and explicit gap warnings. No resampling is performed.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses normalization, source timestamps, explicit gap warnings, and absence of resampling, which are meaningful behavioral traits. It does not cover auth, rate limits, or error behavior, but the disclosed traits are substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly packed sentence with no filler. It front-loads the core purpose and then adds only high-value caveats: source timestamps, gap warnings, and no resampling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is viable but not fully complete for a tool with a nested request schema, many parameters, and no output schema. It mentions return-related characteristics like source timestamps and gap warnings, but omits output structure, pagination, error behavior, and session timezone semantics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter semantics, but it only hints at 'exact UTC' and 'session window.' It does not explain interval units, instrument_type values, timeout_ms, max_candles, or session object details, leaving much of the parameter meaning to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Retrieve normalized historical OHLCV candles from tastytrade DXLink.' It further distinguishes this tool from siblings by emphasizing exact UTC/session windows, source timestamps, gap warnings, and no resampling, making its role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when to use this tool: when an exact UTC and session window is needed, with normalized candles, source timestamps, and gap warnings. It does not explicitly name alternatives or exclusion criteria, but the context is strong enough to guide selection among the backtest/simulation siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_list_backtestsA

List backtests submitted for the authenticated tastytrade API grant.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description must carry behavioral context. It indicates authentication scope and the 'submitted' filter, and 'List' implies a read-only operation, but it does not disclose pagination, ordering, response shape, or whether in-progress backtests are included.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It conveys the action, resource, and scope economically, which is appropriate for a zero-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple parameterless list operation, the description gives enough to know what the tool returns conceptually. It lacks explicit return-field or pagination details, but with no output schema, that is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is fully empty, so there is nothing for the description to explain. The baseline of 4 applies because parameter semantics are simply not applicable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') and resource ('backtests'), and adds a scope qualifier ('submitted for the authenticated tastytrade API grant'). This makes it easy to distinguish from sibling tools such as get_backtest, create_backtest, and get_backtest_logs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this to see all submitted backtests for the grant. However, it does not explicitly state when not to use it or point to a sibling alternative, so an agent has to infer the boundary from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_prepare_spx_spreadB

Normalize an SPX debit vertical, credit vertical, iron condor, or double diagonal into deterministic Backtester and exact-leg simulation requests without submitting it.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description is the sole behavioral source. It reveals that the tool does not submit the order and that it normalizes input for both backtester and exact-leg simulation, which is useful. However, it does not disclose whether it validates inputs, checks symbol existence, or what happens with invalid spread configurations. For a non-annotated tool, this leaves key behavioral aspects implied rather than stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that front-loads the types of spreads and the core output (backtester and exact-leg simulation requests) and ends with the critical non-submitting behavior. It is concise and accomplishes its purpose in one breath, though it could have added a brief list of outputs for even faster scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (deeply nested schema over four spread families, multiple simulation paths, and no output schema), the description is minimal. It does not mention the distinction between backtester and exact-leg simulation, nor the validation or normalization steps, nor what the return value looks like. Siblings like tastytrade_simulate_spx_spread and tastytrade_create_spx_spread_backtest exist, so an agent needs to know what this returns to use it as a prerequisite. The description is not sufficient on its own.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, but it is an extremely detailed schema with nested properties, enums for family/action/price_effect, and even a backtest_selector object. The description only says 'normalize' and lists the spread types, which adds minimal value beyond what the schema already shows. It doesn't explain the distinction between the backtest and exact-leg outputs, nor does it clarify the interplay of fields like backtest_selector and days_until_expiration. The schema carries most of the semantic burden, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool normalizes SPX spread variants (debit/credit vertical, iron condor, double diagonal) into two specific kinds of requests (Backtester and exact-leg simulation), and explicitly notes it does not submit the trade. This distinguishes it from siblings like tastytrade_simulate_spx_spread (which submits) and tastytrade_create_spx_spread_backtest (which creates a backtest).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'normalize... without submitting it' but does not explicitly state when to use it versus alternatives like tastytrade_simulate_spx_spread (which presumably submits) or tastytrade_create_spx_spread_backtest (which creates a backtest). An agent must infer that this is a preprocessing step, but there is no explicit guidance on when to call this first versus directly calling the simulation/backtest creation tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_price_option_packageC

Price a 2-4 leg option package with explicit native-package, synthetic-natural, and midpoint-reference provenance using exact decimal arithmetic.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does mention 'exact decimal arithmetic' and 'provenance', which are useful computational details, but it does not disclose whether the tool performs a read-only calculation, requires live quotes, has side effects, or how it handles missing or stale data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. The main action and scope appear immediately, though the dense provenance phrase adds complexity without much clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a complex nested schema, no annotations, and no output schema, yet the description only covers leg count and arithmetic style. It omits the request wrapper, required fields, family semantics, and the meaning of the provenance modes, leaving the agent reliant on schema inspection for critical context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only hints at '2-4 leg' and 'native-package'. It does not explain the required request object, the family enum, the structure of legs, or the meaning of fields like references, max_quote_age_ms, and max_temporal_skew_ms.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Price'), the resource ('option package'), and the scope ('2-4 leg'), which distinguishes it from sibling backtest and simulation tools. However, the phrase 'explicit native-package, synthetic-natural, and midpoint-reference provenance' is jargon-heavy and may obscure the core purpose for an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for pricing option packages, but it gives no explicit guidance on when to choose it over alternatives like tastytrade_prepare_spx_spread or tastytrade_simulate_trade. There are no when-not-to-use conditions or references to related tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_simulate_spx_spreadC

Run exact-leg historical simulation for a normalized SPX defined-risk spread and return a stable regression result contract.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It states it is a simulation and returns a contract, but does not disclose whether it is read-only, any side effects, authentication needs, rate limits, or what happens to data. The term 'stable' is vague and does not explain behavioral traits beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the action and output. However, it is extremely short given the complexity of the tool, but conciseness is high because every word is purposeful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex nested 'request' parameter with many required fields, and no annotations or output schema, the description is completely inadequate. An agent cannot know what to put in the request, what the contract contains, or any constraints. It is missing almost all contextual information needed to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters. It mentions none of the parameters, including the required 'request' object and its fields like 'family', 'legs', 'entry_at', etc. The description adds no semantic meaning to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Run' and a precise object: 'exact-leg historical simulation for a normalized SPX defined-risk spread'. Also specifies the return: 'stable regression result contract'. This is distinct from sibling tools like tastytrade_simulate_trade and tastytrade_create_spx_spread_backtest by mentioning 'exact-leg' and 'normalized'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives. Does not mention scenarios, prerequisites, or exclusions. Sibling tools like tastytrade_simulate_trade and tastytrade_create_spx_spread_backtest are not referenced, leaving the agent to infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_simulate_tradeA

Simulate a single historical option trade with tastytrade Backtester and return its historical path/results. Research-only; no brokerage order is placed.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYesRaw request body accepted by tastytrade POST /simulate-trade.

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explicitly discloses that no brokerage order is placed and that this is research-only, which is a critical safety-related behavior. It also mentions the return of historical path/results, though it does not cover error behavior, authentication requirements, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. The core action and output are front-loaded, and the research-only safety note is placed clearly. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations, no output schema, and a single opaque 'request' parameter that is essentially a raw API body. The description provides the high-level purpose and safety profile but does not explain how to construct the request, what fields are required, or what the returned path/results will look like. An agent would likely need external API documentation to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds minimal meaning beyond the schema, only indicating the trade is historical and single. The 'request' parameter is still opaque ('Raw request body accepted by tastytrade POST /simulate-trade'), and the description does not compensate with additional parameter-level details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Simulate'), a specific resource ('a single historical option trade with tastytrade Backtester'), and the expected output ('historical path/results'). It also clearly differentiates from sibling tools by emphasizing 'single' trade and 'Research-only', which distinguishes it from broader backtest creation tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage context with 'Research-only; no brokerage order is placed', but it does not explicitly state when to use this tool versus alternatives like tastytrade_create_backtest or tastytrade_simulate_spx_spread. The guidance is mostly implicit and relies on the agent inferring the difference from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

tastytrade_verify_historical_fillB

Verify whether a frozen paper-order limit was touched during a forward historical interval. Returns separate post-session evidence and never rewrites the live paper event.

ParametersJSON Schema
NameRequiredDescriptionDefault
requestYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full burden of behavioral disclosure. It does disclose a key non-mutating behavior: 'never rewrites the live paper event,' which is important for an agent to understand side effects. It also mentions that it returns 'separate post-session evidence.' However, it does not describe other behavioral traits such as rate limits, authentication requirements, or failure modes. The description is partially transparent but leaves gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the primary purpose. The two sentences have no redundancy and effectively convey the core function and a key behavioral guarantee. It is appropriately sized for a tool of moderate complexity, though it could include a bit more guidance without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (a single request object with many nested required fields, no output schema), the description is far from complete. It does not explain what 'frozen paper-order limit' means, what 'forward historical interval' implies, or how to structure the request. It also does not describe the output format or any error conditions. An agent attempting to call this tool correctly would be missing essential context, especially without annotations or an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides zero information about the 'request' parameter or its nested fields. The schema is complex with many required properties, but the description does not explain what any of them mean, how to construct a valid request, or which combinations are valid. With schema description coverage at 0%, the description must compensate, but it fails to do so. The agent would have to rely entirely on the schema, which has minimal field descriptions (only 'provider_symbol' has a description).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('verify') and a precise resource ('frozen paper-order limit') within a defined context ('forward historical interval'). It also clarifies the output ('separate post-session evidence') and the non-mutating behavior ('never rewrites the live paper event'). This clearly differentiates it from sibling tools like backtest creation or simulation, which are about generating or running tests rather than verifying an existing frozen order.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide explicit guidance on when to use this tool versus alternatives. It does not mention conditions that would select this tool over others (e.g., when you need to check if a limit was touched in a past session) nor does it name any sibling tools. The verb 'verify' and the mention of 'frozen paper-order limit' imply a verification use case, but without explicit context or exclusions, the agent is left to infer usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 13 tool updatesv0.2.0
    • First observedtastytrade_cancel_backtest
    • First observedtastytrade_create_backtest
    • First observedtastytrade_create_spx_spread_backtest
    • First observedtastytrade_get_backtest
    • First observedtastytrade_get_backtest_available_dates
    • First observedtastytrade_get_backtest_logs
    • First observedtastytrade_get_historical_candles
    • First observedtastytrade_list_backtests
    • First observedtastytrade_prepare_spx_spread
    • First observedtastytrade_price_option_package
    • First observedtastytrade_simulate_spx_spread
    • First observedtastytrade_simulate_trade
    • First observedtastytrade_verify_historical_fill

TDQS

A3.6/5.0

Scored across 13 tools

Disambiguation4/5

Most tools are clearly distinct by resource and action, but the backtest/simulation cluster (create_backtest, simulate_trade, simulate_spx_spread, create_spx_spread_backtest) risks confusion despite clarifying descriptions. The price/prepare/simulate spread pipeline is well-separated with explicit roles.

Naming Consistency5/5

All tools follow a consistent tastytrade_ prefix with snake_case verb_noun naming (list_backtests, create_backtest, get_backtest_logs, cancel_backtest). Verbs and objects are uniform and predictable across the entire set.

Tool Count5/5

13 tools is well within the ideal 3-15 range for a research-focused server. Each tool contributes to distinct workflows such as backtest management, option pricing, spread simulation, and historical data access without redundancy.

Completeness4/5

The core research workflow is well covered: backtest lifecycle (list, create, get, cancel, logs), historical simulation, option pricing, spread normalization, and candle retrieval. Minor gaps exist such as lack of a delete-backtest or batch operation, but no critical dead ends are apparent.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    exposes a remote MCP endpoint so agents can: run strategy backtests by symbol/timeframe/date range, pass strategy inputs programmatically, receive structured backtest results (trades, win rate, profit, drawdown), keep long-running runs observable via progress notifications, support Binance Futures tickers only, enforce a maximum of 1440 candles per backtest, apply a rate limit of 3 backtests per
    6
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables interaction with the Nubra trading platform for authentication, instrument lookup, quotes, historical data, options analytics, portfolio management, report generation, screening, backtesting, and order placement via UAT environment.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables backtesting of limit-order strategies on Polymarket's BTC 5-minute markets using historical data, with tools to browse markets, get price series, and run simulations.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables interactive access to TradeSearcher strategies and backtests via CLI and MCP, allowing agents to search, backtest, and compare trading strategies.
    MIT