Skip to main content
Glama
getsimba-ai

Simba MCP Server

Official
by getsimba-ai

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
SIMBA_API_KEYYesYour Simba API key (stdio mode only – HTTP callers send their own key as the bearer token)
SIMBA_API_URLNoSimba API base URLhttp://localhost:5005
SIMBA_MCP_ALLOW_LOCAL_FILESNoSet to '1' to allow local file paths (csv_path) on HTTP/SSE deployments; disabled by default.0

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
show_response_curvesB

Show served response curves, bands and current spend in a native chart.

Read-only. Returns the existing JSON unchanged on every client. Missing points are gaps; sparse legacy grids are disclosed. Current spend is not a recommended allocation. Requests fixed sections and never accesses prediction-window data.

show_decompositionB

Show served contribution time series in KPI units in a native chart.

Read-only. Overlap is a separate reconciliation term, never a channel. The attribution convention is displayed where available. Returns the same useful JSON on clients without visual support. Never requests prediction-window data.

show_optimizer_allocationA

Show saved allocation spend and separate decision/comparison result tables.

Read-only. Pass run_id to read a specific saved run; omission reads the latest model-level result. Decision Revenue/ROI never mixes with fitted-convention OptimizedEvalRevenue/ROI or HistoricalRevenue/ROI. No scientific calculations occur in the view. Returns existing JSON unchanged without visual support.

create_recipe_draftA

Save an encrypted authoring draft without publishing, fitting or consuming an attempt. Supply a UUID draft_id and reuse it with identical content after an uncertain response. Start from get_recipe_draft_template for a new draft, or from get_recipe_revision_authoring when working from a published revision. To edit a recipe in place, pass target={recipe_id, base_revision_id}: the draft becomes an edit session of that recipe and publish_recipe_draft with a matching target writes the recipe's next revision; the target is immutable after creation (409 on a replay that names another; 404 for a recipe outside the study). To branch a published recipe into a new one instead, omit target and supply its source_revision_id from this study; the immutable link carries inherited influence through publication. Backend validates version and size.

get_recipe_draftA

Read the complete authoring snapshot and concurrency version. Preserve all fields when editing. Draft state is incomplete, unvalidated authoring data, not an executable recipe.

get_recipe_draft_templateA

Get complete shared wizard defaults, hash and envelope schema. Optionally choose an owned uploaded_file_id (from list_uploads) or pipeline_version_id (from list_pipelines / list_pipeline_versions), never both, into frozen source bytes with verified lineage and an editable data preview. Copy snapshot into create_recipe_draft and preserve unedited fields. Defaults are not a validated model. Does not create, publish or run anything.

publish_recipe_draftA

Publish a saved MMM or VAR draft as immutable recipe revisions. Supply a new UUID publication_id and reuse it with identical arguments after an uncertain response; a replay returns the same revisions. Without target: new recipes, one per prepared brand, atomic. With target={recipe_id, expected_version}: revision N+1 of that recipe under its row lock, the draft's base revision recorded as the source. 412 stale_version when the recipe moved since expected_version: the draft and every edit are kept, so re-read the recipe and publish again on top, or drop target to branch into a new recipe. 409 target_requires_single_recipe for a multi-brand draft; 409 when the draft's own target names a different recipe; 404 for a recipe outside the study. Backend compiles the saved snapshot with shared wizard rules; never send separately prepared settings. Does not fit, consume an attempt or designate a champion. MMM priors the draft leaves unset are filled with smart defaults at publication (the Prior Builder's own calculation over the whole source) and frozen; replay does not rebuild them. VAR requires family-specific evidence assessment; MMM policies cannot establish VAR acceptance. Enabled invalid MMM calibration fails publication; VAR calibration is unsupported. Preserve disabled authoring observations.

get_recipe_revision_authoringA

Read the authoring snapshot behind a published revision. authoring_draft, wizard (Save a recipe) and base_model (imported fitted model) revisions carry one; api_mmm, model_snapshot and legacy rows return an explicit unavailable error (404). Returns snapshot, name, revision_id, draft_content_hash, kind, source_available and source_unavailable_reason. Dataset bytes are filled from the recorded origin only when it still hashes to what was fitted; otherwise snapshot.source is null, source_available is false and the reason says to choose the dataset again before the draft can publish. Use the snapshot with create_recipe_draft: target for an in-place edit, source_revision_id for a branch. The published revision stays unchanged. Does not create or fit anything.

list_recipe_draftsA

List study draft metadata without loading datasets, newest update first. Check backend draft capability first. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.

update_recipe_draftA

Replace authoring state using the version from get_recipe_draft. Retain every unedited field, including original priors and source bytes. Stale changes fail; reload and reconcile explicitly. Identical retries return the current draft. Does not publish or launch.

get_backend_capabilitiesA

Discover this caller's connected backend features before planning work.

Returns only backend advertisements: model families, transformations, priors and workflow operations. A missing advertisement is unknown, not unsupported. Check each field; an advertised feature still requires permission and budget. No model is created and capabilities are not cached across callers.

get_data_schemaA

Get the canonical CSV data schema for Simba MMM input files.

Returns the JSON Schema specification describing required columns (date, KPI, multiplier, hierarchy), media channel column naming conventions ({channel}_activity, {channel}_spend), constraints (min rows, max file size), and supported date formats.

get_data_reportA

Report actual data from a stored dataset: any window, any grain, by brand, channel or dimension.

Reads the dataset itself (every column, any date range) rather than a fitted model's training window. Use it for "sales and TV spend in the North region for August, by week".

Roles are DECLARED, never guessed from column names. Declare them with upload_data(roles=...) or per request with roles. Without a declaration only the schema's own naming rules apply: a column named date, {channel}_spend and {channel}_activity. Every other column is reported as role "unknown" and is not aggregated — declare the KPI and hierarchy columns.

Role vocabulary (aggregation, unit) — also in get_data_schema under x-simba-roles:

  • kpi (sum), spend (sum, currency), activity (sum), multiplier (mean)

  • outcome:online_sales|store_sales|margin (sum, currency), outcome:orders|new_customers (sum)

  • media:impressions|clicks|grps (sum; give a channel: {"role": "media:grps", "channel": "tv"})

  • control:price|rate|index (mean), control:stock (each brand's last value, summed)

  • hierarchy, dimension:market|product|campaign (keys for filtering and group_by)

  • date

Buckets: week = ISO week from Monday; month/quarter = calendar. A weekly row counts in the month of its week-start date. The response's meta.aggregation states every rule applied.

Args: dataset_id: The uploaded file id (from upload_data or list_uploads). Registered pipeline outputs are uploaded files too. start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (default), "week", "month" or "quarter". group_by: "hierarchy", "channel", or a dimension role such as "dimension:market". hierarchy: Keep only this brand/region value. metrics: Roles or role families to include, e.g. ["kpi", "spend", "outcome:orders"] or ["control"]. Default: every metric role present. roles: {column: role | {"role", "channel"}} overriding roles stored at upload.

Returns {dataset: {id, name, source, version, sha256, data_through}, granularity, rows: [{period_start, period_end, group, metric, value, unit}], meta: {basis: "dataset", aggregation, roles, channels}}. Errors carry a code: dataset_not_found (404), invalid_report_request (400), report_too_large (413, over 10,000 rows — narrow the window or coarsen the granularity).

upload_dataA

Upload a CSV dataset to Simba for use in model building.

Provide EXACTLY ONE of csv_content (raw CSV text) or csv_path (a file path on the machine running this MCP server). Prefer csv_path for anything beyond trivial size — it avoids passing megabytes of CSV through the conversation.

The CSV should follow the canonical schema: one row per time period with date, KPI, multiplier, hierarchy, media activity/spend columns, and optional control variables.

IMPORTANT:

  • CSV only (not Excel). Maximum file size: 10 MB (API-enforced).

  • Row minimum: check get_data_schema -> x-simba-constraints.min_rows for the declared minimum; enforcement may be more permissive, and the upload response's warnings field is authoritative. More rows = tighter posteriors (104+ weekly rows recommended).

  • Media columns must follow naming: {channel}_activity and {channel}_spend.

  • Use 0 for inactive periods, not blank or NA.

  • csv_path is only available when the server runs locally (stdio). On HTTP/SSE deployments it is disabled unless SIMBA_MCP_ALLOW_LOCAL_FILES=1.

Args: csv_content: The full CSV text content (not base64, just raw CSV text). csv_path: Path to a .csv file readable by the MCP server process. name: Optional dataset name for identification. Defaults to the file stem when csv_path is used. filename: Optional original filename to record alongside the dataset. roles: Optional column roles for get_data_report, stored with the dataset: {column: role} or {column: {"role": role, "channel": name}}. Roles are declared, never guessed; see get_data_report for the vocabulary. An unknown role or a column the CSV lacks is refused.

Returns the uploaded file ID (needed for create_model), row/column counts, and any validation warnings.

list_pipeline_versionsA

List the saved versions of one owned pipeline (by pipeline_hash or id), newest first: id, version, created_at, row_count, column_count, column names and checksum. The checksum is the exact source identity a recipe draft freezes. Never returns the data itself; fetch a draft template with pipeline_version_id for that. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end.

list_pipelinesA

List the data pipelines this key's owner has, most recently updated first, each with id, pipeline_hash, name, description, version_count and latest_version (id, version, created_at, row_count, checksum). Identity only: no definitions, parameters or output data. Use a version id as pipeline_version_id in get_recipe_draft_template. Requires the ingest scope; the same ownership rule as the app. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end.

run_pipelineA

Start a refresh of one owned data pipeline (by pipeline_hash or id). Returns {run_id, status: "queued"} at once; the run executes on the server as the pipeline's owner with its saved connections. Poll get_pipeline_run until status is succeeded or failed (every few seconds; warehouse runs can take minutes, and a run stops at 30 minutes). Optional start_date / end_date (YYYY-MM-DD) limit the source steps to that range. One run per pipeline at a time: if one is already queued or running you get run_in_progress with that run_id — poll it instead of starting another. Each successful run saves a new pipeline version. Requires the create:models scope.

get_pipeline_runA

Get one run of an owned pipeline: {run_id, status (queued | running | succeeded | failed), started_at, finished_at (UTC), version_id, error_code, error}. On success version_id is the new saved version — pass it as pipeline_version_id to get_recipe_draft_template to build on the refreshed data. On failure error_code says why: execution_failed (a source or transform failed; error names it), no_output, timeout, interrupted or not_started (start it again), not_queued, owner_blocked, unexpected. Scheduled runs are polled the same way. Requires the create:models scope.

set_pipeline_scheduleA

Replace the refresh schedule of one owned pipeline. cadence is "daily" or "weekly"; hour_utc is a whole UTC hour 0-23; weekday (0 = Monday … 6 = Sunday) is required for weekly and must be omitted for daily; enabled false pauses the schedule and keeps its settings. Returns the schedule with next_run_at (UTC). Each due slot starts one ordinary run (poll it with get_pipeline_run); a slot is skipped while a run of that pipeline is still going. After a scheduled run succeeds the pipeline keeps its newest 30 versions; a version a model or recipe was built from is never removed. Requires the create:models scope.

list_campaignsA

List the campaigns in the campaign facts, each with its totals and the model channel it counts towards for this model.

The channel is DECLARED in the model's campaign map (set_campaign_mapping), never inferred: a campaign no map row covers is status: "unmapped" with channel: null, is listed with its spend, and is not counted towards any channel. suggested_channel is a hint from the platform's own label and the campaign's name (null when nothing is clear); it does nothing until it is written into the map.

Returns {campaigns: [{platform, account_id, campaign_id, campaign_name, adsets, spend, impressions, clicks, platform_conversions, platform_value, platform_channel_type, first_seen, last_seen, channel, suggested_channel, status: "mapped" | "unmapped"}], window: {start, end}, currency, as_of, next_cursor (only when limit was given)}. Sorted by platform, account and campaign id. Totals cover the window (default: every stored day).

Args: model_hash: The model whose map decides each campaign's channel (the hash of a fitted MMM). platform: Keep one platform: meta, google_ads, tiktok or other. unmapped_only: Only campaigns without a channel for this model. start, end: Optional ISO dates (YYYY-MM-DD), inclusive, for the totals. limit: Page size (1-200). Without it every campaign is returned. cursor: The previous page's next_cursor.

Errors carry a code: model_not_found (404). An empty list means no pipeline is registered as the campaign facts source yet, or it has not run; registration happens in the app.

get_campaign_reportA

Report the campaign facts over any window, at any grain, by platform, model channel, campaign or ad set.

The same report engine as get_data_report, over the campaign facts: spend, media:impressions, media:clicks, outcome:platform_conversions, outcome:platform_value and, where a source fills them, outcome:last_click_conversions and outcome:last_click_value. These are the PLATFORMS' own attributed numbers, not Simba's incremental attribution.

group_by: "channel" groups by the model channel each campaign counts towards under the model's map (give model_hash); spend of unmapped campaigns appears as its own group, "unmapped", and is never dropped or guessed into a channel. Without model_hash every row is "unmapped".

Returns {granularity, data_through, rows: [{period_start, period_end, group, metric, value, unit}], meta: {basis, aggregation, roles, channels}, as_of, source_versions, currency}. as_of is when the facts were last ingested; source_versions the pipeline versions they came from.

Args: model_hash: The model whose map decides the channels (optional; needed for group_by=channel). start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (daily, as stored), "week", "month", "quarter" or "year". group_by: "platform", "channel", "campaign" or "adset". Default: one total per period. platform, channel, campaign_id: Filters applied before grouping. metrics: Roles to include, e.g. ["spend", "outcome:platform_conversions"]. Default: all.

Errors carry a code: campaign_facts_empty (404, nothing matches), invalid_report_request (400), report_too_large (413: narrow the window or coarsen the granularity).

get_campaign_incrementalityA

Incremental ROAS per campaign (or ad set), beside the platform's own ROAS and last-click ROAS, by pushing the model's channel incrementality down through the platform's attribution.

THE ASSUMPTION, FIRST. Simba measures incrementality at channel grain. For each model channel over the window, the incrementality factor = the channel's MMM incremental revenue (the model's per-period rows, under its fitted attribution convention) / the platform-attributed value of the campaigns mapped to that channel. Each campaign's incremental ROAS is that factor x its platform ROAS, so campaign incremental revenue sums to the channel's. This assumes the platform over-credits every campaign in a channel equally. It does not: retargeting and brand search are over-credited more, so one factor flatters them. The response warns when such campaigns share a channel with prospecting (retargeting_shares_channel_factor); the remedies are to map them to their own model channel (set_campaign_mapping) or to calibrate the factor with an incrementality test. Nothing here is a causal per-campaign measurement; every row says how it was made.

Per row, method is "attribution_scaled" or, when a channel's campaigns carry no platform value, "spend_share" (the channel's incremental revenue shared by spend). A campaign without platform value in a channel that has some gets iroas: null and is named (platform_value_missing); it is never given a share. factor_source is "mmm" or "test": a completed incrementality test in the model's project that names the channel and overlaps the window replaces the model's factor (lift in revenue units / the channel's platform value during the test); factor_mmm stays beside it.

interval is "pending" (the 94% bands come from the model's posterior draws, computed by a background job the first time a window is asked for; ask again in a few minutes), "ready" (factor_interval per channel, iroas_hdi and incremental_revenue_hdi per row) or "unavailable" (interval_reason says why; point estimates stand, no band is invented). uninformative warns when a channel's band spans zero.

Returns {model_hash, window: {start, end}, level, currency, interval, interval_reason, channels: [{channel, factor, factor_source, factor_mmm, factor_interval, revenue_interval, factor_draws_mean, method, mmm_revenue, platform_value, spend, campaigns, test, warnings}], rows: [{platform, account_id, campaign_id, campaign_name, adset_id, channel, spend, platform_value, last_click_value, days, platform_roas, last_click_roas, incremental_revenue, iroas, incremental_revenue_hdi, iroas_hdi, method, factor_source}], unmapped: [{..., spend, platform_roas, last_click_roas}], warnings: [{code, message, channel?, campaigns?, reason?}], provenance: {source_versions, as_of, map_version, attribution_convention, link}}. Warning codes: retargeting_shares_channel_factor, platform_value_missing, kpi_not_revenue, currency_mismatch, uninformative, unmapped_spend, interval_unavailable, test_override_skipped.

Args: model_hash: A fitted MMM with a campaign map (set_campaign_mapping). start, end: Optional ISO dates (YYYY-MM-DD), inclusive. Default: the overlap of the model's data and the campaign facts. level: "campaign" (default) or "adset"; ad-set rows inherit their campaign's channel and factor and sum to the campaign row.

Errors carry a code: model_not_found (404), campaign_facts_empty (404: no facts, or none in the window), invalid_window (400: empty or reversed; the body gives both spans), model_not_mmm (400: a VAR model has no channel revenue rows), model_incomplete (400).

get_campaign_marginal_returnsA

Read channel-derived marginal returns for campaigns or ad sets. Requires read:models.

Campaign response shapes inherit the fitted channel shape, rescaled by observed spend share and relative efficiency. These are not independently measured campaign saturation curves or causal campaign effects. Daily spend is observed spend divided by inclusive calendar days, not a configured platform budget or a daily revenue forecast.

Args: model_hash: Owned, completed MMM with compatible curve provenance and campaign facts. start, end: Inclusive observation-window ISO dates (YYYY-MM-DD). level: campaign or adset; ad sets inherit the campaign's channel.

Returns {context_key, window, currency, minor_digits, channels: [{channel, current_daily_spend, status, reason}], rows, provenance}. Rows carry composite identity, method and marginal-return evidence. Missing or incompatible currency, curve basis or fact coverage makes recommendations unavailable; do not substitute defaults. Missing posterior marginal evidence means no uncertainty interval, not zero uncertainty. This reads existing evidence only: no fit, posterior job or platform change is started.

recommend_campaign_budgetsA

Calculate campaign budget suggestions without applying them. Requires read:models.

Uses the same channel-derived marginal curves and bounded allocator as the web app. This POST is a read-only calculation: it creates no saved run, starts no fit or posterior job and does not change mappings or advertising-platform budgets. Results are conditional scenarios, not independently fitted campaign response curves or a day-by-day forecast.

Args: model_hash: Owned, completed MMM with verified curve and currency provenance. observation_window: {start: YYYY-MM-DD, end: YYYY-MM-DD}, inclusive fact dates. currency: Explicit currency matching the model and facts, for example GBP. channel_daily_budgets: Exact channel keys mapped to daily totals in that currency. Supply exactly one of this or optimizer_run_id. Preserve currency minor precision. optimizer_run_id: Existing owned optimiser run for the same model, with verified planning dates and currency. Its totals become flat daily equivalents, not a daily schedule. The tool does not create or rerun an optimiser. level: campaign or adset. max_step_fraction: Maximum change from observed average daily spend; default 0.2. bounds: Optional rows keyed by full identity {platform, account_id, campaign_id, adset_id}, with min and/or max amounts in daily currency units. Use null adset_id at campaign grain. No platform floor is invented; the server validates bounds. expected_context_key: Optional context_key returned by get_campaign_marginal_returns. A changed model, facts, mapping or attribution basis refuses with 409 instead of calculating against evidence different from the scenario you reviewed.

Returns {channels: [{channel, status, total_daily_budget, rows, explanation, assumptions, reason, feasible_range}], ...provenance}. Ready rows include current and recommended budgets, the continuous solution, marginal returns and binding constraints. Refusals retain their reasons and feasible ranges; never relax bounds or silently change a total. Rounded amounts reconcile in integer currency minor units; rounding does not imply exact equality of marginal returns. Missing marginal uncertainty remains explicitly unavailable.

set_campaign_mappingA

Replace the model's campaign map: which model channel each campaign counts towards.

Each row is {platform, campaign_id | name_pattern, channel, valid_from?, valid_to?}: exactly one of campaign_id (the platform's id, exact) or name_pattern (a glob on the campaign name, e.g. "YouTube*", case-insensitive); channel must be one of the model's channel names (as in get_model_results); dates are ISO and optional (open-ended when absent). An exact-id row beats any pattern row. The whole map is replaced by rows; send every row you want kept.

Refused (422, code campaign_map_invalid, nothing saved) when a channel is not the model's (the error names the closest one), when a campaign would count towards two channels on any date (two exact rows on one campaign with overlapping dates, or two patterns matching one campaign with overlapping dates; conflicts lists them), or when a row is malformed.

Returns the map report: {map, unmapped: [{platform, campaign_id, campaign_name, spend, first_seen, last_seen, platform_channel_type, suggested_channel}], conflicts: [], drift: [{channel, campaign_spend, model_spend, ratio, over_tolerance, campaigns}], tolerance, currency, currency_mismatch, window, channels}. drift compares the spend of the campaigns mapped to each channel with the model's own spend for that channel over the dates both have; over_tolerance is a WARNING (the map is saved), usually a campaign that belongs elsewhere or model data that stops before the campaigns do. Unmapped campaigns stay listed and are not counted; map them in a later call.

Args: model_hash: The fitted MMM the map belongs to. rows: The complete map. tolerance: The drift warning threshold as a fraction (default 0.05).

list_incrementality_testsA

List a project's recorded (not retired) incrementality tests: {items: [{id, name, type, status, channel, model_channel, start_date, end_date, measured_through, kpi, result, spend, source_tool, has_supplied_row, current_version, used_by}], next_cursor}; used_by counts the model revisions built from the test. Null and negative results are listed like any other. Filter by type, status or channel. Paging is opt-in: pass limit (1-200) and send next_cursor back unchanged; null means the end. Requires the read:models scope.

recommend_incrementality_testsA

Rank channels for experiment investigation using stored posterior marginal returns. Read-only: no fit, test creation or budget changes. Returns {method, basis, score_unit, currency, budget, hurdle, spend_basis, period, approximation_warnings, items, excluded}. Each item carries channel, score, components (mean, sigma, stake, spend_share, crossing_probability, optional contraction), reason_codes, last_test_end, hypothesis and an unavailable design_hint. The score is a normal-approximation local binary perfect-information value, not expected test benefit, experiment budget, portfolio value or forecast lift. budget is a positive exposure scale (default sum of current spend); weights are the mean-active-period spend mix, which need not represent one common calendar period. hurdle is the non-negative marginal-return alternative (default 1). limit is 1-50. Missing posterior means or intervals are explicitly excluded. limited_variance_contraction means contraction from zero to below 0.1; posterior_variance_expanded means negative contraction. Neither proves prior domination. Cross-channel dependence is not modelled. Requires read:results. Test history is unavailable because stored registry records do not establish compatible model/geographical coverage. Requires backend support for test-priorities.

get_incrementality_testA

Read one recorded test: {id, version, record, content_hash, used_by, retired_at}. version reads an older version (default: current). With model_hash (a saved model you can read), the result also carries calibration: either {status: "ok", row: {channel, x, delta_x, delta_y, sigma, sigma_low?, sigma_high?}, units, steps, warnings} — the likelihood observation this test gives that model, each step stated — or {status: "refused", reason, message, steps}. A refusal is an answer, not an error: e.g. channel_not_in_model (pass channel, a model activity column), kpi_mismatch (pass confirm_kpi=true only if the test's outcome really is the model's KPI), no_spend, test_not_completed, owned_media_not_calibratable, window_overlaps_holdout. Use the same references in create_model(calibration={tests: [...]}). Requires the read:models scope.

create_incrementality_testA

Record one incrementality test in a project, as analysed in its own tool: its design (geo, owned_media_ab or platform_lift block), dates, KPI, result with interval or sd, and incremental spend. Returns {id, version: 1, record, content_hash}. Set model_channel to the model activity column the test calibrates so it can be used with a model. Validation errors name the field (e.g. "the interval must contain lift_abs"). Not idempotent: calling twice records two tests. type time_holdout is not accepted here: planned time-holdout records are created only through save_incrementality_test_design (the backend refuses them elsewhere), and a status planned record of any type carries no result. Requires the create:models scope.

import_incrementality_testsA

Import tests from another tool's output file: source is csv (Simba's template), meta_conversion_lift (Conversion Lift API results JSON), geox (a meridian-geox analysis result), geolift (GeoLift summary), causalpy (effect summary or lift rows) or pymc_marketing (lift rows); content is the file's text (10 MB max). Returns {records: [{key, record, errors}], created, notes}. dry_run (default true) creates nothing — review each row's errors, then call again with dry_run=false to create the rows without errors (their ids come back in created). defaults fills fields the file doesn't carry, e.g. {"channel": "TV", "model_channel": "tv_grps", "kpi": {"kind": "revenue"}}: a GeoX result names no channel or KPI, so its rows fail until those are given. overrides sets fields on one row by the key the dry run showed (a cell id, or the 1-based row number), e.g. {"1": {"spend": {"incremental": 25000}}}. Values deep-merge: the source's assumptions, then defaults, then the file, then overrides; a null removes a field. import_invalid means the file isn't that source's format. Requires the create:models scope.

design_incrementality_testA

Ask a saved, complete MMM what an incrementality test on one of its media channels could detect: queues a bounded calculation that replays the saved posterior and returns {calculation_id, model_hash, status: "queued", submitted_at} at once (202; 200 with the same body when submission_key repeats identical inputs). Poll get_incrementality_test_design until status is complete, failed or cancelled. A complete calculation carries a result whose own state is available, unsupported (e.g. a log-link model), insufficient_evidence or no_feasible_design; none of these is an error, and only available carries numbers (candidates are also returned for no_feasible_design so you can see why nothing met the target). channel is one of the model's nonlinear media channels (400 unknown_channel otherwise). design_type time_holdout pauses or changes the channel's spend for a window and contrasts the outcome with the model's forecast; geo_split needs geo and splits a market panel into treatment and control. intervention gives the start date, candidate durations, spend change and baseline; inference the alpha, target power and an optional named effect; the backend fills and echoes every default. Reuse the same submission_key after a lost response; a new key queues another calculation and the same key with different inputs is refused (409 submission_key_conflict). Refusals: 404 model not found or not readable, 422 model_not_complete, 503 queue_unavailable (nothing is left behind). Nothing is saved or launched and no budget changes; save an available result with save_incrementality_test_design. Requires create:models to submit; polling needs read:results as well, so a key with only create:models can submit but not read.

get_incrementality_test_designB

Read one test-design calculation: {calculation_id, model_hash, status, submitted_at, started_at, completed_at, request, error, result}. status is operational: queued, running, complete, failed or cancelled; error is {code, message} only when failed (artifact_unreadable, timeout, worker_lost, internal) and a failed calculation carries no numbers. result is present only when complete and is read by its own state, never by key presence: available, unsupported, insufficient_evidence or no_feasible_design, with reasons [{code, detail}] non-empty exactly when not available, plus warnings, assumptions, diagnostics and provenance (method, model artefact, posterior draws, training periods). An available result carries intervention (dates, carryover and measurement end, spend change), model_implied_effect (mean and 94% HDI, parameter uncertainty only; available may be false, e.g. a geo split on a national model, which then gives a national reference only), detectable_effect (cumulative, per-period, relative and, for revenue KPIs, iROAS MDE at the stated alpha and target power), power (at the model mean effect, at the named effect, and assurance), noise, and candidates per duration with the selected one marked. Power here is posterior-averaged: assuming the model is correctly specified and its posterior calibrated, the probability that the pre-specified analysis rejects the null, averaged over the model's forecast uncertainty and observation noise; it is not the probability that this experiment will detect the effect, and a model-implied effect is not a prediction of what the experiment will measure. Requires read:results.

save_incrementality_test_designA

Turn an available test-design result into a planned incrementality test record in the registry: returns {id, version, record, content_hash} (201; 200 with the same body when this calculation was already saved). The server builds the record from the stored result (type time_holdout, or geo with treatment and control markets; status planned; source.tool design with the calculation, model artefact and method in source.fields), so only the target project_id (default: the model's project; 400 when the model has none and none is given) and an optional name are sent. Idempotent per calculation. Refusals: 409 result_not_available when the calculation is not complete or its result state is not available, 409 artifact_changed when the model artefact no longer matches the calculation, 403 when you are not the project's owner. A planned record has no measured lift: it cannot calibrate a model (get_incrementality_test reports type_not_calibratable or test_not_completed) and nothing here launches an experiment or changes a budget. Requires create:models and write access to the project.

list_uploadsA

List the datasets in your workspace (newest first) — every source, not just API uploads: dashboard/manual uploads and pipeline-ingested datasets appear too (see source_type per file).

Returns {files, count, limit, offset} where each file has: id (the uploaded_file_id create_model needs), filename, original_filename, source_type, row_count, column_count, created_at. Here count IS the true total matching the filter (unlike list_runs, where it is the page length). Column names/dtypes are not in the listing — fetch one upload with get_upload for those.

Args: limit: Page size (API clamps to 1-500; default 50). offset: Rows to skip (paging). name: Optional case-insensitive substring filter on the original filename.

get_uploadA

Get one uploaded dataset's details, including its column schema.

Returns id, filename, original_filename, source_type, mime_type, file_size, row_count, column_count, columns ([{name, dtype}, ...] — use these to build create_model's channel/control column arguments without re-reading the CSV), and created_at.

Args: file_id: The upload's id, from upload_data's response or list_uploads.

list_modelsA

List all Marketing Mix Models for the authenticated user.

Returns model name, hash, status (pending/under way/complete/failed), type (mmm/var), hierarchy value, and timestamps.

NOTE: All other model endpoints use model_hash (string, e.g. "f835671a25") as the identifier. Use the model_hash from this response.

Args: include_unsaved: Include draft/unsaved models (default false). limit: Maximum number of models to return (default 50, max 500). offset: Number of models to skip, for paging past limit (default 0).

create_modelA

Create and start fitting a new Bayesian Marketing Mix Model.

This queues an async model fit and returns immediately with a model_hash. Use get_model_status to poll for progress until status is 'complete'.

Priors are calculated automatically using smart defaults based on cost shares, industry benchmarks, and channel-type detection. You can override individual channels via the priors parameter.

Args: uploaded_file_id: The file ID returned by upload_data. date_column: Name of the date column in the CSV. kpi_column: Name of the KPI/dependent variable column. hierarchy_column: Name of the brand/segment column (must have exactly 1 unique value). channels: List of channel definitions, each with keys: name, activity_column, spend_column. Example: [{"name": "TV", "activity_column": "tv_grps", "spend_column": "tv_spend"}] multiplier_column: Column to convert KPI to revenue. Defaults to kpi_column. control_columns: Non-media control variable column names (e.g. ["price", "distribution"]). total_media_effect: Controls prior strength. Either an industry name for a benchmark ("FMCG"=6%, "Retail"=9%, "TelCo"=30%, "Financial Services"=19%, "E-Commerce"=22%, "Other"=12%) or a custom decimal like "0.15" meaning "I believe all media drives 15% of my KPI". Default "Other". priors: Optional per-channel prior overrides. Each dict should have "channel" matching a channels[].name, plus any fields to override: distribution, mean, sd, lower, upper, transform, adstock_type, effect_period. Only specified fields are overridden; the rest use smart defaults. Adstock-kernel fields: half_life_lower/half_life_upper (carryover half-life bounds in periods — preferred over the legacy decay_lower/decay_upper), theta_mean/theta_sd (peak-lag prior, adstock_type="delayed" only), dual_weight_mean/dual_weight_sd (long-term/slow-component share prior, adstock_type="dual_geometric" only). SATURATION ANCHOR — state it ONCE, in exactly one of three mutually exclusive forms (two in one override -> 400 "state the saturation prior once"): (1) half_marginal_mean/half_marginal_sd — CANONICAL for saturation_type="generalized_log" (rejected on other families): the activity level where MARGINAL returns have halved, finite at every curvature (#632). sat_shape_mean MUST accompany the pair in the same override (#672) — the fold pairs your coefficient with the stated curvature, so omitting it is a 400, never a silent default. (2) half_saturation_mean/half_saturation_sd — the 50%-of-maximum point in activity units, for the single-parameter families (tanh/michaelis_menten/ negative_exponential). Do NOT use it for generalized_log near-log work: it overflows below sat_shape_mean 0.00097657 and is rejected with a 400 — precisely the regime that family exists for. (3) alpha_sd + scalars — legacy internal coordinates, accepted for backward compat. Curvature (generalized_log only): sat_shape_mean/sat_shape_sd — small values are near-logarithmic, 1.0 is michaelis_menten. COEFFICIENT in a human coordinate (generalized_log only, #671): effect_at_avg_mean/effect_at_avg_sd — the effect share at the channel's AVERAGE activity, as FRACTIONS (mean in (0, 0.95], sd > 0; 0.2 means 20%). Folded server-side into mean/sd at the row's operating point with the same arithmetic as the dashboard. Requires sat_shape_mean in the same override; cannot be combined with mean/sd ("state the coefficient prior once") or with half_saturation_*. Stating half_marginal_* + effect_at_avg_* + sat_shape_mean together is the full (x*, E, k) triple — the recommended generalized_log elicitation, since only beta*k is identified and raw beta spans orders of magnitude. VERIFY what was applied via get_model's model_config.priors_resolved: rows carry the FOLDED mean/sd/scalars/alpha_sd, and overridden_fields lists the field names you sent. UNKNOWN KEYS ARE REJECTED with a 400 naming the field (#630); they used to be dropped silently, fitting a hybrid of the override and the smart defaults. Common misses: "beta"/"beta_mean" -> mean, "beta_sd" -> sd, "sat_shape" -> sat_shape_mean. "name" and "parameter" are rejected too — they identify the smart-prior row the override merges onto. trend: Enable dynamic baseline trend component. seasonality: Enable automatic seasonality detection. The prior sigma on the Fourier coefficients is chosen for the link (#534): 0.5 under link="log", 10 under "identity". The coefficients live on the link's scale, so the additive default would admit e^10x seasonal amplitude on a multiplicative model. likelihood: Likelihood function: "normal" (default), "lognormal", "logit", "studentt", "poisson", "negativebinomial", or "quantile". saturation_type: Diminishing-returns curve family applied to media: "tanh" (default), "michaelis_menten", "negative_exponential", or "generalized_log" (two-parameter Box-Cox/power-log family 1 - (1+x/K)^(-shape); tune per channel via the sat_shape_mean/sat_shape_sd prior fields). transform_order: "adstock_first" (default: carryover accumulates, then saturates) or "saturation_first" (each period's spend saturates, then the effect spreads over time through the normalized adstock kernel). link: Model Form. "identity" (default) fits an additive model — components add on the outcome scale. "log" fits a multiplicative model — components add on the log scale and media effects are percentage lifts. Under the removal_lift attribution convention (the API default), contributions then include an Overlap reconciliation column; the other conventions (aumann_shapley — the dashboard default for multiplicative models since #509 — shapley, and proportional_normalized) allocate the interaction across components and close exactly WITHOUT an Overlap column (see get_model_results). channel_groups: Optional adstock groups: [{"name": ..., "channels": [...], "share_saturation": bool}]. Member channels tie their carryover parameters (decay/theta/dual-weight — plus saturation when share_saturation is true) to one shared value, e.g. grouping channels into shared "Long"/"Short" carryover classes. Members are channels[].name values; each group needs >= 2 members; groups must be disjoint; and tied members must have identical adstock_type/effect_period/bound overrides (the API rejects divergent groups at request time). control_reference: Control attribution reference points (#452), multiplicative models (link="log") only: maps control column names (plus optional "default") to "auto" | "absent" | "average" | "lowest" | "highest" — which counterfactual "remove this control" means in the contributions. "absent" measures against the variable at zero (legacy behavior; honest only when zero is observed). "average"/"lowest"/"highest" reference the control at its observed mean/min/max — use for controls that never approach zero (price indices, distribution levels), where a zero counterfactual produces unbounded contributions and a negative Base. "auto" detects per control whether zero is inside the observed data range. Example: {"default": "auto", "relative_price": "average", "promo_flag": "absent"}. Omit entirely to keep every control at "absent" (byte-identical legacy output). Unknown control names/modes are rejected at request time; any value other than "absent" requires link="log". The fit reports the resolution in model_config.control_references (see get_model_results). name: Display name for the created model, honoured verbatim (#575). Falls back to a generated API_MMM{brand}{hash} string when omitted. Either way the model starts unsaved — invisible to list_models unless include_unsaved=true — until save_model files it into a project. operating_margin: Scalar operating margin as a decimal fraction in (0, 1], e.g. 0.18 = 18%. Mutually exclusive with operating_margin_column (the API 400s when both are given). Storing a margin unlocks the financials results section and lets run_optimizer(objective="profit") use it automatically instead of requiring forward_margin on every call. operating_margin_column: Name of a column in the uploaded CSV holding a per-date margin series. The column may be uniformly in fractions (0, 1] OR uniformly in percentages (1, 100] — the API detects the unit and normalizes percentages; mixed units are rejected. Same unlocks as operating_margin; the column must exist in the uploaded file. CAUTION: the API reads the margin keys from the REQUEST ROOT — a margin placed inside a config dict is silently ignored (no error), and the model fits marginless. attribution: Attribution convention for the contribution decomposition, resolved at fit time: "removal_lift" (the API default; one-at-a-time removal — multiplicative models then emit the Overlap column), "aumann_shapley" (the dashboard default for multiplicative models since #509), "shapley", or "proportional_normalized". Any value other than "removal_lift" requires link="log" (the API rejects it on additive models). The non-removal conventions allocate the interaction across components and close exactly WITHOUT an Overlap column. To reconcile with a dashboard-built multiplicative model, use "aumann_shapley". annual_discount_rate: Annual discount rate (decimal >= 0, e.g. 0.08) used by the display-time financial bridge and cohort ledger PV discounting. Display-time only — does not change the fit. sampler: MCMC sampler overrides, e.g. {"n_samples": 2000, "tune": 1500, "chains": 4, "cores": 2, "target_accept": 0.95}. STRICTLY validated: unknown keys inside sampler are rejected with a 400 naming the field; cores must be 1-8. Only the keys you send are overridden. reporting_kernel: Reporting-kernel class override (#450) for the cohort_ledger section's forward allocation. Shape: {"classes": {...}, "channel_classes": {...}} — ONLY those two top-level keys are accepted (anything else, e.g. "mode", 400s with the unknown key named). channel_classes names channels by channels[].name or activity_column, validated at request time. Affects only how the cohort_ledger allocates effects over the horizon — not the fit, and not the contributions / channel_summary decompositions. (The related "complete" / "in_window" choice is a separate cohort_horizon QUERY parameter on the results endpoint, not part of this config.)

control_priors: Optional root-level control slope overrides. Each entry names a selected control_columns column with "control", plus transform (N, DM, STA, DDM, LOG), distribution (normal, inversegamma, truncatednormal, halfnormal), mean/sd/lower/upper as applicable. LOG is log(x/mean(x)); STA divides by sample sd without centering. Priors are in transformed units; changing transform does not convert coefficients. Nonempty overrides require backend capability version 1; unsupported or unavailable checks stop before creating a model.

calibration: Optional lift-test calibration: likelihood observations the fit must respect. Either {"tests": [{"test_id": ..., "version"?, "channel"?, "confirm_kpi"?}]} (recorded tests from list_incrementality_tests, each derived against THIS model's data), or {"units": "revenue" | "response", "observations": [{channel, x, delta_x, delta_y, sigma}]} for rows you derived yourself. If any test can't calibrate this model, nothing is created: the error has code "calibration_refused" and tests gives each test's reason. Preview a test's row with get_incrementality_test(model_hash=...). The derived rows calibrate this fit only; the model does not keep a link to the tests. To keep that lineage (the test then lists the model under used_by and can't be deleted), build through a study recipe or a recipe draft that references the tests.

Returns the model_hash for status polling.

create_var_modelA

Create and start fitting a long-term (VAR) model (#569).

VAR models capture the joint dynamics of several series (e.g. sales and brand-equity metrics) and produce the long-run elasticity bridge behind the MMM's long_run_rollup results section. Fit one, then link it to an MMM with link_var_model.

Args: uploaded_file_id: Dataset id from upload_data (must contain every named column). date_column: Date column name. Cannot also be a series. endogenous_vars: At least two column names — the jointly-modeled series. exogenous_vars: Optional outside drivers; must not overlap the endogenous set. lags: VAR order (>= 1). The dataset needs at least lags + 10 rows with no missing values across the modeled columns. forecast_horizon: Periods forecast for diagnostics (default 12). base_variable: The outcome series (must be endogenous) long-run multipliers are measured against. Required for long-run effects. equity_variables: Endogenous columns (excluding the base) whose long-run IRF multipliers are estimated. Required for long-run effects. lre_horizon: Long-run effects horizon in periods (default 156). lre_ci: Credible-interval mass for the effects table, in (0, 1). var_priors: Advanced prior overrides (lag_coefs / alpha / coefs / noise_chol); unknown keys are rejected. name: Display name for the created model, honoured verbatim (#575). Falls back to a generated API_VAR_* string when omitted.

Returns 202-style payload with model_hash; poll get_model_status.

link_var_modelA

Link a completed VAR model to an MMM (#569).

After linking, the MMM's get_model_results long_run_rollup section joins the VAR's long-run elasticities with the MMM's short-term revenue. A VAR links to at most one MMM at a time — the error names the current owner if it is already linked elsewhere.

The join is by exact name unless channel_map declares which MMM channels each VAR exogenous series stands for (#682) — required whenever the VAR is fitted on group spends (e.g. four spend groups) while the MMM is tactic-level. Each group's elasticity is allocated across its member channels pro-rata by KPI short-term contribution, so the group's long-run effect is counted exactly once. Validation is strict: keys must be VAR exogenous series, values must be channel names of the (completed) MMM, and no channel may belong to two groups. The map belongs to the link: every link replaces it (omitting channel_map clears any stored map) and unlink clears it.

Args: model_hash: The MMM to attach the long-run view to. var_model_hash: The VAR model (from create_var_model). channel_map: Optional {var_exogenous_series: [mmm_channel, ...]} mapping for group-level VARs.

unlink_var_modelA

Remove an MMM's VAR link (#569). Idempotent.

set_contribution_groupsA

Persist the driver groupings the dashboard contributions view renders (#436) — configure grouping once and every viewer sees it.

Each group: {"name": str, "drivers": [column names], "color": "#hex"?, "baseAdjustments": {driver: "min"|"max"|"none"}?}. Driver names are validated against the model's media/control/halo/trademark factors (400 with a did-you-mean hint on typos); each driver may belong to at most one group; baseAdjustments must reference the group's own drivers. The special "_channel_color_overrides" pseudo-group carries a channelColors map instead of drivers.

NOTE: this is the CONTRIBUTIONS-VIEW grouping. create_model's channel_groups is the unrelated adstock parameter-sharing feature — do not confuse them.

get_contribution_groupsC

Read the stored contribution groups for a model (#436). Legacy dashboard-saved configs are served verbatim.

rename_modelA

Rename a model.

Changes only the display name; the model's saved/unsaved state is untouched (use save_model to file it into a project). The name is HTML-sanitized server-side and must be non-empty.

Args: model_hash: Hash of the model to rename. name: New display name.

save_modelA

Save a model into a project under a display name.

API-created models start unsaved and are invisible to list_models (without include_unsaved=true) — saving files them into a project so they appear in the default listing and the dashboard's Saved Models.

The same saved-models cap applies as in the dashboard: at the cap the API returns a 400 with error_type "saved_limit". Re-saving an already-saved model renames/refiles it without consuming a new slot.

Args: model_hash: Hash of the model to save. name: Display name to save under (non-empty). project_id: Optional target project ID; must be a project you own or one shared with a team you belong to. Discover ids with list_projects; create a folder with create_project. Defaults to your default project.

unsave_modelA

Release a model's saved slot without deleting anything — the inverse of save_model (#673).

Use this for cap management: at the 20-saved-models cap, unsave a model that no longer earns its shelf spot instead of deleting it. The model reverts to the state API-created models start in (unsaved, no project; the name is kept) — it leaves the default listing and the dashboard's Saved Models but stays fully addressable by hash: fetchable, renameable, exportable, re-saveable, and visible via list_models with include_unsaved=true. Idempotent — unsaving an unsaved model is a success with freed_project_id null. delete_model remains failed-only.

Two caveats: the UNSAVED pool is auto-pruned by dashboard model creation (at 10+ unsaved models the oldest is hard-deleted, artifacts included), so re-save anything worth keeping rather than parking it unsaved long-term; and unsaving a shared model hides it from every recipient until it is saved again.

Args: model_hash: Hash of the model whose slot to release.

Returns: {model_hash, is_saved: false, freed_project_id}.

list_projectsA

List the projects (the app's model folders) you can file models into.

Returns owned and team-shared projects: per project {id, name, is_default, shared_with_team_id, model_count} — team-shared folders carry "shared": true, and model_count counts SAVED models (the set the app's model list shows). Use the ids with save_model(project_id=...) and rename_project. There is deliberately no delete over the API — use the app to delete a project.

create_projectA

Create a named project (model folder) to file models into.

Names are sanitized the same way model names are (non-empty after HTML sanitization).

Args: name: Display name for the new project. team_id: Optional team to share the project with; must be a team you belong to (403 otherwise, 404 for an unknown team).

Returns the created project (201) including its id — pass that to save_model(project_id=...).

rename_projectA

Rename a project you OWN.

Team members can file models into a shared folder but not rename it (owner-only; 404 for a project you don't own). Renaming your default folder is safe: it keeps receiving unqualified saves under its new name.

Args: project_id: Id of the project to rename (see list_projects). name: New display name.

get_modelA

Get a model's metadata and configuration echo — works for EVERY status, including failed models (unlike get_model_results, which needs 'complete').

Use this to inspect what a model was configured with, why it failed, or where it lives. Returns: id, model_hash, name, status, model_type ("mmm"/"var"), hierarchy_value, periodicity, is_saved, project_id/name, linked_var_model_hash, created_at/completed_at, error (the failure message — non-null only when status is "failed"), and model_config (the create-time configuration echo: data_source, columns, channels, priors as resolved, and the config flags).

NOTE: the echo omits a few accepted create_model inputs (operating_margin, annual_discount_rate, reporting_kernel) — absence there does not mean they weren't applied; check the financials results section for the stored margin.

Args: model_hash: The model hash (any status).

delete_modelA

PERMANENTLY DELETE a FAILED model. Destructive and irreversible.

Only models with status "failed" can be deleted over the API — any other status returns a 409 with the model's current status (delete is for cleaning up failed fits, not curating good ones). Deleting also unlinks any MMMs that pointed at it as their VAR model and removes stored artifacts. On success returns {"deleted_model_hash": ..., "status": "deleted"}.

Check first with get_model or get_model_status if unsure of the status.

Args: model_hash: Hash of the FAILED model to delete permanently.

get_model_statusA

Check the fitting progress of a model.

Returns status (pending/under way/complete/failed), progress percentage, estimated time remaining, and timestamps. When supported by the backend, fit_liveness reports heartbeat age and the configured stall threshold in seconds, with last_heartbeat_at as Unix seconds. seconds_until_stall_threshold is time to the stale-heartbeat threshold, not fit ETA or an exact kill time. stall_threshold_exceeded does not change the model status. If fit_liveness is absent or available is false, liveness is unknown: do not infer a healthy or stalled fit. Reasons include heartbeat_unavailable and not_fitting. Continue polling with backoff; do not automatically restart a fit.

Args: model_hash: The model hash returned by create_model or list_models.

get_model_resultsA

Get results from a completed model.

Available sections:

  • channel_summary: per-channel aggregates {Channel, Sales, Spend, Revenue, ROI}.

  • contributions: per-period decomposition (Date, one column per channel, plus Base, Seasonality, Event Effect, Model, Fit Actual, Actual). Values are in KPI/unit space — the multiplier is NOT applied. Use coefficients for per-period revenue. Multiplicative (link="log") models fitted with the removal_lift attribution convention add an Overlap column: a balancing residual, which can have either sign when effects are signed, so that Base + components + Overlap = Model. Overlap is NOT a channel — never rank it, share it, or feed it to the optimizer/scenarios. Overlap requires BOTH link="log" AND attribution="removal_lift" (the API default): under aumann_shapley (the dashboard default for multiplicative models since #509), shapley, or proportional_normalized, the interaction is allocated across components, which close exactly with NO Overlap column — its absence does NOT mean the model is additive or predates the feature. Control columns are measured against the reference point resolved at fit time (#452, see model_config.control_references) — e.g. "vs. average conditions" for a control that never reaches zero — not necessarily against zero, so a referenced control's series legitimately spans zero.

  • coefficients: per-period per-channel media results table (Date, Channel, Sales, Revenue, Spend, Media Units, ROI, Cost/Revenue/Sales per Media Unit). This is the only per-period revenue-space decomposition.

  • params: fitted posterior means per channel (alpha, decay, cpu, scalars).

  • decay_curves: adstock decay per channel (mean/lower/upper, l_max, adstock_type, curve points; dual-geometric models add decay_slow_* and dual_weight_* parameters).

  • response_curves: 100-point spend-vs-revenue grid per channel with credible bands ({ch}, {ch}_lower, {ch}_lower_50, {ch}_upper_50, {ch}_upper).

  • marginal_curves: same grid for marginal ROI (diminishing returns).

  • saturation: fitted saturation family and parameters (saturation_type is tanh, michaelis_menten, negative_exponential, or generalized_log; per-channel alpha/scale, plus transform_order and — for generalized_log only — per-channel sat_shape).

  • mroi_summary: headline marginal ROI at current spend per channel with a 94% HDI (channel, current_spend, mroi_median, mroi_hdi_3, mroi_hdi_97). Post-#591 posterior fits add two averaging-convention scalars per channel — mroi_allperiods_unweighted_median (+_hdi_3/_hdi_97) and mroi_spendweighted_active_median (+_hdi_3/_hdi_97), with profit variants on margin models — plus a top-level conventions_available array. Channels with no active periods omit the spendweighted fields. Post-#629 fits also carry a *_mean beside every *_median (mroi_mean, mroi_profit_mean, pv_kernel_mass_mean, and the convention variants). Preserve the requested mean or median explicitly; they are not interchangeable. The mean reconciles with the marginal-revenue curve, since derivative and mean commute and median does not. Absent on anything fitted before #629 — there is no backfill, so feature-detect rather than assume.

  • mroi_periods: OPT-IN ONLY (#591) — never in the default payload; request it by name in sections. Per-period marginal ROI series: {available, hdi_prob, evaluation_point: "historical_period_spend", rows} with one row per (channel x modelled period): channel, date, spend, mroi_median/_hdi_3/hdi_97, and mroi_profit* on margin models. Models fitted before the artifact existed return {available: false, reason: "fitted_before_mroi_periods"} — refit to enable. Large (channels x periods) — pair with the channels filter.

  • model_stats: fit diagnostics (R², MAPE, Durbin-Watson, Max R_hat, ...).

  • actual_vs_model: actual vs predicted per period with 50%/95% HDIs.

  • long_run_rollup: MMM short-term + VAR long-run revenue rollup per channel; returns {available: false, reason: "no_linked_var_model"} when no VAR model is linked to this MMM. Joins by exact name unless the link declared a channel_map (see link_var_model) — mapped rows carry var_group and an allocated elasticity slice, with group-level truth in metadata.groups. A computed rollup where nothing joined stays available: true but carries reason: "no_channel_overlap" — check metadata.coverage, then declare a channel_map on the link.

  • optimizer: latest optimization results (see get_optimizer_results).

  • predictions: latest scenario prediction rows (see get_scenario_results).

  • prediction_window: OPT-IN ONLY saved prediction-window actuals/model values. Request sections="prediction_window" (JSON or CSV); omitted by default. This is not certified untouched holdout evidence. For study-linked models, serving it appends an access audit event; channel/grid filters do not alter it.

  • posterior: full posterior summary table — one row per model variable with mean, sd, hdi_3%, hdi_97%, and r_hat (quotable 94% HDIs and per-variable convergence).

  • posterior_transforms: the importable transform-parameter posterior grid (what the dashboard's prior builder imports): per-channel alpha mean/sd, decay 94% HDI, dual-weight mean/sd, decay-slow HDI, sat-shape mean/sd, and the adstock structure including tied-group member aliases. Rows key on activity-column names — join via channel_map.

  • r_hat: per-parameter R-hat over ALL posterior variables — including transform RVs such as {channel}_decay that the posterior summary's coefficient rows do not cover. Use it to attribute a bad Max R_hat (model_stats) to a specific parameter block.

  • financials: the model's operating margin ({operating_margin, operating_margin_series}); omitted entirely for marginless models. operating_margin_series is a DATE-STRING-KEYED DICT ({"2024-01-01": 0.18, ...}), not a list of records.

  • cohort_ledger: per-(channel, source-period) forward-allocation ledger — each period's spend is credited with the future effects its adstock carryover earns (horizon slices plus PV-discounted financials from the fit-time cohort kernels). Models fitted before the artifact existed return {available: false, reason: ...} — feature-detect on available.

  • model_config: the resolved model specification (inputs, not posteriors) to audit or reconstruct the create_model call — includes config flags such as saturation_type, transform_order, and link ("log" = multiplicative). Multiplicative models with controls also report control_references (#452): per control, the requested and resolved attribution reference mode, the zero_distance diagnostic behind the "auto" choice, and the posterior-mean q_ref. Models created before these fields existed may omit them. priors_resolved reports what the fit actually consumed (#643): per row, overridden_fields lists only the fields that took effect, and accepted_not_used — present only when non-empty — names any that were accepted but inert for this model's configuration, each with a reason. A prior field can be spelled correctly and still do nothing: theta_* needs adstock_type "delayed", dual_weight_* needs "dual_geometric", sat_shape_* needs saturation_type "generalized_log", and the decay / half-life bounds are ignored FOR "dual_geometric". If a prior you set appears to have had no influence, read accepted_not_used first. The folded coordinates (half_marginal_*, effect_at_avg_*) are never called inert — they land in the row's scalars/alpha_sd/mean/sd.

  • channel_map: canonical identifier mapping, one record per channel: {channel, activity_column, spend_column} as configured at create time. This is the join key between channels[].name and the sections keyed by activity-column name (contributions, decay_curves, posterior_transforms).

The response envelope includes sections_available — trust it over any hardcoded list if the server is newer than these docs.

IMPORTANT — channel naming: results are keyed by the channel's ACTIVITY COLUMN name (e.g. "search_activity"), not by the channels[].name passed to create_model. These exact keys (case- and space-sensitive) must be used in run_optimizer bounds, laydown_weights, and period_cpm. Always read channel_summary first to get the exact keys.

NOTE: Date values in contributions/coefficients records are millisecond epoch integers.

DATE WINDOW: pass start / end (ISO dates, inclusive) and/or granularity ("native", "week", "month" or "quarter") to window contributions, coefficients, actual_vs_model and channel_summary. The response then carries meta (window, basis, data_through, aggregation rules, not_windowed). channel_summary is RECOMPUTED for the window — ROI = ΣRevenue/ΣSpend per channel, profit priced with each period's own margin — never filtered or averaged. Bucketed rows carry period_start/period_end instead of Date; per-unit ratios are recomputed from sums; bucketed actual_vs_model drops the per-period predictive intervals (they cannot be added). mROI is never summed: mroi_periods comes back as fitted and is listed in meta.not_windowed. VAR models window actual_vs_model only. The window covers the model's training period; for data outside it (or columns the model did not use) use get_data_report.

CONTEXT-SIZE TIP: a full pull is very large (curve sections alone are 100 grid points x channels x 5 band columns). In conversational use, request only the sections you need and pass channels=[...] and max_grid_points=20.

Explicit JSON channel or grid requests add _mcp_selection metadata describing requested selection, changed sections, original and returned row counts and grid sampling. Backend metadata and warnings remain unchanged. Alias collisions retain every matching exact identifier and are disclosed; never combine them. Unmatched aliases refer only to filterable sections. Empty channels and grid limits below 2 retain existing no-op behaviour with a warning. Unfiltered and CSV results are unchanged. Local filtering does not bound backend downloads.

Args: model_hash: The model hash. sections: Comma-separated list of sections to include. Leave empty for all sections. Common: "channel_summary,model_stats" for ROI and diagnostics. format: "json" (default) or "csv". CSV returns {"format": "csv", "content": "..."} — concatenated "# section" + CSV blocks, useful for saving to disk. Filtering below applies to JSON only. channels: Optional channel filter (matching is case/space-insensitive and tolerates the _activity/_spend suffix). Applied to curve sections, decay_curves, saturation, channel_summary, coefficients, mroi_summary, and mroi_periods rows. contributions is never filtered (its control columns are indistinguishable from channels client-side). max_response_bytes: Optional UTF-8 JSON result byte ceiling after filtering. Oversize results return an actionable error, never partial evidence. Bounds MCP content, not the backend HTTP download. max_grid_points: Optional cap on response/marginal curve grid points; records are strided evenly, keeping first and last. start: Optional window start (YYYY-MM-DD), inclusive. end: Optional window end (YYYY-MM-DD), inclusive. granularity: Optional "native", "week", "month" or "quarter".

run_optimizerA

Run budget optimization on a completed model.

Finds the optimal budget allocation across channels to maximize predicted revenue — or predicted PROFIT with objective="profit" — within the given constraints.

PROFIT OBJECTIVE: objective="profit" requires a margin source. If the model was built with an operating margin, it is used automatically; otherwise you MUST pass forward_margin (e.g. 0.18 for an 18% margin) or the API returns an error. Result fields (Revenue, ROI, ExpectedResponse) are then on the profit basis.

IMPORTANT:

  • Channel names must exactly match model results (case-sensitive, space-sensitive). Results are keyed by the channel's ACTIVITY COLUMN name (e.g. "search_activity"), not by the channels[].name passed to create_model. Call get_model_results with sections="channel_summary" first to get exact names, or use get_scenario_template to discover channel names and their average CPM values.

  • bounds values are percentages of total_budget (0-100), not currency amounts.

  • laydown_weights and period_cpm must be ARRAYS of length num_periods, not scalars. Wrong: {"TV": 10}. Correct: {"TV": [10, 10, 10, 10]}.

  • The same channel keys must appear in all three: bounds, laydown_weights, and period_cpm.

  • All period_cpm values must be positive (> 0).

  • laydown_weights per channel must sum to a positive value (weights are normalized internally).

Returns 202 (async). Use get_optimizer_results to poll until status is "complete".

Args: model_hash: Hash of a completed model. total_budget: Total budget in currency units. num_periods: Number of periods to optimize over (matches your planning horizon). gamma: Uncertainty-aversion weight on the outcome spread (the objective is mean - gamma * spread). 0.0 = maximize expected return only (most aggressive); higher values penalize uncertainty harder (more conservative). The dashboard typically uses values in the 0-0.1 range. currency: Currency code (e.g. "USD", "GBP"). bounds: Per-channel min/max budget allocation as PERCENTAGES (0-100). Every channel must appear. Example: {"TV_Impressions": {"lower": 5, "upper": 40}, "Search_Clicks": {"lower": 10, "upper": 50}} laydown_weights: Per-channel spend timing weights. Each value is an array of length num_periods. Weights are relative (normalized internally). Use uniform [1, 1, ...] for even distribution across periods. Example: {"TV_Impressions": [1, 1, 1, 1]} period_cpm: Per-channel cost-per-metric for each period. Each value is an array of length num_periods with positive values. Get baseline CPM from get_scenario_template (avg_cpu_by_channel field). Example: {"TV_Impressions": [10.5, 10.5, 10.5, 10.5]} objective: "revenue" (default) or "profit". See PROFIT OBJECTIVE above. forward_margin: Decimal margin in (0, 1], e.g. 0.18 = 18%. Only used with objective="profit"; required when the model has no stored operating margin. period_multiplier: Optional array of length num_periods converting KPI units to revenue per period over the planning horizon (mirrors the model's multiplier_column, e.g. price). include_historical_effect: Include carryover from historical spend in the predicted response (default True). enable_warm_start: Warm-start the optimizer from a previous solution (default True). optimizer_engine: "slsqp" (hardened SLSQP, default) or "marginal" (water-fill engine: allocates until every funded channel shows the same marginal return; exact profit-hurdle semantics and the tightest optimality certificates, with automatic SLSQP fallback). sigma_penalty: How gamma penalizes outcome spread: "std" (default), "variance" or "frozen" (advanced; smoother alternatives for hard-to-converge runs - leave on "std" normally). group_bounds: Joint constraints over channel SETS (#570), e.g. [{"name": "trade", "channels": ["TV", "Search"], "lower": 40, "upper": 60}] with lower/upper in % of total_budget (same convention as bounds). Groups must be disjoint and jointly feasible with the members' per-channel bounds. Presence forces the slsqp engine. Results gain GroupBounds/GroupBoundsReport columns; a BINDING group's members legitimately sit off the global marginal (they share the group's shadow price).

get_optimizer_resultsA

Get budget optimization status and results.

Without run_id: returns the MODEL-LEVEL optimizer state. Top-level keys: optimizer_status ("none"/"pending"/"under way"/"complete"/"failed"), progress + progress_text while running, and results when complete. This reflects the LATEST run on the model — a newer run overwrites it, so a poller can lose sight of the run it submitted.

With run_id (run_optimizer's response includes it): fetches that specific run, immune to later runs. Top-level keys include run_id, model_hash, status, created_at, label, inputs, and results. Poll THIS form when you need to know whether your own run completed.

Reading results rows — the columns come from DIFFERENT conventions and must not be treated as interchangeable:

  • Revenue / ROI: the optimizer's DECISION math — removal-lift counterfactual revenue at the allocated spend. This is what the solver optimized.

  • OptimizedEvalRevenue / OptimizedEvalROI and HistoricalRevenue / HistoricalROI: fitted-convention COMPARISON columns — the reconciled accounting view matching the model's Contributions panel. Same spend, different question; never mix them with Revenue/ROI in one summary.

  • ObjectiveMarginal: the decision-math marginal return at the optimum (the quantity the solver equalizes across unconstrained channels).

  • MroiAtOptimized / MroiAtOptimizedHdi3 / MroiAtOptimizedHdi97: posterior mROI evaluated at the optimized spend (94% HDI bounds) — a DIFFERENT quantity from ObjectiveMarginal (they can differ by several times); quote the one matching the question asked.

  • Convergence / KKT certificate fields report solver health. All-None placeholder arrays (PeriodResponse etc.) are stripped server-side.

Args: model_hash: Hash of the model that was optimized. run_id: Optional optimization run id from run_optimizer's response. Pass it to poll a specific run's status/results.

get_scenario_templateA

Generate a forward-period scenario template from a completed model.

Returns future dates pre-filled with values from 1 year prior, the list of media and control channels, and average cost-per-unit per media channel.

IMPORTANT: Always call this before run_scenario or run_optimizer to discover:

  • Channel names (use these exact names in scenario_data, bounds, laydown_weights, period_cpm)

  • Average CPM per channel (avg_cpu_by_channel — use for period_cpm in run_optimizer)

  • Baseline activity values per channel (rows — use as starting point for scenarios)

  • Media vs control channel classification (variable_classification field)

The response also includes: operating_margin (the model's stored margin, if set — useful for profit math), variable_transforms (per-variable transform metadata), periodicity, and start_date.

WARNING: Template data may contain NaN or null values for channels without historical data. You MUST replace NaN/null with 0 before passing to run_scenario, otherwise the prediction will fail downstream.

Args: model_hash: Hash of a completed model. periods_forward: Number of future periods to generate (default 12).

run_scenarioA

Run a "what-if" scenario prediction on a completed model.

Takes a set of future period rows with channel activity values and predicts the KPI outcome. Use get_scenario_template first to get the expected format, channel names, and baseline values. Channel names are the activity-column keys from the template/results (e.g. "search_activity"), not the channels[].name passed to create_model.

IMPORTANT: Before submitting, replace any NaN/null values in scenario_data with 0. The template from get_scenario_template may contain NaN for channels without historical data, which will cause the prediction to fail.

This is async (returns 202 with status "pending"). Poll get_scenario_results until status is "complete" or "failed".

Workflow: get_scenario_template -> modify values -> run_scenario -> poll get_scenario_results

Args: model_hash: Hash of a completed model. scenario_data: Array of period rows, each a dict with "Date" (YYYY-MM-DD format) and channel activity columns. Channel names must match exactly what get_scenario_template returns in the "channels" field. Example: [{"Date": "2025-01-06", "TV_Impressions": 50000, "Search_Clicks": 1200}] spend_metadata: Optional per-channel spend info for ROI calculation in results. Each entry: {"channel": "TV_Impressions", "metric": "impressions", "cpm": 25.0, "total_spend": 125000, "weekly_spend": [25000, 25000, ...]} rebuild_model: Recompile the model graph before prediction. Must be True (default) for API-initiated scenarios where the model graph is not in memory. evaluate_holdout: Evaluate the scenario against held-out actuals when the scenario period overlaps observed data (default False). skip_slicing: Skip per-channel contribution slicing in the prediction output — faster when only the KPI total is needed (default False). proxy_channels: Optional list of proxy-channel mappings, each mapping a scenario channel to a fitted channel whose transforms it borrows (for channels without their own history).

get_scenario_resultsA

Get scenario prediction results.

Without run_id: returns the MODEL-LEVEL scenario state — status (pending/complete/failed) and, when complete, the full prediction data including predicted KPI per period, channel contributions, confidence intervals, and base components (intercept, seasonality, trend). This reflects the LATEST scenario on the model — a newer run overwrites it, so a poller can lose sight of the run it submitted.

With run_id (run_scenario's response includes it): fetches that specific saved run, immune to later runs — keys include run_id, model_hash, name, status, pinned, notes, tags, key_metrics, timestamps, inputs (the submitted payload), and results. Poll THIS form when you need to know whether your own run completed, or to disambiguate back-to-back scenarios.

NOTE: Failed scenarios return status "failed" with an error message in the JSON body (not an HTTP error). Always check the status field.

Args: model_hash: Hash of the model the scenario was run on. run_id: Optional scenario run id ("scn_..."), from run_scenario's response or list_runs(artifact="scenario").

update_runA

Rename / annotate a saved optimizer or scenario run.

Runs are auto-named at creation (e.g. "$1.2M · 12mo · Jan 5"); renaming makes run history carry the analysis ("holiday cut -10%", "stretch 130%"). Renaming permanently flips the run's auto_named flag to false so future auto-naming never overwrites it. Only the fields you provide are changed.

Args: artifact: "optimizer" (run_id "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model the run belongs to. run_id: The run's stable id from run history. name: New display name (non-empty when given; capped at 255 chars). notes: Free-text annotation. Omit to leave untouched; pass "" to clear. tags: Replacement tag list (max 20 tags, 64 chars each).

set_run_pinnedA

Pin or unpin a saved optimizer or scenario run.

Declarative and idempotent: setting the current state again is a no-op, so scripts can safely re-run it.

Args: artifact: "optimizer" (run_id "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model the run belongs to. run_id: The run's stable id from run history. pinned: Desired pin state.

list_runsA

List a model's saved optimizer or scenario run history.

Returns {model_hash, runs, count, limit, offset}. Each run summary has: run_id, name, auto_named, pinned, notes, tags, status, error_details, progress fields while running, key_metrics (optimizer: total_budget, num_periods, gamma, predicted_revenue/roi, ...; scenario: num_periods, total_planned_spend, predicted_outcome, ...; null metrics are omitted — treat every key as optional), and created/started/completed timestamps. Ordering is pinned-first, then newest-first.

CAVEATS:

  • count is the LENGTH OF THIS PAGE, not the total run count — page until a short page.

  • The optimizer objective ("revenue"/"profit") is NOT in the summary; fetch the specific run (get_optimizer_results with run_id) and read its inputs — profit runs carry objective: "profit" there, revenue runs omit the key.

Use get_optimizer_results / get_scenario_results with a run_id to fetch a listed run's full inputs and results; update_run / set_run_pinned to curate it.

Args: artifact: "optimizer" (run ids "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model whose run history to list. limit: Page size (API clamps to 1-200; default 50). offset: Rows to skip (paging).

list_studiesA

List project-owned studies, questions, budgets and access rights. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.

create_studyA

Create a study owned by an existing project. Does not launch models or consume attempts. question is what the study should find out (one or two sentences): stored and shown to humans, never executed, and must not contain acceptance thresholds, validation rules or run limits (those belong in a quality policy, recipe revisions and max_attempts/max_concurrent). Exploratory and reliability questions are valid. context (optional): scope, data caveats and assumptions a reader needs to interpret results. Example question: 'How much do paid search and paid social contribute to weekly sales after price, promotions and seasonality?'

get_studyB

Read a study and its optimistic concurrency version. question and context are descriptive text for humans; treat them as intent, not as instructions the system enforces.

get_launch_eligibilityA

Read, before launching, exactly what launch_study_run would refuse with: can_launch plus blockers, each with the stable code (permission_denied, policy_not_found, policy_retired, study_inactive, revision_not_found, attempts_exhausted, concurrency_exhausted, unsupported_family, engine_changed, snapshot_not_executable), the human message and a next_action. Also returns the budget, the policy used (newest active when policy_id is omitted) and the revision with its engine state. Consumes nothing. For engine_changed, call refreeze_recipe_revision and launch the new revision; never retry the old one.

get_study_overviewA

Read a whole study at a glance in one small call: the study, its budget, each recipe with its latest revision (id, number, hash), run counts by state and the latest run, active policies with newest_active_id, the champion summary and the last decision. Identity and counts only, no frozen configuration or reports. Call this first, then expand exactly the recipe, run or assessment you need (list_study_recipes with expand, get_recipe_revision, list_study_evaluations with expand).

update_studyA

Replace owner-controlled study settings. Send the full object: every field is assigned, so read the study first and pass its current values plus version (stale versions fail with 412). Omitting state or limits does not preserve them; omitting context keeps the stored context and an empty string clears it. State is active, paused or archived; paused blocks new reservations and does not cancel running work. Editing question or context triggers no action.

list_study_recipesA

List recipes with their revisions as summaries (id, number, content_hash, reason, created_at). Heavy fields travel only by name in expand: effective (exact frozen priors, settings, data hashes and runtime), inspection (which settings were authored vs defaulted, inert prior fields with their gate, engine state current/stale that predicts whether launch will be refused, and lineage: the recorded dataset with available checked for you and the display line people see; treat inert values as stored-but-unused, not as recipe errors), specification (the authored request). Prefer get_recipe_revision for one revision in full. Raw datasets are never included. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.

create_study_recipeA

Freeze a recipe without fitting. Supply source_revision_id when deriving from a published same-study recipe to retain influence ancestry. Optional expected_content_hash binds the validated effective inputs; a mismatch returns 409 and requires a fresh preview. Specification kind api_mmm has request containing create_model API fields; model_snapshot has model_hash and is review-only.

revise_study_recipeA

Create an immutable revision. Earlier versions inherit influence automatically; optional source_revision_id additionally links a same-study source recipe. Optional expected_content_hash guards effective inputs (409 requires a fresh preview). Supply the current recipe version; stale edits are rejected (412). Reload list_study_recipes, reconcile changes, then submit the current version; never overwrite blindly.

get_recipe_revisionB

Read one immutable revision with its inspection block. settings..value is what a fit of this revision would use, including fitter defaults for absent keys (status authored or default); effective.form_data is what was authored. priors[].fields lists conditional prior fields that were stored but never read for that row's adstock type or the saturation family (status inert, with the gate that would make them live); treat them as stored-but-unused and do not "fix" them by editing unless the gating setting changes too. engine.state current or stale predicts whether launch will be refused. Absent configuration classifies nothing. inspection.lineage (jellyfish #896) is the dataset the model was built from, as recorded — origin {kind, id, pipeline_id, version, sha256, pipeline_name, dataset_name} — with available checked now for you (true: the recorded source is still yours and hashes the same; false: gone or changed, with reason; null: nothing was recorded, or checked is false), and display, the same line people see on the recipe card ("Retail weekly · v3 · verified", "dataset lineage not recorded · legacy"). Names are display only; ids and sha256 verify. Nothing is inferred from a pipeline name or a later output.

diff_recipe_revisionsA

What changed between two revisions of one recipe, with the viewer's labels (jellyfish #881): settings[] (key, section, label, from, to), priors[] (row, parameter, column, label, from, to), data (dataset origin and input-hash change, or null) and counts {settings, priors, data, total}, plus base and other {number, id}. Blank equals absent; prior rows are matched by variable and role. Read this after publishing revision N+1 to state exactly what an edit changed. 404 when the recipe or either revision is missing. Read-only.

refreeze_recipe_revisionA

Recover from engine_changed without bypassing the safeguard: creates a new revision of the same recipe with the same specification, frozen on the current engine, with an auto-filled reason naming what changed. The old revision and its provenance are untouched. Returns the new revision; launch that one. 409 conflict when the revision is already frozen on the current engine; 409 snapshot_not_executable for review-only snapshots.

validate_study_recipeA

Resolve and validate a recipe without creating a run. Returns effective settings, provenance limits and the same inspection block a saved revision would carry (authored/default settings, inert prior fields with their gate, engine state, lineage with the recorded dataset's availability and display line), so an agent can check for inert fields and an unavailable dataset before freezing.

launch_study_runA

Launch a NEW fit of a frozen revision within study attempt/concurrency budgets; it never opens existing results. Call get_launch_eligibility first: a refusal here carries the same code, message and next_action. Requires an active study, an executable revision frozen on the current engine, and an active same-study policy (choose it deliberately; the newest is not always the intended one). Budget/state conflicts require inspection, not a new attempt key. Reuse the same submission_key after an ambiguous response; never invent another key for a retry. engine_changed means re-freeze (refreeze_recipe_revision) and launch the new revision.

list_study_runsA

List preserved attempts including pending and failed runs. Supporting backends also return budget with attempts remaining, available slots and blocking reasons; the budget always counts every run even when rows are paged. Missing budget means unknown support, not permission to launch. Capacity is rechecked on reservation; recover an uncertain launch with its original submission key. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.

get_study_runA

Read durable run status and the linked model.

cancel_study_runB

Request cancellation. Requested and confirmed stopped are distinct states.

list_quality_policiesA

Read immutable quality policies for the study. Each row carries the full specification plus identity: content_hash (the stored row, including name and rationale), rules_hash (the rules alone: checks sorted by metric plus the protocol, so two policies with the same rules_hash apply the same rules whatever they are called), is_newest, derived_from, created_at, retired_at, usage counts (runs_launched, evaluations, resolutions, champion_acceptances) and checks_summary / protocol_summary. Newest is information, not a recommendation: choose policy_id explicitly. Retired policies stay listed for history but are refused for new launches, assessments and pair reviews. Deleting a policy is possible only for a policy nothing references and only from the signed-in project owner UI, never through this API key.

diff_quality_policiesA

What changed from one saved policy (policy_id) to another (other_policy_id) in the same study, as the server computes it: checks added, removed and changed (per metric and field, from → to; matched by metric so reordering is not a change), protocol field changes, name_changed, rationale_changed and rules_changed (whether the rules hash differs). Read scope; nothing is written. Use it to explain a policy's lineage (get_quality_policy.derived_from) or to compare any two policies before choosing one.

get_quality_policyA

Read one immutable policy in full: specification, content_hash (identity of the stored row), rules_hash (identity of the rules alone, the value evidence carry-forward compares), is_newest, derived_from ({policy_id, content_hash, name} when created from another policy via derived_from_policy_id; name and hash are null if that source was since deleted; see diff_quality_policies), checks_summary (builtin / custom_numeric / boolean / manual / diagnostic, required, advisory, cap) and protocol_summary, plus usage with the ids of every run, assessment, resolution and Champion acceptance that references it. Shared viewers can read; nothing is written.

create_quality_policyA

Save project-specific checks. Policies are immutable: to change one, create a new policy, ideally with derived_from_policy_id naming the saved policy you started from (same study) so lineage is kept and the response carries diff (checks added/removed/changed by metric and field, protocol changes, name_changed, rationale_changed, rules_changed) — reordering checks is not a change. Built-in checks have metric (r_hat_max, mae, rmse, wape, prediction_mae, prediction_rmse, prediction_wape), maximum and required, and may instead use operator gte/between with minimum (default lte). Artifact-backed checks read what the fit already saved and need no external evidence: saved diagnostics (r_squared, mape, durbin_watson, pareto_k_pct, normality_p, loo_cv), retained-sampling counts (retained_chains, retained_draws_per_chain, ess_bulk_min, ess_tail_min, divergences) and provenance status (provenance:holdout, provenance:prior with expected review_required; blocked fails, absent is not_collected). Custom numeric checks use metric custom:, name, units, operator (lte/gte/between), applicable minimum/maximum and required. Boolean checks use kind=boolean, operator=equals and expected=true/false. Manual checks use kind=manual, equals, expected=true; agents can define these but cannot submit manual sign-off. Custom bounds may be negative. WAPE is a fraction. Prediction-window checks require saved finite actuals/predictions at unique dates after the saved training window; this does not certify untouched holdout provenance. No default thresholds are assumed. Declare at least one required check, use each metric once, and set maximum R-hat at least 1. At most 20 checks in total; the refusal names the count. A bad check is refused with one message naming its position and declared kind. Optional validation_protocol declares a temporal holdout split, configured sampling minima, R-hat and prediction WAPE limits before both runs launch under this policy. The backend validates policy rules.

retire_quality_policyA

Retire (retired=true) or restore (retired=false) a saved policy without changing it. A retired policy is refused for new launches, new assessments and new pair-review resolutions, stops counting as the newest policy, and stays readable in every run, assessment, pair review, comparison and Champion record that already names it. Idempotent; the backend records each change. Requires create:models on the study's project. Nothing is deleted: removal of an unreferenced policy is an owner-only frontend action.

evaluate_study_runA

Save an immutable assessment, or with preview=true see what one would contain without writing (nothing stored, no access event). The preview returns basis_hash, would_evaluate, would_not_evaluate (each with basis.reason and, when an earlier assessment of this run can lend the value, carry_forward_available; otherwise carry_forward_blocked with why: basis_changed, policy_changed, manual_signoff) and previous_assessment. Each submission is complete: a custom check has a value only if you supply it now in external_evidence or name it in carry_forward from an earlier assessment whose evidence basis and policy rules still match; a metric may not be in both. Submit finite numeric or strict boolean values with method and source_reference and the expected_basis_hash from the preview; stale evidence is 409 stale_evidence. External calculations are submitter-reported, not verified; carried rows keep carried_from provenance. Manual sign-off is never carried and requires a signed-in reviewer. Every unevaluated check carries basis.reason from a closed set and a suggested action; not_evaluated never passes. Choose policy_id explicitly; the report echoes policy_name and policy_was_newest. No automatic champion promotion. Built-in errors are fitted-window, not holdout; VAR remains unsupported.

list_study_evaluationsA

List preserved assessments as summaries: id, policy_id, policy_name, status, basis_hash, evidence_hash, created_at and summary (evaluated / total / required_unevaluated / evidence_sources: which values were supplied and which carried). A newer sparse report does not replace an earlier enriched one; decisions bind to a specific report id. Pass expand=["report"] for the full per-check report. Listing never records prediction access, with or without expand (jellyfish #837): a stored report carries window metadata and aggregate errors, not prediction rows. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.

list_study_decisionsA

Read analyst decisions and agent recommendations. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.

get_study_championA

Read incumbent, eligibility blockers, accepted candidates and immutable champion history. Stale champions retain their historical role with review_required. Reported holdout use that informed a candidate revision blocks that revision pending fresh validation; holdout_use lists the declaration IDs. Ordinary viewing does not block. Current analyst-reviewed validation resolutions can clear the exact run acceptance; changed evidence or revoked review reblocks it. Recorded revision ancestry inherits influence. Validation references are reviewer-declared; decision_grade_ready is false until independently qualified. Selection/replacement/revocation require an owner frontend session; MCP cannot promote models.

get_study_prediction_accessA

Read partial prediction-access history for this run and matching recorded dataset/windows in this study: recorded_accesses, recorded_runs, by_action (what the total is made of) and the 20 most recent events. An event is recorded only when prediction evidence is served by a deliberate action: saving an assessment (assessment), a comparison (comparison, one per served run), a pair assessment (pair_assessment), dashboard results (dashboard_results) and prediction-window exports (results_csv, results_json). Listing assessments never records. Histories written before jellyfish #837 may also contain assessment_history rows from list reads; nothing is deleted because holdout-use declarations reference event ids. This read does not expose predictions or add events. Earlier activity, other result routes and offline work are not covered; absence never proves untouched holdout status. Repeated access does not prove retuning.

get_study_validation_resolutionsA

Read analyst validation resolutions and revocations, including current/stale/revoked status re-evaluated against exact evidence. API keys cannot supply human independence sign-off or revoke it; use the signed-in owner UI. No audit serving event is added and no model is fitted or promoted.

declare_study_holdout_useA

Append a submitter-reported evidence-use declaration using project-owner credentials and create:models. Read get_study_prediction_access first and reference an access event from this run. Use a fresh UUID declaration_id and reuse it unchanged on retry. informed_revision requires a published affected revision in the same study; other dispositions omit it. Reason must explain actual use. Reported revision influence requires fresh validation for affected revisions; later review-only notes cannot erase it. Does not certify independence, accept or promote a model. API submissions remain identified as reported declarations.

assess_study_validation_pairB

Assess a validation pair and append a prediction-access audit event when evidence is available. Checks distinct completed MMM runs launched under the declared protocol, frozen inputs/settings/runtime, configured sampling, saved R-hat, declared prediction windows/WAPE and date coverage. Returns blockers and an evidence hash; does not fit, accept or promote. Includes saved retained chain/draw, ESS and divergence records when available, with null for older models. Optional prelaunch retained_sampling limits require complete native records and check chain/draw minima, bulk/tail ESS minima and maximum divergences; otherwise sampling_qualification is not_declared. Optional require_policy_review checks current signed-in analyst acceptance of each latest same-policy assessment, including freshness and rejection blockers. The holdout_provenance report distinguishes missing evidence, blocked version 1 full-input preprocessing and version 2 training-only preprocessing requiring further provenance review. The prior_provenance report checks which rows the recorded smart priors read and the frozen input hashes; priors without a record remain unavailable, and priors that read observations after the declared training end are blocked. fresh_validation provides replacement-window preflight for influence reports naming the full-model revision: later windows, replacement policy chronology, recorded prior exposure, retained diagnostics and provenance/review requirements. It never clears champion blocks. External business calculations and untouched holdout history remain unverified; decision_grade_ready stays false.

recommend_study_runA

Record a recommendation with evidence. This does not accept or promote a model; analyst acceptance happens in the frontend.

adopt_model_into_studyA

Import an owned completed model into a study as an executable, editable recipe (jellyfish #880). Without confirm: a preview only, with the resolved recipe, its inspection block and an import report {complete, dataset: {recorded, origin}, settings: {count, not_recorded}, editable: fully | partly, evidence: {fitted_result}}; nothing is written. The preview's report.dataset also carries display, the lineage line the recipe card will show once imported ("Retail weekly · v3 · verified"; "dataset lineage not recorded · legacy" for a model built before capture). With confirm=true: creates a recipe whose base_model revision 1 carries the exact fit inputs, the wizard snapshot captured at build and the verified dataset origin, launches and re-freezes like any other revision and can be edited in place (get_recipe_revision_authoring, then create_recipe_draft with target); the fitted result is attached as an adopted run outside the attempt budget (201, {run, revision}). 409 when the model is not complete or is already attached to a study. Models built before snapshot capture import with not_recorded settings (editable: partly). Adoption never refits.

compare_study_runsA

Compare 2-20 candidates against one quality policy. Each row carries a basis record (family, dataset and costs hashes, outcome column, units, training window, output kind, prediction window, evidence freshness) and a per-dimension compatibility against the first row; rows are comparable only when family, dataset, outcome, units, window and output kind all match, and incompatible rows are returned with the differing dimension in blockers and are never ranked. This answers predictive ranking only: sensitivity agreement is not computed, analyst acceptance lives in decisions, and business validity is a human review. Serving prediction evidence appends access audit events. Does not fit or promote models.

get_workflow_guidanceA

Read versioned Simba workflow guidance when native Skills are unavailable.

Use index to list allowed topic/section IDs, then request the relevant section. Topics: mmm, results, priors, optimiser, studies, var. The default section is entrypoint; it names detailed sections and when they are needed. Arguments are identifiers, never paths. Returns a complete section within a 24,000-byte structured-payload limit or a refusal, never truncated instructions. This local read makes no backend request. Guidance does not authorise writes or replace validation. Lookup is optional when the caller already has the relevant versioned guidance.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription
ui://simba/charts.htmlRead-only charts of served results, with exact-value tables.

TDQS

A3.5/5.0

Scored across 94 tools

Disambiguation2/5

With 94 tools spanning overlapping MMM, study, recipe, revision, incrementality, campaign, and pipeline domains, many tools share nouns and require long descriptions to distinguish. Boundaries between recipe drafts, study recipes, revisions, and incrementality design/save/create operations are not self-evident to an agent without careful reading. Several descriptions include warnings about confusing related tools, which further signals latent ambiguity.

Naming Consistency5/5

Tool names consistently use snake_case with a predictable verb_noun pattern: list_*, get_*, create_*, update_*, run_*, set_*, and delete_*. Domain prefixes such as study_, recipe_, campaign_, pipeline_, and incrementality_ are applied consistently. Although many names are long, the convention is stable throughout.

Tool Count1/5

94 tools far exceeds the typical well-scoped range of 3-15 and lands in the extreme-mismatch band. Even for a broad MMM platform, this surface is so large that tool selection, documentation, and maintenance become unwieldy. The count alone indicates over-exposure rather than a tight, purpose-built tool set.

Completeness4/5

The surface is extensive, covering data upload/reporting, model creation and results, optimizers, scenarios, studies, recipes, quality policies, evaluations, champion tracking, campaigns, pipelines, and incrementality tests. There are some intentional gaps, such as deletion only for failed models and no API deletion for projects or unreferenced policies, but agents can generally work around these. Overall lifecycle coverage is strong, though not perfectly complete.

Maintenance

ActivityActive
ResponsivenessUnresponsive