Simba MCP Server
OfficialServer Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| SIMBA_API_KEY | Yes | Your Simba API key (stdio mode only – HTTP callers send their own key as the bearer token) | |
| SIMBA_API_URL | No | Simba API base URL | http://localhost:5005 |
| SIMBA_MCP_ALLOW_LOCAL_FILES | No | Set to '1' to allow local file paths (csv_path) on HTTP/SSE deployments; disabled by default. | 0 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| show_response_curvesB | Show served response curves, bands and current spend in a native chart. Read-only. Returns the existing JSON unchanged on every client. Missing points are gaps; sparse legacy grids are disclosed. Current spend is not a recommended allocation. Requests fixed sections and never accesses prediction-window data. |
| show_decompositionB | Show served contribution time series in KPI units in a native chart. Read-only. Overlap is a separate reconciliation term, never a channel. The attribution convention is displayed where available. Returns the same useful JSON on clients without visual support. Never requests prediction-window data. |
| show_optimizer_allocationA | Show saved allocation spend and separate decision/comparison result tables. Read-only. Pass run_id to read a specific saved run; omission reads the latest model-level result. Decision Revenue/ROI never mixes with fitted-convention OptimizedEvalRevenue/ROI or HistoricalRevenue/ROI. No scientific calculations occur in the view. Returns existing JSON unchanged without visual support. |
| create_recipe_draftA | Save an encrypted authoring draft without publishing, fitting or consuming an attempt. Supply a UUID draft_id and reuse it with identical content after an uncertain response. Start from get_recipe_draft_template for a new draft, or from get_recipe_revision_authoring when working from a published revision. To edit a recipe in place, pass target={recipe_id, base_revision_id}: the draft becomes an edit session of that recipe and publish_recipe_draft with a matching target writes the recipe's next revision; the target is immutable after creation (409 on a replay that names another; 404 for a recipe outside the study). To branch a published recipe into a new one instead, omit target and supply its source_revision_id from this study; the immutable link carries inherited influence through publication. Backend validates version and size. |
| get_recipe_draftA | Read the complete authoring snapshot and concurrency version. Preserve all fields when editing. Draft state is incomplete, unvalidated authoring data, not an executable recipe. |
| get_recipe_draft_templateA | Get complete shared wizard defaults, hash and envelope schema. Optionally choose an owned uploaded_file_id (from list_uploads) or pipeline_version_id (from list_pipelines / list_pipeline_versions), never both, into frozen source bytes with verified lineage and an editable data preview. Copy snapshot into create_recipe_draft and preserve unedited fields. Defaults are not a validated model. Does not create, publish or run anything. |
| publish_recipe_draftA | Publish a saved MMM or VAR draft as immutable recipe revisions. Supply a new UUID publication_id and reuse it with identical arguments after an uncertain response; a replay returns the same revisions. Without target: new recipes, one per prepared brand, atomic. With target={recipe_id, expected_version}: revision N+1 of that recipe under its row lock, the draft's base revision recorded as the source. 412 stale_version when the recipe moved since expected_version: the draft and every edit are kept, so re-read the recipe and publish again on top, or drop target to branch into a new recipe. 409 target_requires_single_recipe for a multi-brand draft; 409 when the draft's own target names a different recipe; 404 for a recipe outside the study. Backend compiles the saved snapshot with shared wizard rules; never send separately prepared settings. Does not fit, consume an attempt or designate a champion. MMM priors the draft leaves unset are filled with smart defaults at publication (the Prior Builder's own calculation over the whole source) and frozen; replay does not rebuild them. VAR requires family-specific evidence assessment; MMM policies cannot establish VAR acceptance. Enabled invalid MMM calibration fails publication; VAR calibration is unsupported. Preserve disabled authoring observations. |
| get_recipe_revision_authoringA | Read the authoring snapshot behind a published revision. authoring_draft, wizard (Save a recipe) and base_model (imported fitted model) revisions carry one; api_mmm, model_snapshot and legacy rows return an explicit unavailable error (404). Returns snapshot, name, revision_id, draft_content_hash, kind, source_available and source_unavailable_reason. Dataset bytes are filled from the recorded origin only when it still hashes to what was fitted; otherwise snapshot.source is null, source_available is false and the reason says to choose the dataset again before the draft can publish. Use the snapshot with create_recipe_draft: target for an in-place edit, source_revision_id for a branch. The published revision stays unchanged. Does not create or fit anything. |
| list_recipe_draftsA | List study draft metadata without loading datasets, newest update first. Check backend draft capability first. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised. |
| update_recipe_draftA | Replace authoring state using the version from get_recipe_draft. Retain every unedited field, including original priors and source bytes. Stale changes fail; reload and reconcile explicitly. Identical retries return the current draft. Does not publish or launch. |
| get_backend_capabilitiesA | Discover this caller's connected backend features before planning work. Returns only backend advertisements: model families, transformations, priors and workflow operations. A missing advertisement is unknown, not unsupported. Check each field; an advertised feature still requires permission and budget. No model is created and capabilities are not cached across callers. |
| get_data_schemaA | Get the canonical CSV data schema for Simba MMM input files. Returns the JSON Schema specification describing required columns (date, KPI, multiplier, hierarchy), media channel column naming conventions ({channel}_activity, {channel}_spend), constraints (min rows, max file size), and supported date formats. |
| get_data_reportA | Report actual data from a stored dataset: any window, any grain, by brand, channel or dimension. Reads the dataset itself (every column, any date range) rather than a fitted model's training window. Use it for "sales and TV spend in the North region for August, by week". Roles are DECLARED, never guessed from column names. Declare them with upload_data(roles=...)
or per request with Role vocabulary (aggregation, unit) — also in get_data_schema under x-simba-roles:
Buckets: week = ISO week from Monday; month/quarter = calendar. A weekly row counts in the month of its week-start date. The response's meta.aggregation states every rule applied. Args: dataset_id: The uploaded file id (from upload_data or list_uploads). Registered pipeline outputs are uploaded files too. start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (default), "week", "month" or "quarter". group_by: "hierarchy", "channel", or a dimension role such as "dimension:market". hierarchy: Keep only this brand/region value. metrics: Roles or role families to include, e.g. ["kpi", "spend", "outcome:orders"] or ["control"]. Default: every metric role present. roles: {column: role | {"role", "channel"}} overriding roles stored at upload. Returns {dataset: {id, name, source, version, sha256, data_through}, granularity, rows: [{period_start, period_end, group, metric, value, unit}], meta: {basis: "dataset", aggregation, roles, channels}}. Errors carry a code: dataset_not_found (404), invalid_report_request (400), report_too_large (413, over 10,000 rows — narrow the window or coarsen the granularity). |
| upload_dataA | Upload a CSV dataset to Simba for use in model building. Provide EXACTLY ONE of csv_content (raw CSV text) or csv_path (a file path on the machine running this MCP server). Prefer csv_path for anything beyond trivial size — it avoids passing megabytes of CSV through the conversation. The CSV should follow the canonical schema: one row per time period with date, KPI, multiplier, hierarchy, media activity/spend columns, and optional control variables. IMPORTANT:
Args: csv_content: The full CSV text content (not base64, just raw CSV text). csv_path: Path to a .csv file readable by the MCP server process. name: Optional dataset name for identification. Defaults to the file stem when csv_path is used. filename: Optional original filename to record alongside the dataset. roles: Optional column roles for get_data_report, stored with the dataset: {column: role} or {column: {"role": role, "channel": name}}. Roles are declared, never guessed; see get_data_report for the vocabulary. An unknown role or a column the CSV lacks is refused. Returns the uploaded file ID (needed for create_model), row/column counts, and any validation warnings. |
| list_pipeline_versionsA | List the saved versions of one owned pipeline (by pipeline_hash or id), newest first: id, version, created_at, row_count, column_count, column names and checksum. The checksum is the exact source identity a recipe draft freezes. Never returns the data itself; fetch a draft template with pipeline_version_id for that. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. |
| list_pipelinesA | List the data pipelines this key's owner has, most recently updated first, each with id, pipeline_hash, name, description, version_count and latest_version (id, version, created_at, row_count, checksum). Identity only: no definitions, parameters or output data. Use a version id as pipeline_version_id in get_recipe_draft_template. Requires the ingest scope; the same ownership rule as the app. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. |
| run_pipelineA | Start a refresh of one owned data pipeline (by pipeline_hash or id). Returns {run_id, status: "queued"} at once; the run executes on the server as the pipeline's owner with its saved connections. Poll get_pipeline_run until status is succeeded or failed (every few seconds; warehouse runs can take minutes, and a run stops at 30 minutes). Optional start_date / end_date (YYYY-MM-DD) limit the source steps to that range. One run per pipeline at a time: if one is already queued or running you get run_in_progress with that run_id — poll it instead of starting another. Each successful run saves a new pipeline version. Requires the create:models scope. |
| get_pipeline_runA | Get one run of an owned pipeline: {run_id, status (queued | running | succeeded | failed), started_at, finished_at (UTC), version_id, error_code, error}. On success version_id is the new saved version — pass it as pipeline_version_id to get_recipe_draft_template to build on the refreshed data. On failure error_code says why: execution_failed (a source or transform failed; error names it), no_output, timeout, interrupted or not_started (start it again), not_queued, owner_blocked, unexpected. Scheduled runs are polled the same way. Requires the create:models scope. |
| set_pipeline_scheduleA | Replace the refresh schedule of one owned pipeline. cadence is "daily" or "weekly"; hour_utc is a whole UTC hour 0-23; weekday (0 = Monday … 6 = Sunday) is required for weekly and must be omitted for daily; enabled false pauses the schedule and keeps its settings. Returns the schedule with next_run_at (UTC). Each due slot starts one ordinary run (poll it with get_pipeline_run); a slot is skipped while a run of that pipeline is still going. After a scheduled run succeeds the pipeline keeps its newest 30 versions; a version a model or recipe was built from is never removed. Requires the create:models scope. |
| list_campaignsA | List the campaigns in the campaign facts, each with its totals and the model channel it counts towards for this model. The channel is DECLARED in the model's campaign map (set_campaign_mapping), never inferred:
a campaign no map row covers is Returns {campaigns: [{platform, account_id, campaign_id, campaign_name, adsets, spend,
impressions, clicks, platform_conversions, platform_value, platform_channel_type, first_seen,
last_seen, channel, suggested_channel, status: "mapped" | "unmapped"}], window: {start, end},
currency, as_of, next_cursor (only when Args: model_hash: The model whose map decides each campaign's channel (the hash of a fitted MMM). platform: Keep one platform: meta, google_ads, tiktok or other. unmapped_only: Only campaigns without a channel for this model. start, end: Optional ISO dates (YYYY-MM-DD), inclusive, for the totals. limit: Page size (1-200). Without it every campaign is returned. cursor: The previous page's next_cursor. Errors carry a code: model_not_found (404). An empty list means no pipeline is registered as the campaign facts source yet, or it has not run; registration happens in the app. |
| get_campaign_reportA | Report the campaign facts over any window, at any grain, by platform, model channel, campaign or ad set. The same report engine as get_data_report, over the campaign facts: spend, media:impressions, media:clicks, outcome:platform_conversions, outcome:platform_value and, where a source fills them, outcome:last_click_conversions and outcome:last_click_value. These are the PLATFORMS' own attributed numbers, not Simba's incremental attribution.
Returns {granularity, data_through, rows: [{period_start, period_end, group, metric, value,
unit}], meta: {basis, aggregation, roles, channels}, as_of, source_versions, currency}.
Args: model_hash: The model whose map decides the channels (optional; needed for group_by=channel). start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (daily, as stored), "week", "month", "quarter" or "year". group_by: "platform", "channel", "campaign" or "adset". Default: one total per period. platform, channel, campaign_id: Filters applied before grouping. metrics: Roles to include, e.g. ["spend", "outcome:platform_conversions"]. Default: all. Errors carry a code: campaign_facts_empty (404, nothing matches), invalid_report_request (400), report_too_large (413: narrow the window or coarsen the granularity). |
| get_campaign_incrementalityA | Incremental ROAS per campaign (or ad set), beside the platform's own ROAS and last-click ROAS, by pushing the model's channel incrementality down through the platform's attribution. THE ASSUMPTION, FIRST. Simba measures incrementality at channel grain. For each model channel
over the window, the incrementality factor = the channel's MMM incremental revenue (the
model's per-period rows, under its fitted attribution convention) / the platform-attributed
value of the campaigns mapped to that channel. Each campaign's incremental ROAS is that
factor x its platform ROAS, so campaign incremental revenue sums to the channel's. This
assumes the platform over-credits every campaign in a channel equally. It does not:
retargeting and brand search are over-credited more, so one factor flatters them. The
response warns when such campaigns share a channel with prospecting
( Per row,
Returns {model_hash, window: {start, end}, level, currency, interval, interval_reason, channels: [{channel, factor, factor_source, factor_mmm, factor_interval, revenue_interval, factor_draws_mean, method, mmm_revenue, platform_value, spend, campaigns, test, warnings}], rows: [{platform, account_id, campaign_id, campaign_name, adset_id, channel, spend, platform_value, last_click_value, days, platform_roas, last_click_roas, incremental_revenue, iroas, incremental_revenue_hdi, iroas_hdi, method, factor_source}], unmapped: [{..., spend, platform_roas, last_click_roas}], warnings: [{code, message, channel?, campaigns?, reason?}], provenance: {source_versions, as_of, map_version, attribution_convention, link}}. Warning codes: retargeting_shares_channel_factor, platform_value_missing, kpi_not_revenue, currency_mismatch, uninformative, unmapped_spend, interval_unavailable, test_override_skipped. Args: model_hash: A fitted MMM with a campaign map (set_campaign_mapping). start, end: Optional ISO dates (YYYY-MM-DD), inclusive. Default: the overlap of the model's data and the campaign facts. level: "campaign" (default) or "adset"; ad-set rows inherit their campaign's channel and factor and sum to the campaign row. Errors carry a code: model_not_found (404), campaign_facts_empty (404: no facts, or none in the window), invalid_window (400: empty or reversed; the body gives both spans), model_not_mmm (400: a VAR model has no channel revenue rows), model_incomplete (400). |
| get_campaign_marginal_returnsA | Read channel-derived marginal returns for campaigns or ad sets. Requires read:models. Campaign response shapes inherit the fitted channel shape, rescaled by observed spend share and relative efficiency. These are not independently measured campaign saturation curves or causal campaign effects. Daily spend is observed spend divided by inclusive calendar days, not a configured platform budget or a daily revenue forecast. Args: model_hash: Owned, completed MMM with compatible curve provenance and campaign facts. start, end: Inclusive observation-window ISO dates (YYYY-MM-DD). level: campaign or adset; ad sets inherit the campaign's channel. Returns {context_key, window, currency, minor_digits, channels: [{channel, current_daily_spend, status, reason}], rows, provenance}. Rows carry composite identity, method and marginal-return evidence. Missing or incompatible currency, curve basis or fact coverage makes recommendations unavailable; do not substitute defaults. Missing posterior marginal evidence means no uncertainty interval, not zero uncertainty. This reads existing evidence only: no fit, posterior job or platform change is started. |
| recommend_campaign_budgetsA | Calculate campaign budget suggestions without applying them. Requires read:models. Uses the same channel-derived marginal curves and bounded allocator as the web app. This POST is a read-only calculation: it creates no saved run, starts no fit or posterior job and does not change mappings or advertising-platform budgets. Results are conditional scenarios, not independently fitted campaign response curves or a day-by-day forecast. Args: model_hash: Owned, completed MMM with verified curve and currency provenance. observation_window: {start: YYYY-MM-DD, end: YYYY-MM-DD}, inclusive fact dates. currency: Explicit currency matching the model and facts, for example GBP. channel_daily_budgets: Exact channel keys mapped to daily totals in that currency. Supply exactly one of this or optimizer_run_id. Preserve currency minor precision. optimizer_run_id: Existing owned optimiser run for the same model, with verified planning dates and currency. Its totals become flat daily equivalents, not a daily schedule. The tool does not create or rerun an optimiser. level: campaign or adset. max_step_fraction: Maximum change from observed average daily spend; default 0.2. bounds: Optional rows keyed by full identity {platform, account_id, campaign_id, adset_id}, with min and/or max amounts in daily currency units. Use null adset_id at campaign grain. No platform floor is invented; the server validates bounds. expected_context_key: Optional context_key returned by get_campaign_marginal_returns. A changed model, facts, mapping or attribution basis refuses with 409 instead of calculating against evidence different from the scenario you reviewed. Returns {channels: [{channel, status, total_daily_budget, rows, explanation, assumptions, reason, feasible_range}], ...provenance}. Ready rows include current and recommended budgets, the continuous solution, marginal returns and binding constraints. Refusals retain their reasons and feasible ranges; never relax bounds or silently change a total. Rounded amounts reconcile in integer currency minor units; rounding does not imply exact equality of marginal returns. Missing marginal uncertainty remains explicitly unavailable. |
| set_campaign_mappingA | Replace the model's campaign map: which model channel each campaign counts towards. Each row is {platform, campaign_id | name_pattern, channel, valid_from?, valid_to?}: exactly
one of Refused (422, code campaign_map_invalid, nothing saved) when a channel is not the model's
(the error names the closest one), when a campaign would count towards two channels on any
date (two exact rows on one campaign with overlapping dates, or two patterns matching one
campaign with overlapping dates; Returns the map report: {map, unmapped: [{platform, campaign_id, campaign_name, spend,
first_seen, last_seen, platform_channel_type, suggested_channel}], conflicts: [], drift:
[{channel, campaign_spend, model_spend, ratio, over_tolerance, campaigns}], tolerance,
currency, currency_mismatch, window, channels}. Args: model_hash: The fitted MMM the map belongs to. rows: The complete map. tolerance: The drift warning threshold as a fraction (default 0.05). |
| list_incrementality_testsA | List a project's recorded (not retired) incrementality tests: {items: [{id, name, type, status, channel, model_channel, start_date, end_date, measured_through, kpi, result, spend, source_tool, has_supplied_row, current_version, used_by}], next_cursor}; used_by counts the model revisions built from the test. Null and negative results are listed like any other. Filter by type, status or channel. Paging is opt-in: pass limit (1-200) and send next_cursor back unchanged; null means the end. Requires the read:models scope. |
| recommend_incrementality_testsA | Rank channels for experiment investigation using stored posterior marginal returns. Read-only: no fit, test creation or budget changes. Returns {method, basis, score_unit, currency, budget, hurdle, spend_basis, period, approximation_warnings, items, excluded}. Each item carries channel, score, components (mean, sigma, stake, spend_share, crossing_probability, optional contraction), reason_codes, last_test_end, hypothesis and an unavailable design_hint. The score is a normal-approximation local binary perfect-information value, not expected test benefit, experiment budget, portfolio value or forecast lift. budget is a positive exposure scale (default sum of current spend); weights are the mean-active-period spend mix, which need not represent one common calendar period. hurdle is the non-negative marginal-return alternative (default 1). limit is 1-50. Missing posterior means or intervals are explicitly excluded. limited_variance_contraction means contraction from zero to below 0.1; posterior_variance_expanded means negative contraction. Neither proves prior domination. Cross-channel dependence is not modelled. Requires read:results. Test history is unavailable because stored registry records do not establish compatible model/geographical coverage. Requires backend support for test-priorities. |
| get_incrementality_testA | Read one recorded test: {id, version, record, content_hash, used_by, retired_at}. version reads an older version (default: current). With model_hash (a saved model you can read), the result also carries |
| create_incrementality_testA | Record one incrementality test in a project, as analysed in its own tool: its design (geo, owned_media_ab or platform_lift block), dates, KPI, result with interval or sd, and incremental spend. Returns {id, version: 1, record, content_hash}. Set model_channel to the model activity column the test calibrates so it can be used with a model. Validation errors name the field (e.g. "the interval must contain lift_abs"). Not idempotent: calling twice records two tests. type time_holdout is not accepted here: planned time-holdout records are created only through save_incrementality_test_design (the backend refuses them elsewhere), and a status planned record of any type carries no result. Requires the create:models scope. |
| import_incrementality_testsA | Import tests from another tool's output file: source is csv (Simba's template), meta_conversion_lift (Conversion Lift API results JSON), geox (a meridian-geox analysis result), geolift (GeoLift summary), causalpy (effect summary or lift rows) or pymc_marketing (lift rows); content is the file's text (10 MB max). Returns {records: [{key, record, errors}], created, notes}. dry_run (default true) creates nothing — review each row's errors, then call again with dry_run=false to create the rows without errors (their ids come back in |
| design_incrementality_testA | Ask a saved, complete MMM what an incrementality test on one of its media channels could detect: queues a bounded calculation that replays the saved posterior and returns {calculation_id, model_hash, status: "queued", submitted_at} at once (202; 200 with the same body when submission_key repeats identical inputs). Poll get_incrementality_test_design until status is complete, failed or cancelled. A complete calculation carries a result whose own state is available, unsupported (e.g. a log-link model), insufficient_evidence or no_feasible_design; none of these is an error, and only available carries numbers (candidates are also returned for no_feasible_design so you can see why nothing met the target). channel is one of the model's nonlinear media channels (400 unknown_channel otherwise). design_type time_holdout pauses or changes the channel's spend for a window and contrasts the outcome with the model's forecast; geo_split needs geo and splits a market panel into treatment and control. intervention gives the start date, candidate durations, spend change and baseline; inference the alpha, target power and an optional named effect; the backend fills and echoes every default. Reuse the same submission_key after a lost response; a new key queues another calculation and the same key with different inputs is refused (409 submission_key_conflict). Refusals: 404 model not found or not readable, 422 model_not_complete, 503 queue_unavailable (nothing is left behind). Nothing is saved or launched and no budget changes; save an available result with save_incrementality_test_design. Requires create:models to submit; polling needs read:results as well, so a key with only create:models can submit but not read. |
| get_incrementality_test_designB | Read one test-design calculation: {calculation_id, model_hash, status, submitted_at, started_at, completed_at, request, error, result}. status is operational: queued, running, complete, failed or cancelled; error is {code, message} only when failed (artifact_unreadable, timeout, worker_lost, internal) and a failed calculation carries no numbers. result is present only when complete and is read by its own state, never by key presence: available, unsupported, insufficient_evidence or no_feasible_design, with reasons [{code, detail}] non-empty exactly when not available, plus warnings, assumptions, diagnostics and provenance (method, model artefact, posterior draws, training periods). An available result carries intervention (dates, carryover and measurement end, spend change), model_implied_effect (mean and 94% HDI, parameter uncertainty only; available may be false, e.g. a geo split on a national model, which then gives a national reference only), detectable_effect (cumulative, per-period, relative and, for revenue KPIs, iROAS MDE at the stated alpha and target power), power (at the model mean effect, at the named effect, and assurance), noise, and candidates per duration with the selected one marked. Power here is posterior-averaged: assuming the model is correctly specified and its posterior calibrated, the probability that the pre-specified analysis rejects the null, averaged over the model's forecast uncertainty and observation noise; it is not the probability that this experiment will detect the effect, and a model-implied effect is not a prediction of what the experiment will measure. Requires read:results. |
| save_incrementality_test_designA | Turn an available test-design result into a planned incrementality test record in the registry: returns {id, version, record, content_hash} (201; 200 with the same body when this calculation was already saved). The server builds the record from the stored result (type time_holdout, or geo with treatment and control markets; status planned; source.tool design with the calculation, model artefact and method in source.fields), so only the target project_id (default: the model's project; 400 when the model has none and none is given) and an optional name are sent. Idempotent per calculation. Refusals: 409 result_not_available when the calculation is not complete or its result state is not available, 409 artifact_changed when the model artefact no longer matches the calculation, 403 when you are not the project's owner. A planned record has no measured lift: it cannot calibrate a model (get_incrementality_test reports type_not_calibratable or test_not_completed) and nothing here launches an experiment or changes a budget. Requires create:models and write access to the project. |
| list_uploadsA | List the datasets in your workspace (newest first) — every source, not just API uploads: dashboard/manual uploads and pipeline-ingested datasets appear too (see source_type per file). Returns {files, count, limit, offset} where each file has: id (the
uploaded_file_id create_model needs), filename, original_filename,
source_type, row_count, column_count, created_at. Here Args: limit: Page size (API clamps to 1-500; default 50). offset: Rows to skip (paging). name: Optional case-insensitive substring filter on the original filename. |
| get_uploadA | Get one uploaded dataset's details, including its column schema. Returns id, filename, original_filename, source_type, mime_type, file_size, row_count, column_count, columns ([{name, dtype}, ...] — use these to build create_model's channel/control column arguments without re-reading the CSV), and created_at. Args: file_id: The upload's id, from upload_data's response or list_uploads. |
| list_modelsA | List all Marketing Mix Models for the authenticated user. Returns model name, hash, status (pending/under way/complete/failed), type (mmm/var), hierarchy value, and timestamps. NOTE: All other model endpoints use model_hash (string, e.g. "f835671a25") as the identifier. Use the model_hash from this response. Args:
include_unsaved: Include draft/unsaved models (default false).
limit: Maximum number of models to return (default 50, max 500).
offset: Number of models to skip, for paging past |
| create_modelA | Create and start fitting a new Bayesian Marketing Mix Model. This queues an async model fit and returns immediately with a model_hash. Use get_model_status to poll for progress until status is 'complete'. Priors are calculated automatically using smart defaults based on cost shares, industry benchmarks, and channel-type detection. You can override individual channels via the priors parameter. Args:
uploaded_file_id: The file ID returned by upload_data.
date_column: Name of the date column in the CSV.
kpi_column: Name of the KPI/dependent variable column.
hierarchy_column: Name of the brand/segment column (must have exactly 1 unique value).
channels: List of channel definitions, each with keys: name, activity_column, spend_column.
Example: [{"name": "TV", "activity_column": "tv_grps", "spend_column": "tv_spend"}]
multiplier_column: Column to convert KPI to revenue. Defaults to kpi_column.
control_columns: Non-media control variable column names (e.g. ["price", "distribution"]).
total_media_effect: Controls prior strength. Either an industry name for a benchmark
("FMCG"=6%, "Retail"=9%, "TelCo"=30%, "Financial Services"=19%,
"E-Commerce"=22%, "Other"=12%) or a custom decimal like "0.15"
meaning "I believe all media drives 15% of my KPI". Default "Other".
priors: Optional per-channel prior overrides. Each dict should have "channel" matching
a channels[].name, plus any fields to override: distribution, mean, sd, lower,
upper, transform, adstock_type, effect_period.
Only specified fields are overridden; the rest use smart defaults.
Adstock-kernel fields: half_life_lower/half_life_upper (carryover half-life
bounds in periods — preferred over the legacy decay_lower/decay_upper),
theta_mean/theta_sd (peak-lag prior, adstock_type="delayed" only),
dual_weight_mean/dual_weight_sd (long-term/slow-component share prior,
adstock_type="dual_geometric" only).
SATURATION ANCHOR — state it ONCE, in exactly one of three
mutually exclusive forms (two in one override -> 400 "state
the saturation prior once"):
(1) half_marginal_mean/half_marginal_sd — CANONICAL for
saturation_type="generalized_log" (rejected on other
families): the activity level where MARGINAL returns have
halved, finite at every curvature (#632).
sat_shape_mean MUST accompany the pair in the same override
(#672) — the fold pairs your coefficient with the
stated curvature, so omitting it is a 400, never a silent
default.
(2) half_saturation_mean/half_saturation_sd — the
50%-of-maximum point in activity units, for the
single-parameter families (tanh/michaelis_menten/
negative_exponential). Do NOT use it for generalized_log
near-log work: it overflows below sat_shape_mean 0.00097657
and is rejected with a 400 — precisely the regime that
family exists for.
(3) alpha_sd + scalars — legacy internal coordinates,
accepted for backward compat.
Curvature (generalized_log only): sat_shape_mean/sat_shape_sd
— small values are near-logarithmic, 1.0 is michaelis_menten.
COEFFICIENT in a human coordinate (generalized_log only,
#671): effect_at_avg_mean/effect_at_avg_sd — the
effect share at the channel's AVERAGE activity, as FRACTIONS
(mean in (0, 0.95], sd > 0; 0.2 means 20%). Folded
server-side into mean/sd at the row's operating point with
the same arithmetic as the dashboard. Requires sat_shape_mean
in the same override; cannot be combined with mean/sd
("state the coefficient prior once") or with
half_saturation_*. Stating half_marginal_* + effect_at_avg_*
+ sat_shape_mean together is the full (x*, E, k) triple —
the recommended generalized_log elicitation, since only
beta*k is identified and raw beta spans orders of magnitude.
VERIFY what was applied via get_model's
model_config.priors_resolved: rows carry the FOLDED
mean/sd/scalars/alpha_sd, and overridden_fields lists the
field names you sent.
UNKNOWN KEYS ARE REJECTED with a 400 naming the field
(#630); they used to be dropped silently, fitting a
hybrid of the override and the smart defaults. Common misses:
"beta"/"beta_mean" -> mean, "beta_sd" -> sd, "sat_shape" ->
sat_shape_mean. "name" and "parameter" are rejected too — they
identify the smart-prior row the override merges onto.
trend: Enable dynamic baseline trend component.
seasonality: Enable automatic seasonality detection. The prior sigma on
the Fourier coefficients is chosen for the link (#534):
0.5 under link="log", 10 under "identity". The coefficients
live on the link's scale, so the additive default would
admit e^10x seasonal amplitude on a multiplicative model.
likelihood: Likelihood function: "normal" (default), "lognormal", "logit",
"studentt", "poisson", "negativebinomial", or "quantile".
saturation_type: Diminishing-returns curve family applied to media:
"tanh" (default), "michaelis_menten", "negative_exponential",
or "generalized_log" (two-parameter Box-Cox/power-log family
1 - (1+x/K)^(-shape); tune per channel via the
sat_shape_mean/sat_shape_sd prior fields).
transform_order: "adstock_first" (default: carryover accumulates, then
saturates) or "saturation_first" (each period's spend
saturates, then the effect spreads over time through the
normalized adstock kernel).
link: Model Form. "identity" (default) fits an additive model — components
add on the outcome scale. "log" fits a multiplicative model —
components add on the log scale and media effects are percentage
lifts. Under the removal_lift attribution convention (the API
default), contributions then include an Overlap reconciliation
column; the other conventions (aumann_shapley — the dashboard
default for multiplicative models since #509 —
shapley, and proportional_normalized) allocate the interaction
across components and close exactly WITHOUT an Overlap column
(see get_model_results).
channel_groups: Optional adstock groups: [{"name": ..., "channels":
[...], "share_saturation": bool}]. Member channels tie
their carryover parameters (decay/theta/dual-weight —
plus saturation when share_saturation is true) to one
shared value, e.g. grouping channels into shared
"Long"/"Short" carryover classes. Members are
channels[].name values; each group needs >= 2 members;
groups must be disjoint; and tied members must have
identical adstock_type/effect_period/bound overrides
(the API rejects divergent groups at request time).
control_reference: Control attribution reference points (#452),
multiplicative models (link="log") only: maps control
column names (plus optional "default") to
"auto" | "absent" | "average" | "lowest" | "highest" —
which counterfactual "remove this control" means in
the contributions. "absent" measures against the
variable at zero (legacy behavior; honest only when
zero is observed). "average"/"lowest"/"highest"
reference the control at its observed mean/min/max —
use for controls that never approach zero (price
indices, distribution levels), where a zero
counterfactual produces unbounded contributions and a
negative Base. "auto" detects per control whether
zero is inside the observed data range. Example:
{"default": "auto", "relative_price": "average",
"promo_flag": "absent"}. Omit entirely to keep every
control at "absent" (byte-identical legacy output).
Unknown control names/modes are rejected at request
time; any value other than "absent" requires
link="log". The fit reports the resolution in
model_config.control_references (see
get_model_results).
name: Display name for the created model, honoured verbatim (#575).
Falls back to a generated API_MMM{brand}{hash} string when
omitted. Either way the model starts unsaved — invisible to
list_models unless include_unsaved=true — until save_model
files it into a project.
operating_margin: Scalar operating margin as a decimal fraction in
(0, 1], e.g. 0.18 = 18%. Mutually exclusive with
operating_margin_column (the API 400s when both are given).
Storing a margin unlocks the control_priors: Optional root-level control slope overrides. Each entry names a selected control_columns column with "control", plus transform (N, DM, STA, DDM, LOG), distribution (normal, inversegamma, truncatednormal, halfnormal), mean/sd/lower/upper as applicable. LOG is log(x/mean(x)); STA divides by sample sd without centering. Priors are in transformed units; changing transform does not convert coefficients. Nonempty overrides require backend capability version 1; unsupported or unavailable checks stop before creating a model. calibration: Optional lift-test calibration: likelihood observations the
fit must respect. Either {"tests": [{"test_id": ..., "version"?,
"channel"?, "confirm_kpi"?}]} (recorded tests from
list_incrementality_tests, each derived against THIS model's data),
or {"units":
"revenue" | "response", "observations": [{channel, x, delta_x,
delta_y, sigma}]} for rows you derived yourself. If any test can't
calibrate this model, nothing is created: the error has
code "calibration_refused" and Returns the model_hash for status polling. |
| create_var_modelA | Create and start fitting a long-term (VAR) model (#569). VAR models capture the joint dynamics of several series (e.g. sales and
brand-equity metrics) and produce the long-run elasticity bridge behind
the MMM's Args: uploaded_file_id: Dataset id from upload_data (must contain every named column). date_column: Date column name. Cannot also be a series. endogenous_vars: At least two column names — the jointly-modeled series. exogenous_vars: Optional outside drivers; must not overlap the endogenous set. lags: VAR order (>= 1). The dataset needs at least lags + 10 rows with no missing values across the modeled columns. forecast_horizon: Periods forecast for diagnostics (default 12). base_variable: The outcome series (must be endogenous) long-run multipliers are measured against. Required for long-run effects. equity_variables: Endogenous columns (excluding the base) whose long-run IRF multipliers are estimated. Required for long-run effects. lre_horizon: Long-run effects horizon in periods (default 156). lre_ci: Credible-interval mass for the effects table, in (0, 1). var_priors: Advanced prior overrides (lag_coefs / alpha / coefs / noise_chol); unknown keys are rejected. name: Display name for the created model, honoured verbatim (#575). Falls back to a generated API_VAR_* string when omitted. Returns 202-style payload with model_hash; poll get_model_status. |
| link_var_modelA | Link a completed VAR model to an MMM (#569). After linking, the MMM's get_model_results The join is by exact name unless channel_map declares which MMM channels each VAR exogenous series stands for (#682) — required whenever the VAR is fitted on group spends (e.g. four spend groups) while the MMM is tactic-level. Each group's elasticity is allocated across its member channels pro-rata by KPI short-term contribution, so the group's long-run effect is counted exactly once. Validation is strict: keys must be VAR exogenous series, values must be channel names of the (completed) MMM, and no channel may belong to two groups. The map belongs to the link: every link replaces it (omitting channel_map clears any stored map) and unlink clears it. Args: model_hash: The MMM to attach the long-run view to. var_model_hash: The VAR model (from create_var_model). channel_map: Optional {var_exogenous_series: [mmm_channel, ...]} mapping for group-level VARs. |
| unlink_var_modelA | Remove an MMM's VAR link (#569). Idempotent. |
| set_contribution_groupsA | Persist the driver groupings the dashboard contributions view renders (#436) — configure grouping once and every viewer sees it. Each group: {"name": str, "drivers": [column names], "color": "#hex"?, "baseAdjustments": {driver: "min"|"max"|"none"}?}. Driver names are validated against the model's media/control/halo/trademark factors (400 with a did-you-mean hint on typos); each driver may belong to at most one group; baseAdjustments must reference the group's own drivers. The special "_channel_color_overrides" pseudo-group carries a channelColors map instead of drivers. NOTE: this is the CONTRIBUTIONS-VIEW grouping. create_model's channel_groups is the unrelated adstock parameter-sharing feature — do not confuse them. |
| get_contribution_groupsC | Read the stored contribution groups for a model (#436). Legacy dashboard-saved configs are served verbatim. |
| rename_modelA | Rename a model. Changes only the display name; the model's saved/unsaved state is untouched (use save_model to file it into a project). The name is HTML-sanitized server-side and must be non-empty. Args: model_hash: Hash of the model to rename. name: New display name. |
| save_modelA | Save a model into a project under a display name. API-created models start unsaved and are invisible to list_models (without include_unsaved=true) — saving files them into a project so they appear in the default listing and the dashboard's Saved Models. The same saved-models cap applies as in the dashboard: at the cap the API returns a 400 with error_type "saved_limit". Re-saving an already-saved model renames/refiles it without consuming a new slot. Args: model_hash: Hash of the model to save. name: Display name to save under (non-empty). project_id: Optional target project ID; must be a project you own or one shared with a team you belong to. Discover ids with list_projects; create a folder with create_project. Defaults to your default project. |
| unsave_modelA | Release a model's saved slot without deleting anything — the inverse of save_model (#673). Use this for cap management: at the 20-saved-models cap, unsave a model that no longer earns its shelf spot instead of deleting it. The model reverts to the state API-created models start in (unsaved, no project; the name is kept) — it leaves the default listing and the dashboard's Saved Models but stays fully addressable by hash: fetchable, renameable, exportable, re-saveable, and visible via list_models with include_unsaved=true. Idempotent — unsaving an unsaved model is a success with freed_project_id null. delete_model remains failed-only. Two caveats: the UNSAVED pool is auto-pruned by dashboard model creation (at 10+ unsaved models the oldest is hard-deleted, artifacts included), so re-save anything worth keeping rather than parking it unsaved long-term; and unsaving a shared model hides it from every recipient until it is saved again. Args: model_hash: Hash of the model whose slot to release. Returns: {model_hash, is_saved: false, freed_project_id}. |
| list_projectsA | List the projects (the app's model folders) you can file models into. Returns owned and team-shared projects: per project {id, name, is_default, shared_with_team_id, model_count} — team-shared folders carry "shared": true, and model_count counts SAVED models (the set the app's model list shows). Use the ids with save_model(project_id=...) and rename_project. There is deliberately no delete over the API — use the app to delete a project. |
| create_projectA | Create a named project (model folder) to file models into. Names are sanitized the same way model names are (non-empty after HTML sanitization). Args: name: Display name for the new project. team_id: Optional team to share the project with; must be a team you belong to (403 otherwise, 404 for an unknown team). Returns the created project (201) including its id — pass that to save_model(project_id=...). |
| rename_projectA | Rename a project you OWN. Team members can file models into a shared folder but not rename it (owner-only; 404 for a project you don't own). Renaming your default folder is safe: it keeps receiving unqualified saves under its new name. Args: project_id: Id of the project to rename (see list_projects). name: New display name. |
| get_modelA | Get a model's metadata and configuration echo — works for EVERY status, including failed models (unlike get_model_results, which needs 'complete'). Use this to inspect what a model was configured with, why it failed, or where it lives. Returns: id, model_hash, name, status, model_type ("mmm"/"var"), hierarchy_value, periodicity, is_saved, project_id/name, linked_var_model_hash, created_at/completed_at, error (the failure message — non-null only when status is "failed"), and model_config (the create-time configuration echo: data_source, columns, channels, priors as resolved, and the config flags). NOTE: the echo omits a few accepted create_model inputs (operating_margin, annual_discount_rate, reporting_kernel) — absence there does not mean they weren't applied; check the financials results section for the stored margin. Args: model_hash: The model hash (any status). |
| delete_modelA | PERMANENTLY DELETE a FAILED model. Destructive and irreversible. Only models with status "failed" can be deleted over the API — any other status returns a 409 with the model's current status (delete is for cleaning up failed fits, not curating good ones). Deleting also unlinks any MMMs that pointed at it as their VAR model and removes stored artifacts. On success returns {"deleted_model_hash": ..., "status": "deleted"}. Check first with get_model or get_model_status if unsure of the status. Args: model_hash: Hash of the FAILED model to delete permanently. |
| get_model_statusA | Check the fitting progress of a model. Returns status (pending/under way/complete/failed), progress percentage, estimated time remaining, and timestamps. When supported by the backend, fit_liveness reports heartbeat age and the configured stall threshold in seconds, with last_heartbeat_at as Unix seconds. seconds_until_stall_threshold is time to the stale-heartbeat threshold, not fit ETA or an exact kill time. stall_threshold_exceeded does not change the model status. If fit_liveness is absent or available is false, liveness is unknown: do not infer a healthy or stalled fit. Reasons include heartbeat_unavailable and not_fitting. Continue polling with backoff; do not automatically restart a fit. Args: model_hash: The model hash returned by create_model or list_models. |
| get_model_resultsA | Get results from a completed model. Available sections:
The response envelope includes IMPORTANT — channel naming: results are keyed by the channel's ACTIVITY
COLUMN name (e.g. "search_activity"), not by the NOTE: Date values in contributions/coefficients records are millisecond epoch integers. DATE WINDOW: pass start / end (ISO dates, inclusive) and/or granularity
("native", "week", "month" or "quarter") to window contributions,
coefficients, actual_vs_model and channel_summary. The response then
carries CONTEXT-SIZE TIP: a full pull is very large (curve sections alone are 100 grid points x channels x 5 band columns). In conversational use, request only the sections you need and pass channels=[...] and max_grid_points=20. Explicit JSON channel or grid requests add Args:
model_hash: The model hash.
sections: Comma-separated list of sections to include.
Leave empty for all sections.
Common: "channel_summary,model_stats" for ROI and diagnostics.
format: "json" (default) or "csv". CSV returns
{"format": "csv", "content": "..."} — concatenated
"# section" + CSV blocks, useful for saving to disk.
Filtering below applies to JSON only.
channels: Optional channel filter (matching is case/space-insensitive
and tolerates the _activity/_spend suffix). Applied to curve
sections, decay_curves, saturation, channel_summary,
coefficients, mroi_summary, and mroi_periods rows.
|
| run_optimizerA | Run budget optimization on a completed model. Finds the optimal budget allocation across channels to maximize predicted revenue — or predicted PROFIT with objective="profit" — within the given constraints. PROFIT OBJECTIVE: objective="profit" requires a margin source. If the model was built with an operating margin, it is used automatically; otherwise you MUST pass forward_margin (e.g. 0.18 for an 18% margin) or the API returns an error. Result fields (Revenue, ROI, ExpectedResponse) are then on the profit basis. IMPORTANT:
Returns 202 (async). Use get_optimizer_results to poll until status is "complete". Args: model_hash: Hash of a completed model. total_budget: Total budget in currency units. num_periods: Number of periods to optimize over (matches your planning horizon). gamma: Uncertainty-aversion weight on the outcome spread (the objective is mean - gamma * spread). 0.0 = maximize expected return only (most aggressive); higher values penalize uncertainty harder (more conservative). The dashboard typically uses values in the 0-0.1 range. currency: Currency code (e.g. "USD", "GBP"). bounds: Per-channel min/max budget allocation as PERCENTAGES (0-100). Every channel must appear. Example: {"TV_Impressions": {"lower": 5, "upper": 40}, "Search_Clicks": {"lower": 10, "upper": 50}} laydown_weights: Per-channel spend timing weights. Each value is an array of length num_periods. Weights are relative (normalized internally). Use uniform [1, 1, ...] for even distribution across periods. Example: {"TV_Impressions": [1, 1, 1, 1]} period_cpm: Per-channel cost-per-metric for each period. Each value is an array of length num_periods with positive values. Get baseline CPM from get_scenario_template (avg_cpu_by_channel field). Example: {"TV_Impressions": [10.5, 10.5, 10.5, 10.5]} objective: "revenue" (default) or "profit". See PROFIT OBJECTIVE above. forward_margin: Decimal margin in (0, 1], e.g. 0.18 = 18%. Only used with objective="profit"; required when the model has no stored operating margin. period_multiplier: Optional array of length num_periods converting KPI units to revenue per period over the planning horizon (mirrors the model's multiplier_column, e.g. price). include_historical_effect: Include carryover from historical spend in the predicted response (default True). enable_warm_start: Warm-start the optimizer from a previous solution (default True). optimizer_engine: "slsqp" (hardened SLSQP, default) or "marginal" (water-fill engine: allocates until every funded channel shows the same marginal return; exact profit-hurdle semantics and the tightest optimality certificates, with automatic SLSQP fallback). sigma_penalty: How gamma penalizes outcome spread: "std" (default), "variance" or "frozen" (advanced; smoother alternatives for hard-to-converge runs - leave on "std" normally). group_bounds: Joint constraints over channel SETS (#570), e.g. [{"name": "trade", "channels": ["TV", "Search"], "lower": 40, "upper": 60}] with lower/upper in % of total_budget (same convention as bounds). Groups must be disjoint and jointly feasible with the members' per-channel bounds. Presence forces the slsqp engine. Results gain GroupBounds/GroupBoundsReport columns; a BINDING group's members legitimately sit off the global marginal (they share the group's shadow price). |
| get_optimizer_resultsA | Get budget optimization status and results. Without run_id: returns the MODEL-LEVEL optimizer state. Top-level keys:
With run_id (run_optimizer's response includes it): fetches that specific
run, immune to later runs. Top-level keys include Reading
Args: model_hash: Hash of the model that was optimized. run_id: Optional optimization run id from run_optimizer's response. Pass it to poll a specific run's status/results. |
| get_scenario_templateA | Generate a forward-period scenario template from a completed model. Returns future dates pre-filled with values from 1 year prior, the list of media and control channels, and average cost-per-unit per media channel. IMPORTANT: Always call this before run_scenario or run_optimizer to discover:
The response also includes: operating_margin (the model's stored margin, if set — useful for profit math), variable_transforms (per-variable transform metadata), periodicity, and start_date. WARNING: Template data may contain NaN or null values for channels without historical data. You MUST replace NaN/null with 0 before passing to run_scenario, otherwise the prediction will fail downstream. Args: model_hash: Hash of a completed model. periods_forward: Number of future periods to generate (default 12). |
| run_scenarioA | Run a "what-if" scenario prediction on a completed model. Takes a set of future period rows with channel activity values and
predicts the KPI outcome. Use get_scenario_template first to get
the expected format, channel names, and baseline values. Channel names are
the activity-column keys from the template/results (e.g. "search_activity"),
not the IMPORTANT: Before submitting, replace any NaN/null values in scenario_data with 0. The template from get_scenario_template may contain NaN for channels without historical data, which will cause the prediction to fail. This is async (returns 202 with status "pending"). Poll get_scenario_results until status is "complete" or "failed". Workflow: get_scenario_template -> modify values -> run_scenario -> poll get_scenario_results Args: model_hash: Hash of a completed model. scenario_data: Array of period rows, each a dict with "Date" (YYYY-MM-DD format) and channel activity columns. Channel names must match exactly what get_scenario_template returns in the "channels" field. Example: [{"Date": "2025-01-06", "TV_Impressions": 50000, "Search_Clicks": 1200}] spend_metadata: Optional per-channel spend info for ROI calculation in results. Each entry: {"channel": "TV_Impressions", "metric": "impressions", "cpm": 25.0, "total_spend": 125000, "weekly_spend": [25000, 25000, ...]} rebuild_model: Recompile the model graph before prediction. Must be True (default) for API-initiated scenarios where the model graph is not in memory. evaluate_holdout: Evaluate the scenario against held-out actuals when the scenario period overlaps observed data (default False). skip_slicing: Skip per-channel contribution slicing in the prediction output — faster when only the KPI total is needed (default False). proxy_channels: Optional list of proxy-channel mappings, each mapping a scenario channel to a fitted channel whose transforms it borrows (for channels without their own history). |
| get_scenario_resultsA | Get scenario prediction results. Without run_id: returns the MODEL-LEVEL scenario state — status (pending/complete/failed) and, when complete, the full prediction data including predicted KPI per period, channel contributions, confidence intervals, and base components (intercept, seasonality, trend). This reflects the LATEST scenario on the model — a newer run overwrites it, so a poller can lose sight of the run it submitted. With run_id (run_scenario's response includes it): fetches that specific
saved run, immune to later runs — keys include NOTE: Failed scenarios return status "failed" with an error message in the JSON body (not an HTTP error). Always check the status field. Args: model_hash: Hash of the model the scenario was run on. run_id: Optional scenario run id ("scn_..."), from run_scenario's response or list_runs(artifact="scenario"). |
| update_runA | Rename / annotate a saved optimizer or scenario run. Runs are auto-named at creation (e.g. "$1.2M · 12mo · Jan 5"); renaming makes run history carry the analysis ("holiday cut -10%", "stretch 130%"). Renaming permanently flips the run's auto_named flag to false so future auto-naming never overwrites it. Only the fields you provide are changed. Args: artifact: "optimizer" (run_id "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model the run belongs to. run_id: The run's stable id from run history. name: New display name (non-empty when given; capped at 255 chars). notes: Free-text annotation. Omit to leave untouched; pass "" to clear. tags: Replacement tag list (max 20 tags, 64 chars each). |
| set_run_pinnedA | Pin or unpin a saved optimizer or scenario run. Declarative and idempotent: setting the current state again is a no-op, so scripts can safely re-run it. Args: artifact: "optimizer" (run_id "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model the run belongs to. run_id: The run's stable id from run history. pinned: Desired pin state. |
| list_runsA | List a model's saved optimizer or scenario run history. Returns {model_hash, runs, count, limit, offset}. Each run summary has: run_id, name, auto_named, pinned, notes, tags, status, error_details, progress fields while running, key_metrics (optimizer: total_budget, num_periods, gamma, predicted_revenue/roi, ...; scenario: num_periods, total_planned_spend, predicted_outcome, ...; null metrics are omitted — treat every key as optional), and created/started/completed timestamps. Ordering is pinned-first, then newest-first. CAVEATS:
Use get_optimizer_results / get_scenario_results with a run_id to fetch a listed run's full inputs and results; update_run / set_run_pinned to curate it. Args: artifact: "optimizer" (run ids "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model whose run history to list. limit: Page size (API clamps to 1-200; default 50). offset: Rows to skip (paging). |
| list_studiesA | List project-owned studies, questions, budgets and access rights. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised. |
| create_studyA | Create a study owned by an existing project. Does not launch models or consume attempts. |
| get_studyB | Read a study and its optimistic concurrency version. |
| get_launch_eligibilityA | Read, before launching, exactly what launch_study_run would refuse with: can_launch plus blockers, each with the stable code (permission_denied, policy_not_found, policy_retired, study_inactive, revision_not_found, attempts_exhausted, concurrency_exhausted, unsupported_family, engine_changed, snapshot_not_executable), the human message and a next_action. Also returns the budget, the policy used (newest active when policy_id is omitted) and the revision with its engine state. Consumes nothing. For engine_changed, call refreeze_recipe_revision and launch the new revision; never retry the old one. |
| get_study_overviewA | Read a whole study at a glance in one small call: the study, its budget, each recipe with its latest revision (id, number, hash), run counts by state and the latest run, active policies with newest_active_id, the champion summary and the last decision. Identity and counts only, no frozen configuration or reports. Call this first, then expand exactly the recipe, run or assessment you need (list_study_recipes with expand, get_recipe_revision, list_study_evaluations with expand). |
| update_studyA | Replace owner-controlled study settings. Send the full object: every field is assigned, so read the study first and pass its current values plus |
| list_study_recipesA | List recipes with their revisions as summaries (id, number, content_hash, reason, created_at). Heavy fields travel only by name in expand: effective (exact frozen priors, settings, data hashes and runtime), inspection (which settings were authored vs defaulted, inert prior fields with their gate, engine state current/stale that predicts whether launch will be refused, and lineage: the recorded dataset with available checked for you and the display line people see; treat inert values as stored-but-unused, not as recipe errors), specification (the authored request). Prefer get_recipe_revision for one revision in full. Raw datasets are never included. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised. |
| create_study_recipeA | Freeze a recipe without fitting. Supply source_revision_id when deriving from a published same-study recipe to retain influence ancestry. Optional expected_content_hash binds the validated effective inputs; a mismatch returns 409 and requires a fresh preview. Specification kind api_mmm has request containing create_model API fields; model_snapshot has model_hash and is review-only. |
| revise_study_recipeA | Create an immutable revision. Earlier versions inherit influence automatically; optional source_revision_id additionally links a same-study source recipe. Optional expected_content_hash guards effective inputs (409 requires a fresh preview). Supply the current recipe version; stale edits are rejected (412). Reload list_study_recipes, reconcile changes, then submit the current version; never overwrite blindly. |
| get_recipe_revisionB | Read one immutable revision with its inspection block. settings..value is what a fit of this revision would use, including fitter defaults for absent keys (status authored or default); effective.form_data is what was authored. priors[].fields lists conditional prior fields that were stored but never read for that row's adstock type or the saturation family (status inert, with the gate that would make them live); treat them as stored-but-unused and do not "fix" them by editing unless the gating setting changes too. engine.state current or stale predicts whether launch will be refused. Absent configuration classifies nothing. inspection.lineage (jellyfish #896) is the dataset the model was built from, as recorded — origin {kind, id, pipeline_id, version, sha256, pipeline_name, dataset_name} — with available checked now for you (true: the recorded source is still yours and hashes the same; false: gone or changed, with reason; null: nothing was recorded, or checked is false), and display, the same line people see on the recipe card ("Retail weekly · v3 · verified", "dataset lineage not recorded · legacy"). Names are display only; ids and sha256 verify. Nothing is inferred from a pipeline name or a later output. |
| diff_recipe_revisionsA | What changed between two revisions of one recipe, with the viewer's labels (jellyfish #881): settings[] (key, section, label, from, to), priors[] (row, parameter, column, label, from, to), data (dataset origin and input-hash change, or null) and counts {settings, priors, data, total}, plus base and other {number, id}. Blank equals absent; prior rows are matched by variable and role. Read this after publishing revision N+1 to state exactly what an edit changed. 404 when the recipe or either revision is missing. Read-only. |
| refreeze_recipe_revisionA | Recover from engine_changed without bypassing the safeguard: creates a new revision of the same recipe with the same specification, frozen on the current engine, with an auto-filled reason naming what changed. The old revision and its provenance are untouched. Returns the new revision; launch that one. 409 conflict when the revision is already frozen on the current engine; 409 snapshot_not_executable for review-only snapshots. |
| validate_study_recipeA | Resolve and validate a recipe without creating a run. Returns effective settings, provenance limits and the same inspection block a saved revision would carry (authored/default settings, inert prior fields with their gate, engine state, lineage with the recorded dataset's availability and display line), so an agent can check for inert fields and an unavailable dataset before freezing. |
| launch_study_runA | Launch a NEW fit of a frozen revision within study attempt/concurrency budgets; it never opens existing results. Call get_launch_eligibility first: a refusal here carries the same code, message and next_action. Requires an active study, an executable revision frozen on the current engine, and an active same-study policy (choose it deliberately; the newest is not always the intended one). Budget/state conflicts require inspection, not a new attempt key. Reuse the same submission_key after an ambiguous response; never invent another key for a retry. engine_changed means re-freeze (refreeze_recipe_revision) and launch the new revision. |
| list_study_runsA | List preserved attempts including pending and failed runs. Supporting backends also return budget with attempts remaining, available slots and blocking reasons; the budget always counts every run even when rows are paged. Missing budget means unknown support, not permission to launch. Capacity is rechecked on reservation; recover an uncertain launch with its original submission key. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised. |
| get_study_runA | Read durable run status and the linked model. |
| cancel_study_runB | Request cancellation. Requested and confirmed stopped are distinct states. |
| list_quality_policiesA | Read immutable quality policies for the study. Each row carries the full specification plus identity: content_hash (the stored row, including name and rationale), rules_hash (the rules alone: checks sorted by metric plus the protocol, so two policies with the same rules_hash apply the same rules whatever they are called), is_newest, derived_from, created_at, retired_at, usage counts (runs_launched, evaluations, resolutions, champion_acceptances) and checks_summary / protocol_summary. Newest is information, not a recommendation: choose policy_id explicitly. Retired policies stay listed for history but are refused for new launches, assessments and pair reviews. Deleting a policy is possible only for a policy nothing references and only from the signed-in project owner UI, never through this API key. |
| diff_quality_policiesA | What changed from one saved policy (policy_id) to another (other_policy_id) in the same study, as the server computes it: checks added, removed and changed (per metric and field, from → to; matched by metric so reordering is not a change), protocol field changes, name_changed, rationale_changed and rules_changed (whether the rules hash differs). Read scope; nothing is written. Use it to explain a policy's lineage (get_quality_policy.derived_from) or to compare any two policies before choosing one. |
| get_quality_policyA | Read one immutable policy in full: specification, content_hash (identity of the stored row), rules_hash (identity of the rules alone, the value evidence carry-forward compares), is_newest, derived_from ({policy_id, content_hash, name} when created from another policy via derived_from_policy_id; name and hash are null if that source was since deleted; see diff_quality_policies), checks_summary (builtin / custom_numeric / boolean / manual / diagnostic, required, advisory, cap) and protocol_summary, plus usage with the ids of every run, assessment, resolution and Champion acceptance that references it. Shared viewers can read; nothing is written. |
| create_quality_policyA | Save project-specific checks. Policies are immutable: to change one, create a new policy, ideally with derived_from_policy_id naming the saved policy you started from (same study) so lineage is kept and the response carries diff (checks added/removed/changed by metric and field, protocol changes, name_changed, rationale_changed, rules_changed) — reordering checks is not a change. Built-in checks have metric (r_hat_max, mae, rmse, wape, prediction_mae, prediction_rmse, prediction_wape), maximum and required, and may instead use operator gte/between with minimum (default lte). Artifact-backed checks read what the fit already saved and need no external evidence: saved diagnostics (r_squared, mape, durbin_watson, pareto_k_pct, normality_p, loo_cv), retained-sampling counts (retained_chains, retained_draws_per_chain, ess_bulk_min, ess_tail_min, divergences) and provenance status (provenance:holdout, provenance:prior with expected review_required; blocked fails, absent is not_collected). Custom numeric checks use metric custom:, name, units, operator (lte/gte/between), applicable minimum/maximum and required. Boolean checks use kind=boolean, operator=equals and expected=true/false. Manual checks use kind=manual, equals, expected=true; agents can define these but cannot submit manual sign-off. Custom bounds may be negative. WAPE is a fraction. Prediction-window checks require saved finite actuals/predictions at unique dates after the saved training window; this does not certify untouched holdout provenance. No default thresholds are assumed. Declare at least one required check, use each metric once, and set maximum R-hat at least 1. At most 20 checks in total; the refusal names the count. A bad check is refused with one message naming its position and declared kind. Optional validation_protocol declares a temporal holdout split, configured sampling minima, R-hat and prediction WAPE limits before both runs launch under this policy. The backend validates policy rules. |
| retire_quality_policyA | Retire (retired=true) or restore (retired=false) a saved policy without changing it. A retired policy is refused for new launches, new assessments and new pair-review resolutions, stops counting as the newest policy, and stays readable in every run, assessment, pair review, comparison and Champion record that already names it. Idempotent; the backend records each change. Requires create:models on the study's project. Nothing is deleted: removal of an unreferenced policy is an owner-only frontend action. |
| evaluate_study_runA | Save an immutable assessment, or with preview=true see what one would contain without writing (nothing stored, no access event). The preview returns basis_hash, would_evaluate, would_not_evaluate (each with basis.reason and, when an earlier assessment of this run can lend the value, carry_forward_available; otherwise carry_forward_blocked with why: basis_changed, policy_changed, manual_signoff) and previous_assessment. Each submission is complete: a custom check has a value only if you supply it now in external_evidence or name it in carry_forward from an earlier assessment whose evidence basis and policy rules still match; a metric may not be in both. Submit finite numeric or strict boolean values with method and source_reference and the expected_basis_hash from the preview; stale evidence is 409 stale_evidence. External calculations are submitter-reported, not verified; carried rows keep carried_from provenance. Manual sign-off is never carried and requires a signed-in reviewer. Every unevaluated check carries basis.reason from a closed set and a suggested action; not_evaluated never passes. Choose policy_id explicitly; the report echoes policy_name and policy_was_newest. No automatic champion promotion. Built-in errors are fitted-window, not holdout; VAR remains unsupported. |
| list_study_evaluationsA | List preserved assessments as summaries: id, policy_id, policy_name, status, basis_hash, evidence_hash, created_at and summary (evaluated / total / required_unevaluated / evidence_sources: which values were supplied and which carried). A newer sparse report does not replace an earlier enriched one; decisions bind to a specific report id. Pass expand=["report"] for the full per-check report. Listing never records prediction access, with or without expand (jellyfish #837): a stored report carries window metadata and aggregate errors, not prediction rows. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised. |
| list_study_decisionsA | Read analyst decisions and agent recommendations. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised. |
| get_study_championA | Read incumbent, eligibility blockers, accepted candidates and immutable champion history. Stale champions retain their historical role with review_required. Reported holdout use that informed a candidate revision blocks that revision pending fresh validation; holdout_use lists the declaration IDs. Ordinary viewing does not block. Current analyst-reviewed validation resolutions can clear the exact run acceptance; changed evidence or revoked review reblocks it. Recorded revision ancestry inherits influence. Validation references are reviewer-declared; decision_grade_ready is false until independently qualified. Selection/replacement/revocation require an owner frontend session; MCP cannot promote models. |
| get_study_prediction_accessA | Read partial prediction-access history for this run and matching recorded dataset/windows in this study: recorded_accesses, recorded_runs, by_action (what the total is made of) and the 20 most recent events. An event is recorded only when prediction evidence is served by a deliberate action: saving an assessment (assessment), a comparison (comparison, one per served run), a pair assessment (pair_assessment), dashboard results (dashboard_results) and prediction-window exports (results_csv, results_json). Listing assessments never records. Histories written before jellyfish #837 may also contain assessment_history rows from list reads; nothing is deleted because holdout-use declarations reference event ids. This read does not expose predictions or add events. Earlier activity, other result routes and offline work are not covered; absence never proves untouched holdout status. Repeated access does not prove retuning. |
| get_study_validation_resolutionsA | Read analyst validation resolutions and revocations, including current/stale/revoked status re-evaluated against exact evidence. API keys cannot supply human independence sign-off or revoke it; use the signed-in owner UI. No audit serving event is added and no model is fitted or promoted. |
| declare_study_holdout_useA | Append a submitter-reported evidence-use declaration using project-owner credentials and create:models. Read get_study_prediction_access first and reference an access event from this run. Use a fresh UUID declaration_id and reuse it unchanged on retry. informed_revision requires a published affected revision in the same study; other dispositions omit it. Reason must explain actual use. Reported revision influence requires fresh validation for affected revisions; later review-only notes cannot erase it. Does not certify independence, accept or promote a model. API submissions remain identified as reported declarations. |
| assess_study_validation_pairB | Assess a validation pair and append a prediction-access audit event when evidence is available. Checks distinct completed MMM runs launched under the declared protocol, frozen inputs/settings/runtime, configured sampling, saved R-hat, declared prediction windows/WAPE and date coverage. Returns blockers and an evidence hash; does not fit, accept or promote. Includes saved retained chain/draw, ESS and divergence records when available, with null for older models. Optional prelaunch retained_sampling limits require complete native records and check chain/draw minima, bulk/tail ESS minima and maximum divergences; otherwise sampling_qualification is not_declared. Optional require_policy_review checks current signed-in analyst acceptance of each latest same-policy assessment, including freshness and rejection blockers. The holdout_provenance report distinguishes missing evidence, blocked version 1 full-input preprocessing and version 2 training-only preprocessing requiring further provenance review. The prior_provenance report checks which rows the recorded smart priors read and the frozen input hashes; priors without a record remain unavailable, and priors that read observations after the declared training end are blocked. fresh_validation provides replacement-window preflight for influence reports naming the full-model revision: later windows, replacement policy chronology, recorded prior exposure, retained diagnostics and provenance/review requirements. It never clears champion blocks. External business calculations and untouched holdout history remain unverified; decision_grade_ready stays false. |
| recommend_study_runA | Record a recommendation with evidence. This does not accept or promote a model; analyst acceptance happens in the frontend. |
| adopt_model_into_studyA | Import an owned completed model into a study as an executable, editable recipe (jellyfish #880). Without confirm: a preview only, with the resolved recipe, its inspection block and an import report {complete, dataset: {recorded, origin}, settings: {count, not_recorded}, editable: fully | partly, evidence: {fitted_result}}; nothing is written. The preview's report.dataset also carries display, the lineage line the recipe card will show once imported ("Retail weekly · v3 · verified"; "dataset lineage not recorded · legacy" for a model built before capture). With confirm=true: creates a recipe whose base_model revision 1 carries the exact fit inputs, the wizard snapshot captured at build and the verified dataset origin, launches and re-freezes like any other revision and can be edited in place (get_recipe_revision_authoring, then create_recipe_draft with target); the fitted result is attached as an adopted run outside the attempt budget (201, {run, revision}). 409 when the model is not complete or is already attached to a study. Models built before snapshot capture import with not_recorded settings (editable: partly). Adoption never refits. |
| compare_study_runsA | Compare 2-20 candidates against one quality policy. Each row carries a basis record (family, dataset and costs hashes, outcome column, units, training window, output kind, prediction window, evidence freshness) and a per-dimension compatibility against the first row; rows are comparable only when family, dataset, outcome, units, window and output kind all match, and incompatible rows are returned with the differing dimension in blockers and are never ranked. This answers predictive ranking only: sensitivity agreement is not computed, analyst acceptance lives in decisions, and business validity is a human review. Serving prediction evidence appends access audit events. Does not fit or promote models. |
| get_workflow_guidanceA | Read versioned Simba workflow guidance when native Skills are unavailable. Use index to list allowed topic/section IDs, then request the relevant section. Topics: mmm, results, priors, optimiser, studies, var. The default section is entrypoint; it names detailed sections and when they are needed. Arguments are identifiers, never paths. Returns a complete section within a 24,000-byte structured-payload limit or a refusal, never truncated instructions. This local read makes no backend request. Guidance does not authorise writes or replace validation. Lookup is optional when the caller already has the relevant versioned guidance. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| ui://simba/charts.html | Read-only charts of served results, with exact-value tables. |
TDQS
Scored across 94 tools
With 94 tools spanning overlapping MMM, study, recipe, revision, incrementality, campaign, and pipeline domains, many tools share nouns and require long descriptions to distinguish. Boundaries between recipe drafts, study recipes, revisions, and incrementality design/save/create operations are not self-evident to an agent without careful reading. Several descriptions include warnings about confusing related tools, which further signals latent ambiguity.
Tool names consistently use snake_case with a predictable verb_noun pattern: list_*, get_*, create_*, update_*, run_*, set_*, and delete_*. Domain prefixes such as study_, recipe_, campaign_, pipeline_, and incrementality_ are applied consistently. Although many names are long, the convention is stable throughout.
94 tools far exceeds the typical well-scoped range of 3-15 and lands in the extreme-mismatch band. Even for a broad MMM platform, this surface is so large that tool selection, documentation, and maintenance become unwieldy. The count alone indicates over-exposure rather than a tight, purpose-built tool set.
The surface is extensive, covering data upload/reporting, model creation and results, optimizers, scenarios, studies, recipes, quality policies, evaluations, champion tracking, campaigns, pipelines, and incrementality tests. There are some intentional gaps, such as deletion only for failed models and no API deletion for projects or unreferenced policies, but agents can generally work around these. Overall lifecycle coverage is strong, though not perfectly complete.