Simba MCP Server
OfficialSimba MCP Server lets an AI assistant run the full Bayesian Marketing Mix Modeling workflow — data in, models fitted, results read, budgets optimized, experiments designed, and studies governed — against a Simba backend.
Data & pipelines — Upload CSVs and inspect uploads/schemas, report dataset slices by window/grain/role, and list, run, poll and schedule data pipelines.
Model building — Create and fit MMMs and long-run VAR models, poll fit status/liveness, rename, save/unsave, delete failed models, organize projects, and link VAR models or contribution groups.
Results & charts — Read model results by section (channel summary, contributions, coefficients, curves, saturation, mROI, diagnostics, posterior, financials, model config, etc.) with windowing and size controls, plus native read-only charts for response curves, decomposition and optimizer allocations.
Optimization & scenarios — Run bounded budget optimizations (revenue or profit), build and run what-if scenarios, and curate saved run history (list, rename, annotate, pin).
Campaign analysis — Map campaigns to model channels, report campaign facts, derive incremental ROAS/marginal returns, and calculate campaign budget recommendations without applying them.
Incrementality tests — Record, import and list geo/owned-media/platform-lift tests, design detectability studies, and save designs as planned tests for calibration.
Studies & governance — Manage studies, recipe drafts and immutable revisions, quality policies, launches and cancellations, evaluations, run comparisons, champion/validation reads, prediction-access and holdout-use declarations, and validation-pair assessment (acceptance stays with a signed-in human).
Discovery & guidance — Probe backend capabilities and read versioned workflow guidance (or packaged Skills) to plan work correctly.
Imports Meta Conversion Lift experiment results (Conversion Lift API JSON) and records Meta platform-lift incrementality tests, which can then be attached to and calibrate Bayesian Marketing Mix Models in Simba.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Simba MCP ServerShow me the channel contributions and ROI for my latest model"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Simba MCP Server
Simba is a Bayesian Marketing Mix Modeling (MMM) platform. This Marketing Mix Modeling MCP server lets AI assistants interact with your models directly — upload data, build models, check results, and run budget optimizations through natural language in Claude, Cursor, or Claude Code.
Installation
For job-specific startup catalogues, see tool profiles. For independently installed native Skills or bounded MCP fallback, see workflow guidance. Connecting MCP does not install Skills. Current campaign, experiment-screening, chart and actual-data examples are packaged with that guidance; candidate acceptance records tested configurations and remaining rollout checks.
pip install simba-mcpOr run directly without installing:
uvx simba-mcpRelated MCP server: Ads MCP
Quick Start
Cursor IDE
Add to your Cursor MCP settings (.cursor/mcp.json in the workspace or global settings):
{
"mcpServers": {
"simba": {
"command": "uvx",
"args": ["simba-mcp"],
"env": {
"SIMBA_API_URL": "https://demo.simba-mmm.com",
"SIMBA_API_KEY": "simba_sk_..."
}
}
}
}Claude Code
Add to your Claude Code MCP config:
{
"mcpServers": {
"simba": {
"command": "uvx",
"args": ["simba-mcp"],
"env": {
"SIMBA_API_URL": "https://demo.simba-mmm.com",
"SIMBA_API_KEY": "simba_sk_..."
}
}
}
}Claude API (MCP Connector)
Use the remote Streamable HTTP transport with the Anthropic MCP connector:
import anthropic
client = anthropic.Anthropic()
response = client.beta.messages.create(
model="claude-sonnet-4-6",
max_tokens=4096,
messages=[{"role": "user", "content": "List my Simba models"}],
mcp_servers=[
{
"type": "url",
"url": "https://demo.simba-mmm.com/mcp",
"name": "simba",
"authorization_token": "simba_sk_...",
}
],
tools=[{"type": "mcp_toolset", "mcp_server_name": "simba"}],
betas=["mcp-client-2025-11-20"],
)Configuration
Operator settings, defaults and the non-sensitive effective check are in docs/configuration.md. Inspect the current process without printing credentials:
python -m simba_mcp.configurationA development .env configures only the process that loads it. It does not configure
a remote server or every MCP host. This release does not add an application settings UI.
Native result charts
Three read-only native chart tools present response curves, decomposition and saved allocations through MCP Apps. Structured and text results remain available without visual support. The compatibility page distinguishes fixture tests from actual client acceptance.
Available Tools
The full, generated reference — every tool, whether it reads or writes, its description and its
parameters — is docs/tools.md. It is rendered from the running server
(python -m simba_mcp.reference) and a test keeps it current, so it never drifts from what the
server serves. The tools cover:
Data and pipelines — the canonical CSV schema, uploads, pipeline versions, backend capabilities
Projects and models — create, fit (MMM and long-run VAR), poll, save, rename, link VAR models, contribution groups
Incrementality tests — record, list and import geo, owned-media and platform-lift results; the calibration row each test gives a model; calibrated fits
Results — ROI, contributions, response curves, diagnostics and more, with section and size controls
Scenarios and optimization — scenario templates and runs, budget optimization, saved-run history
Studies — questions, recipes and revisions, drafts, launches, quality policies, evaluation, validation pairs, comparisons and recommendations (acceptance stays with a signed-in person)
Example Prompts
Try these with any connected AI assistant:
Explore your models:
"List my Simba models and show me the channel ROI summary for the most recent complete model."
Build a model:
"Upload this CSV data to Simba and create a new MMM model with TV, Search, and Social as media channels. Use 'revenue' as the KPI and 'date' as the date column."
Check progress:
"What's the fitting status of model a1b2c3d4?"
Get results:
"Show me the model diagnostics and channel contributions for model a1b2c3d4."
Optimize budget:
"Run a budget optimization on model a1b2c3d4 with $1M total budget over 12 months. Set TV bounds to 5-40% and Search to 10-50%. Use uniform laydown weights."
Response curves:
"Show me the response curves for model a1b2c3d4. At what spend level does TV hit diminishing returns?"
Scenario planning:
"Get a scenario template for model a1b2c3d4 for the next 12 weeks. Then run a scenario where I increase TV by 20% and cut Search by 10%. What happens to revenue?"
Full workflow:
"I have marketing data I want to analyze. First get the schema so I know what format is needed, then upload my data, create a model, and once it's done show me the ROI by channel."
Agent Skills
Workflow guidance is available as self-contained native Skills and through
get_workflow_guidance for clients without native Skills. Both use the same versioned,
packaged source. Connecting the MCP endpoint does not install native Skills.
Skill | Task |
Upload, create and recover a fit | |
Saved results, channel identity, units, intervals and windows | |
Media/control priors, anchor pairing and resolved settings | |
Budget constraints, objectives and saved runs | |
Drafts, immutable revisions, launch recovery and review | |
VAR creation, linking and long-run rollups |
See installation, fallback lookup and maintenance.
Tool descriptions default to the full legacy catalogue. Set
SIMBA_TOOL_DESCRIPTIONS=compact before starting the server to opt into shorter
descriptions for model creation, results and optimisation. Names, schemas and
execution stay the same. See configuration, rollback and comparison evidence.
Optional tool profiles provide marketer and reviewer views.
On a supporting Simba backend, the account's choice under Profile > Connected apps
also narrows hosted tool discovery and dispatch within the operator's catalogue.
Use simba-mcp --profile marketer or SIMBA_TOOL_PROFILE=reviewer at startup.
full remains the default; data_scientist includes every tool and the entire Studies
lifecycle. Profiles are starting views, not backend permissions.
An experimental native-discovery host example compares provider tool search against the compact eager catalogue using synthetic backends. It is a paid, explicitly invoked evaluation command and does not enable discovery in your connected MCP client. Its evidence includes task failures and fallback limits. Existing-result questions can start with selected result sections directly; capability, schema and guidance discovery are conditional on the task. Backend validation and human approval remain authoritative.
Gotchas & Tips
Things that commonly trip up both AI agents and humans:
Hosted server: your bearer token IS your login
On HTTP deployments each request is authenticated with the caller's own
Authorization: Bearer simba_sk_... token — there is no server-side shared
key. If tool calls return "No API key on this request", your MCP client
isn't sending the token (check the authorization_token / headers setting
in its config).
OAuth (optional, from 0.12.0). A deployment can also run in OAuth resource-server mode
(MCP_OAUTH_ENABLED=1). Clients whose connector UI expects OAuth — claude.ai custom connectors and
ChatGPT connectors — then sign in through the Simba backend and get a token; the server verifies it
against the backend and forwards it unchanged. An API key still works everywhere, and it is the
method for Claude Code, Cursor and the Claude API connector. Step-by-step connection guides per
client are in the product documentation.
Channel names are exact-match
Model results are keyed by the channel's activity column name (e.g. "search_activity", "TV_impressions"), not by the channels[].name you passed to create_model. Keys can contain spaces and matching is case-sensitive and space-sensitive — the optimizer and scenario tools use them as dictionary keys.
Always call get_model_results with sections="channel_summary" first to see exact channel keys, then use those verbatim in optimizer/scenario payloads.
Results sections
get_model_results serves these sections (request only what you need via sections=):
channel_summary, contributions (KPI/unit space — multiplier not applied), coefficients (per-period per-channel revenue table), params, decay_curves, response_curves, marginal_curves, saturation, mroi_summary (marginal ROI at current spend with 94% HDI; post-#591 fits add the allperiods_unweighted / spendweighted_active convention scalars, and post-#629 fits add a *_mean beside every *_median — the median is displayed, the mean is what reconciles with the marginal-revenue curve), mroi_periods (opt-in only — the per-period marginal ROI series; never in the default payload, request it by name), model_stats, actual_vs_model, long_run_rollup, optimizer, predictions, posterior, financials, model_config. The response's sections_available field is authoritative if the server is newer than these docs.
Models are identified by model_hash
All model endpoints use the string model_hash (e.g. "f835671a25") returned by create_model and list_models.
API-key management is deliberately not exposed
The /api/v1/keys endpoints (create/list/revoke API keys) are session-auth only and have no MCP tools by design: a server holding one key must not be able to mint or revoke keys. Manage keys in the Simba UI (Profile → API Keys).
Optimizer arrays, not scalars
laydown_weights and period_cpm must be objects of arrays, each array having exactly num_periods elements:
// Wrong
"period_cpm": {"TV": 10}
// Correct
"period_cpm": {"TV": [10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10, 10]}The same channel keys must appear in bounds, laydown_weights, and period_cpm. Bounds values are percentages (0-100) of total_budget, not currency amounts.
Clean NaN from scenario templates
The template from get_scenario_template may contain NaN/null for channels without historical data. Replace them with 0 before passing to run_scenario:
import math
for row in scenario_data:
for key, val in row.items():
if val is None or (isinstance(val, float) and math.isnan(val)):
row[key] = 0Three endpoints are async
These return 202 and require polling:
Action | Start | Poll |
Fit model |
|
|
Optimize |
|
|
Scenario |
|
|
Poll every 5-10 seconds. Check the status field for "complete" or "failed".
Data upload requirements
CSV only (not Excel). Maximum 10 MB (API-enforced).
Row minimum: check
get_data_schema→x-simba-constraints.min_rows; the upload response'swarningsfield is authoritative. More rows = tighter posteriors (104+ weekly rows recommended).Media columns:
{channel}_activityand{channel}_spendper channel.Use
0for inactive periods, not blank or NA.Large file? Pass
csv_path(a local file path) instead ofcsv_content— the server reads it directly instead of the CSV going through the conversation. Local (stdio) servers only; disabled on HTTP/SSE deployments unlessSIMBA_MCP_ALLOW_LOCAL_FILES=1.
Common Errors
Error | Cause | Fix |
| No API key or expired key | Check |
| Key doesn't have the needed scope | Create a key with all scopes |
| Payload missing required keys | Check the tool's parameter list |
| Model still fitting or failed | Poll |
| Scalar instead of array, or wrong length | Use arrays matching |
| Zero or negative CPM | All CPM values must be > 0 |
| Mismatched channel names | Same keys in bounds, laydown_weights, and period_cpm |
| Column name typo | Check CSV headers match exactly |
| CSV too large | Reduce file size or aggregate data |
Direct API Access
The MCP server wraps the Simba REST API. For scripting, CI/CD, or environments without MCP, you can call the API directly.
When to use MCP vs direct API
MCP (via AI assistant) | Direct API (curl / Python) | |
Best for | Exploratory analysis, conversational workflows | Automated pipelines, scheduled jobs, scripts |
Async polling | Assistant handles it automatically | You implement poll-until-complete logic |
Data cleaning | Assistant cleans NaN/null, builds payloads | You write the data prep code |
Reproducibility | Conversational | Scriptable, version-controlled |
Both use the same API keys with the same scopes.
Quick start (Python)
import requests, time
BASE = "https://demo.simba-mmm.com"
HEADERS = {"Authorization": "Bearer simba_sk_..."}
# Upload data
with open("marketing_data.csv", "rb") as f:
r = requests.post(f"{BASE}/api/v1/ingest",
headers={**HEADERS, "Content-Type": "text/csv"},
data=f.read(), params={"name": "q1_data"})
file_id = r.json()["id"]
# Create model
r = requests.post(f"{BASE}/api/v1/models", headers=HEADERS, json={
"data_source": {"uploaded_file_id": file_id},
"date_column": "date",
"kpi_column": "revenue",
"hierarchy_column": "brand",
"channels": [
{"name": "TV", "activity_column": "tv_grps", "spend_column": "tv_spend"},
{"name": "Search", "activity_column": "search_impressions", "spend_column": "search_spend"},
],
"total_media_effect": "Retail",
})
model_hash = r.json()["model_hash"]
# Poll until complete
while True:
status = requests.get(f"{BASE}/api/v1/models/{model_hash}/status",
headers=HEADERS).json()
if status["status"] in ("complete", "failed"):
break
print(f"Fitting... {status.get('progress', '?')}%")
time.sleep(10)
# Get results
results = requests.get(f"{BASE}/api/v1/models/{model_hash}/results",
headers=HEADERS,
params={"sections": "channel_summary,model_stats"}).json()
for ch in results["results"]["channel_summary"]:
print(f"{ch['Channel']}: ROI {ch['ROI']:.1f}")Quick start (curl)
API_KEY="simba_sk_..."
BASE="https://demo.simba-mmm.com"
# Upload data
curl -X POST "$BASE/api/v1/ingest?name=q1_data" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: text/csv" \
--data-binary @marketing_data.csv
# Create model (replace uploaded_file_id with id from upload)
curl -X POST "$BASE/api/v1/models" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"data_source": {"uploaded_file_id": 1}, "date_column": "date", "kpi_column": "revenue", "hierarchy_column": "brand", "channels": [{"name": "TV", "activity_column": "tv_grps", "spend_column": "tv_spend"}]}'
# Poll status (replace MODEL_HASH)
curl "$BASE/api/v1/models/MODEL_HASH/status" -H "Authorization: Bearer $API_KEY"
# Get results
curl "$BASE/api/v1/models/MODEL_HASH/results?sections=channel_summary,model_stats" \
-H "Authorization: Bearer $API_KEY"API Key Setup
The MCP server authenticates with the same API keys used by the Simba REST API. Create a key with the required scopes:
Go to Profile > API Keys in the Simba UI
Click Create Key
Set scopes:
ingest,read:models,read:results,create:models,optimize,scenarioCopy the key (shown only once)
How the key is supplied depends on where the server runs:
Local (stdio — Cursor, Claude Code): set it as the
SIMBA_API_KEYenvironment variable in your MCP config (the examples above).Hosted (
https://demo.simba-mmm.com/mcp): send it as the HTTPAuthorization: Bearerheader — theauthorization_tokenfield in the Claude MCP connector config. Every caller uses their own key (v0.2.2+): the server never shares an identity between callers, a request without a key gets a structured 401 with guidance, and you only ever see your own account's models.
Configuration
Environment Variable | Description | Default |
| Simba API base URL |
|
| Total/phase timeouts and process/caller admission policy. Unset disables it; a complete JSON object enables one deployment's values. See request budgets | Disabled |
| Optional maximum response entity bytes before decompression | Unset (uncapped for compatibility) |
| Optional maximum response bytes supplied to JSON/CSV parsing | Unset (uncapped for compatibility) |
| Your Simba API key (stdio mode only; HTTP callers send their own key as the bearer token) | (required for stdio) |
Set response byte ceilings to positive integer byte counts. If you set only one,
the same ceiling applies before and after decompression. These opt-in limits refuse
oversized responses without returning partial evidence. Review typical full-result
sizes before choosing limits; the results tool's max_response_bytes option is a
separate later cap on selected output.
Transport Modes
The server supports all MCP transport modes:
# stdio (default) — for Cursor, Claude Code
simba-mcp
# Streamable HTTP — for remote deployment
simba-mcp --transport streamable-http --port 8100
# SSE — legacy transport
simba-mcp --transport sse --port 8100
# Or via uvicorn directly
uvicorn simba_mcp.server:app --host 0.0.0.0 --port 8100License
MIT
Fit heartbeat visibility
On backends that support it, get_model_status also returns fit_liveness.
The MCP server passes the response through unchanged; older backends may omit
this field. With available: true, heartbeat_age_seconds,
stall_timeout_seconds, and seconds_until_stall_threshold are in seconds;
last_heartbeat_at is a Unix timestamp in seconds. stall_threshold_exceeded
reports whether heartbeat age is strictly greater than the configured threshold.
The countdown is to a stale-heartbeat threshold, not completion ETA or an exact
termination time: the watchdog runs periodically. It does not change status.
With available: false, reason is heartbeat_unavailable (including unavailable
heartbeat storage) or not_fitting. Missing or unavailable metadata is unknown,
not evidence that a fit is healthy or stalled. Continue polling with backoff and
use the reported model status; do not automatically restart or duplicate a fit.
Explicit control priors and transforms
Use control_columns to include controls and optional control_priors to
configure them. Overrides select an exact column via control; media priors
continue to select a channel via channel.
{"control_columns": ["price", "discount_depth"], "control_priors": [
{"control": "price", "transform": "LOG", "distribution": "normal", "mean": -1, "sd": 0.5},
{"control": "discount_depth", "transform": "N", "distribution": "normal", "mean": 1, "sd": 0.5}
]}These are illustrative coefficients, not fitted estimates or recommended priors. Native transforms: N = x; DM = x/mean(x); STA = x/sample_sd(x) without centering; DDM = x/mean(KPI); LOG = log(x/mean(x)), requiring positive x. Under a log link, a LOG coefficient is an elasticity; N is a semi-elasticity. Changing a transform does not convert prior units: explicitly choose suitable coefficient priors. Normal and inversegamma use mean/sd (positive mean for inversegamma); truncatednormal also requires ordered lower/upper bounds; halfnormal uses only sd. Unused fields, unknown keys, null field values and nonselected/duplicate controls are rejected. Omitted fields preserve applicable smart defaults; changing family clears inapplicable bounds/mean.
Nonempty overrides preflight get_data_schema for
x-simba-model-capabilities.control_priors version 1 and all five transforms.
Unsupported, malformed or unavailable capability checks stop before creation;
never retry by removing requested settings. Omitted/None/empty overrides keep
legacy calls unchanged. Backend support must be deployed before using this
option. Direct REST clients must perform the same capability check and put
control_priors at the request root. Check get_model's resolved priors and
overridden_fields after creation. Prediction uses existing fit-time constants;
this feature does not change preprocessing split order or fit a model for you.
Study workflows
The study tools require a Simba backend with the workflow API deployed. Studies belong to projects and provide a shared record for analysts and agents:
Create/read/update studies with a declared question (what to learn), optional context (scope and assumptions), attempt limits and concurrency limits. Question and context are descriptive text; no launch or evaluation path reads them, and acceptance rules live only in quality policies.
Validate, save and inspect immutable recipe revisions; list recipes captured by the wizard; edit a recipe in place through a draft that names its target, and diff two revisions.
Launch a revision with a declared quality policy and caller-generated submission key.
Inspect progress, request cancellation, evaluate saved evidence and compare candidates.
Preview/import existing models and record recommendations with a rationale.
Reuse the same submission key when retrying an uncertain launch; changing the recipe or policy requires a new key. Workflow writes are sent once, without automatic HTTP retries. A study does not autonomously launch its budget of fits. Its runs use the existing models, workers and progress records.
Project sharing permits reading studies; mutations require ownership. The MCP
can recommend but cannot record analyst acceptance. adopt_model_into_study
imports a fitted model as an executable, editable revision 1 (legacy snapshot
imports stay review-only). Wizard-captured and imported recipes can be inspected,
launched and edited in place: read the snapshot with
get_recipe_revision_authoring, save a draft with target, publish with
target for revision N+1.
Quality reports distinguish failed checks, missing evidence and required analyst review. Current saved-window error metrics and R-hat are not held-out validation or proof of business validity. No tool automatically promotes a winning model.
Discovery and reliable Studies workflows
Start with get_backend_capabilities to read the connected backend's model,
transform/prior and workflow advertisements. Missing fields mean unknown;
installing this package does not upgrade the backend. All tool annotations are
informational hints, never permission checks.
For Studies: inspect the project/study budget, validate a recipe, freeze a revision,
declare a quality policy, then launch with an explicit submission key. Preserve
that key and the exact inputs after an uncertain response. Poll the shared run;
requested cancellation is not confirmed completion. Reload and reconcile on 412;
revalidate on an input-hash conflict. Optional expected_content_hash on recipe
create/revise binds the validated effective inputs. Evaluate existing evidence and
recommend with limitations; analyst acceptance remains in the frontend. Missing
evidence never passes, and fitted-window metrics are not holdout validation.
Writes are sent once, without automatic retries. Reconcile uncertain mutations
before repeating them. Reads retain bounded transient retries. Existing error
objects retain error / _status_code, with additive _error_code and
_next_action guidance. Backend additive fields remain intact in structured output.
For bounded results, request sections="channel_summary,model_stats" first and use
channels / max_grid_points where appropriate. Optional max_response_bytes
returns an actionable 413 instead of partial evidence when the filtered JSON
payload is too large. It bounds payload serialization, not backend download or MCP
envelope overhead. Existing defaults remain unchanged.
Every tool, its parameters and whether it reads or writes: docs/tools.md (generated from the
running server — python -m simba_mcp.reference).
See architecture and compatibility for ownership, transport/authentication boundaries, known limits and validation.
Draft authoring discovery: get_recipe_draft_template(family="mmm" | "var") returns complete defaults generated by the shared wizard, a template hash and the draft envelope schema. Start new drafts from this snapshot and preserve unedited fields. Defaults are not validated models; check publication capabilities before publishing. Requires the corresponding backend capability and read:models.
Pass either uploaded_file_id or pipeline_version_id to get_recipe_draft_template to copy an owned uploaded dataset or saved pipeline output version into the new authoring snapshot and populate its bounded preview. Optional source origin is verified by the backend against exact bytes; detail responses include a source manifest. Reopening uses frozen data even if the original upload disappears. This does not enable publication or model fitting.
Draft publication: publish_recipe_draft freezes a saved MMM or VAR draft as an atomic batch, using expected version and caller UUID recovery; it never launches a fit. With target={recipe_id, expected_version} it writes revision N+1 of an existing recipe instead of a new one (412 stale_version keeps the draft and its edits; 409 target_requires_single_recipe for a multi-brand draft); create_recipe_draft with target={recipe_id, base_revision_id} opens that edit session. get_recipe_revision_authoring retrieves the authoring snapshot behind any revision that carries one (authoring_draft, wizard, base_model) and says whether its dataset is still available; diff_recipe_revisions states what changed between two revisions. Check backend publication capabilities. When publication_constraints.automatic_prior_resolution is freeze_at_publication, automatic MMM priors are resolved once and replayed without rebuilding. Original authoring choices remain recoverable for editing a new draft. This applies to draft publication; legacy api_mmm recipe resolution still requires fixed priors. VAR preserves its raw input and engine manifest; MMM quality policies cannot establish VAR acceptance.
Calibration: check the backend calibration capability. MMM draft publication validates active likelihood observations and returns their count, units, channels and hash in recipe provenance. Enabled invalid or unapplied observations fail explicitly; VAR calibration is unsupported. Preserve disabled authoring rows when editing. Imported wizard JSON is retained as editable rows; multipart wizard CSV capture retains its original bytes. No new MCP route is needed.
Pipeline sources retain the exact version ID, pipeline ID, version number and verified content hash. No pipeline is executed. Existing exported uploads are not assigned inferred pipeline lineage.
Draft source.history preserves up to 100 recorded column transformations/removals with parameters and before/after data hashes. The backend checks chain continuity and the terminal data hash. These are client-reported authoring records, not independently replayed operations or quality evidence. Edited data must omit an unchanged source origin. Unknown nested fields remain preserved.
Optional calibration_import retains an original JSON file (1 MB maximum) in the authorized authoring snapshot. Preserve it independently of current editable observations. The backend verifies its bytes/hash; published provenance includes filename/hash and explicitly states current observations may differ. Ordinary recipe responses omit the raw attachment. This reference is not scientific validation.
Custom numeric quality checks: create_quality_policy accepts checks such as
{"metric":"custom:benchmark_deviation","name":"Benchmark deviation","units":"%","operator":"lte","maximum":10,"required":true}.
Use gte with minimum, or between with both inclusive bounds. Read backend capability discovery before using this additive contract.
For externally calculated metrics, call evaluate_study_run(run_id, policy_id) first and retain report.basis_hash. Calculate from that run's saved outputs, then call the same tool with expected_basis_hash and external_evidence=[{"metric":"custom:benchmark_deviation","value":8,"method":"Absolute deviation as percentage of benchmark","source_reference":"Versioned model export and benchmark"}]. An optional source_sha256 records a reported source digest. The backend determines pass/fail and rejects stale output bases. Every submission creates a new assessment and must supply all intended custom values; omitted values remain unevaluated. Source references and calculations are submitter-reported, not independently verified. This numeric MMM contract does not execute agent code, accept a model, support VAR acceptance, or designate a champion.
Boolean checks use kind: "boolean", operator: "equals" and strict boolean expected. Submit a JSON boolean in external_evidence.value; numeric or string substitutes are rejected. Manual checks use kind: "manual", operator: "equals", expected: true. MCP can define these rules and read evidence, but API keys cannot submit manual sign-off: a signed-in reviewer must supply confirmation, rationale and source through the frontend. Missing answers stay unevaluated; sign-off is not automatic model acceptance or champion selection.
get_study_champion(study_id) reads the incumbent, accepted candidates/eligibility blockers and immutable selection/replacement/revocation history. The backend requires migration workflow_champion_001. Writes require the project owner's frontend session; MCP cannot promote or revoke. Stale evidence retains the incumbent with review_required. Validation references are reviewer-declared and decision_grade_ready remains false until scientific protocol qualification is implemented. Champion designation does not deploy or fit a model.
Native prediction-window gates are available as prediction_mae, prediction_rmse and prediction_wape (WAPE is a fraction). They use the existing create_quality_policy/evaluate_study_run tools. The backend reads saved actual/prediction rows, requires unique prediction dates after the saved training window and leaves missing/malformed evidence unevaluated. Both windows are bound into the assessment hash. This does not prove untouched holdout provenance, leakage-free preprocessing or full sampling intent; decision-grade champion qualification remains separate.
create_quality_policy(..., validation_protocol=...) can declare a temporal holdout before launching both runs. Required protocol fields are training_end, prediction_start/end, min_draws, min_tune, min_chains, max_r_hat and max_prediction_wape. Use kind: "temporal_holdout"; dates are ISO and WAPE a fraction. No defaults are recommended. Both runs must launch under that exact policy.
assess_study_validation_pair(study_id, full_run_id, validation_run_id, policy_id) assesses saved evidence without fitting or writing a model decision; serving available prediction evidence appends an access audit event. It checks distinct completed MMM tasks, launch binding, matching frozen files/settings/runtime, configured sampling minima, R-hat, prediction-window WAPE/dates and full-date coverage. Missing evidence blocks. The response fingerprint identifies the assessed records. Passing compatibility does not verify retained draws/ESS/divergences, preprocessing or untouched holdout history, and decision_grade_ready remains false.
Validation-pair responses include sampling_evidence for the full and validation runs: retained chain/draw counts, bulk/tail ESS minima and divergences when saved by the fitting engine. Missing or partial records are explicit. These values are not yet checked against acceptance limits and do not certify scientific readiness.
Protocols may now opt into retained_sampling: {min_ess_bulk, min_ess_tail, max_divergences} before launching both models. All limits are explicit, with positive ESS minima and a nonnegative integer divergence maximum. Pair assessments also apply declared chain/draw minima to retained counts. Missing/partial native records block; sampling_qualification distinguishes pass, blocked and not_declared. Business validity and holdout provenance still prevent scientific certification.
Set validation_protocol.require_policy_review before launching to require current analyst acceptance of each latest launch-policy assessment in pair checks. Stale evidence, subsequent rejection, missing acceptance or ambiguous ordering blocks. This reuses required policy gates; external business calculations remain reported evidence. MCP can declare and read these requirements but cannot supply analyst acceptance.
Pair responses expose holdout_provenance separately from compatibility. Version 1 full-input preprocessing records remain blocked. Version 2 records identify training-only transformation/scaling and report review_required with preprocessing_status: training_only. Missing or unsupported records are unavailable; no record is silently certified. Prior-source independence and holdout access/reuse evidence remain required; decision_grade_ready stays false. Existing routes and tool arguments are unchanged.
Pair responses also expose prior_provenance: native frozen automatic-prior source windows are checked against the declared training end and bound to recipe input hashes. A later source window or mismatched hashes is blocked; missing historical/uploaded provenance is unavailable. A valid window still requires review of assumptions, external calibration and holdout reuse. This does not record access history or certify independence. No tool arguments or routes changed.
get_study_prediction_access(run_id) reads partial Studies prediction-access history (totals and latest 20 events). Pair responses include prediction_access with the same partial-coverage summary, independent of the assessment hash. Assessment creation/history reads, comparison and pair checks append audit events when serving available prediction evidence. The three previously read-only inspection tools are now annotated non-destructive/additive and non-idempotent. Reading access metadata itself does not add events.
History is scoped to the study and matching recorded dataset/windows. Older activity, other results routes/exports, other studies and offline work are not covered. Repeated serving does not prove retuning; zero events never proves an untouched holdout. The backend requires its additive prediction-access migration; scientific qualification remains incomplete.
get_model_results(model_hash, sections="prediction_window") explicitly exports saved prediction-window actuals/model values in JSON or CSV; it is omitted from default exports and distinct from scenario predictions. Serving it for a study-linked model appends a native access event, so this tool is conservatively annotated additive/non-destructive. Dashboard results now record access to the same saved window. Filters do not truncate or channel-filter prediction evidence. This broadens instrumentation without certifying untouched holdout status; other routes, source files, earlier activity and offline work remain outside coverage.
declare_study_holdout_use(run_id, declaration_id, source_access_id, disposition, reason, affected_revision_id=None) appends a reported evidence-use declaration with project-owner credentials and create:models. Read access history first and select an event from that run. Use review_only, informed_revision (requires a same-study published revision) or uncertain. Reuse the same UUID/content on retry; conflicting content is rejected. Declarations retain submitting interface and do not certify independence, accept or promote models.
Access/pair responses include holdout_use; reported revision influence retains fresh_validation_required for affected revisions even after later review-only notes. Other declarations remain review_required, and no declarations means not_recorded. Declaration history changes the pair fingerprint; additional viewing alone does not. The backend requires its additive holdout-use migration. Champion reads expose candidate-specific holdout_use: an informed-revision declaration blocks selection of that named revision and marks an existing champion review_required. Ordinary access, review-only notes and unrelated revisions do not block. Later acceptance or review-only notes cannot clear an earlier influence report. This applies to named revisions and their recorded descendants, including ordered recipe predecessors and source-revision ancestry; current analyst-reviewed validation resolutions may clear the exact accepted run; changed or revoked evidence reblocks it. No report is not proof of an untouched holdout.
Replacement holdout preflight: assess_study_validation_pair returns fresh_validation for reports naming the full-model revision. It checks that a replacement policy follows the influence declarations, precedes both runs, and uses prediction dates strictly after the used holdouts. Recorded same-study access overlapping the new dates before policy creation blocks preflight, conservatively across dataset versions. Complete retained sampling, accepted policy evidence and supported training-only preprocessing/prior records are required. Status is not_required, blocked, or review_required; resolves_champion_block is always false. This is not a durable resolution or proof of offline independence. No extra tool or request field is needed.
Reviewed resolutions and recipe ancestry
get_study_validation_resolutions(study_id) returns append-only analyst resolutions and revocations with re-evaluated current, stale or revoked status. The backend requires migration workflow_resolution_001. API keys cannot sign off independence or revoke reviews: the signed-in owner uses the UI, which binds the exact pair fingerprint and requires a rationale plus an explicit independence review. Current resolution clears only the exact accepted full run; new outputs, declarations or reviews invalidate it. Champion reads expose resolved_by_review and the resolution identity. No automatic scientific certification or promotion is performed.
When creating a derived recipe or draft, supply source_revision_id from the same study. create_recipe_draft, create_study_recipe and revise_study_recipe forward this optional field. Draft source linkage is immutable and survives full-editor publication; previous versions of a recipe inherit influence automatically. Unrecorded/off-platform copies remain outside recorded ancestry. Resolution does not transfer to new runs or descendants.
Release note: these tools require the corresponding MCP package release and updated backend. Draft-branch tests do not establish package publication or application deployment.
Experiment priorities
See What to test next for the read-only experiment screening tool, its inputs, approximation and limitations. Requires compatible backend support.
Available Tools
94 toolsadopt_model_into_studyAdopt Model Into StudyA
Import an owned completed model into a study as an executable, editable recipe (jellyfish #880). Without confirm: a preview only, with the resolved recipe, its inspection block and an import report {complete, dataset: {recorded, origin}, settings: {count, not_recorded}, editable: fully | partly, evidence: {fitted_result}}; nothing is written. The preview's report.dataset also carries display, the lineage line the recipe card will show once imported ("Retail weekly · v3 · verified"; "dataset lineage not recorded · legacy" for a model built before capture). With confirm=true: creates a recipe whose base_model revision 1 carries the exact fit inputs, the wizard snapshot captured at build and the verified dataset origin, launches and re-freezes like any other revision and can be edited in place (get_recipe_revision_authoring, then create_recipe_draft with target); the fitted result is attached as an adopted run outside the attempt budget (201, {run, revision}). 409 when the model is not complete or is already attached to a study. Models built before snapshot capture import with not_recorded settings (editable: partly). Adoption never refits.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| confirm | No | ||
| study_id | Yes | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only flag it as non-read-only, non-destructive, non-idempotent and open-world; the description goes far beyond by detailing the dry-run preview contents, the exact report shape, the 201/409 outcomes, that nothing is written without confirm, that adoption never refits, and that the run sits outside the attempt budget.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, which is good, but the body is a dense run-on sentence with nested braces, semicolons and parenthetical issue codes that is hard to parse on a single read; every clause carries information, yet the structure could be split for scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, yet the description still supplies return details (201 {run, revision}, report fields, 409 conditions) plus the legacy/pre-snapshot edge case. It is near-complete for a mutation tool, with the unexplained 'reason' parameter the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate: confirm is thoroughly explained (default false, preview semantics) and study_id/model_hash are implied by the surrounding narrative, but the required 'reason' parameter is never mentioned at all, leaving one of four parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Import an owned completed model into a study as an executable, editable recipe') and immediately distinguishes itself from sibling write tools like create_study_recipe/create_recipe_draft by describing what it actually produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the two modes (preview without confirm vs. actual creation with confirm=true) and the 409 preconditions (model not complete, already attached). It names get_recipe_revision_authoring/create_recipe_draft as the follow-on editing path, but never explicitly says when to prefer this over create_study_recipe.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_study_validation_pairAssess Study Validation PairB
Assess a validation pair and append a prediction-access audit event when evidence is available. Checks distinct completed MMM runs launched under the declared protocol, frozen inputs/settings/runtime, configured sampling, saved R-hat, declared prediction windows/WAPE and date coverage. Returns blockers and an evidence hash; does not fit, accept or promote. Includes saved retained chain/draw, ESS and divergence records when available, with null for older models. Optional prelaunch retained_sampling limits require complete native records and check chain/draw minima, bulk/tail ESS minima and maximum divergences; otherwise sampling_qualification is not_declared. Optional require_policy_review checks current signed-in analyst acceptance of each latest same-policy assessment, including freshness and rejection blockers. The holdout_provenance report distinguishes missing evidence, blocked version 1 full-input preprocessing and version 2 training-only preprocessing requiring further provenance review. The prior_provenance report checks which rows the recorded smart priors read and the frozen input hashes; priors without a record remain unavailable, and priors that read observations after the declared training end are blocked. fresh_validation provides replacement-window preflight for influence reports naming the full-model revision: later windows, replacement policy chronology, recorded prior exposure, retained diagnostics and provenance/review requirements. It never clears champion blocks. External business calculations and untouched holdout history remain unverified; decision_grade_ready stays false.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes | ||
| policy_id | Yes | ||
| full_run_id | Yes | ||
| validation_run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations declaring readOnlyHint=false, openWorldHint=true and idempotentHint=false, the description adds substantial behavioral context well beyond the annotations: it appends an audit event, returns blockers and an evidence hash, substitutes null for older models, gates optional checks (retained_sampling, require_policy_review) with fallbacks like 'not_declared', and explicitly states decision_grade_ready stays false. These disclosures are consistent with the non-read-only, non-idempotent annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is an unusually long, dense block of run-on sentences mixing many distinct topics (audit event, R-hat/WAPE, retained diagnostics, policy review, two provenance reports, fresh_validation). It is not front-loaded or scannable, and key routing/scoping information is buried mid-paragraph, so it exceeds what an agent needs to select and invoke the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be restated, and the description does cover behavioral outcome states and gating logic richly. However, it leaves parameters entirely unexplained and gives no explicit alternative-tool routing, so it is adequate but with real gaps for a 4-required-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and all four parameters (study_id, policy_id, full_run_id, validation_run_id) are undocumented in both schema and description. The prose references concepts like the 'declared protocol,' 'same-policy assessment,' and 'full-model revision' that gesture at the parameters, but it never defines what each id represents or how they interrelate, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Assess a validation pair and append a prediction-access audit event') and enumerates exactly what it inspects (MMM runs, frozen inputs, sampling, R-hat, WAPE, provenance). It also distinguishes itself from adjacent actions by stating it 'does not fit, accept or promote' and 'never clears champion blocks,' which separates it from launch/promote tools. It stops short of naming a sibling tool directly, so differentiation is inferred rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage conditions are implied heavily through the feature list (validation pairing, retainment, policy review, holdout/prior provenance, fresh_validation), but there is no explicit 'use this when / use X instead' routing to alternatives such as evaluate_study_run or get_study_validation_resolutions. The negative scoping ('does not fit, accept or promote') gives some when-not guidance, but an agent must infer the primary trigger from context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_study_runCancel Study RunBDestructive
Request cancellation. Requested and confirmed stopped are distinct states.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey that this is destructive and non-idempotent. The description adds valuable behavioral nuance by clarifying that cancellation is a request and that 'requested' and 'confirmed stopped' are distinct states, implying an asynchronous or multi-step lifecycle. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two short sentences, with the key action upfront and the important state distinction immediately after. There is no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with an output schema and rich annotations, the description is mostly adequate. However, it does not explain the cancellation flow, how to confirm the eventual stopped state, or what side effects occur. The state distinction is a useful start but could be more actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention run_id at all. The parameter name is self-explanatory to a degree, but the description does not compensate for the lack of schema documentation by explaining what run_id refers to, how it should be obtained, or any constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('request') and a clear resource ('cancellation' of a study run), and it distinguishes the requested state from a confirmed stopped state. It does not explicitly differentiate it from sibling tools like update_run or get_study_run, but the core intent is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives such as launch_study_run, get_study_run, or update_run. The state distinction hints at usage context but does not explain prerequisites, follow-up steps, or which sibling handles confirmed stopping.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_study_runsCompare Study RunsA
Compare 2-20 candidates against one quality policy. Each row carries a basis record (family, dataset and costs hashes, outcome column, units, training window, output kind, prediction window, evidence freshness) and a per-dimension compatibility against the first row; rows are comparable only when family, dataset, outcome, units, window and output kind all match, and incompatible rows are returned with the differing dimension in blockers and are never ranked. This answers predictive ranking only: sensitivity agreement is not computed, analyst acceptance lives in decisions, and business validity is a human review. Serving prediction evidence appends access audit events. Does not fit or promote models.
| Name | Required | Description | Default |
|---|---|---|---|
| run_ids | Yes | ||
| study_id | Yes | ||
| policy_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far beyond the annotations, the description discloses crucial behavior: rows are comparable only when multiple dimensions match, incompatible rows are returned with the differing dimension in blockers and are never ranked, and serving prediction evidence appends access audit events. It also clarifies that no model fitting or promotion occurs, adding meaningful non-obvious context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but purposeful, with the core purpose front-loaded and supporting behavior details following. It is long, though each clause adds real value; a slight structural breakdown could improve readability, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description need not explain return values, and it does a strong job explaining behavior, exclusions, and ranking semantics. The main gap is that study_id's meaning is left implicit, and since the schema gives no parameter descriptions, that missing piece prevents full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It implicitly explains run_ids as the 2-20 candidates and policy_id as the quality policy, but it never mentions study_id, its role, or any constraints on it. The description partially compensates for the bare schema but does not fully document all three required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Compare 2-20 candidates against one quality policy,' which is unambiguous about what the tool does. It further defines the output rows, comparability conditions, and explicitly scopes the tool to predictive ranking, making it distinguishable from nearby siblings like evaluate_study_run or assess_study_validation_pair even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use the tool ('answers predictive ranking only') and gives explicit exclusions: sensitivity agreement is not computed, analyst acceptance lives in decisions, business validity is a human review, and it does not fit or promote models. It does not name alternative sibling tools, but the when-not guidance is strong enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_incrementality_testCreate Incrementality TestA
Record one incrementality test in a project, as analysed in its own tool: its design (geo, owned_media_ab or platform_lift block), dates, KPI, result with interval or sd, and incremental spend. Returns {id, version: 1, record, content_hash}. Set model_channel to the model activity column the test calibrates so it can be used with a model. Validation errors name the field (e.g. "the interval must contain lift_abs"). Not idempotent: calling twice records two tests. type time_holdout is not accepted here: planned time-holdout records are created only through save_incrementality_test_design (the backend refuses them elsewhere), and a status planned record of any type carries no result. Requires the create:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| record | Yes | One incrementality test as analysed in its own tool (Simba fits nothing here). type picks the design block: geo, owned_media_ab or platform_lift. Dates are YYYY-MM-DD; measured_through is the end of the carryover window (default end_date). result.lift_abs is the total lift in KPI units over the window; give its interval (two-sided, level e.g. 0.9) and/or sd. Null and negative lifts are valid. spend.incremental is the extra spend the test caused (signed). The time_holdout type and block appear on records the server builds from a test-design calculation (save_incrementality_test_design); they are read and listed here but not accepted on create or import, and a status planned record carries no result. | |
| project_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes beyond annotations by naming the required auth scope, describing validation error behavior ('name the field'), and spelling out the consequence of non-idempotency ('calling twice records two tests', consistent with idempotentHint=false). It also notes that planned records carry no result. These add real behavioral context beyond the safety annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the primary action and record contents, then progressively adds return shape, validation, idempotency, exclusion and scope. Dense but every clause carries information; only slight redundancy with annotations on idempotency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested-schema mutation with an output schema present, the description covers the design-block semantics, validation, scope, and the excluded time_holdout path. Return values are documented by the output schema, so omitting them is fine; a few minor behaviors (e.g. limits on record size) are not addressed but not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%, so the description must compensate, and it does: it explains what type selects (geo/owned_media_ab/platform_lift block), that dates are YYYY-MM-DD, that measured_through defaults to end_date, that lift_abs is total lift in KPI units, that null/negative lifts are valid, and that spend.incremental is signed. This meaningfully supplements the nested schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Record one incrementality test in a project') and immediately scopes what a record contains (design block, dates, KPI, result, spend). It also names a sibling it is NOT: 'type time_holdout is not accepted here... created only through save_incrementality_test_design', letting an agent distinguish it from nearby tools without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-not guidance: time_holdout is refused here and must go through save_incrementality_test_design. It also states the required scope ('create:models') and the setup condition ('Set model_channel to the model activity column the test calibrates'). No alternatives are left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_modelCreate ModelA
Create and start fitting a new Bayesian Marketing Mix Model.
This queues an async model fit and returns immediately with a model_hash. Use get_model_status to poll for progress until status is 'complete'.
Priors are calculated automatically using smart defaults based on cost shares, industry benchmarks, and channel-type detection. You can override individual channels via the priors parameter.
Args:
uploaded_file_id: The file ID returned by upload_data.
date_column: Name of the date column in the CSV.
kpi_column: Name of the KPI/dependent variable column.
hierarchy_column: Name of the brand/segment column (must have exactly 1 unique value).
channels: List of channel definitions, each with keys: name, activity_column, spend_column.
Example: [{"name": "TV", "activity_column": "tv_grps", "spend_column": "tv_spend"}]
multiplier_column: Column to convert KPI to revenue. Defaults to kpi_column.
control_columns: Non-media control variable column names (e.g. ["price", "distribution"]).
total_media_effect: Controls prior strength. Either an industry name for a benchmark
("FMCG"=6%, "Retail"=9%, "TelCo"=30%, "Financial Services"=19%,
"E-Commerce"=22%, "Other"=12%) or a custom decimal like "0.15"
meaning "I believe all media drives 15% of my KPI". Default "Other".
priors: Optional per-channel prior overrides. Each dict should have "channel" matching
a channels[].name, plus any fields to override: distribution, mean, sd, lower,
upper, transform, adstock_type, effect_period.
Only specified fields are overridden; the rest use smart defaults.
Adstock-kernel fields: half_life_lower/half_life_upper (carryover half-life
bounds in periods — preferred over the legacy decay_lower/decay_upper),
theta_mean/theta_sd (peak-lag prior, adstock_type="delayed" only),
dual_weight_mean/dual_weight_sd (long-term/slow-component share prior,
adstock_type="dual_geometric" only).
SATURATION ANCHOR — state it ONCE, in exactly one of three
mutually exclusive forms (two in one override -> 400 "state
the saturation prior once"):
(1) half_marginal_mean/half_marginal_sd — CANONICAL for
saturation_type="generalized_log" (rejected on other
families): the activity level where MARGINAL returns have
halved, finite at every curvature (#632).
sat_shape_mean MUST accompany the pair in the same override
(#672) — the fold pairs your coefficient with the
stated curvature, so omitting it is a 400, never a silent
default.
(2) half_saturation_mean/half_saturation_sd — the
50%-of-maximum point in activity units, for the
single-parameter families (tanh/michaelis_menten/
negative_exponential). Do NOT use it for generalized_log
near-log work: it overflows below sat_shape_mean 0.00097657
and is rejected with a 400 — precisely the regime that
family exists for.
(3) alpha_sd + scalars — legacy internal coordinates,
accepted for backward compat.
Curvature (generalized_log only): sat_shape_mean/sat_shape_sd
— small values are near-logarithmic, 1.0 is michaelis_menten.
COEFFICIENT in a human coordinate (generalized_log only,
#671): effect_at_avg_mean/effect_at_avg_sd — the
effect share at the channel's AVERAGE activity, as FRACTIONS
(mean in (0, 0.95], sd > 0; 0.2 means 20%). Folded
server-side into mean/sd at the row's operating point with
the same arithmetic as the dashboard. Requires sat_shape_mean
in the same override; cannot be combined with mean/sd
("state the coefficient prior once") or with
half_saturation_*. Stating half_marginal_* + effect_at_avg_*
+ sat_shape_mean together is the full (x*, E, k) triple —
the recommended generalized_log elicitation, since only
beta*k is identified and raw beta spans orders of magnitude.
VERIFY what was applied via get_model's
model_config.priors_resolved: rows carry the FOLDED
mean/sd/scalars/alpha_sd, and overridden_fields lists the
field names you sent.
UNKNOWN KEYS ARE REJECTED with a 400 naming the field
(#630); they used to be dropped silently, fitting a
hybrid of the override and the smart defaults. Common misses:
"beta"/"beta_mean" -> mean, "beta_sd" -> sd, "sat_shape" ->
sat_shape_mean. "name" and "parameter" are rejected too — they
identify the smart-prior row the override merges onto.
trend: Enable dynamic baseline trend component.
seasonality: Enable automatic seasonality detection. The prior sigma on
the Fourier coefficients is chosen for the link (#534):
0.5 under link="log", 10 under "identity". The coefficients
live on the link's scale, so the additive default would
admit e^10x seasonal amplitude on a multiplicative model.
likelihood: Likelihood function: "normal" (default), "lognormal", "logit",
"studentt", "poisson", "negativebinomial", or "quantile".
saturation_type: Diminishing-returns curve family applied to media:
"tanh" (default), "michaelis_menten", "negative_exponential",
or "generalized_log" (two-parameter Box-Cox/power-log family
1 - (1+x/K)^(-shape); tune per channel via the
sat_shape_mean/sat_shape_sd prior fields).
transform_order: "adstock_first" (default: carryover accumulates, then
saturates) or "saturation_first" (each period's spend
saturates, then the effect spreads over time through the
normalized adstock kernel).
link: Model Form. "identity" (default) fits an additive model — components
add on the outcome scale. "log" fits a multiplicative model —
components add on the log scale and media effects are percentage
lifts. Under the removal_lift attribution convention (the API
default), contributions then include an Overlap reconciliation
column; the other conventions (aumann_shapley — the dashboard
default for multiplicative models since #509 —
shapley, and proportional_normalized) allocate the interaction
across components and close exactly WITHOUT an Overlap column
(see get_model_results).
channel_groups: Optional adstock groups: [{"name": ..., "channels":
[...], "share_saturation": bool}]. Member channels tie
their carryover parameters (decay/theta/dual-weight —
plus saturation when share_saturation is true) to one
shared value, e.g. grouping channels into shared
"Long"/"Short" carryover classes. Members are
channels[].name values; each group needs >= 2 members;
groups must be disjoint; and tied members must have
identical adstock_type/effect_period/bound overrides
(the API rejects divergent groups at request time).
control_reference: Control attribution reference points (#452),
multiplicative models (link="log") only: maps control
column names (plus optional "default") to
"auto" | "absent" | "average" | "lowest" | "highest" —
which counterfactual "remove this control" means in
the contributions. "absent" measures against the
variable at zero (legacy behavior; honest only when
zero is observed). "average"/"lowest"/"highest"
reference the control at its observed mean/min/max —
use for controls that never approach zero (price
indices, distribution levels), where a zero
counterfactual produces unbounded contributions and a
negative Base. "auto" detects per control whether
zero is inside the observed data range. Example:
{"default": "auto", "relative_price": "average",
"promo_flag": "absent"}. Omit entirely to keep every
control at "absent" (byte-identical legacy output).
Unknown control names/modes are rejected at request
time; any value other than "absent" requires
link="log". The fit reports the resolution in
model_config.control_references (see
get_model_results).
name: Display name for the created model, honoured verbatim (#575).
Falls back to a generated API_MMM{brand}{hash} string when
omitted. Either way the model starts unsaved — invisible to
list_models unless include_unsaved=true — until save_model
files it into a project.
operating_margin: Scalar operating margin as a decimal fraction in
(0, 1], e.g. 0.18 = 18%. Mutually exclusive with
operating_margin_column (the API 400s when both are given).
Storing a margin unlocks the financials results section and
lets run_optimizer(objective="profit") use it automatically
instead of requiring forward_margin on every call.
operating_margin_column: Name of a column in the uploaded CSV holding
a per-date margin series. The column may be uniformly in
fractions (0, 1] OR uniformly in percentages (1, 100] — the
API detects the unit and normalizes percentages; mixed units
are rejected. Same unlocks as operating_margin; the column
must exist in the uploaded file. CAUTION: the API reads the
margin keys from the REQUEST ROOT — a margin placed inside a
config dict is silently ignored (no error), and the model fits
marginless.
attribution: Attribution convention for the contribution decomposition,
resolved at fit time: "removal_lift" (the API default;
one-at-a-time removal — multiplicative models then emit the
Overlap column), "aumann_shapley" (the dashboard default for
multiplicative models since #509), "shapley", or
"proportional_normalized". Any value other than "removal_lift"
requires link="log" (the API rejects it on additive models).
The non-removal conventions allocate the interaction across
components and close exactly WITHOUT an Overlap column. To
reconcile with a dashboard-built multiplicative model, use
"aumann_shapley".
annual_discount_rate: Annual discount rate (decimal >= 0, e.g. 0.08)
used by the display-time financial bridge and cohort ledger PV
discounting. Display-time only — does not change the fit.
sampler: MCMC sampler overrides, e.g. {"n_samples": 2000,
"tune": 1500, "chains": 4, "cores": 2, "target_accept": 0.95}.
STRICTLY validated: unknown keys inside sampler are rejected
with a 400 naming the field; cores must be 1-8. Only the keys
you send are overridden.
reporting_kernel: Reporting-kernel class override (#450) for
the cohort_ledger section's forward allocation. Shape:
{"classes": {...}, "channel_classes": {...}} — ONLY those two
top-level keys are accepted (anything else, e.g. "mode", 400s
with the unknown key named). channel_classes names channels by
channels[].name or activity_column, validated at request time.
Affects only how the cohort_ledger allocates effects over the
horizon — not the fit, and not the contributions /
channel_summary decompositions. (The related "complete" /
"in_window" choice is a separate cohort_horizon QUERY parameter
on the results endpoint, not part of this config.)
control_priors: Optional root-level control slope overrides. Each entry names a selected control_columns column with "control", plus transform (N, DM, STA, DDM, LOG), distribution (normal, inversegamma, truncatednormal, halfnormal), mean/sd/lower/upper as applicable. LOG is log(x/mean(x)); STA divides by sample sd without centering. Priors are in transformed units; changing transform does not convert coefficients. Nonempty overrides require backend capability version 1; unsupported or unavailable checks stop before creating a model.
calibration: Optional lift-test calibration: likelihood observations the
fit must respect. Either {"tests": [{"test_id": ..., "version"?,
"channel"?, "confirm_kpi"?}]} (recorded tests from
list_incrementality_tests, each derived against THIS model's data),
or {"units":
"revenue" | "response", "observations": [{channel, x, delta_x,
delta_y, sigma}]} for rows you derived yourself. If any test can't
calibrate this model, nothing is created: the error has
code "calibration_refused" and tests gives each test's reason.
Preview a test's row with get_incrementality_test(model_hash=...).
The derived rows calibrate this fit only; the model does not keep a
link to the tests. To keep that lineage (the test then lists the
model under used_by and can't be deleted), build through a study
recipe or a recipe draft that references the tests.
Returns the model_hash for status polling.
| Name | Required | Description | Default |
|---|---|---|---|
| link | No | identity | |
| name | No | ||
| trend | No | ||
| priors | No | ||
| sampler | No | ||
| channels | Yes | ||
| kpi_column | Yes | ||
| likelihood | No | normal | |
| attribution | No | ||
| calibration | No | ||
| date_column | Yes | ||
| seasonality | No | ||
| channel_groups | No | ||
| control_priors | No | ||
| control_columns | No | ||
| saturation_type | No | tanh | |
| transform_order | No | adstock_first | |
| hierarchy_column | Yes | ||
| operating_margin | No | ||
| reporting_kernel | No | ||
| uploaded_file_id | Yes | ||
| control_reference | No | ||
| multiplier_column | No | ||
| total_media_effect | No | Other | |
| annual_discount_rate | No | ||
| operating_margin_column | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is a non-readonly, non-idempotent, open-world write. The description goes far beyond: async return contract, the model starting 'unsaved — invisible to list_models,' the silent-failure trap where a margin key placed in a config dict is ignored, strict 400 rejection of unknown keys, and calibration_refused behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose, async contract, and polling guidance are correctly front-loaded, but the Args section runs to roughly two thousand words with internal issue numbers (#632, #672, #630) and dense register detail that a caller must wade through. Much is justified by genuine complexity, yet it is not scannable and could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 26-parameter fitting tool with an output schema, this is complete: required inputs, defaults, cross-parameter constraints (operating_margin vs operating_margin_column, generalized_log anchor combinations), and downstream verification paths are all covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0% and 26 parameters, the description carries the entire burden and does so: it documents every argument, including enum-like values for likelihood/saturation_type/link/attribution, total_media_effect industry benchmarks, the three mutually exclusive saturation-anchor forms, transform units, and rejected key aliases.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a precise verb+resource: 'Create and start fitting a new Bayesian Marketing Mix Model,' which distinguishes it from sibling creation tools (create_var_model, create_study). The async semantics ('queues an async model fit and returns immediately with a model_hash') further pin down what this tool uniquely does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear next-step context ('Use get_model_status to poll for progress until status is complete') and workflow prerequisites via upload_data, save_model, and calibration. It doesn't explicitly state when to choose this over sibling creation paths (e.g., a study recipe), so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectCreate ProjectA
Create a named project (model folder) to file models into.
Names are sanitized the same way model names are (non-empty after HTML sanitization).
Args: name: Display name for the new project. team_id: Optional team to share the project with; must be a team you belong to (403 otherwise, 404 for an unknown team).
Returns the created project (201) including its id — pass that to save_model(project_id=...).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| team_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavior beyond annotations: name sanitization rules, team membership requirements with distinct 403/404 errors, and the 201 response including the created project id. Annotations already indicate a mutating, non-idempotent operation, and the description is consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the core purpose, and organized into clear sections: purpose, sanitization note, Args, and Returns. Every sentence adds distinct value with no repetition of schema or annotation information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter creation tool with an output schema present, the description covers the essential behavior: display name, sanitation, team-sharing constraints, return status, and how the result connects to save_model. No critical calling information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates by explaining both parameters: 'name' as a display name and 'team_id' as optional sharing with an owned team, including error semantics. This adds real meaning beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Create a named project (model folder) to file models into.' This clearly distinguishes create_project from sibling tools like create_model and create_recipe_draft by establishing that a project is a container for models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: projects are for filing models into, and the returned id is meant to be passed to save_model(project_id=...). It does not explicitly name alternatives or when-not-to-use conditions, but the workflow hint is strong enough to guide appropriate invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_quality_policyCreate Quality PolicyA
Save project-specific checks. Policies are immutable: to change one, create a new policy, ideally with derived_from_policy_id naming the saved policy you started from (same study) so lineage is kept and the response carries diff (checks added/removed/changed by metric and field, protocol changes, name_changed, rationale_changed, rules_changed) — reordering checks is not a change. Built-in checks have metric (r_hat_max, mae, rmse, wape, prediction_mae, prediction_rmse, prediction_wape), maximum and required, and may instead use operator gte/between with minimum (default lte). Artifact-backed checks read what the fit already saved and need no external evidence: saved diagnostics (r_squared, mape, durbin_watson, pareto_k_pct, normality_p, loo_cv), retained-sampling counts (retained_chains, retained_draws_per_chain, ess_bulk_min, ess_tail_min, divergences) and provenance status (provenance:holdout, provenance:prior with expected review_required; blocked fails, absent is not_collected). Custom numeric checks use metric custom:, name, units, operator (lte/gte/between), applicable minimum/maximum and required. Boolean checks use kind=boolean, operator=equals and expected=true/false. Manual checks use kind=manual, equals, expected=true; agents can define these but cannot submit manual sign-off. Custom bounds may be negative. WAPE is a fraction. Prediction-window checks require saved finite actuals/predictions at unique dates after the saved training window; this does not certify untouched holdout provenance. No default thresholds are assumed. Declare at least one required check, use each metric once, and set maximum R-hat at least 1. At most 20 checks in total; the refusal names the count. A bad check is refused with one message naming its position and declared kind. Optional validation_protocol declares a temporal holdout split, configured sampling minima, R-hat and prediction WAPE limits before both runs launch under this policy. The backend validates policy rules.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| checks | Yes | ||
| study_id | Yes | ||
| rationale | Yes | ||
| validation_protocol | No | ||
| derived_from_policy_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses far beyond sparse annotations: immutability of saved policies, refusal behavior ("the refusal names the count", "refused with one message naming its position and declared kind"), backend validation of rules, agent restrictions ("cannot submit manual sign-off"), provenance edge semantics ("blocked fails, absent is not_collected"), and response contents (diff carrying checks/protocol changes). No contradiction with readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose ("Save project-specific checks") followed by the most decision-relevant constraint (immutability). The description is long (~400 words) but each sentence carries distinct rules for a genuinely complex tool — check taxonomy, constraints, error behavior, protocol. Minor redundancy exists with the checks item description already embedded in the schema (e.g., "At least one required gate per policy; at most 20 checks"), so it is not maximally lean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with zero schema descriptions, the description is remarkably complete: every check kind and its expression, validation invariants (at least one required check, each metric once, max R-hat ≥ 1, ≤20 checks, no default thresholds), prediction-window semantics, validation protocol semantics, lineage/diff behavior, and error/refusal format. The output schema covers return values, so nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, putting the full burden on the description, and it delivers: categorizes the metric enum into built-in/artifact-backed/custom with per-kind parameter requirements (operator defaults like lte, custom:<slug> syntax, boolean vs manual shapes, expected semantics), explains validation_protocol as a temporal holdout split with configured sampling intent, and clarifies derived_from_policy_id lineage. This far exceeds the bare schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Save project-specific checks" states a specific verb (save/create) and resource (project-specific quality policy), and the immutability clause immediately establishes this as the creation tool in a create/list/get/diff/retire family. An agent can distinguish it from diff_quality_policies, list_quality_policies, get_quality_policy, and retire_quality_policy without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use context: "Policies are immutable: to change one, create a new policy, ideally with derived_from_policy_id naming the saved policy you started from (same study)" — this defines the derivation workflow clearly. However, it never names sibling alternatives (retire_quality_policy for retiring, diff_quality_policies for comparing) nor states explicit when-not-to-use conditions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_recipe_draftCreate Recipe DraftAIdempotent
Save an encrypted authoring draft without publishing, fitting or consuming an attempt. Supply a UUID draft_id and reuse it with identical content after an uncertain response. Start from get_recipe_draft_template for a new draft, or from get_recipe_revision_authoring when working from a published revision. To edit a recipe in place, pass target={recipe_id, base_revision_id}: the draft becomes an edit session of that recipe and publish_recipe_draft with a matching target writes the recipe's next revision; the target is immutable after creation (409 on a replay that names another; 404 for a recipe outside the study). To branch a published recipe into a new one instead, omit target and supply its source_revision_id from this study; the immutable link carries inherited influence through publication. Backend validates version and size.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| target | No | ||
| draft_id | Yes | ||
| snapshot | Yes | Lossless editor authoring state. Backend is authoritative; preserve unknown nested fields. Source bytes are copied into encrypted draft storage (10 MB source limit); metadata limit is 5 MB. No local filesystem paths or executable code. | |
| study_id | Yes | ||
| source_revision_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations (readOnly=false, idempotent=true, destructive=false, openWorld=true): the idempotent replay contract (reuse identical draft_id), target immutability with specific 409/404 semantics, and backend version/size validation. The description's write/idempotent behavior is fully consistent with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose before extending into the mode-switching rules, and every sentence carries load (idempotency, target vs branch, error codes). It is dense and long, which slightly raises parsing effort, but there is little pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers the full complexity of the two draft modes, the template entry points, and replay semantics. Minor gaps remain around name/study_id scoping, but nothing that would cause a mis-invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 17% schema description coverage, the description compensates well for the complex parameters: target becomes an edit session with an immutable recipe link, source_revision_id branches into a new recipe scoped to the study, and draft_id is a reusable UUID. It leaves name, study_id, and snapshot body undocumented, but those are self-evident or covered by nested schema field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save an encrypted authoring draft') plus the distinguishing constraint that it neither publishes nor consumes an attempt. It explicitly differentiates itself from siblings like get_recipe_draft_template, get_recipe_revision_authoring, and publish_recipe_draft, so an agent can tell the modes apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use routing: start from get_recipe_draft_template for a new draft, from get_recipe_revision_authoring when working from a published revision, pass target to edit in place, or omit target and supply source_revision_id to branch. The alternative paths and their selecting conditions are spelled out, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_studyCreate StudyA
Create a study owned by an existing project. Does not launch models or consume attempts. question is what the study should find out (one or two sentences): stored and shown to humans, never executed, and must not contain acceptance thresholds, validation rules or run limits (those belong in a quality policy, recipe revisions and max_attempts/max_concurrent). Exploratory and reliability questions are valid. context (optional): scope, data caveats and assumptions a reader needs to interpret results. Example question: 'How much do paid search and paid social contribute to weekly sales after price, promotions and seasonality?'
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| context | No | Optional scope, data caveats and assumptions a reader needs to interpret results. Not rules; see question. Omit to leave unchanged on update; send an empty string to clear. | |
| question | Yes | What the study should find out, in one or two sentences. Stored and shown to humans; never executed or used as an acceptance rule. Do not put thresholds, validation rules or fit limits here: those belong in a quality policy, recipe revisions and max_attempts/max_concurrent. Exploratory and reliability questions are valid. | |
| project_id | Yes | ||
| max_attempts | No | ||
| max_concurrent | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag this as a non-read-only, non-idempotent create operation, and the description adds valuable behavioral detail: creation has no launch side effects, does not consume attempts, and the question is stored for humans rather than executed. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is front-loaded with the main purpose and scoping, and the example question is useful. It is a single dense paragraph without filler, though some content repeats what is already in the schema descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with an output schema and annotations covering safety, the description covers purpose, ownership, side-effect absence, and question constraints well enough to invoke correctly. It does not fully explain max_attempts/max_concurrent semantics, but those have defaults and are named in the run-limits guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only 2 of 6 parameters have schema descriptions (33% coverage), so the description needed to compensate; it does well for question and context, including what belongs elsewhere. However, max_attempts and max_concurrent are only referenced as places for run limits, and name/project_id receive no semantic detail, leaving meaningful gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Create'), the resource ('a study'), and the ownership constraint ('owned by an existing project'), and immediately differentiates itself from related actions by noting it does not launch models or consume attempts. This makes it clearly distinct from siblings like launch_study_run and update_study.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear guidance on what this tool is not for: it does not launch models or consume attempts, and it explicitly routes acceptance thresholds and validation rules away from the question field to quality policies, recipe revisions, and max_attempts/max_concurrent. It stops short of naming sibling tools like update_study or launch_study_run as explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_study_recipeCreate Study RecipeA
Freeze a recipe without fitting. Supply source_revision_id when deriving from a published same-study recipe to retain influence ancestry. Optional expected_content_hash binds the validated effective inputs; a mismatch returns 409 and requires a fresh preview. Specification kind api_mmm has request containing create_model API fields; model_snapshot has model_hash and is review-only.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| reason | Yes | ||
| study_id | Yes | ||
| specification | Yes | Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, unknown config keys are rejected, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact. | |
| source_revision_id | No | ||
| expected_content_hash | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the write/non-idempotent/non-destructive profile, and the description adds real behavior beyond them: a hash mismatch returns 409 and forces a fresh preview, and model_snapshot is review-only (launch refused, lineage unknown). It does not state permission/auth requirements, but the mutation semantics and failure mode are well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, each carrying distinct information, with the core purpose front-loaded. No filler, though the spec-kind sentence packs multiple constraints into one line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with a nested spec object, an output schema (so returns need no explanation), and rich annotations, the description covers the key branching (kind selection), the ancestry concern, and the 409 failure path. Only the untouched scalar params keep it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17%, so the description must compensate, and it does add meaning for source_revision_id and expected_content_hash that the schema lacks (titles only). It leaves name, reason, and study_id unexplained, so roughly half the parameters remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Freeze a recipe without fitting', which pins down the operation and distinguishes it from fitting/authoring. It doesn't name the sibling tools it replaces (create_recipe_draft, publish_recipe_draft), but 'freeze' vs 'draft'/'fit' is inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional guidance for source_revision_id ('when deriving from a published same-study recipe') and for choosing between api_mmm and model_snapshot kinds. However, it never states when to use this tool versus create_recipe_draft, revise_study_recipe, or refreeze_recipe_revision, so tool-selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_var_modelCreate Var ModelA
Create and start fitting a long-term (VAR) model (#569).
VAR models capture the joint dynamics of several series (e.g. sales and
brand-equity metrics) and produce the long-run elasticity bridge behind
the MMM's long_run_rollup results section. Fit one, then link it to an
MMM with link_var_model.
Args: uploaded_file_id: Dataset id from upload_data (must contain every named column). date_column: Date column name. Cannot also be a series. endogenous_vars: At least two column names — the jointly-modeled series. exogenous_vars: Optional outside drivers; must not overlap the endogenous set. lags: VAR order (>= 1). The dataset needs at least lags + 10 rows with no missing values across the modeled columns. forecast_horizon: Periods forecast for diagnostics (default 12). base_variable: The outcome series (must be endogenous) long-run multipliers are measured against. Required for long-run effects. equity_variables: Endogenous columns (excluding the base) whose long-run IRF multipliers are estimated. Required for long-run effects. lre_horizon: Long-run effects horizon in periods (default 156). lre_ci: Credible-interval mass for the effects table, in (0, 1). var_priors: Advanced prior overrides (lag_coefs / alpha / coefs / noise_chol); unknown keys are rejected. name: Display name for the created model, honoured verbatim (#575). Falls back to a generated API_VAR_* string when omitted.
Returns 202-style payload with model_hash; poll get_model_status.
| Name | Required | Description | Default |
|---|---|---|---|
| lags | No | ||
| name | No | ||
| lre_ci | No | ||
| var_priors | No | ||
| date_column | Yes | ||
| lre_horizon | No | ||
| base_variable | No | ||
| exogenous_vars | No | ||
| endogenous_vars | Yes | ||
| equity_variables | No | ||
| forecast_horizon | No | ||
| uploaded_file_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, indicating mutation but no destruction. The description adds important behavioral details: it returns a 202-style payload with model_hash, requires polling get_model_status, and notes that unknown keys in var_priors are rejected. It also specifies that fitting starts asynchronously, which is beyond what annotations provide. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured with a purpose statement, contextual background, a detailed Args section, and a Returns note. Each sentence serves a purpose, explaining either the tool's role or parameter constraints. It's slightly verbose but justified given the 12 parameters and the need to compensate for zero schema coverage. The key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (12 params, no schema coverage, mutation, asynchronous behavior), the description is remarkably complete. It covers prerequisites (dataset content, row count), parameter relationships, defaults for optional params, and the return type with follow-up action. Nothing an agent needs to invoke correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter documentation. It explains every parameter in the Args section, adding meaning well beyond the schema: e.g., uploaded_file_id must contain every named column, date_column cannot also be a series, endogenous_vars must have at least two, lags requires enough rows, base_variable must be endogenous, etc. This is essential for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Create and start fitting a long-term (VAR) model'. It distinguishes itself from the sibling link_var_model by explaining that this tool creates the model while link_var_model connects it to an MMM. The reference to the MMM's long_run_rollup gives domain context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: for modeling joint dynamics of series and producing long-run elasticity. It explicitly mentions the follow-up step of linking with link_var_model, which guides the agent on the workflow. However, it doesn't explicitly state when NOT to use this tool or list alternative model-creation tools (like create_model), so it's not fully explicit about exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
declare_study_holdout_useDeclare Study Holdout UseAIdempotent
Append a submitter-reported evidence-use declaration using project-owner credentials and create:models. Read get_study_prediction_access first and reference an access event from this run. Use a fresh UUID declaration_id and reuse it unchanged on retry. informed_revision requires a published affected revision in the same study; other dispositions omit it. Reason must explain actual use. Reported revision influence requires fresh validation for affected revisions; later review-only notes cannot erase it. Does not certify independence, accept or promote a model. API submissions remain identified as reported declarations.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| run_id | Yes | ||
| disposition | Yes | ||
| declaration_id | Yes | ||
| source_access_id | Yes | ||
| affected_revision_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide idempotentHint and readOnlyHint=false, and the description adds behavioral context beyond them: write-scope requires project-owner credentials and create:models, declaration_id idempotency mechanics, revocation of stale revision influence, and the fact that API declarations stay marked as reported. This is rich behavioral disclosure with no contradiction to the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place by adding an operational constraint or clarification. Core purpose is front-loaded, followed by prerequisite steps, retry behavior, parameter conditions, and exclusions, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a non-trivial six-parameter mutation with zero schema descriptions, the description covers prerequisites, auth, idempotent retry, conditional parameter requirements, validation implications, and non-goals. An agent has enough context to call this tool correctly and to understand follow-up obligations such as fresh validation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it succeeds: declaration_id is tied to fresh-UUID/retry semantics, source_access_id to an access event from the run, disposition and affected_revision_id to the informed_revision conditional, and reason to explaining actual use. Every parameter receives meaningful semantic guidance beyond its bare title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Append a submitter-reported evidence-use declaration' names a specific verb, resource, and scope, and the final sentence explicitly excludes certification/promotion duties, differentiating this from sibling study actions. It also names get_study_prediction_access as a prerequisite, which helps distinguish the read path from this write path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use instructions: read get_study_prediction_access first, reference an access event from this run, and use a fresh UUID reused unchanged on retry. It also states conditional usage rules for informed_revision versus other dispositions, giving clear operational guidance without requiring the agent to guess.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_modelDelete ModelADestructiveIdempotent
PERMANENTLY DELETE a FAILED model. Destructive and irreversible.
Only models with status "failed" can be deleted over the API — any other status returns a 409 with the model's current status (delete is for cleaning up failed fits, not curating good ones). Deleting also unlinks any MMMs that pointed at it as their VAR model and removes stored artifacts. On success returns {"deleted_model_hash": ..., "status": "deleted"}.
Check first with get_model or get_model_status if unsure of the status.
Args: model_hash: Hash of the FAILED model to delete permanently.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this destructive and non-read-only, but the description adds meaningful context: deleting is irreversible, unlinks MMMs pointing to the model, removes stored artifacts, and returns a specific success payload. It also discloses the 409 error behavior for non-failed models. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: scope, irreversibility, precondition, side effects, return value, and pre-check advice are all included. The most critical warning is front-loaded, and the Args section is clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive one-parameter API, the description covers allowed inputs, error conditions, side effects, success response, and recommended pre-checks. Even with an output schema present, it explicitly shows the success payload, leaving little for an agent to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden. It defines model_hash as the hash of the FAILED model to delete permanently, adding the critical status constraint and purpose beyond the schema's bare 'Model Hash' title. It could mention how to obtain the hash or confirm its format, but for a single parameter this is adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('PERMANENTLY DELETE a FAILED model') and clearly scopes the operation to failed models only. This distinguishes it from sibling tools like get_model, rename_model, or unlink_var_model, even without inspecting their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool: only for failed models, and explicitly states that any other status returns a 409. It also advises checking with get_model or get_model_status first, giving an agent clear decision guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
design_incrementality_testDesign Incrementality TestAIdempotent
Ask a saved, complete MMM what an incrementality test on one of its media channels could detect: queues a bounded calculation that replays the saved posterior and returns {calculation_id, model_hash, status: "queued", submitted_at} at once (202; 200 with the same body when submission_key repeats identical inputs). Poll get_incrementality_test_design until status is complete, failed or cancelled. A complete calculation carries a result whose own state is available, unsupported (e.g. a log-link model), insufficient_evidence or no_feasible_design; none of these is an error, and only available carries numbers (candidates are also returned for no_feasible_design so you can see why nothing met the target). channel is one of the model's nonlinear media channels (400 unknown_channel otherwise). design_type time_holdout pauses or changes the channel's spend for a window and contrasts the outcome with the model's forecast; geo_split needs geo and splits a market panel into treatment and control. intervention gives the start date, candidate durations, spend change and baseline; inference the alpha, target power and an optional named effect; the backend fills and echoes every default. Reuse the same submission_key after a lost response; a new key queues another calculation and the same key with different inputs is refused (409 submission_key_conflict). Refusals: 404 model not found or not readable, 422 model_not_complete, 503 queue_unavailable (nothing is left behind). Nothing is saved or launched and no budget changes; save an available result with save_incrementality_test_design. Requires create:models to submit; polling needs read:results as well, so a key with only create:models can submit but not read.
| Name | Required | Description | Default |
|---|---|---|---|
| geo | No | ||
| channel | Yes | ||
| inference | No | ||
| model_hash | Yes | ||
| design_type | Yes | ||
| intervention | Yes | What changes in the channel's spend and when. start_date (YYYY-MM-DD) is on or after the model's last training period plus one, at most 13 weekly, 91 daily or 3 monthly periods ahead. durations lists 1-8 distinct candidate lengths in the model's cadence (1-52 periods); the backend evaluates each and selects one. spend_change is {mode: pause} (spend to zero), {mode: percent, pct} (pct in [-100, 500], not 0) or {mode: schedule, baseline: [...], intervention: [...]} (equal-length currency arrays covering the longest duration). baseline is {mode: recent_average, periods} (mean spend over the last 4-52 training periods) or {mode: schedule}; required for pause and percent. | |
| submission_key | Yes | Caller-generated 8-128 character identity for one intentional attempt. Reuse identical key and inputs after a lost response; a new key may consume another attempt. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations already covering idempotency and safety, the description still adds substantial behavior: 202 vs 200 response semantics, submission_key replay rules, the 409 submission_key_conflict, refusal codes (404, 422, 503 queue_unavailable with nothing left behind), and the permission model (create:models to submit, read:results to poll). It also enumerates result states (available/unsupported/insufficient_evidence/no_feasible_design) and clarifies none are errors. No annotation is contradicted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The return shape and the queue/poll model are front-loaded, and almost every sentence carries operational content. It is dense and long, with error and permission details packed into trailing run-ons, so it is information-rich but not maximally scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 7-parameter nested schema with low coverage, an output schema, and mutation-adjacent semantics, the description covers return body, status lifecycle, result states, defaults, error codes, and auth requirements. An agent has everything needed to submit, poll, and handle failures.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 29%, so the description must carry the load, and it does: channel must be a nonlinear media channel (400 unknown_channel), design_type semantics for time_holdout vs geo_split, the full intervention contract (start_date window, 1-8 durations, spend_change modes, baseline requirement), and inference defaults that are echoed. This adds meaning well beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: ask a saved, complete MMM what an incrementality test on one of its media channels could detect, and immediately frames it as a bounded queued calculation. It is clearly distinguished from siblings like get_incrementality_test_design (polling) and save_incrementality_test_design (persisting), so an agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the follow-up tool (get_incrementality_test_design) and the persistence path (save_incrementality_test_design), plus the explicit constraint that nothing is saved or launched and no budget changes. Usage context is strong but it never explicitly contrasts this with the sibling create_incrementality_test, leaving the design-vs-launch boundary inferential.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_quality_policiesDiff Quality PoliciesARead-onlyIdempotent
What changed from one saved policy (policy_id) to another (other_policy_id) in the same study, as the server computes it: checks added, removed and changed (per metric and field, from → to; matched by metric so reordering is not a change), protocol field changes, name_changed, rationale_changed and rules_changed (whether the rules hash differs). Read scope; nothing is written. Use it to explain a policy's lineage (get_quality_policy.derived_from) or to compare any two policies before choosing one.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes | ||
| policy_id | Yes | ||
| other_policy_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal readOnly and non-destructive; the description reinforces this with 'Read scope; nothing is written' and adds substantive behavioral detail: matching is by metric so reordering is not a change, rules comparison is hash-based, and the direction from → to is meaningful. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-load the operation and diff semantics, then add lineage/use-case context. No filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the operation, directionality, diff semantics, read-only behavior, and use cases. Output schema exists and annotations cover safety, so nothing needed for calling the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning, and it does: it defines policy_id as the source saved policy, other_policy_id as the target, and notes they must be in the same study. It does not explicitly describe study_id, but the phrase 'same study' and the parameter name make it clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with the precise operation: computing changes between two saved policies in the same study. It enumerates concrete diff dimensions (checks, protocol fields, name_changed, rationale_changed, rules_changed), which clearly distinguishes it from get_quality_policy and list_quality_policies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit use cases: explaining lineage via get_quality_policy.derived_from and comparing policies before selecting one. It names a sibling (get_quality_policy) as related, but does not explicitly state when this tool should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diff_recipe_revisionsDiff Recipe RevisionsARead-onlyIdempotent
What changed between two revisions of one recipe, with the viewer's labels (jellyfish #881): settings[] (key, section, label, from, to), priors[] (row, parameter, column, label, from, to), data (dataset origin and input-hash change, or null) and counts {settings, priors, data, total}, plus base and other {number, id}. Blank equals absent; prior rows are matched by variable and role. Read this after publishing revision N+1 to state exactly what an edit changed. 404 when the recipe or either revision is missing. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| base | Yes | ||
| other | Yes | ||
| recipe_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds genuine operational context: a 404 when the recipe or either revision is missing, the "blank equals absent" convention, and that prior rows are matched by variable and role. Return-format details duplicate the output schema, but the error and matching semantics are real additions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core answer, but a large share of the text enumerates the return payload (settings[], priors[], data, counts) that the output schema already defines, which is wasted space. Dense but not efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers error behavior and edge semantics for a read-only three-parameter tool, and the output schema handles return values. The main remaining gap is that the three required input parameters are never described, leaving their meaning to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the parameter burden. It implies base/other are the two revisions being compared and that recipe_id scopes them, but gives no types, ordering rules, or which is "from" vs "to". Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (diff) and resource (two revisions of one recipe) with the exact scope of comparison. An agent can distinguish it from diff_quality_policies and compare_study_runs without reading the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Read this after publishing revision N+1 to state exactly what an edit changed" gives clear timing/context for use. It does not name alternatives (e.g. get_recipe_revision for a single revision) or state when not to use it, so it stops short of the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evaluate_study_runEvaluate Study RunA
Save an immutable assessment, or with preview=true see what one would contain without writing (nothing stored, no access event). The preview returns basis_hash, would_evaluate, would_not_evaluate (each with basis.reason and, when an earlier assessment of this run can lend the value, carry_forward_available; otherwise carry_forward_blocked with why: basis_changed, policy_changed, manual_signoff) and previous_assessment. Each submission is complete: a custom check has a value only if you supply it now in external_evidence or name it in carry_forward from an earlier assessment whose evidence basis and policy rules still match; a metric may not be in both. Submit finite numeric or strict boolean values with method and source_reference and the expected_basis_hash from the preview; stale evidence is 409 stale_evidence. External calculations are submitter-reported, not verified; carried rows keep carried_from provenance. Manual sign-off is never carried and requires a signed-in reviewer. Every unevaluated check carries basis.reason from a closed set and a suggested action; not_evaluated never passes. Choose policy_id explicitly; the report echoes policy_name and policy_was_newest. No automatic champion promotion. Built-in errors are fitted-window, not holdout; VAR remains unsupported.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| preview | No | ||
| policy_id | Yes | ||
| carry_forward | No | ||
| external_evidence | No | ||
| expected_basis_hash | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations. It discloses that preview writes nothing and creates no access event, that stale evidence returns 409 stale_evidence, that external calculations are submitter-reported and not verified, that carried rows keep carried_from provenance, that manual sign-off is never carried, and that built-in errors are fitted-window not holdout. These are behavioral traits an agent needs to know and are not present in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with the most important behavioral facts front-loaded (immutable save, preview mode, no write). Every sentence adds a distinct fact, but the density is high and some sentences are packed with multiple clauses that require careful parsing. It is appropriately sized for the tool's complexity, though slightly less structured than ideal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, preview vs. write modes, carry-forward rules, error conditions, provenance semantics), the description covers the critical operational context: what happens on write, what preview returns, what causes 409, what is never carried, and what is unsupported. The output schema exists, so return values need not be described in detail. The description is complete enough for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the key parameter semantics: preview=true means no write, expected_basis_hash comes from the preview, carry_forward requires the metric to be listed under carry_forward_available, and external_evidence requires finite numeric or strict boolean values with method and source_reference. However, it doesn't explicitly map every parameter (e.g., run_id, policy_id) and doesn't explain the policy_id choice beyond 'Choose policy_id explicitly; the report echoes policy_name and policy_was_newest.' Still, the description adds substantial meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Save an immutable assessment, or with preview=true see what one would contain without writing.' This clearly distinguishes the two modes of the tool and names the resource (assessment of a study run). It also differentiates from siblings like list_study_evaluations and recommend_study_run by describing the write/preview behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: use preview=true to see what an assessment would contain without writing, and use the full submission to save. It also states exclusions: 'Manual sign-off is never carried and requires a signed-in reviewer,' 'VAR remains unsupported,' and 'No automatic champion promotion.' These exclusions help an agent decide when not to use this tool or what not to expect.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_backend_capabilitiesGet Backend CapabilitiesARead-onlyIdempotent
Discover this caller's connected backend features before planning work.
Returns only backend advertisements: model families, transformations, priors and workflow operations. A missing advertisement is unknown, not unsupported. Check each field; an advertised feature still requires permission and budget. No model is created and capabilities are not cached across callers.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations: it explains open-world semantics ('missing advertisement is unknown, not unsupported'), warns that advertised features still require permission and budget, and discloses that no model is created and capabilities are not cached across callers. This meaningfully informs how the agent should treat the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four compact sentences, each earning its place: when to use, what it returns, how to interpret missing/advertised features, and what side effects to expect. The key guidance is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters, rich annotations, and an output schema present, the description is fully sufficient for an agent to select and invoke the tool correctly. It covers timing, interpretation, permission nuance, and non-caching behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the description carries no parameter burden. The baseline of 4 applies, and the description still helps set expectations about what the returned result represents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Discover this caller's connected backend features' and clarifies it returns only backend advertisements for model families, transformations, priors, and workflow operations. This clearly differentiates it from the many sibling tools by scope and intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use this before planning work and explains how to interpret missing advertisements ('unknown, not unsupported'). It does not name exclusions or alternative tools, but none of the siblings compete with this capability-discovery role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaign_incrementalityGet Campaign IncrementalityARead-onlyIdempotent
Incremental ROAS per campaign (or ad set), beside the platform's own ROAS and last-click ROAS, by pushing the model's channel incrementality down through the platform's attribution.
THE ASSUMPTION, FIRST. Simba measures incrementality at channel grain. For each model channel
over the window, the incrementality factor = the channel's MMM incremental revenue (the
model's per-period rows, under its fitted attribution convention) / the platform-attributed
value of the campaigns mapped to that channel. Each campaign's incremental ROAS is that
factor x its platform ROAS, so campaign incremental revenue sums to the channel's. This
assumes the platform over-credits every campaign in a channel equally. It does not:
retargeting and brand search are over-credited more, so one factor flatters them. The
response warns when such campaigns share a channel with prospecting
(retargeting_shares_channel_factor); the remedies are to map them to their own model
channel (set_campaign_mapping) or to calibrate the factor with an incrementality test.
Nothing here is a causal per-campaign measurement; every row says how it was made.
Per row, method is "attribution_scaled" or, when a channel's campaigns carry no platform
value, "spend_share" (the channel's incremental revenue shared by spend). A campaign without
platform value in a channel that has some gets iroas: null and is named
(platform_value_missing); it is never given a share. factor_source is "mmm" or "test":
a completed incrementality test in the model's project that names the channel and overlaps
the window replaces the model's factor (lift in revenue units / the channel's platform value
during the test); factor_mmm stays beside it.
interval is "pending" (the 94% bands come from the model's posterior draws, computed by a
background job the first time a window is asked for; ask again in a few minutes), "ready"
(factor_interval per channel, iroas_hdi and incremental_revenue_hdi per row) or
"unavailable" (interval_reason says why; point estimates stand, no band is invented).
uninformative warns when a channel's band spans zero.
Returns {model_hash, window: {start, end}, level, currency, interval, interval_reason, channels: [{channel, factor, factor_source, factor_mmm, factor_interval, revenue_interval, factor_draws_mean, method, mmm_revenue, platform_value, spend, campaigns, test, warnings}], rows: [{platform, account_id, campaign_id, campaign_name, adset_id, channel, spend, platform_value, last_click_value, days, platform_roas, last_click_roas, incremental_revenue, iroas, incremental_revenue_hdi, iroas_hdi, method, factor_source}], unmapped: [{..., spend, platform_roas, last_click_roas}], warnings: [{code, message, channel?, campaigns?, reason?}], provenance: {source_versions, as_of, map_version, attribution_convention, link}}. Warning codes: retargeting_shares_channel_factor, platform_value_missing, kpi_not_revenue, currency_mismatch, uninformative, unmapped_spend, interval_unavailable, test_override_skipped.
Args: model_hash: A fitted MMM with a campaign map (set_campaign_mapping). start, end: Optional ISO dates (YYYY-MM-DD), inclusive. Default: the overlap of the model's data and the campaign facts. level: "campaign" (default) or "adset"; ad-set rows inherit their campaign's channel and factor and sum to the campaign row.
Errors carry a code: model_not_found (404), campaign_facts_empty (404: no facts, or none in the window), invalid_window (400: empty or reversed; the body gives both spans), model_not_mmm (400: a VAR model has no channel revenue rows), model_incomplete (400).
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | ||
| level | No | campaign | |
| start | No | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the read-only/idempotent safety profile, and the description adds a great deal beyond them: the core grain-assumption and its bias direction, method variants (attribution_scaled vs spend_share), factor_source override semantics, the 'pending' interval background-job behavior ('ask again in a few minutes'), nulling rules for platform_value_missing, and error codes. This is unusually rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and the crucial assumption are front-loaded, but the description then enumerates the entire return payload (channels/rows/unmapped/provenance fields) and the full warning-code list even though an output schema exists and already documents return values. That redundancy makes it considerably longer than necessary for its job.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a methodology-heavy analytical tool, the definition covers the assumption, method provenance, interval lifecycle, warning semantics, error codes, and parameter defaults. Nothing an agent needs to call it correctly or interpret its output caveats is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden and does: model_hash is defined as 'A fitted MMM with a campaign map (set_campaign_mapping)', start/end as inclusive ISO dates defaulting to the data/facts overlap, and level's campaign default with ad-set inheritance and summing behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line names a specific verb, resource and scope: 'Incremental ROAS per campaign (or ad set), beside the platform's own ROAS and last-click ROAS.' It further distinguishes the tool from siblings like get_campaign_report and get_incrementality_test by explaining the attribution-scaling methodology that is unique to this endpoint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong contextual guidance and even routes the agent elsewhere when the assumption is violated ('map them to their own model channel (set_campaign_mapping) or to calibrate the factor with an incrementality test'), plus the key caveat 'Nothing here is a causal per-campaign measurement.' It stops short of an explicit when-to-use-this-vs-that comparison against get_campaign_report, so it is clear context rather than full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaign_marginal_returnsGet Campaign Marginal ReturnsARead-onlyIdempotent
Read channel-derived marginal returns for campaigns or ad sets. Requires read:models.
Campaign response shapes inherit the fitted channel shape, rescaled by observed spend share and relative efficiency. These are not independently measured campaign saturation curves or causal campaign effects. Daily spend is observed spend divided by inclusive calendar days, not a configured platform budget or a daily revenue forecast.
Args: model_hash: Owned, completed MMM with compatible curve provenance and campaign facts. start, end: Inclusive observation-window ISO dates (YYYY-MM-DD). level: campaign or adset; ad sets inherit the campaign's channel.
Returns {context_key, window, currency, minor_digits, channels: [{channel, current_daily_spend, status, reason}], rows, provenance}. Rows carry composite identity, method and marginal-return evidence. Missing or incompatible currency, curve basis or fact coverage makes recommendations unavailable; do not substitute defaults. Missing posterior marginal evidence means no uncertainty interval, not zero uncertainty. This reads existing evidence only: no fit, posterior job or platform change is started.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | ||
| level | No | campaign | |
| start | Yes | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly/idempotent/openWorld), the description discloses substantive behavior: how campaign shapes are derived (rescaled by spend share and relative efficiency), that daily spend is observed spend over inclusive calendar days rather than a configured budget, that missing currency/curve/fact coverage yields unavailable recommendations, and that absent posterior evidence means no interval rather than zero uncertainty. It also confirms no fit or platform change is triggered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, then Args, then Returns, so structure is strong. It is somewhat long and the 'no fit, posterior job or platform change is started' clause partially overlaps the readOnly annotation, but most sentences carry genuine, non-redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytical tool with four params and an output schema, the definition is complete: prerequisites, input semantics, caveats on what the numbers do and do not mean, and failure behavior are all covered, and it even summarizes the return shape despite an output schema existing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden and delivers: model_hash must be an owned, completed MMM with compatible provenance; start/end are inclusive ISO dates (YYYY-MM-DD); level is campaign or adset with ad sets inheriting the campaign's channel. All four parameters get meaning the bare schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource (read channel-derived marginal returns for campaigns or ad sets) and delimits scope against adjacent concepts by stating these are 'not independently measured campaign saturation curves or causal campaign effects.' It does not name a sibling tool directly, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a precondition ('Requires read:models') and constrains model_hash to an 'Owned, completed MMM with compatible curve provenance and campaign facts,' plus a 'do not substitute defaults' rule on missing data. However, it never says when to prefer this over neighbors like get_campaign_report, get_campaign_incrementality, or recommend_campaign_budgets, so routing guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_campaign_reportGet Campaign ReportARead-onlyIdempotent
Report the campaign facts over any window, at any grain, by platform, model channel, campaign or ad set.
The same report engine as get_data_report, over the campaign facts: spend, media:impressions, media:clicks, outcome:platform_conversions, outcome:platform_value and, where a source fills them, outcome:last_click_conversions and outcome:last_click_value. These are the PLATFORMS' own attributed numbers, not Simba's incremental attribution.
group_by: "channel" groups by the model channel each campaign counts towards under the
model's map (give model_hash); spend of unmapped campaigns appears as its own group,
"unmapped", and is never dropped or guessed into a channel. Without model_hash every row is
"unmapped".
Returns {granularity, data_through, rows: [{period_start, period_end, group, metric, value,
unit}], meta: {basis, aggregation, roles, channels}, as_of, source_versions, currency}.
as_of is when the facts were last ingested; source_versions the pipeline versions they
came from.
Args: model_hash: The model whose map decides the channels (optional; needed for group_by=channel). start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (daily, as stored), "week", "month", "quarter" or "year". group_by: "platform", "channel", "campaign" or "adset". Default: one total per period. platform, channel, campaign_id: Filters applied before grouping. metrics: Roles to include, e.g. ["spend", "outcome:platform_conversions"]. Default: all.
Errors carry a code: campaign_facts_empty (404, nothing matches), invalid_report_request (400), report_too_large (413: narrow the window or coarsen the granularity).
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | ||
| start | No | ||
| channel | No | ||
| metrics | No | ||
| group_by | No | ||
| platform | No | ||
| model_hash | No | ||
| campaign_id | No | ||
| granularity | No | native |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), and the description adds substantial context beyond them: the 'unmapped' group guarantee that spend is 'never dropped or guessed into a channel,' the meaning of as_of and source_versions, and the error codes with their HTTP statuses and remediation. This is exactly the extra behavioral detail annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and sibling contrast, then a Returns block, then an Args block that mirrors the schema. Every sentence carries information, but the return-shape paragraph is partially redundant given an output schema exists, and the Args block, while necessary at 0% coverage, makes the definition long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, zero-required aggregation tool this is complete: purpose, metric availability caveats ('where a source fills them'), grouping edge case, return shape, defaults, and error taxonomy are all present. An agent has everything needed to call it correctly without further exploration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all nine parameters, and it does: model_hash (and its dependency for group_by=channel), start/end inclusivity, granularity semantics, group_by values with the 'unmapped' behavior, filters applied before grouping, and a metrics example with default. It adds meaning the schema cannot express, such as what happens without model_hash.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('report the campaign facts over any window, at any grain') and pins down the exact metric set (spend, media:impressions, media:clicks, platform conversions/value). It explicitly differentiates itself from siblings by naming get_data_report as the same engine over a different fact table, and from get_campaign_incrementality by noting these are 'the PLATFORMS' own attributed numbers, not Simba's incremental attribution.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to reach for this tool (platform-attributed campaign facts reporting) and contrasts it against incremental attribution, which routes the agent away from get_campaign_incrementality. It also documents the error cases and the remediation ('narrow the window or coarsen the granularity'). It stops short of an explicit 'use X instead when…' rule for get_data_report beyond the fact-table distinction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contribution_groupsGet Contribution GroupsCRead-onlyIdempotent
Read the stored contribution groups for a model (#436). Legacy dashboard-saved configs are served verbatim.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, idempotentHint, and destructiveHint false, so the safety profile is disclosed. The description adds 'Legacy dashboard-saved configs are served verbatim,' which is a behavioral nuance beyond the annotations, but it is terse and cryptic ('#436') and does not explain implications like formatting or transformation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short (two sentences) and front-loads the purpose, which is good. However, the second sentence is cryptic and not well structured; it references '#436' and a legacy behavior without context, making it less clear and slightly wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one parameter and an output schema, the description is minimal. It omits any explanation of the model_hash parameter and does not mention potential error cases or prerequisites. The cryptic legacy note adds confusion rather than completeness. The output schema exists, so return values are not required, but parameter description and edge cases are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the model_hash parameter at all. The parameter is named model_hash, which is somewhat self-explanatory, but the description adds no meaning about its format, requiredness, or how to obtain it. With zero coverage, the description must compensate, and it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads stored contribution groups for a model, using the verb 'Read' and naming the resource. It distinguishes from the sibling set_contribution_groups, though it does not explicitly name that sibling. The reference to '#436' is cryptic and not helpful.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (reading contribution groups) but provides no explicit guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. It does not mention that set_contribution_groups is the write counterpart or any conditions that would route an agent here.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_data_reportGet Data ReportARead-onlyIdempotent
Report actual data from a stored dataset: any window, any grain, by brand, channel or dimension.
Reads the dataset itself (every column, any date range) rather than a fitted model's training window. Use it for "sales and TV spend in the North region for August, by week".
Roles are DECLARED, never guessed from column names. Declare them with upload_data(roles=...)
or per request with roles. Without a declaration only the schema's own naming rules apply:
a column named date, {channel}_spend and {channel}_activity. Every other column is
reported as role "unknown" and is not aggregated — declare the KPI and hierarchy columns.
Role vocabulary (aggregation, unit) — also in get_data_schema under x-simba-roles:
kpi (sum), spend (sum, currency), activity (sum), multiplier (mean)
outcome:online_sales|store_sales|margin (sum, currency), outcome:orders|new_customers (sum)
media:impressions|clicks|grps (sum; give a channel: {"role": "media:grps", "channel": "tv"})
control:price|rate|index (mean), control:stock (each brand's last value, summed)
hierarchy, dimension:market|product|campaign (keys for filtering and group_by)
date
Buckets: week = ISO week from Monday; month/quarter = calendar. A weekly row counts in the month of its week-start date. The response's meta.aggregation states every rule applied.
Args: dataset_id: The uploaded file id (from upload_data or list_uploads). Registered pipeline outputs are uploaded files too. start, end: Optional ISO dates (YYYY-MM-DD), inclusive. granularity: "native" (default), "week", "month" or "quarter". group_by: "hierarchy", "channel", or a dimension role such as "dimension:market". hierarchy: Keep only this brand/region value. metrics: Roles or role families to include, e.g. ["kpi", "spend", "outcome:orders"] or ["control"]. Default: every metric role present. roles: {column: role | {"role", "channel"}} overriding roles stored at upload.
Returns {dataset: {id, name, source, version, sha256, data_through}, granularity, rows: [{period_start, period_end, group, metric, value, unit}], meta: {basis: "dataset", aggregation, roles, channels}}. Errors carry a code: dataset_not_found (404), invalid_report_request (400), report_too_large (413, over 10,000 rows — narrow the window or coarsen the granularity).
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | ||
| roles | No | ||
| start | No | ||
| metrics | No | ||
| group_by | No | ||
| hierarchy | No | ||
| dataset_id | Yes | ||
| granularity | No | native |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnly/idempotent/non-destructive): it documents aggregation rules, bucket semantics (ISO week from Monday, weekly rows counted in the month of week-start), the meta.aggregation disclosure, and specific error codes with statuses (404, 400, 413 with the 10,000-row cap and remediation). This is exactly the extra behavioral context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose and routes the reader through clearly labeled sections (roles, buckets, args, returns). It is dense and long, but given 0% schema coverage most of the length is load-bearing; only minor tightening is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the full lifecycle an agent needs: correct invocation, role declaration fallback, aggregation semantics, return shape, and error handling. Since an output schema exists it wisely only summarizes the return structure rather than fully restating it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it does: all 8 parameters are explained in an Args block (dataset_id origin, inclusive ISO start/end, granularity values, group_by options, hierarchy filter, metrics defaults, roles override format). It additionally supplies the role vocabulary and channel-qualified role syntax that the schema does not encode.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Report actual data from a stored dataset') with explicit scope ('any window, any grain, by brand, channel or dimension'). It also distinguishes itself from the model-based alternative by noting it 'Reads the dataset itself ... rather than a fitted model's training window', which separates it from siblings like get_model_results and get_data_schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete usage example ('sales and TV spend in the North region for August, by week') and implies the read-raw-data-vs-model distinction, plus the prerequisite that roles must be declared via upload_data(roles=...) or per request. It never explicitly names a sibling to avoid the way a 5 would, so guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_data_schemaGet Data SchemaARead-onlyIdempotent
Get the canonical CSV data schema for Simba MMM input files.
Returns the JSON Schema specification describing required columns (date, KPI, multiplier, hierarchy), media channel column naming conventions ({channel}_activity, {channel}_spend), constraints (min rows, max file size), and supported date formats.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds value by disclosing the tool's return contract: a JSON Schema specification listing required columns, naming conventions, constraints, and date formats. No contradiction exists, and the behavior is fully consistent with a read-only metadata query.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the action and resource, and the second sentence enumerates the return contents with no wasted words. Every sentence contributes useful information, and the structure is easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only schema tool with an output schema, the description is complete. It tells the agent exactly what to expect in the response, lists the schema's key contents, and needs no additional context to select or invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the input schema is trivially documented at 100% coverage. There are no parameter semantics for the description to clarify, and the baseline for zero-parameter tools is 4; the description appropriately focuses on the return value instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the canonical CSV data schema for Simba MMM input files.' It further clarifies the exact content returned (required columns, naming conventions, constraints, date formats), leaving no ambiguity about what the tool does. It is clearly distinguishable from all siblings, none of which offer a data schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: this is the canonical schema for Simba MMM input files, which signals when an agent should call it for constructing or validating input data. It does not explicitly name exclusions or alternatives, but no sibling tool serves the same schema-retrieval purpose, so the implicit use-case guidance is adequate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_incrementality_testGet Incrementality TestARead-onlyIdempotent
Read one recorded test: {id, version, record, content_hash, used_by, retired_at}. version reads an older version (default: current). With model_hash (a saved model you can read), the result also carries calibration: either {status: "ok", row: {channel, x, delta_x, delta_y, sigma, sigma_low?, sigma_high?}, units, steps, warnings} — the likelihood observation this test gives that model, each step stated — or {status: "refused", reason, message, steps}. A refusal is an answer, not an error: e.g. channel_not_in_model (pass channel, a model activity column), kpi_mismatch (pass confirm_kpi=true only if the test's outcome really is the model's KPI), no_spend, test_not_completed, owned_media_not_calibratable, window_overlaps_holdout. Use the same references in create_model(calibration={tests: [...]}). Requires the read:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| channel | No | ||
| test_id | Yes | ||
| version | No | ||
| model_hash | No | ||
| confirm_kpi | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial behavioral context beyond them: the calibration success vs. refusal shape, the enumerable refusal reasons, and the required read:models scope. This is unusually rich disclosure for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose first, then version, then the model_hash/calibration branch, then refusal handling and scope. Every clause carries meaning, though the inline brace-listing of base record fields is somewhat redundant against the output schema and adds bulk.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with a required scope, a rich conditional output (calibration), and five params at 0% schema coverage, the description supplies everything needed: id/version semantics, the model_hash branch, refusal enumeration, and cross-tool linkage. With an output schema present, it need not do more.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it does: version (older version, default current), model_hash (a saved model you can read, triggers calibration), channel (a model activity column, passed for channel_not_in_model), and confirm_kpi (pass true only if the test's outcome really is the model's KPI). Four of five params are given real semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read one recorded test') and enumerates the record shape ({id, version, record, content_hash, used_by, retired_at}), which cleanly distinguishes it from the sibling list_incrementality_tests (singular fetch vs. list). An agent can pick this tool without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear usage context: version reads older versions (default current), model_hash enriches the result, and refusals are answers not errors with named remediations. It cross-links to create_model(calibration={tests: [...]}). It stops short of an explicit 'when-not/what-alternative' statement (e.g. pointing to list_incrementality_tests to find ids), so not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_incrementality_test_designGet Incrementality Test DesignBRead-onlyIdempotent
Read one test-design calculation: {calculation_id, model_hash, status, submitted_at, started_at, completed_at, request, error, result}. status is operational: queued, running, complete, failed or cancelled; error is {code, message} only when failed (artifact_unreadable, timeout, worker_lost, internal) and a failed calculation carries no numbers. result is present only when complete and is read by its own state, never by key presence: available, unsupported, insufficient_evidence or no_feasible_design, with reasons [{code, detail}] non-empty exactly when not available, plus warnings, assumptions, diagnostics and provenance (method, model artefact, posterior draws, training periods). An available result carries intervention (dates, carryover and measurement end, spend change), model_implied_effect (mean and 94% HDI, parameter uncertainty only; available may be false, e.g. a geo split on a national model, which then gives a national reference only), detectable_effect (cumulative, per-period, relative and, for revenue KPIs, iROAS MDE at the stated alpha and target power), power (at the model mean effect, at the named effect, and assurance), noise, and candidates per duration with the selected one marked. Power here is posterior-averaged: assuming the model is correctly specified and its posterior calibrated, the probability that the pre-specified analysis rejects the null, averaged over the model's forecast uncertainty and observation noise; it is not the probability that this experiment will detect the effect, and a model-implied effect is not a prediction of what the experiment will measure. Requires read:results.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes | ||
| calculation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so safety is covered. The description adds genuinely useful behavioral context beyond the schema: the read:results permission requirement and the non-obvious rule that result must be read by its own state rather than key presence, with defined state values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense run-on paragraph that buries the operation's purpose behind an exhaustive payload enumeration. It is information-rich but poorly structured and far longer than needed given that an output schema already exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity and the presence of an output schema, the description is thorough about state semantics, error codes, auth requirements and the caveats around power/model-implied effect. The main gap is that it never grounds the two required identifiers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never defines model_hash or calculation_id beyond echoing them as output fields. It does not say where these identifiers come from or how they relate, so the description fails to compensate for the coverage gap on the two required params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one test-design calculation') and lists the identifying fields and payload shape, so an agent can tell it is a singular retrieval of a design calculation. It does not, however, explicitly differentiate itself from siblings like design_incrementality_test or get_incrementality_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no mention of alternatives. The only usage signal is the trailing 'Requires read:results' permission note; the agent must infer the workflow position from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_launch_eligibilityGet Launch EligibilityARead-onlyIdempotent
Read, before launching, exactly what launch_study_run would refuse with: can_launch plus blockers, each with the stable code (permission_denied, policy_not_found, policy_retired, study_inactive, revision_not_found, attempts_exhausted, concurrency_exhausted, unsupported_family, engine_changed, snapshot_not_executable), the human message and a next_action. Also returns the budget, the policy used (newest active when policy_id is omitted) and the revision with its engine state. Consumes nothing. For engine_changed, call refreeze_recipe_revision and launch the new revision; never retry the old one.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes | ||
| policy_id | No | ||
| revision_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description reinforces this with 'Read' and 'Consumes nothing.' It adds genuinely non-obvious behavior beyond annotations: the policy selection rule ('newest active when policy_id is omitted') and the disclosure that the revision comes back with its engine state. It doesn't fully cover auth requirements or rate limits, but for a read-only tool with full annotation coverage, the description adds meaningful value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first clause, followed logically by returns, safety, and the special-case workflow. The enumeration of ten blocker codes is long, but each is a stable contract value agents need for downstream branching; it borders on redundancy with the output schema but earns its place overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read tool with a rich annotation set and an output schema, the description is unusually complete: it covers purpose, when to call it, what it returns, safety ('Consumes nothing'), the policy default, and the one critical post-call workflow branch (engine_changed). Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the only non-obvious parameter, policy_id, including its omission behavior ('newest active when policy_id is omitted'), and clarifies what revision_id yields ('the revision with its engine state'). study_id is left to name inference, but given the tool's pre-launch context, it is trivially inferable. This is strong compensation for a 0%-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read, before launching, exactly what launch_study_run would refuse with.' It enumerates the precise deliverable (can_launch plus blockers with stable codes, messages, next_action), which distinguishes it from its siblings launch_study_run (the mutation it preflights), get_study_run, and get_recipe_revision. The first clause alone would differentiate it; the detail makes the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit temporal trigger ('before launching'), names the sibling it pairs with (launch_study_run), and provides a concrete conditional routing: 'For engine_changed, call refreeze_recipe_revision and launch the new revision; never retry the old one.' The 'never retry' clause is an explicit when-not, which is rare and valuable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_modelGet ModelARead-onlyIdempotent
Get a model's metadata and configuration echo — works for EVERY status, including failed models (unlike get_model_results, which needs 'complete').
Use this to inspect what a model was configured with, why it failed, or where it lives. Returns: id, model_hash, name, status, model_type ("mmm"/"var"), hierarchy_value, periodicity, is_saved, project_id/name, linked_var_model_hash, created_at/completed_at, error (the failure message — non-null only when status is "failed"), and model_config (the create-time configuration echo: data_source, columns, channels, priors as resolved, and the config flags).
NOTE: the echo omits a few accepted create_model inputs (operating_margin, annual_discount_rate, reporting_kernel) — absence there does not mean they weren't applied; check the financials results section for the stored margin.
Args: model_hash: The model hash (any status).
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only and idempotent. The description adds substantive behavioral detail beyond that: it works for any status, the error field is non-null only for failed models, and it lists the configuration echo and its known omissions. This materially helps an agent predict behavior without over-claiming.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized: opening purpose, use cases, returns, and a note about omissions. Every sentence carries useful information, including the critical caveat about omitted create_model inputs. The most important differentiator (works for any status) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple one-parameter input, the existing output schema, and the annotations, the description covers all decision-relevant details: why to use this over siblings, what the return shape means, and an important limitation of the echo. No critical gap remains for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides the parameter name and type with 0% description coverage. The description's Args section adds 'The model hash (any status),' clarifying that this identifier is valid even for failed models. For a single parameter this is adequate, though more detail on how to obtain the hash could be helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'Get a model's metadata and configuration echo.' It clearly states what the tool returns and explicitly contrasts itself with get_model_results, making its scope unambiguous and differentiating it from a close sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: 'works for EVERY status, including failed models' and contrasts with get_model_results which 'needs complete.' It also explains common use cases: inspect configuration, understand failure, or locate a model. This is direct and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_resultsGet Model ResultsA
Get results from a completed model.
Available sections:
channel_summary: per-channel aggregates {Channel, Sales, Spend, Revenue, ROI}.
contributions: per-period decomposition (Date, one column per channel, plus Base, Seasonality, Event Effect, Model, Fit Actual, Actual). Values are in KPI/unit space — the multiplier is NOT applied. Use
coefficientsfor per-period revenue. Multiplicative (link="log") models fitted with the removal_lift attribution convention add anOverlapcolumn: a balancing residual, which can have either sign when effects are signed, so that Base + components + Overlap = Model. Overlap is NOT a channel — never rank it, share it, or feed it to the optimizer/scenarios. Overlap requires BOTH link="log" AND attribution="removal_lift" (the API default): under aumann_shapley (the dashboard default for multiplicative models since #509), shapley, or proportional_normalized, the interaction is allocated across components, which close exactly with NO Overlap column — its absence does NOT mean the model is additive or predates the feature. Control columns are measured against the reference point resolved at fit time (#452, see model_config.control_references) — e.g. "vs. average conditions" for a control that never reaches zero — not necessarily against zero, so a referenced control's series legitimately spans zero.coefficients: per-period per-channel media results table (Date, Channel, Sales, Revenue, Spend, Media Units, ROI, Cost/Revenue/Sales per Media Unit). This is the only per-period revenue-space decomposition.
params: fitted posterior means per channel (alpha, decay, cpu, scalars).
decay_curves: adstock decay per channel (mean/lower/upper, l_max, adstock_type, curve points; dual-geometric models add decay_slow_* and dual_weight_* parameters).
response_curves: 100-point spend-vs-revenue grid per channel with credible bands ({ch}, {ch}_lower, {ch}_lower_50, {ch}_upper_50, {ch}_upper).
marginal_curves: same grid for marginal ROI (diminishing returns).
saturation: fitted saturation family and parameters (saturation_type is tanh, michaelis_menten, negative_exponential, or generalized_log; per-channel alpha/scale, plus transform_order and — for generalized_log only — per-channel sat_shape).
mroi_summary: headline marginal ROI at current spend per channel with a 94% HDI (channel, current_spend, mroi_median, mroi_hdi_3, mroi_hdi_97). Post-#591 posterior fits add two averaging-convention scalars per channel — mroi_allperiods_unweighted_median (+_hdi_3/_hdi_97) and mroi_spendweighted_active_median (+_hdi_3/_hdi_97), with profit variants on margin models — plus a top-level conventions_available array. Channels with no active periods omit the spendweighted fields. Post-#629 fits also carry a *_mean beside every *_median (mroi_mean, mroi_profit_mean, pv_kernel_mass_mean, and the convention variants). Preserve the requested mean or median explicitly; they are not interchangeable. The mean reconciles with the marginal-revenue curve, since derivative and mean commute and median does not. Absent on anything fitted before #629 — there is no backfill, so feature-detect rather than assume.
mroi_periods: OPT-IN ONLY (#591) — never in the default payload; request it by name in
sections. Per-period marginal ROI series: {available, hdi_prob, evaluation_point: "historical_period_spend", rows} with one row per (channel x modelled period): channel, date, spend, mroi_median/_hdi_3/hdi_97, and mroi_profit* on margin models. Models fitted before the artifact existed return {available: false, reason: "fitted_before_mroi_periods"} — refit to enable. Large (channels x periods) — pair with the channels filter.model_stats: fit diagnostics (R², MAPE, Durbin-Watson, Max R_hat, ...).
actual_vs_model: actual vs predicted per period with 50%/95% HDIs.
long_run_rollup: MMM short-term + VAR long-run revenue rollup per channel; returns {available: false, reason: "no_linked_var_model"} when no VAR model is linked to this MMM. Joins by exact name unless the link declared a channel_map (see link_var_model) — mapped rows carry var_group and an allocated elasticity slice, with group-level truth in metadata.groups. A computed rollup where nothing joined stays available: true but carries reason: "no_channel_overlap" — check metadata.coverage, then declare a channel_map on the link.
optimizer: latest optimization results (see get_optimizer_results).
predictions: latest scenario prediction rows (see get_scenario_results).
prediction_window: OPT-IN ONLY saved prediction-window actuals/model values. Request sections="prediction_window" (JSON or CSV); omitted by default. This is not certified untouched holdout evidence. For study-linked models, serving it appends an access audit event; channel/grid filters do not alter it.
posterior: full posterior summary table — one row per model variable with mean, sd, hdi_3%, hdi_97%, and r_hat (quotable 94% HDIs and per-variable convergence).
posterior_transforms: the importable transform-parameter posterior grid (what the dashboard's prior builder imports): per-channel alpha mean/sd, decay 94% HDI, dual-weight mean/sd, decay-slow HDI, sat-shape mean/sd, and the adstock structure including tied-group member aliases. Rows key on activity-column names — join via channel_map.
r_hat: per-parameter R-hat over ALL posterior variables — including transform RVs such as {channel}_decay that the posterior summary's coefficient rows do not cover. Use it to attribute a bad Max R_hat (model_stats) to a specific parameter block.
financials: the model's operating margin ({operating_margin, operating_margin_series}); omitted entirely for marginless models. operating_margin_series is a DATE-STRING-KEYED DICT ({"2024-01-01": 0.18, ...}), not a list of records.
cohort_ledger: per-(channel, source-period) forward-allocation ledger — each period's spend is credited with the future effects its adstock carryover earns (horizon slices plus PV-discounted financials from the fit-time cohort kernels). Models fitted before the artifact existed return {available: false, reason: ...} — feature-detect on
available.model_config: the resolved model specification (inputs, not posteriors) to audit or reconstruct the create_model call — includes config flags such as saturation_type, transform_order, and link ("log" = multiplicative). Multiplicative models with controls also report control_references (#452): per control, the requested and resolved attribution reference mode, the zero_distance diagnostic behind the "auto" choice, and the posterior-mean q_ref. Models created before these fields existed may omit them. priors_resolved reports what the fit actually consumed (#643): per row, overridden_fields lists only the fields that took effect, and accepted_not_used — present only when non-empty — names any that were accepted but inert for this model's configuration, each with a reason. A prior field can be spelled correctly and still do nothing: theta_* needs adstock_type "delayed", dual_weight_* needs "dual_geometric", sat_shape_* needs saturation_type "generalized_log", and the decay / half-life bounds are ignored FOR "dual_geometric". If a prior you set appears to have had no influence, read accepted_not_used first. The folded coordinates (half_marginal_*, effect_at_avg_*) are never called inert — they land in the row's scalars/alpha_sd/mean/sd.
channel_map: canonical identifier mapping, one record per channel: {channel, activity_column, spend_column} as configured at create time. This is the join key between channels[].name and the sections keyed by activity-column name (contributions, decay_curves, posterior_transforms).
The response envelope includes sections_available — trust it over any
hardcoded list if the server is newer than these docs.
IMPORTANT — channel naming: results are keyed by the channel's ACTIVITY
COLUMN name (e.g. "search_activity"), not by the channels[].name passed to
create_model. These exact keys (case- and space-sensitive) must be used in
run_optimizer bounds, laydown_weights, and period_cpm. Always read
channel_summary first to get the exact keys.
NOTE: Date values in contributions/coefficients records are millisecond epoch integers.
DATE WINDOW: pass start / end (ISO dates, inclusive) and/or granularity
("native", "week", "month" or "quarter") to window contributions,
coefficients, actual_vs_model and channel_summary. The response then
carries meta (window, basis, data_through, aggregation rules,
not_windowed). channel_summary is RECOMPUTED for the window — ROI =
ΣRevenue/ΣSpend per channel, profit priced with each period's own margin —
never filtered or averaged. Bucketed rows carry period_start/period_end
instead of Date; per-unit ratios are recomputed from sums; bucketed
actual_vs_model drops the per-period predictive intervals (they cannot be
added). mROI is never summed: mroi_periods comes back as fitted and is
listed in meta.not_windowed. VAR models window actual_vs_model only. The
window covers the model's training period; for data outside it (or
columns the model did not use) use get_data_report.
CONTEXT-SIZE TIP: a full pull is very large (curve sections alone are 100 grid points x channels x 5 band columns). In conversational use, request only the sections you need and pass channels=[...] and max_grid_points=20.
Explicit JSON channel or grid requests add _mcp_selection metadata describing
requested selection, changed sections, original and returned row counts and grid
sampling. Backend metadata and warnings remain unchanged. Alias collisions
retain every matching exact identifier and are disclosed; never combine them.
Unmatched aliases refer only to filterable sections. Empty channels and grid
limits below 2 retain existing no-op behaviour with a warning. Unfiltered and
CSV results are unchanged. Local filtering does not bound backend downloads.
Args:
model_hash: The model hash.
sections: Comma-separated list of sections to include.
Leave empty for all sections.
Common: "channel_summary,model_stats" for ROI and diagnostics.
format: "json" (default) or "csv". CSV returns
{"format": "csv", "content": "..."} — concatenated
"# section" + CSV blocks, useful for saving to disk.
Filtering below applies to JSON only.
channels: Optional channel filter (matching is case/space-insensitive
and tolerates the _activity/_spend suffix). Applied to curve
sections, decay_curves, saturation, channel_summary,
coefficients, mroi_summary, and mroi_periods rows.
contributions is never filtered (its control columns are
indistinguishable from channels client-side).
max_response_bytes: Optional UTF-8 JSON result byte ceiling after filtering.
Oversize results return an actionable error, never partial evidence.
Bounds MCP content, not the backend HTTP download.
max_grid_points: Optional cap on response/marginal curve grid points;
records are strided evenly, keeping first and last.
start: Optional window start (YYYY-MM-DD), inclusive.
end: Optional window end (YYYY-MM-DD), inclusive.
granularity: Optional "native", "week", "month" or "quarter".
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | ||
| start | No | ||
| format | No | json | |
| channels | No | ||
| sections | No | ||
| model_hash | Yes | ||
| granularity | No | ||
| max_grid_points | No | ||
| max_response_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and idempotentHint=false; the description explains the side effect ('serving it appends an access audit event' for study-linked models), which is consistent and clarifying. It also discloses non-obvious behavior well beyond annotations: opt-in sections never in the default payload, feature-detect patterns ({available:false, reason:...}), no backfill for pre-#629 fits, oversize results returning an actionable error rather than partial evidence, and that local filtering does not bound backend downloads.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well organized with headers, bullets, and front-loaded purpose, but grossly over-long for a tool description — the per-section prose plus issue-number archaeology (#509, #452, #591, #629, #643) and edge-case narrations push it well past what an agent needs to select and invoke correctly. Structure is good; size is not.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with ~20 sections, the description covers everything needed: channel-key naming (activity column vs channels[].name), date windowing semantics and meta.not_windowed, filter scope, feature detection for absent artifacts, and the trust-sections_available-over-docs rule. An output schema exists, yet the section shapes are also described, leaving no ambiguity about what comes back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does: sections (comma-separated, empty = all, common combos), format (json/csv and the exact CSV return shape), channels (case/space-insensitive matching, suffix tolerance, which sections it applies to, and that contributions is never filtered), max_response_bytes, max_grid_points striding rules, start/end inclusivity, and granularity values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Get results from a completed model') and then enumerates the exact sections available, so an agent knows precisely what surface this covers. It does not, however, explicitly differentiate itself from close siblings like get_model, show_response_curves, or show_decomposition — it only cross-references get_optimizer_results and get_scenario_results for two sub-sections.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance throughout: which sections are opt-in only (mroi_periods, prediction_window), 'Always read channel_summary first to get the exact keys', 'Use coefficients for per-period revenue', and an explicit routing rule to get_data_report for data outside the training window. The context-size tip tells the agent how to call it in conversational use (subset sections, channels filter, max_grid_points).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_statusGet Model StatusARead-onlyIdempotent
Check the fitting progress of a model.
Returns status (pending/under way/complete/failed), progress percentage, estimated time remaining, and timestamps. When supported by the backend, fit_liveness reports heartbeat age and the configured stall threshold in seconds, with last_heartbeat_at as Unix seconds. seconds_until_stall_threshold is time to the stale-heartbeat threshold, not fit ETA or an exact kill time. stall_threshold_exceeded does not change the model status. If fit_liveness is absent or available is false, liveness is unknown: do not infer a healthy or stalled fit. Reasons include heartbeat_unavailable and not_fitting. Continue polling with backoff; do not automatically restart a fit.
Args: model_hash: The model hash returned by create_model or list_models.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the annotations (which already declare readOnly, idempotent, etc.) by explaining subtle behaviors: fit_liveness heartbeat semantics, stall_threshold meaning, that stall_threshold_exceeded does not change status, and how to handle missing liveness data. This level of detail is exceptional and eliminates ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with a clear purpose statement, then details on return values, liveness behavior, and polling guidance. Every sentence adds value; there is no filler. The critical information is front-loaded, and the parameter explanation is appended logically. It is concise for the complexity it covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (liveness, stall thresholds, edge cases) and that an output schema exists (so return format is covered elsewhere), the description is fully complete. It addresses all possible scenarios an agent might encounter: supported vs unsupported backend, absent liveness, stall threshold semantics, and appropriate polling behavior. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully explain the parameter. It does: 'model_hash: The model hash returned by create_model or list_models.' This adds clear meaning about where to obtain the value, which is essential for correct usage. The description fully compensates for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check the fitting progress of a model.' It specifies the resource (model) and the action (check status). It also distinguishes itself from sibling tools like get_model_results or get_optimizer_results by focusing on fitting progress, not results. The description is explicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it's for polling fitting progress, and explicitly instructs to 'Continue polling with backoff; do not automatically restart a fit.' It also warns about how to interpret liveness data and when not to infer health. While it doesn't name specific alternative tools, the guidance on when to use and what to avoid is sufficient for an agent to decide correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_optimizer_resultsGet Optimizer ResultsARead-onlyIdempotent
Get budget optimization status and results.
Without run_id: returns the MODEL-LEVEL optimizer state. Top-level keys:
optimizer_status ("none"/"pending"/"under way"/"complete"/"failed"),
progress + progress_text while running, and results when complete.
This reflects the LATEST run on the model — a newer run overwrites it, so
a poller can lose sight of the run it submitted.
With run_id (run_optimizer's response includes it): fetches that specific
run, immune to later runs. Top-level keys include run_id, model_hash,
status, created_at, label, inputs, and results. Poll THIS form
when you need to know whether your own run completed.
Reading results rows — the columns come from DIFFERENT conventions and
must not be treated as interchangeable:
Revenue/ROI: the optimizer's DECISION math — removal-lift counterfactual revenue at the allocated spend. This is what the solver optimized.OptimizedEvalRevenue/OptimizedEvalROIandHistoricalRevenue/HistoricalROI: fitted-convention COMPARISON columns — the reconciled accounting view matching the model's Contributions panel. Same spend, different question; never mix them withRevenue/ROIin one summary.ObjectiveMarginal: the decision-math marginal return at the optimum (the quantity the solver equalizes across unconstrained channels).MroiAtOptimized/MroiAtOptimizedHdi3/MroiAtOptimizedHdi97: posterior mROI evaluated at the optimized spend (94% HDI bounds) — a DIFFERENT quantity from ObjectiveMarginal (they can differ by several times); quote the one matching the question asked.Convergence / KKT certificate fields report solver health. All-None placeholder arrays (PeriodResponse etc.) are stripped server-side.
Args: model_hash: Hash of the model that was optimized. run_id: Optional optimization run id from run_optimizer's response. Pass it to poll a specific run's status/results.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds substantial behavioral context: the overwrite semantics of the latest run, the distinction between decision-math and fitted-convention columns, the stripping of all-None placeholder arrays, and the warning that ObjectiveMarginal and MroiAtOptimized are different quantities. This goes far beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every section earns its place: the two-mode distinction, the column-convention warning, and the parameter explanations are all necessary for correct use. It is front-loaded with the core purpose and mode distinction, then details. Slightly dense, but not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity — two modes, multiple result conventions, and a rich output schema — the description is remarkably complete. It explains the top-level keys for both modes, warns about column incompatibility, and clarifies which quantity answers which question. The output schema exists, so return-value details are not the description's job, and the description still adds the interpretive context the schema cannot.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning. It does: model_hash is 'Hash of the model that was optimized,' and run_id is 'Optional optimization run id from run_optimizer's response' with guidance to pass it to poll a specific run. The description could add a bit more about the format of model_hash, but it fully compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('budget optimization status and results'), and immediately distinguishes two modes: without run_id returns model-level state, with run_id returns a specific run. This clearly differentiates it from siblings like run_optimizer, get_model_results, and get_scenario_results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use each form: poll without run_id for the latest model-level state, and pass run_id to poll a specific run's completion. It also warns that a newer run overwrites the latest state, so a poller can lose sight of its own run — a clear when-not-to-use signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pipeline_runGet Pipeline RunARead-onlyIdempotent
Get one run of an owned pipeline: {run_id, status (queued | running | succeeded | failed), started_at, finished_at (UTC), version_id, error_code, error}. On success version_id is the new saved version — pass it as pipeline_version_id to get_recipe_draft_template to build on the refreshed data. On failure error_code says why: execution_failed (a source or transform failed; error names it), no_output, timeout, interrupted or not_started (start it again), not_queued, owner_blocked, unexpected. Scheduled runs are polled the same way. Requires the create:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| pipeline_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds genuinely new behavioral context: the create:models scope requirement, the meaning of each status value, and actionable error_code semantics (re-start on interrupted/not_started, error names the failing source/transform). This is far beyond what structured fields convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the returned shape, then next-step routing, error semantics, and scope. Dense but each clause earns its place; slightly packed with parenthetical enumerations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description needn't explain return values but does so helpfully, and it adds scope and error handling. The remaining gap is the undocumented parameter formats, which leaves the 0%-coverage schema slightly under-supported.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for pipeline_ref and run_id, but it never defines pipeline_ref's format (id vs slug) or where to obtain it. It only obliquely references run_id via the returned-object notation, so it partially but not fully compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get one run of an owned pipeline') and enumerates the returned fields, so the agent knows exactly what it retrieves. It does not explicitly differentiate itself from nearby siblings such as list_pipelines or list_pipeline_versions, leaving that gap to the reader.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear workflow context: poll scheduled runs the same way, and on success route version_id into get_recipe_draft_template. It stops short of stating when to prefer this over alternatives like list_runs or list_pipeline_versions, so there are no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_quality_policyGet Quality PolicyARead-onlyIdempotent
Read one immutable policy in full: specification, content_hash (identity of the stored row), rules_hash (identity of the rules alone, the value evidence carry-forward compares), is_newest, derived_from ({policy_id, content_hash, name} when created from another policy via derived_from_policy_id; name and hash are null if that source was since deleted; see diff_quality_policies), checks_summary (builtin / custom_numeric / boolean / manual / diagnostic, required, advisory, cap) and protocol_summary, plus usage with the ids of every run, assessment, resolution and Champion acceptance that references it. Shared viewers can read; nothing is written.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes | ||
| policy_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint and destructiveHint, so the safety profile is covered. The description adds valuable behavioral context beyond that: the meaning of content_hash vs rules_hash, the nullability of derived_from name/hash when the source has been deleted, and the fact that 'nothing is written.' This helps the agent understand side effects and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense run-on sentence packed with field lists and parentheticals. It is front-loaded with the core purpose, and every piece adds information, but the lack of punctuation breaks makes it harder to parse. It is appropriately sized for the complexity but could be structured better with separate sentences or bullet-like breaks.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return format is covered structurally. The description enriches context by explaining hash semantics, derived_from edge cases, and checks_summary composition. It includes the access note ('Shared viewers can read') and explicitly asserts no writes. For a read-only fetch with two params, this is quite complete, though it omits error behavior for missing policies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It does not. It never explains study_id or policy_id, nor whether policy_id is globally unique or scoped to the study. The only clue is the tool name and schema titles, which are generic. This is a significant gap for a tool with two required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read one immutable policy in full' and enumerates the exact fields returned, including hashes and usage references. It differentiates from siblings like list_quality_policies (listing) and diff_quality_policies (comparing) through its focus on a single policy's full detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the detailed single-policy reader but never explicitly contrasts it with alternatives. The only sibling mention is 'see diff_quality_policies' in the context of the derived_from field, which is a cross-reference, not a when-to-use guide. There is no statement like 'for a list, use list_quality_policies' or 'for comparisons, use diff_quality_policies'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recipe_draftGet Recipe DraftARead-onlyIdempotent
Read the complete authoring snapshot and concurrency version. Preserve all fields when editing. Draft state is incomplete, unvalidated authoring data, not an executable recipe.
| Name | Required | Description | Default |
|---|---|---|---|
| draft_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world, and non-destructive behavior, so the bar is lower. The description adds useful context: the response is a complete snapshot with a concurrency version, edits should preserve all fields, and the draft is not an executable recipe.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each adding distinct information: action, editing directive, and state warning. It is front-loaded with the core read operation and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single obvious parameter, an output schema, and rich annotations, the definition covers the important non-obvious semantics: completeness, concurrency, and draft validity. It could go further by explaining how the concurrency version should be used when updating, but this is not essential for a read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no property description and the tool description never mentions draft_id, so it doesn't compensate for 0% schema description coverage. The parameter name is self-explanatory, but the description adds no meaning about how to obtain or format the ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence uses a clear verb ('Read') with a specific resource ('complete authoring snapshot and concurrency version'), going well beyond the tool's title. It also distinguishes the draft-read operation from siblings like get_recipe_draft_template or publish_recipe_draft by emphasizing completeness and draft state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction 'Preserve all fields when editing' provides real guidance for a read-then-edit workflow, and the warning that drafts are unvalidated authoring data sets expectations. However, it never names sibling tools or explicitly states when to choose get_recipe_draft over get_recipe_draft_template, get_recipe_revision_authoring, or list_recipe_drafts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recipe_draft_templateGet Recipe Draft TemplateARead-onlyIdempotent
Get complete shared wizard defaults, hash and envelope schema. Optionally choose an owned uploaded_file_id (from list_uploads) or pipeline_version_id (from list_pipelines / list_pipeline_versions), never both, into frozen source bytes with verified lineage and an editable data preview. Copy snapshot into create_recipe_draft and preserve unedited fields. Defaults are not a validated model. Does not create, publish or run anything.
| Name | Required | Description | Default |
|---|---|---|---|
| family | No | mmm | |
| uploaded_file_id | No | ||
| pipeline_version_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only, idempotent, and non-destructive. The description adds non-obvious behavior: defaults are not a validated model, retrieved data is frozen with verified lineage, and the preview is editable. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core deliverable, then parameter options, then caveats. No filler or repetition of schema titles.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return format needs no explanation. Combined with schema, the description covers the deliverable, parameter sourcing, constraints, downstream workflow, and non-mutating behavior sufficiently for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions are 0% so the description must carry parameter meaning. It meaningfully explains uploaded_file_id and pipeline_version_id by naming their source tools and exclusivity, but family is left to its enum values with no explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Uses a specific verb ('get') and names the resource: shared wizard defaults, hash, and envelope schema. Also states what it is not ('Does not create, publish or run anything'), separating it from create_recipe_draft and publish_recipe_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete sourcing instructions: uploaded_file_id comes from list_uploads and pipeline_version_id from list_pipelines/list_pipeline_versions, with an explicit 'never both' constraint. It describes downstream use (copy snapshot into create_recipe_draft) but does not explicitly route to a sibling like get_recipe_draft when a saved draft is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recipe_revisionGet Recipe RevisionBRead-onlyIdempotent
Read one immutable revision with its inspection block. settings..value is what a fit of this revision would use, including fitter defaults for absent keys (status authored or default); effective.form_data is what was authored. priors[].fields lists conditional prior fields that were stored but never read for that row's adstock type or the saturation family (status inert, with the gate that would make them live); treat them as stored-but-unused and do not "fix" them by editing unless the gating setting changes too. engine.state current or stale predicts whether launch will be refused. Absent configuration classifies nothing. inspection.lineage (jellyfish #896) is the dataset the model was built from, as recorded — origin {kind, id, pipeline_id, version, sha256, pipeline_name, dataset_name} — with available checked now for you (true: the recorded source is still yours and hashes the same; false: gone or changed, with reason; null: nothing was recorded, or checked is false), and display, the same line people see on the recipe card ("Retail weekly · v3 · verified", "dataset lineage not recorded · legacy"). Names are display only; ids and sha256 verify. Nothing is inferred from a pipeline name or a later output.
| Name | Required | Description | Default |
|---|---|---|---|
| number | Yes | ||
| recipe_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: immutability of the revision, that engine.state current/stale predicts whether launch will be refused, and that lineage availability is checked live with true/false/null meanings. It is let down only by being heavily oriented toward return-value semantics the output schema already carries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The most important statement is front-loaded, and individual sentences do carry information rather than filler. However, the body is a dense, jargon-heavy wall of text with nested parentheticals and quoted example strings that makes it hard to parse for an agent scanning for the operative facts.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with an output schema, explaining return fields is largely redundant, and the description nonetheless spends most of its length there. It is fairly complete on what the response means but omits the arguments entirely, which is the more important gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither of the two required parameters (recipe_id, number) is mentioned anywhere in the description. 'Read one immutable revision' hints that some identifier designates the revision, but the description adds essentially no meaning about what these arguments are or their format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Read one immutable revision with its inspection block.' That clearly distinguishes it from mutating siblings, but the description never differentiates it from close siblings such as get_recipe_revision_authoring or diff_recipe_revisions, leaving the agent to infer which revision-reading tool to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated — it is the tool for inspecting a single stored revision. The description does give actionable interpretation guidance ('treat them as stored-but-unused and do not "fix" them by editing unless the gating setting changes too'), but there is no explicit when-to-use/when-not or routing to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recipe_revision_authoringGet Recipe Revision AuthoringARead-onlyIdempotent
Read the authoring snapshot behind a published revision. authoring_draft, wizard (Save a recipe) and base_model (imported fitted model) revisions carry one; api_mmm, model_snapshot and legacy rows return an explicit unavailable error (404). Returns snapshot, name, revision_id, draft_content_hash, kind, source_available and source_unavailable_reason. Dataset bytes are filled from the recorded origin only when it still hashes to what was fitted; otherwise snapshot.source is null, source_available is false and the reason says to choose the dataset again before the draft can publish. Use the snapshot with create_recipe_draft: target for an in-place edit, source_revision_id for a branch. The published revision stays unchanged. Does not create or fit anything.
| Name | Required | Description | Default |
|---|---|---|---|
| number | Yes | ||
| recipe_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (which already declare read-only, idempotent, non-destructive), the description discloses a specific error contract (404 for api_mmm/model_snapshot/legacy), the exact fallback semantics when dataset bytes no longer hash to the fitted origin (source is null, source_available false, reason), and the reassuring guarantee that 'the published revision stays unchanged' and it 'does not create or fit anything'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences that are front-loaded with the core purpose, followed by return semantics, the dataset-hash caveat, and downstream usage. Every sentence earns its place, though it runs long enough that the parameter meaning never gets a sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description needn't enumerate return fields — yet it explains the non-obvious one (source_available / source_unavailable_reason) rather than just listing names. Combined with the error and downstream-usage context, the definition is complete except for the unexplained 'number' parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no meaning for either parameter. 'recipe_id' is self-evident from the name, but 'number' is ambiguous (revision number vs ordinal) and is never explained; the description only mentions revision_id as a return field, not as input semantics. With 0% coverage the description should have compensated and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the authoring snapshot behind a published revision') and immediately narrows scope by enumerating which revision kinds (authoring_draft, wizard, base_model) carry a snapshot and which (api_mmm, model_snapshot, legacy) do not. This distinguishes it cleanly from the sibling get_recipe_revision and get_recipe_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditions for success (three revision kinds) and failure (404 for the others), and routes the agent onward to create_recipe_draft with the correct argument ('target for an in-place edit, source_revision_id for a branch'). It does not explicitly compare against the sibling get_recipe_revision, so a small inference remains, but the when-to-use context is otherwise clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scenario_resultsGet Scenario ResultsARead-onlyIdempotent
Get scenario prediction results.
Without run_id: returns the MODEL-LEVEL scenario state — status (pending/complete/failed) and, when complete, the full prediction data including predicted KPI per period, channel contributions, confidence intervals, and base components (intercept, seasonality, trend). This reflects the LATEST scenario on the model — a newer run overwrites it, so a poller can lose sight of the run it submitted.
With run_id (run_scenario's response includes it): fetches that specific
saved run, immune to later runs — keys include run_id, model_hash,
name, status, pinned, notes, tags, key_metrics, timestamps,
inputs (the submitted payload), and results. Poll THIS form when you
need to know whether your own run completed, or to disambiguate
back-to-back scenarios.
NOTE: Failed scenarios return status "failed" with an error message in the JSON body (not an HTTP error). Always check the status field.
Args: model_hash: Hash of the model the scenario was run on. run_id: Optional scenario run id ("scn_..."), from run_scenario's response or list_runs(artifact="scenario").
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and non-destructive behavior. The description adds meaningful behavioral context beyond these: a newer run overwrites the latest scenario, and failed scenarios return an HTTP 200 with status 'failed' plus an error message in the body, so callers must check the status field.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but earns its length: mode-by-mode behavior, a critical failure-mode note, and an Args block. The core distinction between the two forms is front-loaded, and every section adds necessary information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so detailed return-value documentation is unnecessary. For a two-mode polling tool, the description covers overwrite semantics, failure representation, run_id provenance, and when to use each form, leaving no practical gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains model_hash as the hash of the model the scenario was run on, and defines run_id's format ('scn_...') plus where to obtain it: from run_scenario's response or list_runs(artifact='scenario'). This fully covers both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pairing ('Get scenario prediction results') and then defines two distinct behaviors: model-level latest state versus a specific saved run identified by run_id. This clearly distinguishes it from run_scenario and list_runs, which are the most related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the condition for each invocation form: without run_id returns the latest scenario state with overwrite risk, while with run_id fetches a specific saved run. It also says directly, 'Poll THIS form when you need to know whether your own run completed, or to disambiguate back-to-back scenarios,' which is unambiguous usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scenario_templateGet Scenario TemplateARead-onlyIdempotent
Generate a forward-period scenario template from a completed model.
Returns future dates pre-filled with values from 1 year prior, the list of media and control channels, and average cost-per-unit per media channel.
IMPORTANT: Always call this before run_scenario or run_optimizer to discover:
Channel names (use these exact names in scenario_data, bounds, laydown_weights, period_cpm)
Average CPM per channel (avg_cpu_by_channel — use for period_cpm in run_optimizer)
Baseline activity values per channel (rows — use as starting point for scenarios)
Media vs control channel classification (variable_classification field)
The response also includes: operating_margin (the model's stored margin, if set — useful for profit math), variable_transforms (per-variable transform metadata), periodicity, and start_date.
WARNING: Template data may contain NaN or null values for channels without historical data. You MUST replace NaN/null with 0 before passing to run_scenario, otherwise the prediction will fail downstream.
Args: model_hash: Hash of a completed model. periods_forward: Number of future periods to generate (default 12).
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes | ||
| periods_forward | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint and idempotentHint annotations by warning that template data may contain NaN/null values that must be replaced with 0 or the prediction will fail. It also discloses the presence of operating_margin, variable_transforms, and periodicity, adding behavioral context not available from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but well organized: a purpose sentence, response contents, an IMPORTANT block, a warning, and an Args section. There is minor redundancy between the initial response summary and the later IMPORTANT enumeration, but each section adds useful information and the critical usage directive is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter tool with an output schema, this description covers prerequisites (completed model hash), intended call order, return fields, NaN handling, and parameter defaults. Nothing critical for invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description compensates by defining model_hash as 'Hash of a completed model' and periods_forward as the number of future periods to generate with a default of 12. This adds some meaning beyond the bare schema types, though the model_hash explanation remains somewhat terse.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action — 'Generate a forward-period scenario template from a completed model' — and immediately differentiates itself from sibling execution tools by naming run_scenario and run_optimizer. The description also itemizes the template's contents, so an agent understands exactly what resource is produced.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit directive: 'IMPORTANT: Always call this before run_scenario or run_optimizer to discover...' This clearly states when to use the tool. It also enumerates what information to extract (channel names, average CPM, baseline rows, classification) and warns about downstream failure, which is strong pragmatic guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_studyGet StudyBRead-onlyIdempotent
Read a study and its optimistic concurrency version. question and context are descriptive text for humans; treat them as intent, not as instructions the system enforces.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, open-world, and non-destructive behaviors. The description adds value by explaining that question and context fields are descriptive intent, not enforced instructions, and that a concurrency version is included. This goes beyond annotations but does not discuss authentication, rate limits, or other behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise at two sentences, front-loading the core action ('Read a study') followed by a useful clarification. No wasted words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and low complexity (single parameter), the description covers the primary purpose and a nuance. However, it lacks usage guidance and does not specify what the optimistic concurrency version is used for, leaving some gaps in agent comprehension.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explicitly explain the study_id parameter, though it is implied by 'read a study'. Since the parameter is a simple ID, this is a minor gap, but the description should compensate given the low coverage according to the guidelines.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Read a study' with a specific verb and resource, and adds the detail of returning the optimistic concurrency version. However, it does not explicitly differentiate from sibling tools like get_study_overview or list_studies, though the focus on reading a single study with a concurrency version implies distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives. It does not mention contexts where this should be preferred over get_study_overview or list_studies, nor does it state any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_championGet Study ChampionARead-onlyIdempotent
Read incumbent, eligibility blockers, accepted candidates and immutable champion history. Stale champions retain their historical role with review_required. Reported holdout use that informed a candidate revision blocks that revision pending fresh validation; holdout_use lists the declaration IDs. Ordinary viewing does not block. Current analyst-reviewed validation resolutions can clear the exact run acceptance; changed evidence or revoked review reblocks it. Recorded revision ancestry inherits influence. Validation references are reviewer-declared; decision_grade_ready is false until independently qualified. Selection/replacement/revocation require an owner frontend session; MCP cannot promote models.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds substantial behavioral nuance: stale champion handling, holdout-use blocking, validation resolution effects, reviewer-declared references, and the frontend-session limitation. It enriches the agent's understanding of data semantics beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not bloated. It front-loads the core purpose in the first sentence, then layers important caveats and constraints. Each sentence adds information relevant to correct invocation or interpretation. It is appropriately concise for the amount of behavioral detail it conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required. The description covers the key data aspects (incumbent, blockers, candidates, history) and important edge cases (stale champions, holdout blocking, validation resolutions, session limitations). It is complete enough for an agent to understand what this tool returns and the conditions that affect its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the tool description must compensate. It does not mention study_id at all, leaving the agent to infer its meaning from the parameter name. For a read tool with a single obvious parameter, the lack of explicit explanation is a minor gap, but given the low coverage, the description fails to add any value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Read incumbent, eligibility blockers, accepted candidates and immutable champion history.' This precisely states what the tool returns and distinguishes it from other study-related tools like get_study or get_study_run by focusing on the champion concept. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides behavioral context (e.g., 'Ordinary viewing does not block', 'Selection/replacement/revocation require an owner frontend session') that implies this is a read-only query tool. However, it never explicitly names alternative tools or states when to prefer this over a sibling like get_study. Usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_overviewGet Study OverviewARead-onlyIdempotent
Read a whole study at a glance in one small call: the study, its budget, each recipe with its latest revision (id, number, hash), run counts by state and the latest run, active policies with newest_active_id, the champion summary and the last decision. Identity and counts only, no frozen configuration or reports. Call this first, then expand exactly the recipe, run or assessment you need (list_study_recipes with expand, get_recipe_revision, list_study_evaluations with expand).
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral context beyond the annotations: it is a lightweight 'identity and counts only' call that deliberately excludes frozen configuration and reports, and it returns the latest revision/run/decision rather than full history. This helps an agent predict cost and scope of the call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized: it front-loads the purpose, lists the included content in a compact series, states exclusions, and ends with a clear next-step routing. Every sentence earns its place, and the length is appropriate for the amount of useful routing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a single required parameter, a rich output schema, and safety annotations. The description covers what the call returns, what it deliberately omits, and how to follow up. There is no missing information an agent needs to decide whether to call this tool or to interpret its scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the single study_id parameter. It does so implicitly by stating the tool reads 'a whole study' and by naming the study-related entities returned. However, it does not explicitly state that study_id is the identifier to pass, nor does it give format guidance. With only one obvious parameter, the baseline is high, and the description is mostly adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'Read a whole study at a glance in one small call' and then enumerates exactly what is included (study, budget, recipes with latest revisions, run counts, active policies, champion summary, last decision). It also explicitly states what is NOT included ('no frozen configuration or reports'), which distinguishes it from heavier study-reading tools like get_study or list_study_recipes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Call this first, then expand exactly the recipe, run or assessment you need' and names the specific sibling tools to use next (list_study_recipes with expand, get_recipe_revision, list_study_evaluations with expand). This is a clear when-to-use and what-to-use-instead directive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_prediction_accessGet Study Prediction AccessARead-onlyIdempotent
Read partial prediction-access history for this run and matching recorded dataset/windows in this study: recorded_accesses, recorded_runs, by_action (what the total is made of) and the 20 most recent events. An event is recorded only when prediction evidence is served by a deliberate action: saving an assessment (assessment), a comparison (comparison, one per served run), a pair assessment (pair_assessment), dashboard results (dashboard_results) and prediction-window exports (results_csv, results_json). Listing assessments never records. Histories written before jellyfish #837 may also contain assessment_history rows from list reads; nothing is deleted because holdout-use declarations reference event ids. This read does not expose predictions or add events. Earlier activity, other result routes and offline work are not covered; absence never proves untouched holdout status. Repeated access does not prove retuning.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations. It discloses that the read does not expose predictions or add events, that listing assessments never records, that histories before jellyfish #837 may contain extra rows, and that nothing is deleted because holdout-use declarations reference event ids. These are critical behavioral traits not captured by readOnlyHint/openWorldHint/idempotentHint/destructiveHint. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It front-loads the core purpose and return fields, then adds necessary caveats about event recording, historical anomalies, and limitations. It is longer than average, but the complexity of the tool justifies the length. Slight redundancy in the caveats could be trimmed, but overall it is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the single parameter, the rich annotations, and the presence of an output schema, the description is complete. It explains what is returned, what is not covered, how events are recorded, and the implications for holdout status. An agent has everything needed to decide when to call this tool and how to interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the single parameter run_id. The description repeatedly references 'this run' and 'matching recorded dataset/windows in this study', implying run_id identifies the run, but it does not explicitly state that run_id is the run identifier or explain its format. However, with only one parameter and the title 'Run Id' in the schema, the meaning is largely inferable. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('partial prediction-access history for this run and matching recorded dataset/windows in this study'), and enumerates the exact fields returned. It clearly distinguishes itself from sibling tools by focusing on prediction-access history rather than study metadata, runs, or models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when this tool is appropriate: to read prediction-access history for a run and its recorded dataset/windows. It also provides exclusions ('Earlier activity, other result routes and offline work are not covered; absence never proves untouched holdout status. Repeated access does not prove retuning.'), which tells an agent when NOT to rely on this tool. This is strong usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_runGet Study RunARead-onlyIdempotent
Read durable run status and the linked model.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds that the tool returns durable status and the linked model, which is useful context, but it does not clarify what 'durable' means or describe other behavioral constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence with no filler. The core action and key output are front-loaded, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with strong annotations and an output schema, this is nearly complete. The only slight gap is the ambiguity of 'durable' and the lack of any link to the broader study context, but the agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter, run_id, with 0% description coverage, and the tool description does not compensate by explaining the parameter's meaning, format, or source. The property name is self-evident, but the description adds no value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read') with a clear object ('durable run status') and a distinct output ('the linked model'). This differentiates it from sibling getters like get_model_status or list_study_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives. The sibling list includes several related read tools, but the description does not name them or state any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_validation_resolutionsGet Study Validation ResolutionsARead-onlyIdempotent
Read analyst validation resolutions and revocations, including current/stale/revoked status re-evaluated against exact evidence. API keys cannot supply human independence sign-off or revoke it; use the signed-in owner UI. No audit serving event is added and no model is fitted or promoted.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint false, but the description adds valuable behavioral details beyond that: no audit serving event is added, no model is fitted or promoted, and statuses are re-evaluated against exact evidence. It also clarifies the human-only nature of sign-off and revocation. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three purposeful sentences with the primary action front-loaded. Each sentence adds distinct value: the read scope, the human-only limitation, and the side-effect guarantees. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, rich annotations, and a single obvious parameter, the description covers purpose, side effects, and usage limitations well. The only meaningful gap is that study_id is undocumented in both the schema and the description, which slightly reduces completeness for an agent encountering this tool cold.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and there is one required parameter, study_id. The description does not mention study_id at all, nor does it explain what values are valid or how the parameter scopes the returned resolutions. The parameter name is self-evident, but the description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read analyst validation resolutions and revocations'. It also clarifies the scope with 'current/stale/revoked status re-evaluated against exact evidence', which distinguishes it from nearby assessment or evaluation tools like assess_study_validation_pair or list_study_evaluations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear when-not boundary: 'API keys cannot supply human independence sign-off or revoke it; use the signed-in owner UI.' This tells an agent not to attempt sign-off or revocation through this API. However, it does not explicitly compare this tool to sibling MCP tools for reading validation data, so it stops short of full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_uploadGet UploadARead-onlyIdempotent
Get one uploaded dataset's details, including its column schema.
Returns id, filename, original_filename, source_type, mime_type, file_size, row_count, column_count, columns ([{name, dtype}, ...] — use these to build create_model's channel/control column arguments without re-reading the CSV), and created_at.
Args: file_id: The upload's id, from upload_data's response or list_uploads.
| Name | Required | Description | Default |
|---|---|---|---|
| file_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, covering the safety profile. The description adds useful non-obvious context: it returns a structured column schema and that this avoids re-reading the CSV for model building. This goes beyond what annotations alone provide without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose, followed by a compact return list and a clear Args section. Despite listing many fields, every sentence earns its place because it helps the agent know exactly what will be returned and how to use it. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the single parameter, the safe-read annotations, and the rich output schema, the description is complete for tool selection and invocation. It explains the parameter source, the return contents, and the downstream use case, leaving no practical gap for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides only 'integer' and 'File Id' with 0% description coverage, so the description must carry the semantics. It does: 'The upload's id, from upload_data's response or list_uploads.' This fully explains the parameter's meaning and provenance, which is exactly what an agent needs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'Get one uploaded dataset's details, including its column schema.' This clearly distinguishes it from list_uploads (which lists uploads) and upload_data (which creates an upload). The resource and output are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when this tool is useful: it explains that the returned columns can be used to build create_model's channel/control arguments without re-reading the CSV. It also specifies where to obtain file_id (from upload_data or list_uploads). It stops short of naming alternatives or explicit when-not-to-use cases, so it does not earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workflow_guidanceGet Workflow GuidanceARead-onlyIdempotent
Read versioned Simba workflow guidance when native Skills are unavailable.
Use index to list allowed topic/section IDs, then request the relevant section. Topics: mmm, results, priors, optimiser, studies, var. The default section is entrypoint; it names detailed sections and when they are needed. Arguments are identifiers, never paths. Returns a complete section within a 24,000-byte structured-payload limit or a refusal, never truncated instructions. This local read makes no backend request. Guidance does not authorise writes or replace validation. Lookup is optional when the caller already has the relevant versioned guidance.
| Name | Required | Description | Default |
|---|---|---|---|
| topic | No | index | |
| section | No | entrypoint |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantial context beyond them: a 24,000-byte structured-payload limit, a refusal rather than truncation, no backend request, and the caveat that guidance does not authorise writes or replace validation. These are exactly the behavioral traits an agent needs and cannot get from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and usage, then behavioral constraints. Most sentences earn their place, though the topic enumeration and the final optional-lookup sentence could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the description still covers the payload limit and refusal semantics. Combined with the usage pattern, topic list, and identifier caveat, an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it does: it enumerates the valid topics (mmm, results, priors, optimiser, studies, var), explains that the default section is 'entrypoint' and what it contains, and warns that arguments are identifiers, never paths. It stops short of documenting section ID syntax, but that is delegated to the index lookup.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read versioned Simba workflow guidance') plus the exact condition under which it applies ('when native Skills are unavailable'). An agent can distinguish it from sibling tools like get_recipe_revision_authoring or get_backend_capabilities without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit two-step usage pattern ('Use index to list allowed topic/section IDs, then request the relevant section') and an explicit when-not ('Lookup is optional when the caller already has the relevant versioned guidance'). The alternative to this tool (native Skills) is named directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_incrementality_testsImport Incrementality TestsA
Import tests from another tool's output file: source is csv (Simba's template), meta_conversion_lift (Conversion Lift API results JSON), geox (a meridian-geox analysis result), geolift (GeoLift summary), causalpy (effect summary or lift rows) or pymc_marketing (lift rows); content is the file's text (10 MB max). Returns {records: [{key, record, errors}], created, notes}. dry_run (default true) creates nothing — review each row's errors, then call again with dry_run=false to create the rows without errors (their ids come back in created). defaults fills fields the file doesn't carry, e.g. {"channel": "TV", "model_channel": "tv_grps", "kpi": {"kind": "revenue"}}: a GeoX result names no channel or KPI, so its rows fail until those are given. overrides sets fields on one row by the key the dry run showed (a cell id, or the 1-based row number), e.g. {"1": {"spend": {"incremental": 25000}}}. Values deep-merge: the source's assumptions, then defaults, then the file, then overrides; a null removes a field. import_invalid means the file isn't that source's format. Requires the create:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| content | Yes | ||
| dry_run | No | ||
| defaults | No | ||
| overrides | No | ||
| project_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses the dry-run safety workflow, the create:models scope requirement, the return shape, and the precise deep-merge precedence (source assumptions, then defaults, then file, then overrides) including that a null removes a field. That is rich behavioral context an agent cannot get from the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is correctly front-loaded and every clause carries information, but it is delivered as one dense semicolon-chained paragraph, which slightly hurts scannability for a six-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, return values need not be explained, yet the description still summarizes them; combined with the scope, merge-order, and per-source failure notes, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does: it defines each source enum value, the 10 MB content limit, dry_run's default, and gives concrete examples for defaults and overrides including key formats (cell id or 1-based row number). This is more than the schema conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Import tests from another tool's output file') and enumerates each supported source format, which clearly distinguishes it from siblings like create_incrementality_test and list_incrementality_tests.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit workflow guidance: dry_run defaults to true and creates nothing, so the agent should review each row's errors then re-call with dry_run=false. It also names the failure condition (import_invalid) and the scope required (create:models), leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
launch_study_runLaunch Study RunAIdempotent
Launch a NEW fit of a frozen revision within study attempt/concurrency budgets; it never opens existing results. Call get_launch_eligibility first: a refusal here carries the same code, message and next_action. Requires an active study, an executable revision frozen on the current engine, and an active same-study policy (choose it deliberately; the newest is not always the intended one). Budget/state conflicts require inspection, not a new attempt key. Reuse the same submission_key after an ambiguous response; never invent another key for a retry. engine_changed means re-freeze (refreeze_recipe_revision) and launch the new revision.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes | ||
| policy_id | Yes | ||
| revision_id | Yes | ||
| submission_key | Yes | Caller-generated 8-128 character identity for one intentional attempt. Reuse identical key and inputs after a lost response; a new key may consume another attempt. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description richly supplements the annotations by explaining retry behavior with submission_key reuse, warning against inventing new keys, and noting that budget conflicts require inspection rather than a new attempt. It also clarifies that the tool never opens existing results, which is useful beyond readOnlyHint/idempotentHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence carries operational value. It front-loads the core purpose, then adds prerequisites, retry protocol, conflict handling, and engine-change remediation in a logical order without unnecessary filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations cover safety/mutation semantics, the description completes the picture: prerequisites, budget behavior, retry key discipline, and engine-change re-freeze workflow. An agent has enough information to invoke this tool correctly and to know when to route to sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 25% schema description coverage, the description compensates by contextualizing all four required parameters: study_id must be active, revision_id must be frozen and executable, policy_id must be an active same-study policy, and submission_key is the caller-generated retry identity. This is meaningful semantic guidance beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action and resource: 'Launch a NEW fit of a frozen revision within study attempt/concurrency budgets; it never opens existing results.' This clearly distinguishes the tool from sibling read-only tools like get_study_run or list_study_runs by emphasizing new launches rather than result inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: call get_launch_eligibility first, ensure an active study, a frozen revision on the current engine, and an active same-study policy. It also gives when-not-to-use guidance for budget/state conflicts and points to refreeze_recipe_revision when engine_changed occurs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_var_modelLink Var ModelADestructive
Link a completed VAR model to an MMM (#569).
After linking, the MMM's get_model_results long_run_rollup section
joins the VAR's long-run elasticities with the MMM's short-term revenue.
A VAR links to at most one MMM at a time — the error names the current
owner if it is already linked elsewhere.
The join is by exact name unless channel_map declares which MMM channels each VAR exogenous series stands for (#682) — required whenever the VAR is fitted on group spends (e.g. four spend groups) while the MMM is tactic-level. Each group's elasticity is allocated across its member channels pro-rata by KPI short-term contribution, so the group's long-run effect is counted exactly once. Validation is strict: keys must be VAR exogenous series, values must be channel names of the (completed) MMM, and no channel may belong to two groups. The map belongs to the link: every link replaces it (omitting channel_map clears any stored map) and unlink clears it.
Args: model_hash: The MMM to attach the long-run view to. var_model_hash: The VAR model (from create_var_model). channel_map: Optional {var_exogenous_series: [mmm_channel, ...]} mapping for group-level VARs.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes | ||
| channel_map | No | ||
| var_model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal destructive and non-idempotent behavior, but the description goes far beyond them: it explains the post-link effect on get_model_results, the at-most-one-MMM constraint, the error behavior naming the current owner, strict validation rules, channel_map replacement semantics, and clearing on unlink. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but structured and front-loaded with the core purpose before diving into details. Every sentence adds operational value, though the depth of detail around channel_map and validation makes it denser than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a linking tool with destructive semantics and non-trivial validation, the description covers the full context: prerequisites, result effects, constraints, error behavior, parameter semantics, and state lifecycle. The existing output schema covers return value details, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries full responsibility. It clearly defines model_hash as the MMM to attach to, var_model_hash via create_var_model, and channel_map with its type, purpose, validation rules, and lifecycle. The description compensates completely for the sparse schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a specific verb and resource: 'Link a completed VAR model to an MMM'. This clearly distinguishes it from the sibling unlink_var_model and other model-management tools. The subsequent explanation of what linking does to results makes the tool's role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when linking is appropriate, including the one-link-per-VAR constraint and the channel_map requirement for group-level VARs. It mentions 'unlink clears it', implicitly pointing to the sibling unlink tool, though it never explicitly instructs the agent to use unlink_var_model for removal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_campaignsList CampaignsARead-onlyIdempotent
List the campaigns in the campaign facts, each with its totals and the model channel it counts towards for this model.
The channel is DECLARED in the model's campaign map (set_campaign_mapping), never inferred:
a campaign no map row covers is status: "unmapped" with channel: null, is listed with its
spend, and is not counted towards any channel. suggested_channel is a hint from the
platform's own label and the campaign's name (null when nothing is clear); it does nothing
until it is written into the map.
Returns {campaigns: [{platform, account_id, campaign_id, campaign_name, adsets, spend,
impressions, clicks, platform_conversions, platform_value, platform_channel_type, first_seen,
last_seen, channel, suggested_channel, status: "mapped" | "unmapped"}], window: {start, end},
currency, as_of, next_cursor (only when limit was given)}. Sorted by platform, account and
campaign id. Totals cover the window (default: every stored day).
Args: model_hash: The model whose map decides each campaign's channel (the hash of a fitted MMM). platform: Keep one platform: meta, google_ads, tiktok or other. unmapped_only: Only campaigns without a channel for this model. start, end: Optional ISO dates (YYYY-MM-DD), inclusive, for the totals. limit: Page size (1-200). Without it every campaign is returned. cursor: The previous page's next_cursor.
Errors carry a code: model_not_found (404). An empty list means no pipeline is registered as the campaign facts source yet, or it has not run; registration happens in the app.
| Name | Required | Description | Default |
|---|---|---|---|
| end | No | ||
| limit | No | ||
| start | No | ||
| cursor | No | ||
| platform | No | ||
| model_hash | Yes | ||
| unmapped_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read/idempotent/destructive safety, but the description adds pagination behavior (next_cursor only when limit is given), sort order, default window ('every stored day'), the meaning of an empty result, and the model_not_found error code — all beyond structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and dense, but the detailed Returns block is redundant given an output schema exists, making it longer than strictly necessary. Still, the prose mostly earns its place through channel/status semantics not documented elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a filtered, paginated read with an output schema, the description covers every arg, the channel-resolution model, pagination, empty-state cause, and errors. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the full parameter burden — and it does, documenting every arg including model_hash semantics, platform enum values, unmapped_only, inclusive ISO start/end, limit range, and cursor. Excellent compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) and resource (campaigns) with scope: campaign facts, totals, and the model channel. The channel-semantics explanation ('declared in the model's campaign map, never inferred') makes clear how this differs from a plain report tool like get_campaign_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when unmapped campaigns appear and that set_campaign_mapping is what assigns channels, giving the agent solid context. However, it never states when to prefer this over sibling report tools such as get_campaign_report or get_campaign_incrementality.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_incrementality_testsList Incrementality TestsARead-onlyIdempotent
List a project's recorded (not retired) incrementality tests: {items: [{id, name, type, status, channel, model_channel, start_date, end_date, measured_through, kpi, result, spend, source_tool, has_supplied_row, current_version, used_by}], next_cursor}; used_by counts the model revisions built from the test. Null and negative results are listed like any other. Filter by type, status or channel. Paging is opt-in: pass limit (1-200) and send next_cursor back unchanged; null means the end. Requires the read:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| limit | No | ||
| cursor | No | ||
| status | No | ||
| channel | No | ||
| project_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, open-world, so the safety profile is covered. The description adds real value beyond them: the paging contract (opt-in, cursor echoed back unchanged, null = end), the limit range 1-200, and the required read:models scope. It stops short of describing error behavior or ordering guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, which is good, but it then enumerates the full field list of the response object even though an output schema exists — that enumeration is largely redundant. The remaining sentences about filtering, paging and scope do earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a filtered, paginated list tool, the description covers scope, filtering, paging contract and auth requirement, and an output schema already documents the return shape (making the field enumeration unnecessary rather than a gap). It is essentially complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage the description must carry the load, and it does for most parameters: type/status/channel as filters, limit with its 1-200 bound, and cursor semantics (echo next_cursor back, null means end). Only project_id is left to inference, and the description's 'next_cursor' naming does not exactly match the 'cursor' input param.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (a project's incrementality tests) with a scope qualifier ('recorded, not retired') that separates it from get_incrementality_test, create_incrementality_test, import_incrementality_tests and recommend_incrementality_tests without needing the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives operational context (filter by type/status/channel, paging is opt-in, pass limit and echo next_cursor), but never states when to choose this over get_incrementality_test or the other incrementality siblings, nor any exclusions. Usage is implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsList ModelsARead-onlyIdempotent
List all Marketing Mix Models for the authenticated user.
Returns model name, hash, status (pending/under way/complete/failed), type (mmm/var), hierarchy value, and timestamps.
NOTE: All other model endpoints use model_hash (string, e.g. "f835671a25") as the identifier. Use the model_hash from this response.
Args:
include_unsaved: Include draft/unsaved models (default false).
limit: Maximum number of models to return (default 50, max 500).
offset: Number of models to skip, for paging past limit (default 0).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| include_unsaved | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral detail beyond the annotations: it lists the returned fields, statuses, type values, hierarchy value, and timestamps, and explains the hash convention across related endpoints. This gives the agent a meaningful picture of what to expect without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well structured and front-loaded: purpose first, then return values, then the critical identifier note, then parameter documentation. Every sentence adds value, and the Args block is concise and directly useful. No redundant fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list endpoint with an output schema and comprehensive annotations, the description is complete. It explains the authenticated-user scope, the key fields returned, the model_hash convention that matters for all downstream calls, and all pagination and filtering parameters. Nothing an agent needs to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter semantics. It fully compensates by explaining include_unsaved as 'draft/unsaved models', limit as 'maximum number to return' with a max of 500, and offset as 'number of models to skip, for paging past limit.' This goes well beyond the bare schema properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List all Marketing Mix Models for the authenticated user.' This clearly identifies the operation and scope, and the resource name differentiates it from sibling list_* tools such as list_studies, list_uploads, and list_projects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by explaining that all other model endpoints use model_hash and instructing the agent to use the model_hash from this response. It does not explicitly name alternative tools or when-not-to-use conditions, but the hash-identifier guidance strongly implies this tool is the entry point for model operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pipelinesList PipelinesARead-onlyIdempotent
List the data pipelines this key's owner has, most recently updated first, each with id, pipeline_hash, name, description, version_count and latest_version (id, version, created_at, row_count, checksum). Identity only: no definitions, parameters or output data. Use a version id as pipeline_version_id in get_recipe_draft_template. Requires the ingest scope; the same ownership rule as the app. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint false, covering the safety profile. The description adds substantial behavioral detail: ordering (most recently updated first), content exclusions (identity only), auth requirement (ingest scope), and a full explanation of opt-in pagination (limit range, cursor semantics, null termination). This goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded with the core purpose. Every sentence adds necessary information—scope, fields, ordering, auth, paging—with no fluff. It is appropriately sized for the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the description explicitly lists the returned fields (id, pipeline_hash, name, etc.), the description covers everything an agent needs: purpose, prerequisites, pagination, and output structure. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It clearly explains both parameters: limit (1-200) triggers paging, and cursor must be sent back unchanged for the next page. This provides complete meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and the resource 'data pipelines' with specific fields returned and ordering. It does not explicitly differentiate from sibling tools like list_pipeline_versions, though the resource name and 'Identity only' scope imply a distinction. This meets the 'clear but no sibling differentiation' criterion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage context: requires the ingest scope, same ownership rule as the app, and mentions how to use the result (version id as pipeline_version_id). It does not explicitly state when not to use this tool or name alternatives, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pipeline_versionsList Pipeline VersionsARead-onlyIdempotent
List the saved versions of one owned pipeline (by pipeline_hash or id), newest first: id, version, created_at, row_count, column_count, column names and checksum. The checksum is the exact source identity a recipe draft freezes. Never returns the data itself; fetch a draft template with pipeline_version_id for that. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| pipeline_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly, non‑destructive, idempotent), the description discloses several important behaviors: newest-first ordering, the exact list of returned attributes, the semantics of checksum as recipe-draft source identity, the fact that data is never returned, and detailed paging rules (limit range, cursor reuse, null cursor meaning end). This is rich, non-obvious context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then packs field list, critical exclusions, and paging rules into one short paragraph. Every sentence serves a purpose; there is no redundant or filler language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations cover safety, the description still supplies all necessary operational details: owned-pipeline prerequisite, ordering, paging contract, and the crucial warning that data is not included. An agent can correctly call this tool without further lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates: it explains pipeline_ref can be a hash or id, defines the limit range (1-200), and specifies cursor semantics (send back unchanged, null means end). All three parameters receive meaningful guidance beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List the saved versions of one owned pipeline' and enumerates the exact fields returned. It differentiates from siblings by focusing on pipeline versions rather than pipelines themselves, and it notes the distinction from data-fetching tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it lists metadata, never data, and tells the user to 'fetch a draft template with pipeline_version_id' for the actual data. It also explains paging behavior. However, it does not explicitly name the sibling tool (e.g., get_recipe_draft_template) or mention how to obtain a pipeline_ref, leaving slight ambiguity about alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsList ProjectsARead-onlyIdempotent
List the projects (the app's model folders) you can file models into.
Returns owned and team-shared projects: per project {id, name, is_default, shared_with_team_id, model_count} — team-shared folders carry "shared": true, and model_count counts SAVED models (the set the app's model list shows). Use the ids with save_model(project_id=...) and rename_project. There is deliberately no delete over the API — use the app to delete a project.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly and idempotent, so the safety profile is covered. The description adds semantic detail: team-shared folders carry 'shared': true, and model_count counts only SAVED models (the set the app's model list shows). It also discloses an intentional API limitation (no deletion), going beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact block that front-loads the core purpose in the first line, then efficiently provides output shape, field meanings, usage pointers, and a limitation. Each sentence adds distinct information and none are redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless listing tool with an output schema, the description explains the return fields and their interpretation (e.g., model_count counting SAVED models), and connects to related tools. The only remaining context lies in the output schema itself, which is provided separately, so nothing needed by an agent is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema describes all parameters (none) and there is nothing for the description to add about inputs. Per the 0-param baseline, a score of 4 is appropriate; the description instead clarifies the return payload semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List the projects (the app's model folders) you can file models into.' It distinguishes projects from other entities by framing them as model folders, and the return format clarifies this is a read-only list tool, separate from list_models and similar siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states that the returned ids should be used with save_model and rename_project, giving an immediate use case. It also notes there is no delete over the API, telling agents to use the app instead. It doesn't explicitly contrast with sibling list tools, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_quality_policiesList Quality PoliciesARead-onlyIdempotent
Read immutable quality policies for the study. Each row carries the full specification plus identity: content_hash (the stored row, including name and rationale), rules_hash (the rules alone: checks sorted by metric plus the protocol, so two policies with the same rules_hash apply the same rules whatever they are called), is_newest, derived_from, created_at, retired_at, usage counts (runs_launched, evaluations, resolutions, champion_acceptances) and checks_summary / protocol_summary. Newest is information, not a recommendation: choose policy_id explicitly. Retired policies stay listed for history but are refused for new launches, assessments and pair reviews. Deleting a policy is possible only for a policy nothing references and only from the signed-in project owner UI, never through this API key.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only and non-destructive, but the description adds substantial behavioral context: it explains the immutability, the inclusion of retired policies and their refusal status, the meaning of content_hash versus rules_hash, and that deletion is impossible via the API. This goes well beyond the annotation hints and gives the agent a precise model of the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: the first sentence states the core purpose, and the subsequent sentences explain row semantics and behavioral caveats. Every sentence adds meaningful information, though the level of detail around content_hash and rules_hash is arguably more than needed for a list endpoint. It remains well-organized and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists, the description need not enumerate return fields; instead it adds semantic context that the schema cannot convey—what the hashes mean, how to treat newest, and the retired-policy behavior. It does not cover pagination ordering or iteration, but these are likely inferable from the cursor/limit parameters. Overall it is sufficiently complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter explanation. It never mentions limit or cursor, and only indirectly identifies study_id as the study being listed. The pagination parameters are left entirely to inference, and the description adds no format, constraints, or interaction details beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Read immutable quality policies for the study', a specific verb-resource pair that clearly indicates a list operation. It further differentiates from sibling tools like get_quality_policy and diff_quality_policies by describing the full row contents and the fact that it lists all policies, including retired ones. The scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context about how to interpret the list: 'Newest is information, not a recommendation' and notes that retired policies are refused for new launches, assessments, and pair reviews. However, it does not explicitly state when to prefer this tool over get_quality_policy or diff_quality_policies, leaving some inference to the agent. The deletion note clarifies a limitation but not alternative tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recipe_draftsList Recipe DraftsARead-onlyIdempotent
List study draft metadata without loading datasets, newest update first. Check backend draft capability first. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds meaningful behavioral context: paging semantics (opt-in limit, cursor flow, null means end), the 'no totals promised' disclosure, and the absence of rows you cannot see – all beyond what annotations provide. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries value: the primary action, the capability prerequisite, paging mechanics, and the absence-of-totals warning. The main action is front-loaded, and there is no filler or repetition – appropriate length for the amount of behavior it must clarify.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return format is covered. The description addresses capability requirements, paging behavior, ordering, and the open-world constraint. Nothing needed for correct invocation is missing; the tool's complexity is limited to paging and capability checks, both fully described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It thoroughly explains limit and cursor with precise paging behavior, and study_id is inherently clear from the tool name and context. It does not explicitly document study_id meaning, but it is unambiguous given the description's opening line.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('study draft metadata') with a key differentiator ('without loading datasets') and ordering ('newest update first'). It clearly distinguishes from get_recipe_draft (single draft) and list_study_recipes (likely published recipes) without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a prerequisite ('Check backend draft capability first') and detailed paging instructions, but does not explicitly name alternative tools or state when not to use this tool. The capability check implies using get_backend_capabilities, but the guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_runsList RunsARead-onlyIdempotent
List a model's saved optimizer or scenario run history.
Returns {model_hash, runs, count, limit, offset}. Each run summary has: run_id, name, auto_named, pinned, notes, tags, status, error_details, progress fields while running, key_metrics (optimizer: total_budget, num_periods, gamma, predicted_revenue/roi, ...; scenario: num_periods, total_planned_spend, predicted_outcome, ...; null metrics are omitted — treat every key as optional), and created/started/completed timestamps. Ordering is pinned-first, then newest-first.
CAVEATS:
countis the LENGTH OF THIS PAGE, not the total run count — page until a short page.The optimizer objective ("revenue"/"profit") is NOT in the summary; fetch the specific run (get_optimizer_results with run_id) and read its
inputs— profit runs carryobjective: "profit"there, revenue runs omit the key.
Use get_optimizer_results / get_scenario_results with a run_id to fetch a listed run's full inputs and results; update_run / set_run_pinned to curate it.
Args: artifact: "optimizer" (run ids "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model whose run history to list. limit: Page size (API clamps to 1-200; default 50). offset: Rows to skip (paging).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| artifact | Yes | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint, idempotentHint, destructiveHint), the description discloses important behavior: count is only the page length, ordering is pinned-first then newest-first, null metrics are omitted, and the optimizer objective is intentionally absent from summaries. These details prevent incorrect assumptions during use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed but tightly organized: a clear purpose, a structured return-value breakdown, explicit caveats, and a compact args section. Every sentence provides actionable information, and the caveats are prominently separated rather than buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the operation's scope, return shape, field semantics, paging pitfalls, missing-data caveats, and relationships to sibling tools. Even with an output schema present, it supplies enough context for an agent to call the tool correctly without further exploration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates: artifact is explained with expected prefixes ('optimizer' for 'opt_...', 'scenario' for 'scn_...'), model_hash is defined as the model whose history to list, limit includes clamping behavior and default, and offset explains paging. This adds meaningful semantic value absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List a model's saved optimizer or scenario run history.' It clarifies this is a listing operation that returns run summaries, not full results, which distinguishes it from sibling tools like get_optimizer_results and get_scenario_results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly routes the agent to alternatives: 'Use get_optimizer_results / get_scenario_results with a run_id to fetch a listed run's full inputs and results; update_run / set_run_pinned to curate it.' It also provides paging guidance with the caveat to page until a short page, making the intended usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_studiesList StudiesARead-onlyIdempotent
List project-owned studies, questions, budgets and access rights. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| project_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint, covering safety. The description adds valuable behavioral context beyond those: the opt-in paging mechanism (limit 1-200, cursor flow, null ends), the fact that missing rows are simply absent, and that no totals are promised. This aligns with and enriches the openWorldHint and is highly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, front-loaded with the core purpose and then the paging details. Every sentence adds unique value—scoping, paging mechanics, and absence semantics. No wasted words, perfectly sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a paged list operation with an output schema present. The description covers the paging protocol, the open-world behavior, and the absence of totals, which is exactly what an agent needs to call it correctly. Return values are presumably in the output schema, so the description is complete without repeating them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for parameter meaning. It explicitly explains limit (1-200, paging) and cursor (return as-is for next page, null ends). It does not describe project_id, but its name and requirement make its purpose self-evident from context. The description compensates well for two of three params, with the third being obvious.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('List') and a specific resource scope ('project-owned studies, questions, budgets and access rights'). It distinguishes this from other list tools by enumerating the entity types, but it does not explicitly name a sibling alternative (e.g., get_study or list_study_runs), so it doesn't fully meet the 5-bar for explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains paging behavior and that rows may be absent, but it does not explicitly state when to use this tool versus alternatives like get_study or list_study_runs. The usage is implied (list everything owned by a project) but no clear exclusions or conditions are given, so it meets the baseline for implied usage without explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_study_decisionsList Study DecisionsARead-onlyIdempotent
Read analyst decisions and agent recommendations. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish read-only, idempotent, and open-world behavior, and the description goes well beyond them by explaining opt-in paging, the 1-200 limit, cursor mechanics, null next_cursor termination, and the absence of promises about totals. It also clarifies that invisible rows are simply omitted, matching the openWorldHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences front-load the core purpose and then efficiently pack all necessary paging semantics. Every sentence earns its place; there is no fluff or repetition of schema defaults.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations already cover safety and open-world behavior, the description supplies the remaining operational details an agent needs: how to paginate, what null cursor means, and how to interpret missing rows. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does this thoroughly for limit and cursor, explaining the allowed range, default behavior, and exact cursor usage. study_id is not described, but its meaning is self-evident from the tool name and title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a clear resource ('analyst decisions and agent recommendations'), which distinguishes it from sibling tools like list_study_runs and list_study_evaluations. It leaves no doubt about what domain of data this tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for reading decisions and recommendations for a study, but it does not explicitly state when to choose it over siblings or when not to use it. No alternatives are named and no exclusions are given, so the usage context is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_study_evaluationsList Study EvaluationsA
List preserved assessments as summaries: id, policy_id, policy_name, status, basis_hash, evidence_hash, created_at and summary (evaluated / total / required_unevaluated / evidence_sources: which values were supplied and which carried). A newer sparse report does not replace an earlier enriched one; decisions bind to a specific report id. Pass expand=["report"] for the full per-check report. Listing never records prediction access, with or without expand (jellyfish #837): a stored report carries window metadata and aggregate errors, not prediction rows. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| expand | No | ||
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing that newer sparse reports do not replace enriched ones, that decisions bind to a specific report id, and that listing never records prediction access. It also clarifies visibility semantics: 'Rows you cannot see are simply absent; no totals are promised.' This is rich, non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and somewhat long, but every sentence carries meaningful information about return shape, side effects, paging, or visibility. It is front-loaded with the summary fields and then covers caveats. The internal reference 'jellyfish #837' is slightly noisy but not harmful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and the description already covers return fields, paging, expand behavior, side effects, and visibility semantics, an agent has everything needed to invoke this tool correctly. The only minor gap is run_id semantics, but that is inferable and not a real obstacle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains limit, cursor, and expand in detail, including ranges, defaults, and cursor behavior. The required run_id parameter is not explicitly described, but its meaning is strongly implied by the tool name and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List preserved assessments as summaries' and enumerates the exact fields returned. This clearly distinguishes it from sibling tools like list_study_decisions or get_study_run by focusing on preserved evaluation summaries rather than decisions or full runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear operational guidance: when to pass expand, how to page with limit and cursor, and that listing never records prediction access. It does not explicitly name alternatives or state when not to use this tool, but the context is strong enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_study_recipesList Study RecipesARead-onlyIdempotent
List recipes with their revisions as summaries (id, number, content_hash, reason, created_at). Heavy fields travel only by name in expand: effective (exact frozen priors, settings, data hashes and runtime), inspection (which settings were authored vs defaulted, inert prior fields with their gate, engine state current/stale that predicts whether launch will be refused, and lineage: the recorded dataset with available checked for you and the display line people see; treat inert values as stored-but-unused, not as recipe errors), specification (the authored request). Prefer get_recipe_revision for one revision in full. Raw datasets are never included. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| expand | No | ||
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover only the safety profile (readOnly/idempotent/non-destructive), and the description adds substantially more: raw datasets are never included, invisible rows are simply absent with no promised totals, cursor must be sent back unchanged, null next_cursor signals the end, and inert prior fields mean stored-but-unused rather than errors. It even flags that engine state predicts whether launch will be refused.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and every clause carries information, but the inspection parenthetical is very long and dense, making the middle hard to scan. It earns its length with content, yet some sub-clauses (e.g. 'the display line people see') could be trimmed without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a filtered, expandable list tool with an output schema already present, the description covers the remaining unknowns an agent needs: which fields are summaries, what expand adds, dataset exclusion, paging protocol, and absence semantics. Nothing required to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so: limit is bounded (1-200) and returns a page with next_cursor, cursor is a verbatim round-trip token, and each expand enum value is defined in domain terms (effective, inspection, specification). This is meaning well beyond the bare enum list in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recipes with their revisions as summaries') and immediately fixes the shape of what is returned (id, number, content_hash, reason, created_at). It also explicitly routes the agent away from itself toward get_recipe_revision for single-revision reads, so it is distinguishable from siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative ('Prefer get_recipe_revision for one revision in full') and the condition that selects it, and gives the condition for each expand value. It also discloses the optional nature of paging ('Paging is opt-in') so the agent knows that omitting limit is a valid, non-error path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_study_runsList Study RunsARead-onlyIdempotent
List preserved attempts including pending and failed runs. Supporting backends also return budget with attempts remaining, available slots and blocking reasons; the budget always counts every run even when rows are paged. Missing budget means unknown support, not permission to launch. Capacity is rechecked on reservation; recover an uncertain launch with its original submission key. Paging is opt-in: pass limit (1-200) to receive a page and next_cursor; send that cursor back unchanged for the next page; null next_cursor means the end. Without limit every row is returned. Rows you cannot see are simply absent; no totals are promised.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| study_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds significant behavioral context beyond annotations: budget is always counted even when paged, missing budget means unknown support, rows you cannot see are simply absent, no totals promised. It doesn't fully describe return object shape but output schema exists and annotations carry the read-only/idempotent load, so a 4 is warranted rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and efficient, front-loading the core purpose before diving into budget and paging details. Every sentence adds information, and the length is justified by the three parameter behaviors and the budget caveat. It could be trimmed slightly—"Supporting backends" and "recover an uncertain launch" are slightly awkward—but overall it earns a high score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering readOnly/openWorld/idempotent, the description covers the remaining high-stakes context: paging semantics, budget interpretation, capacity recheck, and pagination end condition. It doesn't explicitly state error cases or cursor expiry, but for a list tool with an output schema, this is near-complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for the parameters, and it does: limit is explained as opt-in with range 1-200, cursor is explained as opaque token to pass back unchanged, and study_id is implied as the scope of the listing. The description lacks a precise statement that limit without cursor starts from the beginning, but the paging semantics are otherwise clearly tied to the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "List preserved attempts including pending and failed runs." This clearly distinguishes it from siblings like list_runs, get_study_run, and launch_study_run by framing study runs as preserved attempts with statuses included. It goes beyond a vague restatement of the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit paging guidance (opt-in limit 1-200, next_cursor semantics, null means end, no limit returns all rows) and explains when the budget field appears and how to interpret its absence. It warns about capacity recheck and how to recover a launch, which is exactly the kind of when-to-use vs when-not-to-use information an agent needs. No exclusions needed because the sibling set is broad but the scope is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_uploadsList UploadsARead-onlyIdempotent
List the datasets in your workspace (newest first) — every source, not just API uploads: dashboard/manual uploads and pipeline-ingested datasets appear too (see source_type per file).
Returns {files, count, limit, offset} where each file has: id (the
uploaded_file_id create_model needs), filename, original_filename,
source_type, row_count, column_count, created_at. Here count IS the
true total matching the filter (unlike list_runs, where it is the page
length). Column names/dtypes are not in the listing — fetch one upload
with get_upload for those.
Args: limit: Page size (API clamps to 1-500; default 50). offset: Rows to skip (paging). name: Optional case-insensitive substring filter on the original filename.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| limit | No | ||
| offset | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false. The description goes well beyond these by revealing behavioral specifics: the ordering, the inclusion of non-API sources, the true count semantics, pagination behavior (limit/offset), and the omission of column details. No contradiction exists; the description adds significant value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: it opens with the core purpose, then covers the return shape and key behavioral differences, and finally lists parameters. Every sentence adds value—there is no fluff. The contrast with list_runs and the pointer to get_upload are concise and purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema is present (has_output_schema: true) and the description already details the return structure, pagination, filtering, and caveats, the definition is complete for an agent to call this tool correctly. It even points to get_upload for missing details, covering all likely follow-up needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining parameters. It does so explicitly: 'limit: Page size (API clamps to 1-500; default 50)', 'offset: Rows to skip (paging)', and 'name: Optional case-insensitive substring filter on the original filename.' This fully compensates for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a specific resource ('datasets in your workspace'), and an ordering ('newest first'). It also clarifies the scope ('every source, not just API uploads') and explicitly contrasts with list_runs regarding count semantics, which distinguishes it from a sibling tool without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (to list uploads) and provides a clear alternative for a specific need: 'fetch one upload with get_upload for those' when column names/dtypes are required. It also contrasts with list_runs on the meaning of 'count', which is helpful for routing. However, it does not explicitly state 'use this when you need to list uploads' or list exclusions for other scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_recipe_draftPublish Recipe DraftAIdempotent
Publish a saved MMM or VAR draft as immutable recipe revisions. Supply a new UUID publication_id and reuse it with identical arguments after an uncertain response; a replay returns the same revisions. Without target: new recipes, one per prepared brand, atomic. With target={recipe_id, expected_version}: revision N+1 of that recipe under its row lock, the draft's base revision recorded as the source. 412 stale_version when the recipe moved since expected_version: the draft and every edit are kept, so re-read the recipe and publish again on top, or drop target to branch into a new recipe. 409 target_requires_single_recipe for a multi-brand draft; 409 when the draft's own target names a different recipe; 404 for a recipe outside the study. Backend compiles the saved snapshot with shared wizard rules; never send separately prepared settings. Does not fit, consume an attempt or designate a champion. MMM priors the draft leaves unset are filled with smart defaults at publication (the Prior Builder's own calculation over the whole source) and frozen; replay does not rebuild them. VAR requires family-specific evidence assessment; MMM policies cannot establish VAR acceptance. Enabled invalid MMM calibration fails publication; VAR calibration is unsupported. Preserve disabled authoring observations.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| target | No | ||
| draft_id | Yes | ||
| publication_id | Yes | ||
| expected_version | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (idempotentHint=true, openWorldHint=true, destructiveHint=false), and the description goes well beyond them: it explains the publication_id replay contract, atomic per-brand creation, row-lock semantics, the 412/409/404 error taxonomy, and what survives a failed publish (draft and edits kept). It also discloses non-obvious side effects like unset MMM priors being filled and frozen, which the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and idempotency contract are front-loaded, and nearly every sentence carries operational weight (error codes, defaults, MMM/VAR distinctions). The prose is very dense and the mid-section error/rules sentences run long, which slightly hurts scannability, but little is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description still supplies the failure modes, atomicity guarantee, idempotency/replay behavior, and MMM vs VAR constraints an agent needs to call this safely. Given the tool's complexity, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0% for the top-level params, the description must compensate, and it does for the important ones: publication_id (new UUID, reuse on uncertain responses), target (recipe_id + expected_version, 412 semantics), and the interaction with expected_version. It leaves reason and draft_id uninterpreted, so the coverage is strong but not total.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb (Publish), resource (saved MMM or VAR draft), and result (immutable recipe revisions), which cleanly separates it from siblings like create_recipe_draft, update_recipe_draft, and revise_study_recipe. The two-mode framing (new recipe vs. next revision) further sharpens what the tool concretely does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly prescribes the two operating modes: omit target for new recipes (one per prepared brand), or pass target={recipe_id, expected_version} to publish revision N+1, and explicitly tells the agent to 'drop target to branch into a new recipe' on a 412. It does not name sibling alternatives for the create-vs-revise decision, so it stops short of the explicit when-not/alternative routing that earns a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_campaign_budgetsRecommend Campaign BudgetsARead-onlyIdempotent
Calculate campaign budget suggestions without applying them. Requires read:models.
Uses the same channel-derived marginal curves and bounded allocator as the web app. This POST is a read-only calculation: it creates no saved run, starts no fit or posterior job and does not change mappings or advertising-platform budgets. Results are conditional scenarios, not independently fitted campaign response curves or a day-by-day forecast.
Args: model_hash: Owned, completed MMM with verified curve and currency provenance. observation_window: {start: YYYY-MM-DD, end: YYYY-MM-DD}, inclusive fact dates. currency: Explicit currency matching the model and facts, for example GBP. channel_daily_budgets: Exact channel keys mapped to daily totals in that currency. Supply exactly one of this or optimizer_run_id. Preserve currency minor precision. optimizer_run_id: Existing owned optimiser run for the same model, with verified planning dates and currency. Its totals become flat daily equivalents, not a daily schedule. The tool does not create or rerun an optimiser. level: campaign or adset. max_step_fraction: Maximum change from observed average daily spend; default 0.2. bounds: Optional rows keyed by full identity {platform, account_id, campaign_id, adset_id}, with min and/or max amounts in daily currency units. Use null adset_id at campaign grain. No platform floor is invented; the server validates bounds. expected_context_key: Optional context_key returned by get_campaign_marginal_returns. A changed model, facts, mapping or attribution basis refuses with 409 instead of calculating against evidence different from the scenario you reviewed.
Returns {channels: [{channel, status, total_daily_budget, rows, explanation, assumptions, reason, feasible_range}], ...provenance}. Ready rows include current and recommended budgets, the continuous solution, marginal returns and binding constraints. Refusals retain their reasons and feasible ranges; never relax bounds or silently change a total. Rounded amounts reconcile in integer currency minor units; rounding does not imply exact equality of marginal returns. Missing marginal uncertainty remains explicitly unavailable.
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | campaign | |
| bounds | No | ||
| currency | Yes | ||
| model_hash | Yes | ||
| optimizer_run_id | No | ||
| max_step_fraction | No | ||
| observation_window | Yes | ||
| expected_context_key | No | ||
| channel_daily_budgets | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: explains that this POST is a read-only calculation, creates no saved run, starts no fit/posterior job, and changes no mappings or platform budgets. It also discloses refusal semantics (409 on changed provenance), rounding reconciliation behavior, and that bounds are never silently relaxed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and permission, then a dense Args block that earns its length given 9 undocumented parameters. The Returns section is somewhat redundant since an output schema exists, though it adds refusal/rounding semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, nested-object, POST-based calculation tool with 0% schema coverage, the description covers auth, mutual exclusivity, defaults, provenance-refusal behavior, and output shape. Nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so thoroughly: format for observation_window, currency examples, the mutually exclusive channel_daily_budgets/optimizer_run_id, the default for max_step_fraction, and the exact key shape for bounds. Each parameter's meaning and constraints are clarified beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Calculate campaign budget suggestions') and immediately scopes it ('without applying them'), which distinguishes it from run_optimizer and get_campaign_marginal_returns. The agent can tell what the tool produces without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong usage conditions: 'Requires read:models', 'Supply exactly one of this or optimizer_run_id', and points to get_campaign_marginal_returns as the source for expected_context_key. It does not explicitly state when to prefer this over run_optimizer or show_optimizer_allocation, leaving that contrast to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_incrementality_testsRecommend Incrementality TestsARead-onlyIdempotent
Rank channels for experiment investigation using stored posterior marginal returns. Read-only: no fit, test creation or budget changes. Returns {method, basis, score_unit, currency, budget, hurdle, spend_basis, period, approximation_warnings, items, excluded}. Each item carries channel, score, components (mean, sigma, stake, spend_share, crossing_probability, optional contraction), reason_codes, last_test_end, hypothesis and an unavailable design_hint. The score is a normal-approximation local binary perfect-information value, not expected test benefit, experiment budget, portfolio value or forecast lift. budget is a positive exposure scale (default sum of current spend); weights are the mean-active-period spend mix, which need not represent one common calendar period. hurdle is the non-negative marginal-return alternative (default 1). limit is 1-50. Missing posterior means or intervals are explicitly excluded. limited_variance_contraction means contraction from zero to below 0.1; posterior_variance_expanded means negative contraction. Neither proves prior domination. Cross-channel dependence is not modelled. Requires read:results. Test history is unavailable because stored registry records do not establish compatible model/geographical coverage. Requires backend support for test-priorities.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| budget | No | ||
| hurdle | No | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description still adds substantial behavioral context: the auth scope required, the backend capability gate, explicit exclusion of missing posteriors, the meaning of limited_variance_contraction vs posterior_variance_expanded, and the disclaimer that cross-channel dependence is not modelled and test history is unavailable. This is well beyond what the annotations provide and contradicts nothing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded correctly, but the middle is a dense run-on enumerating return keys and warning codes that the output schema already carries, producing a wall of text where a shorter semantic gloss would suffice. Every clause is informative but the listing is not optimally sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers auth requirements, backend gating, exclusion behavior, score interpretation and known modelling limitations, so an agent can judge appropriateness. It leans on prose to restate return fields already in the output schema and omits any meaning for model_hash, which slightly undercuts completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage the description must compensate, and it does for three of four params: limit is bounded 1-50, budget is a positive exposure scale defaulting to current spend, and hurdle is the non-negative marginal-return alternative defaulting to 1. model_hash, the only required parameter, is never explained, leaving one gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+basis: 'Rank channels for experiment investigation using stored posterior marginal returns.' It also contrasts itself against siblings by stating it does 'no fit, test creation or budget changes', which separates it from create_incrementality_test and run_pipeline without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the read-only framing and that it requires read:results and backend support for test-priorities, which tells the agent when the call can succeed. It implies but never explicitly names the alternative (create_incrementality_test) or the condition that selects it, so routing is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommend_study_runRecommend Study RunA
Record a recommendation with evidence. This does not accept or promote a model; analyst acceptance happens in the frontend.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| run_id | Yes | ||
| study_id | Yes | ||
| evaluation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by clarifying that the write operation only records a recommendation and does not itself accept or promote, which is a non-obvious behavioral trait. Annotations already flag readOnly=false, idempotent=false, and destructive=false, and the description does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences deliver the core action and the key exclusion with no filler. The most decision-relevant constraint—'does not accept or promote'—is clearly stated after the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering side-effect flags, the tool is not wildly under-specified: an agent can tell it records a recommendation for a study run with evidence. However, because the parameters are undocumented, the workflow context—such as which evaluation/run IDs are valid and how this relates to adoption—is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description only loosely hints that 'reason' should carry evidence. The three identifier parameters (study_id, run_id, evaluation_id) and their relationships are left entirely to name inference, so the description does not sufficiently compensate for the lack of schema property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a concrete action—record a recommendation—with the object 'with evidence,' and explicitly marks what the tool does not do (accept or promote a model), which separates it from adoption/promotion siblings. It relies on the title for the 'study run' context, but the verb-resource pairing is specific enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The statement 'analyst acceptance happens in the frontend' gives an explicit when-not: this tool is for recording, not for accepting or promoting. It does not name a sibling alternative such as adopt_model_into_study, so an agent must infer which tool to use instead, but the boundary is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refreeze_recipe_revisionRefreeze Recipe RevisionA
Recover from engine_changed without bypassing the safeguard: creates a new revision of the same recipe with the same specification, frozen on the current engine, with an auto-filled reason naming what changed. The old revision and its provenance are untouched. Returns the new revision; launch that one. 409 conflict when the revision is already frozen on the current engine; 409 snapshot_not_executable for review-only snapshots.
| Name | Required | Description | Default |
|---|---|---|---|
| number | Yes | ||
| recipe_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly=false, destructive=false, and idempotent=false. The description adds non-obvious behavior: a new revision is created, the old revision and its provenance are untouched, an auto-filled reason is generated, the return value should be launched, and two specific 409 error cases are disclosed. This goes well beyond the annotation signals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler. Each sentence earns its place: purpose and effect, non-destructive behavior with return guidance, and error conditions. The most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, exact effect, non-destructive behavior, return value, and failure modes, and the tool has an output schema plus only two simple parameters. It is nearly complete, but an explicit mapping of 'recipe_id' and 'number' to the operation would remove the last ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never explicitly names 'recipe_id' or 'number'. It implies the recipe through 'same recipe' and the source revision through 'the revision' in the 409 clause, but parameter semantics are largely left to inference from property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete recovery scenario ('Recover from engine_changed'), then states the specific operation: creates a new revision of the same recipe with the same specification, frozen on the current engine. This clearly differentiates it from sibling revision/draft tools by emphasizing refreezing rather than editing or creating drafts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The first phrase 'Recover from engine_changed' gives an explicit condition for use, and 'without bypassing the safeguard' signals the intent. It does not name an alternative sibling or state an explicit when-not-to-use rule, though the 409 cases implicitly delimit invalid calls.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_modelRename ModelADestructiveIdempotent
Rename a model.
Changes only the display name; the model's saved/unsaved state is untouched (use save_model to file it into a project). The name is HTML-sanitized server-side and must be non-empty.
Args: model_hash: Hash of the model to rename. name: New display name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It adds meaningful non-obvious behavior beyond the annotations: only the display name changes, saved/unsaved state is untouched, and the name is HTML-sanitized and non-empty. This does not contradict the destructiveHint because overwriting the display name is still a mutating update.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, side-effect boundary, input handling rule, then parameter semantics. It is front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with an output schema and annotations, the description is nearly complete: it covers purpose, parameter meaning, side-effect scope, and validation. A minor gap is the lack of explicit mention of what happens when the model hash does not exist or how the new name appears in references.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the Args block compensates by clarifying that model_hash identifies the model to rename and name is the new display name. It also carries the non-empty constraint from the description, though it does not describe hash format or sources.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first line 'Rename a model' states the exact verb and resource, and the next line 'Changes only the display name' sharply scopes the operation against siblings like save_model, delete_model, and rename_project. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly warns not to expect saved/unsaved state changes and directs the agent to save_model for filing a model into a project. It does not enumerate every alternative, but for a narrow rename operation the when-to-use context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_projectRename ProjectADestructiveIdempotent
Rename a project you OWN.
Team members can file models into a shared folder but not rename it (owner-only; 404 for a project you don't own). Renaming your default folder is safe: it keeps receiving unqualified saves under its new name.
Args: project_id: Id of the project to rename (see list_projects). name: New display name.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| project_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as destructive and not read-only, but the description adds valuable behavioral context beyond them: owner-only enforcement with 404 for non-owned projects, and the reassuring detail that renaming the default folder preserves its role for unqualified saves. This meaningfully informs an agent's decision and expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core statement. Every sentence adds useful information: ownership requirement, error behavior, default-folder side effect, and parameter explanations. There is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter, no-nested-object tool with an output schema and annotations, the description covers the essential context: who can call it, error semantics, side effects, and how to find the parameter value. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. project_id is explained as the ID of the project to rename with a pointer to list_projects, and name is described as the new display name. This is sufficient for a simple two-parameter tool, though it doesn't add constraints like uniqueness or length.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific action and resource: 'Rename a project you OWN.' It clearly distinguishes the operation from model-level operations by emphasizing ownership and the project scope, and it adds the critical condition that only owners can rename.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use context: only for projects you own, with a 404 if you don't. It references list_projects for obtaining the project_id. It does not explicitly name sibling alternatives like rename_model, but the ownership and shared-folder distinction makes the appropriate use case clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retire_quality_policyRetire Quality PolicyADestructiveIdempotent
Retire (retired=true) or restore (retired=false) a saved policy without changing it. A retired policy is refused for new launches, new assessments and new pair-review resolutions, stops counting as the newest policy, and stays readable in every run, assessment, pair review, comparison and Champion record that already names it. Idempotent; the backend records each change. Requires create:models on the study's project. Nothing is deleted: removal of an unreferenced policy is an owner-only frontend action.
| Name | Required | Description | Default |
|---|---|---|---|
| retired | No | ||
| study_id | Yes | ||
| policy_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry idempotentHint and destructiveHint, and the description adds meaningful behavioral detail: the policy 'stops counting as the newest policy, and stays readable in every run, assessment, pair review, comparison and Champion record that already names it.' It also mitigates the destructiveHint by stating nothing is deleted, and it discloses the required permission, create:models. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler: the core action is first, then the lifecycle effects, then permissions and deletion boundaries. Every sentence adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating policy operation, the description covers the action, the affected workflows, idempotency behavior, permission requirements, and the explicit non-deletion guarantee. An output schema is present, so return-value details are not missing. Nothing essential is left out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It fully explains the retired parameter with 'Retire (retired=true) or restore (retired=false),' but study_id and policy_id are only anchored by the phrases 'saved policy' and 'the study's project.' It does not explain how to obtain these IDs or clarify their relationship beyond natural inference from their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource pair: 'Retire (retired=true) or restore (retired=false) a saved policy without changing it.' It then clarifies what retirement means in practice, so an agent can distinguish this from creating, updating, or deleting a policy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it affects new launches, assessments, and pair-review resolutions, and it explicitly excludes deletion with 'Nothing is deleted: removal of an unreferenced policy is an owner-only frontend action.' However, it does not name sibling tools such as create_quality_policy as alternatives, so it stops short of explicit when-to-use versus alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revise_study_recipeRevise Study RecipeA
Create an immutable revision. Earlier versions inherit influence automatically; optional source_revision_id additionally links a same-study source recipe. Optional expected_content_hash guards effective inputs (409 requires a fresh preview). Supply the current recipe version; stale edits are rejected (412). Reload list_study_recipes, reconcile changes, then submit the current version; never overwrite blindly.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| reason | Yes | ||
| recipe_id | Yes | ||
| specification | Yes | Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, unknown config keys are rejected, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact. | |
| expected_version | Yes | ||
| source_revision_id | No | ||
| expected_content_hash | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent, non-destructive write (openWorldHint=true), and the description adds substantial behavior beyond that: immutability, automatic influence inheritance from earlier versions, the source_revision_id same-study linkage, the expected_content_hash guard with 409 semantics, and 412 on stale versions. These are exactly the concurrency and lineage traits an agent needs and that structured fields do not carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose first, then optional params, then conflict codes, then the recommended workflow. Every sentence carries signal; the only cost is that the semicolon-heavy style makes it slightly harder to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the nested specification envelope is documented in the schema itself. The description fully covers the concurrency contract and workflow, leaving only the plain-identity parameters (recipe_id, name, reason) unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 14%, so the description must carry the load. It meaningfully explains source_revision_id ('links a same-study source recipe') and expected_content_hash ('guards effective inputs'), but says nothing about recipe_id, name, reason, or expected_version beyond the stale-edit hint. Partial compensation for a low-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create an immutable revision.' Combined with the behavior notes (immutability, version guarding), an agent can tell this apart from create_study_recipe and refreeze_recipe_revision, though the siblings are never named explicitly. The purpose is unambiguous even without that routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete workflow: 'Reload list_study_recipes, reconcile changes, then submit the current version; never overwrite blindly,' plus the stale-edit condition (412). This gives clear when-to-use context, but it never states when NOT to use it versus siblings like update_recipe_draft or refreeze_recipe_revision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_optimizerRun OptimizerA
Run budget optimization on a completed model.
Finds the optimal budget allocation across channels to maximize predicted revenue — or predicted PROFIT with objective="profit" — within the given constraints.
PROFIT OBJECTIVE: objective="profit" requires a margin source. If the model was built with an operating margin, it is used automatically; otherwise you MUST pass forward_margin (e.g. 0.18 for an 18% margin) or the API returns an error. Result fields (Revenue, ROI, ExpectedResponse) are then on the profit basis.
IMPORTANT:
Channel names must exactly match model results (case-sensitive, space-sensitive). Results are keyed by the channel's ACTIVITY COLUMN name (e.g. "search_activity"), not by the
channels[].namepassed to create_model. Call get_model_results with sections="channel_summary" first to get exact names, or use get_scenario_template to discover channel names and their average CPM values.bounds values are percentages of total_budget (0-100), not currency amounts.
laydown_weights and period_cpm must be ARRAYS of length num_periods, not scalars. Wrong: {"TV": 10}. Correct: {"TV": [10, 10, 10, 10]}.
The same channel keys must appear in all three: bounds, laydown_weights, and period_cpm.
All period_cpm values must be positive (> 0).
laydown_weights per channel must sum to a positive value (weights are normalized internally).
Returns 202 (async). Use get_optimizer_results to poll until status is "complete".
Args: model_hash: Hash of a completed model. total_budget: Total budget in currency units. num_periods: Number of periods to optimize over (matches your planning horizon). gamma: Uncertainty-aversion weight on the outcome spread (the objective is mean - gamma * spread). 0.0 = maximize expected return only (most aggressive); higher values penalize uncertainty harder (more conservative). The dashboard typically uses values in the 0-0.1 range. currency: Currency code (e.g. "USD", "GBP"). bounds: Per-channel min/max budget allocation as PERCENTAGES (0-100). Every channel must appear. Example: {"TV_Impressions": {"lower": 5, "upper": 40}, "Search_Clicks": {"lower": 10, "upper": 50}} laydown_weights: Per-channel spend timing weights. Each value is an array of length num_periods. Weights are relative (normalized internally). Use uniform [1, 1, ...] for even distribution across periods. Example: {"TV_Impressions": [1, 1, 1, 1]} period_cpm: Per-channel cost-per-metric for each period. Each value is an array of length num_periods with positive values. Get baseline CPM from get_scenario_template (avg_cpu_by_channel field). Example: {"TV_Impressions": [10.5, 10.5, 10.5, 10.5]} objective: "revenue" (default) or "profit". See PROFIT OBJECTIVE above. forward_margin: Decimal margin in (0, 1], e.g. 0.18 = 18%. Only used with objective="profit"; required when the model has no stored operating margin. period_multiplier: Optional array of length num_periods converting KPI units to revenue per period over the planning horizon (mirrors the model's multiplier_column, e.g. price). include_historical_effect: Include carryover from historical spend in the predicted response (default True). enable_warm_start: Warm-start the optimizer from a previous solution (default True). optimizer_engine: "slsqp" (hardened SLSQP, default) or "marginal" (water-fill engine: allocates until every funded channel shows the same marginal return; exact profit-hurdle semantics and the tightest optimality certificates, with automatic SLSQP fallback). sigma_penalty: How gamma penalizes outcome spread: "std" (default), "variance" or "frozen" (advanced; smoother alternatives for hard-to-converge runs - leave on "std" normally). group_bounds: Joint constraints over channel SETS (#570), e.g. [{"name": "trade", "channels": ["TV", "Search"], "lower": 40, "upper": 60}] with lower/upper in % of total_budget (same convention as bounds). Groups must be disjoint and jointly feasible with the members' per-channel bounds. Presence forces the slsqp engine. Results gain GroupBounds/GroupBoundsReport columns; a BINDING group's members legitimately sit off the global marginal (they share the group's shadow price).
| Name | Required | Description | Default |
|---|---|---|---|
| gamma | Yes | ||
| bounds | Yes | ||
| currency | Yes | ||
| objective | No | revenue | |
| model_hash | Yes | ||
| period_cpm | Yes | ||
| num_periods | Yes | ||
| group_bounds | No | ||
| total_budget | Yes | ||
| sigma_penalty | No | std | |
| forward_margin | No | ||
| laydown_weights | Yes | ||
| optimizer_engine | No | slsqp | |
| enable_warm_start | No | ||
| period_multiplier | No | ||
| include_historical_effect | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only generic booleans (readOnlyHint false, etc.), so the description must disclose all behavioral traits. It does: returns 202 asynchronously, requires polling, mandates exact channel names, treats bounds as percentages, enforces array lengths, requires positive CPM, and explains the objective function (mean - gamma*spread). It also covers error conditions like missing forward_margin. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but justifiably so for a 16-parameter tool. It is well-structured: opening purpose, an 'IMPORTANT' section for critical gotchas, then an Args list. The most critical details (channel name exactness, bounds as percentages, array lengths) are front-loaded. Every sentence adds value, and redundancy between sections is minimal and serves emphasis. The length is proportionate to complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex with many interacting constraints and a schema that provides zero descriptions. The description covers all prerequisites, parameter semantics, error conditions, engine choices, and the async flow. It even explains the effect of group_bounds on results. Since an output schema exists, the description appropriately defers return-value details to get_optimizer_results. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries full responsibility for parameter meaning. It does so comprehensively: each of the 16 parameters gets a detailed explanation with examples, constraints, defaults, and relationships. For instance, gamma is defined as uncertainty-aversion weight with a range suggestion, bounds are clarified as percentages, and group_bounds includes a full example. No parameter is left vague.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Run budget optimization on a completed model.' It clearly states the goal (optimal budget allocation to maximize revenue or profit) and distinguishes itself from sibling tools like get_optimizer_results (polling) and run_scenario (scenario runs). The purpose is unambiguous and immediately actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs when to use related tools: call get_model_results first to obtain exact channel names, use get_scenario_template to discover CPM values, and poll get_optimizer_results after the 202 response. It also explains when the profit objective requires forward_margin and when the 'marginal' engine is appropriate. Clear prerequisites and follow-ups leave no room for confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_pipelineRun PipelineA
Start a refresh of one owned data pipeline (by pipeline_hash or id). Returns {run_id, status: "queued"} at once; the run executes on the server as the pipeline's owner with its saved connections. Poll get_pipeline_run until status is succeeded or failed (every few seconds; warehouse runs can take minutes, and a run stops at 30 minutes). Optional start_date / end_date (YYYY-MM-DD) limit the source steps to that range. One run per pipeline at a time: if one is already queued or running you get run_in_progress with that run_id — poll it instead of starting another. Each successful run saves a new pipeline version. Requires the create:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| start_date | No | ||
| pipeline_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations (readOnlyHint=false, idempotentHint=false) by disclosing that the run executes server-side as the owner with saved connections, returns {run_id, status: \"queued\"} immediately, enforces one-run-per-pipeline, produces a new pipeline version, and requires the create:models scope. These are exactly the operational traits an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense block that is front-loaded with the action and return shape, with each subsequent clause adding a distinct operational fact (polling, duration cap, concurrency, versioning, scope). It is long for one paragraph but essentially waste-free rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an async trigger tool with annotations and a sibling poller, everything needed to invoke correctly is present: inputs, immediate return, concurrency conflict handling, duration bound, and auth scope. An output schema exists, so return-value detail is covered structurally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does: pipeline_ref is explained as accepting a hash or id, and start_date/end_date are explained as YYYY-MM-DD values that limit source steps to that range. It stops just short of stating the date behavior when only one bound is supplied or the default full-refresh range.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a refresh of one owned data pipeline') and immediately scopes it by identifier type (pipeline_hash or id). It is clearly distinguishable from siblings like get_pipeline_run and list_pipelines, which are referenced as polling/listing operations rather than triggers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: poll get_pipeline_run until succeeded/failed, with cadence guidance ('every few seconds'), expected duration, and the 30-minute stop. It also names the alternative behavior when a run is already active ('poll it instead of starting another').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_scenarioRun ScenarioA
Run a "what-if" scenario prediction on a completed model.
Takes a set of future period rows with channel activity values and
predicts the KPI outcome. Use get_scenario_template first to get
the expected format, channel names, and baseline values. Channel names are
the activity-column keys from the template/results (e.g. "search_activity"),
not the channels[].name passed to create_model.
IMPORTANT: Before submitting, replace any NaN/null values in scenario_data with 0. The template from get_scenario_template may contain NaN for channels without historical data, which will cause the prediction to fail.
This is async (returns 202 with status "pending"). Poll get_scenario_results until status is "complete" or "failed".
Workflow: get_scenario_template -> modify values -> run_scenario -> poll get_scenario_results
Args: model_hash: Hash of a completed model. scenario_data: Array of period rows, each a dict with "Date" (YYYY-MM-DD format) and channel activity columns. Channel names must match exactly what get_scenario_template returns in the "channels" field. Example: [{"Date": "2025-01-06", "TV_Impressions": 50000, "Search_Clicks": 1200}] spend_metadata: Optional per-channel spend info for ROI calculation in results. Each entry: {"channel": "TV_Impressions", "metric": "impressions", "cpm": 25.0, "total_spend": 125000, "weekly_spend": [25000, 25000, ...]} rebuild_model: Recompile the model graph before prediction. Must be True (default) for API-initiated scenarios where the model graph is not in memory. evaluate_holdout: Evaluate the scenario against held-out actuals when the scenario period overlaps observed data (default False). skip_slicing: Skip per-channel contribution slicing in the prediction output — faster when only the KPI total is needed (default False). proxy_channels: Optional list of proxy-channel mappings, each mapping a scenario channel to a fitted channel whose transforms it borrows (for channels without their own history).
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes | ||
| skip_slicing | No | ||
| rebuild_model | No | ||
| scenario_data | Yes | ||
| proxy_channels | No | ||
| spend_metadata | No | ||
| evaluate_holdout | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses critical behavioral details beyond the annotations: the operation is async returning 202 with status pending, NaN values must be replaced with 0, and rebuild_model must be True for API-initiated scenarios. No contradiction with annotations is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured and dense with useful information, including an IMPORTANT warning and a workflow line. Each section earns its place, especially given the parameter count and zero schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers all seven parameters, prerequisites, async behavior, failure mode, and post-invocation polling. Since an output schema exists, not detailing return values is acceptable, and nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates by explaining every parameter with types, defaults, examples, and constraints. It even clarifies channel naming provenance and provides a concrete scenario_data example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: running a what-if scenario prediction on a completed model. It also differentiates itself from the related siblings by referencing the get_scenario_template -> run_scenario -> get_scenario_results workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to call get_scenario_template first to obtain the expected format and channel names, and to poll get_scenario_results until completion or failure. Provides a clear workflow and prerequisite ordering, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_incrementality_test_designSave Incrementality Test DesignAIdempotent
Turn an available test-design result into a planned incrementality test record in the registry: returns {id, version, record, content_hash} (201; 200 with the same body when this calculation was already saved). The server builds the record from the stored result (type time_holdout, or geo with treatment and control markets; status planned; source.tool design with the calculation, model artefact and method in source.fields), so only the target project_id (default: the model's project; 400 when the model has none and none is given) and an optional name are sent. Idempotent per calculation. Refusals: 409 result_not_available when the calculation is not complete or its result state is not available, 409 artifact_changed when the model artefact no longer matches the calculation, 403 when you are not the project's owner. A planned record has no measured lift: it cannot calibrate a model (get_incrementality_test reports type_not_calibratable or test_not_completed) and nothing here launches an experiment or changes a budget. Requires create:models and write access to the project.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| model_hash | Yes | ||
| project_id | No | ||
| calculation_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far exceeds the annotations: it discloses idempotency semantics, HTTP status codes (201/200), the specific 409/403 refusal reasons, required scopes (create:models, project write/owner), and the crucial behavioral caveat that a planned record has no measured lift and cannot calibrate a model. No contradiction with the declared hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded, with the return shape and idempotency up front and refusals/permissions trailing. It is a single long paragraph, so a reader must parse carefully, but nearly every clause carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a mutation tool: prerequisites, defaults, refusal modes, permission requirements, idempotency and the no-lift limitation are all covered. Return values are described even though an output schema exists, and the sibling relationship to get_incrementality_test is called out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage the description must carry the load, and it does: it explains that only project_id (defaulting to the model's project, 400 when neither exists) and an optional name are sent, and that the record body is built server-side from model_hash/calculation_id. It stops short of documenting the format of model_hash and calculation_id explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Turn an available test-design result into a planned incrementality test record in the registry'), and the surrounding detail (type time_holdout/geo, status planned, source.tool design) makes it clearly distinct from create_incrementality_test, design_incrementality_test and get_incrementality_test_design.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear precondition (an available test-design result) and explicitly bounds scope ('nothing here launches an experiment or changes a budget'), plus refusal conditions. It stops short of naming which sibling to use when the result is not yet available, so an agent must infer the alternative from the refusals.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_modelSave ModelADestructive
Save a model into a project under a display name.
API-created models start unsaved and are invisible to list_models (without include_unsaved=true) — saving files them into a project so they appear in the default listing and the dashboard's Saved Models.
The same saved-models cap applies as in the dashboard: at the cap the API returns a 400 with error_type "saved_limit". Re-saving an already-saved model renames/refiles it without consuming a new slot.
Args: model_hash: Hash of the model to save. name: Display name to save under (non-empty). project_id: Optional target project ID; must be a project you own or one shared with a team you belong to. Discover ids with list_projects; create a folder with create_project. Defaults to your default project.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| model_hash | Yes | ||
| project_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavior beyond the annotations: unsaved models are invisible to list_models without include_unsaved=true, the saved-model cap returns a 400 with error_type 'saved_limit', re-saving does not consume a new slot, and project ownership is required. This is exactly the kind of contextual detail an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a one-sentence purpose, then a useful lifecycle explanation, then an Args section. It is slightly long but every section contributes necessary behavioral or parameter guidance, so it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core action, the unsaved-to-saved transition, error behavior, parameter constraints, ownership rules, and defaults. With an output schema already present, no important operational detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains model_hash, name (non-empty), and project_id (optional, must be owned or team-shared, defaults to default project), and even suggests how to discover project ids. All three parameters receive meaningful semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Save a model into a project under a display name.' It then clarifies the unsaved vs. saved lifecycle and visibility in list_models, which makes the tool's role distinct from siblings like create_model or rename_model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context: API-created models start unsaved and need this call to become visible in default listings. It also explains project_id constraints and points to list_projects/create_project. However, it does not explicitly state when not to use it or name alternative tools like rename_model or unsave_model.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_campaign_mappingSet Campaign MappingADestructiveIdempotent
Replace the model's campaign map: which model channel each campaign counts towards.
Each row is {platform, campaign_id | name_pattern, channel, valid_from?, valid_to?}: exactly
one of campaign_id (the platform's id, exact) or name_pattern (a glob on the campaign
name, e.g. "YouTube*", case-insensitive); channel must be one of the model's channel names
(as in get_model_results); dates are ISO and optional (open-ended when absent). An exact-id
row beats any pattern row. The whole map is replaced by rows; send every row you want kept.
Refused (422, code campaign_map_invalid, nothing saved) when a channel is not the model's
(the error names the closest one), when a campaign would count towards two channels on any
date (two exact rows on one campaign with overlapping dates, or two patterns matching one
campaign with overlapping dates; conflicts lists them), or when a row is malformed.
Returns the map report: {map, unmapped: [{platform, campaign_id, campaign_name, spend,
first_seen, last_seen, platform_channel_type, suggested_channel}], conflicts: [], drift:
[{channel, campaign_spend, model_spend, ratio, over_tolerance, campaigns}], tolerance,
currency, currency_mismatch, window, channels}. drift compares the spend of the campaigns
mapped to each channel with the model's own spend for that channel over the dates both have;
over_tolerance is a WARNING (the map is saved), usually a campaign that belongs elsewhere
or model data that stops before the campaigns do. Unmapped campaigns stay listed and are not
counted; map them in a later call.
Args: model_hash: The fitted MMM the map belongs to. rows: The complete map. tolerance: The drift warning threshold as a fraction (default 0.05).
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| tolerance | No | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint/idempotentHint/openWorldHint, and the description adds substantial context beyond them: the full-replacement destruction of the previous map, the 422 refusal codes and that nothing is saved, conflict detection with `conflicts`, and that drift `over_tolerance` is only a warning while the map is still saved. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but well-structured and front-loaded with the core action, then the row format, then failure modes. Given the complexity of the row schema and error semantics, the length is largely justified, though the return-value enumeration is longer than strictly needed given an output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, schema-heavy tool with an output schema, the description covers row semantics, collisions, precedence, failure modes, and even the return report meaning. It is complete enough that an agent could invoke it correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden and does so: it fully specifies the row shape ({platform, campaign_id | name_pattern, channel, valid_from?, valid_to?}), the mutual exclusivity of campaign_id vs name_pattern, glob semantics with an example, ISO date rules, and precedence (exact-id beats pattern). Args section also clarifies model_hash and tolerance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Replace the model's campaign map: which model channel each campaign counts towards.' This is unambiguous and clearly distinct from siblings like list_campaigns or get_campaign_report. The agent knows exactly what operation this is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states replacement semantics ('The whole map is replaced by `rows`; send every row you want kept') and gives precise refusal conditions (422 campaign_map_invalid) with the cause of each. It also describes what to do with unmapped campaigns ('map them in a later call'), which is genuine when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_contribution_groupsSet Contribution GroupsADestructiveIdempotent
Persist the driver groupings the dashboard contributions view renders (#436) — configure grouping once and every viewer sees it.
Each group: {"name": str, "drivers": [column names], "color": "#hex"?, "baseAdjustments": {driver: "min"|"max"|"none"}?}. Driver names are validated against the model's media/control/halo/trademark factors (400 with a did-you-mean hint on typos); each driver may belong to at most one group; baseAdjustments must reference the group's own drivers. The special "_channel_color_overrides" pseudo-group carries a channelColors map instead of drivers.
NOTE: this is the CONTRIBUTIONS-VIEW grouping. create_model's channel_groups is the unrelated adstock parameter-sharing feature — do not confuse them.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes | ||
| contribution_groups | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false, idempotentHint=true, destructiveHint=true), the description adds rich behavioral context: validation against model factors, 400 responses with did-you-mean hints, one-group-per-driver constraints, baseAdjustments scoping, and the special _channel_color_overrides pseudo-group. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact despite covering complex validation rules. It front-loads the core purpose, then gives the group shape, constraints, and a disambiguation note—each sentence adds necessary information without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is complete: it explains the object format, validation behavior, persistence semantics, and the unrelated sibling feature. An output schema exists, so return-value documentation is not required from the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It supplies the exact structure of each contribution group, including optional fields, accepted value types, and validation semantics. This is far more than the input schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Persist the driver groupings the dashboard contributions view renders.' It also explicitly contrasts this tool with create_model's channel_groups, making the purpose unmistakable and distinguishing it from a plausible sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states this is the CONTRIBUTIONS-VIEW grouping and explicitly warns not to confuse it with create_model's channel_groups, naming the alternative feature as unrelated. This gives the agent a concrete when-not-to-use signal and prevents cross-tool confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_pipeline_scheduleSet Pipeline ScheduleADestructiveIdempotent
Replace the refresh schedule of one owned pipeline. cadence is "daily" or "weekly"; hour_utc is a whole UTC hour 0-23; weekday (0 = Monday … 6 = Sunday) is required for weekly and must be omitted for daily; enabled false pauses the schedule and keeps its settings. Returns the schedule with next_run_at (UTC). Each due slot starts one ordinary run (poll it with get_pipeline_run); a slot is skipped while a run of that pipeline is still going. After a scheduled run succeeds the pipeline keeps its newest 30 versions; a version a model or recipe was built from is never removed. Requires the create:models scope.
| Name | Required | Description | Default |
|---|---|---|---|
| cadence | Yes | ||
| enabled | Yes | ||
| weekday | No | ||
| hour_utc | Yes | ||
| pipeline_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds rich behavior beyond the annotations: slot-skipping while a run is in progress, one ordinary run per due slot, the newest-30-versions retention rule with protection for versions a model/recipe was built from, and the required create:models scope. This is substantial context the annotations alone do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then layers validation rules and operational behavior. Every sentence adds value, though the version-retention detail is slightly tangential to setting a schedule and makes the description on the longer side.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive/idempotent write with an output schema present, the description still covers inputs, side effects, retention consequences, scope requirement, and the resulting next_run_at. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden and largely delivers: cadence is "daily"/"weekly", hour_utc is a whole UTC hour 0-23, weekday is 0=Monday…6=Sunday, required for weekly and omitted for daily, and enabled false pauses while preserving settings. Only pipeline_ref is left unexplained, which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Replace the refresh schedule of one owned pipeline.' An agent can immediately tell this sets/replaces a schedule rather than polling (get_pipeline_run) or running a pipeline, and 'owned' scopes who it applies to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: 'enabled false pauses the schedule and keeps its settings', weekday required for weekly / omitted for daily, and points to get_pipeline_run for polling the resulting runs. No explicit when-not guidance or named alternative scheduler tool, but no sibling competes for this task either.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_run_pinnedSet Run PinnedADestructiveIdempotent
Pin or unpin a saved optimizer or scenario run.
Declarative and idempotent: setting the current state again is a no-op, so scripts can safely re-run it.
Args: artifact: "optimizer" (run_id "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model the run belongs to. run_id: The run's stable id from run history. pinned: Desired pin state.
| Name | Required | Description | Default |
|---|---|---|---|
| pinned | Yes | ||
| run_id | Yes | ||
| artifact | Yes | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint, readOnlyHint, and idempotentHint. The description adds specificity by explaining that setting the current pin state again is a no-op and that scripts can safely re-run it, which goes beyond the bare hint. It does not detail what happens when unpinning, but the annotation covers the destructive nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the purpose, and organized into a short declarative statement and a clear Args list. Every sentence contributes either to use intent, behavioral guarantees, or parameter understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-scalar-parameter tool with an output schema and idempotency/destructive annotations, the description covers all the information needed to invoke it correctly: accepted artifact types, id semantics, model association, and the pin state. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden, and it succeeds. Each parameter is explained: artifact values with run_id prefix examples, model_hash as the hash of the owning model, run_id as the stable id from run history, and pinned as the desired pin state. This adds the meaning the schema omits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence uses a specific verb and resource: 'Pin or unpin a saved optimizer or scenario run.' It clearly identifies the tool's action and object, making it distinguishable from generic siblings like update_run, and the artifact examples reinforce the domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it by stating the action ('Pin or unpin'), and it notes that idempotency makes re-running safe in scripts. However, it does not explicitly contrast this tool with alternatives or give exclusion criteria, leaving usage guidance to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_decompositionShow DecompositionBRead-onlyIdempotent
Show served contribution time series in KPI units in a native chart.
Read-only. Overlap is a separate reconciliation term, never a channel. The attribution convention is displayed where available. Returns the same useful JSON on clients without visual support. Never requests prediction-window data.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive behavior, so the redundancy of 'Read-only' is low value; however the description adds real context beyond structured data: 'Returns the same useful JSON on clients without visual support', 'attribution convention is displayed where available', and the prediction-window constraint. These disclose return behavior and data-scope limits not captured by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and compact overall. The clipped sentence fragments ('Read-only.', 'Never requests prediction-window data.') are terse but each carries distinct information, so little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, return values needn't be explained, and behavioral context is reasonably covered. The gap is the undocumented required model_hash, plus no framing of how the tool relates to its many visualization siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole required parameter, model_hash, is never mentioned in the description. Unlike a zero-parameter tool, the description leaves the required identifier's meaning entirely undocumented in both structured and unstructured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Show served contribution time series in KPI units in a native chart.' An agent can understand the operation, though the description never contrasts it with closely-named siblings such as show_response_curves or show_optimizer_allocation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no routing to alternatives. 'Never requests prediction-window data' hints at a boundary but does not tell the agent when to prefer this tool over its many show_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_optimizer_allocationShow Optimizer AllocationARead-onlyIdempotent
Show saved allocation spend and separate decision/comparison result tables.
Read-only. Pass run_id to read a specific saved run; omission reads the latest model-level result. Decision Revenue/ROI never mixes with fitted-convention OptimizedEvalRevenue/ROI or HistoricalRevenue/ROI. No scientific calculations occur in the view. Returns existing JSON unchanged without visual support.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent/destructive profile, but the description adds real behavioral context beyond them: no scientific calculations occur, JSON is returned unchanged, there is no visual support, and the Revenue/ROI conventions are never mixed. This is meaningful disclosure an agent couldn't infer from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then the read-only note, then parameter behavior, then data-convention constraints. Dense but every sentence carries information; no filler. The Revenue/ROI sentence is technical but load-bearing for correct interpretation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be detailed, and the description usefully characterizes the return ('existing JSON unchanged'). But it omits any explanation of the required model_hash, leaving a gap for a required input on a moderately complex tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It explains run_id's behavior well (specific run vs. latest model-level result) but says nothing about the required model_hash parameter, leaving one of two parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Show saved allocation spend') and clarifies the output shape ('separate decision/comparison result tables'). It's clearer than a bare 'show' but doesn't explicitly differentiate itself from the sibling get_optimizer_results, leaving some ambiguity about which optimizer view to pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides conditional guidance for run_id ('pass run_id to read a specific saved run; omission reads the latest model-level result'), which is genuine usage context. However, it offers no when-to-use/when-not guidance relative to siblings like get_optimizer_results or run_optimizer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_response_curvesShow Response CurvesBRead-onlyIdempotent
Show served response curves, bands and current spend in a native chart.
Read-only. Returns the existing JSON unchanged on every client. Missing points are gaps; sparse legacy grids are disclosed. Current spend is not a recommended allocation. Requests fixed sections and never accesses prediction-window data.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the 'Read-only' restatement is redundant. The description does add genuine context beyond the annotations: mean-null gaps, legacy sparse grid disclosure, the caveat that current spend is not a recommended allocation, and the fact that fixed sections are requested. But it omits output shape details and any rate or auth context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The fronted sentence does the work, and the following four short lines each add a distinct behavioral note. The phrase 'Returns the existing JSON unchanged' is slightly cryptic, but overall the description earns its space with little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail need not be spelled out. Yet for a one-param tool with 0% schema coverage, the description gives no hint about what model_hash identifies or how the chart is requested. Useful behavioral notes are present, but the description is not complete for an unfamiliar agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single 'model_hash' parameter, so the description could have compensated but never mentions it. The baseline for one param with a bare schema is a 3, and no additional syntax or format information is offered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Show served response curves, bands and current spend in a native chart.' An agent can identify the tool's function, but the description never anchors it against siblings like show_decomposition or show_optimizer_allocation, which also render model charts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternative sibling is named despite a large sibling set containing closely related visualization tools. The mention of 'never accesses prediction-window data' hints at a constraint but gives no positive routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlink_var_modelUnlink Var ModelADestructiveIdempotent
Remove an MMM's VAR link (#569). Idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the description's 'Idempotent' note is somewhat redundant but reinforces the key behavioral trait. The description adds the specific scope ('an MMM's VAR link') and references issue #569, which gives context. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The core action is front-loaded, and the idempotency note is a valuable addition. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive operation, the description is mostly adequate. The output schema exists, so return values are covered. However, the description doesn't clarify what 'VAR link' means in this domain or what happens to the model after unlinking, which could matter for an agent deciding whether to call this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameter meaning. However, the description doesn't explain what model_hash is or how to obtain it. The parameter name is fairly self-explanatory, but the description adds no semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Remove an MMM's VAR link') with a clear resource, and the idempotency note adds useful precision. It doesn't explicitly distinguish from sibling tools like link_var_model, but the verb 'unlink' and the resource 'VAR link' make the purpose clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when you need to remove a VAR link from an MMM. It doesn't explicitly state when not to use it or name alternatives, but the sibling context (link_var_model) makes the contrast fairly obvious. No explicit exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unsave_modelUnsave ModelADestructive
Release a model's saved slot without deleting anything — the inverse of save_model (#673).
Use this for cap management: at the 20-saved-models cap, unsave a model that no longer earns its shelf spot instead of deleting it. The model reverts to the state API-created models start in (unsaved, no project; the name is kept) — it leaves the default listing and the dashboard's Saved Models but stays fully addressable by hash: fetchable, renameable, exportable, re-saveable, and visible via list_models with include_unsaved=true. Idempotent — unsaving an unsaved model is a success with freed_project_id null. delete_model remains failed-only.
Two caveats: the UNSAVED pool is auto-pruned by dashboard model creation (at 10+ unsaved models the oldest is hard-deleted, artifacts included), so re-save anything worth keeping rather than parking it unsaved long-term; and unsaving a shared model hides it from every recipient until it is saved again.
Args: model_hash: Hash of the model whose slot to release.
Returns: {model_hash, is_saved: false, freed_project_id}.
| Name | Required | Description | Default |
|---|---|---|---|
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is exemplary in disclosing side effects, idempotency, retention of hash addressability, and auto-pruning risks. However, the annotations state idempotentHint=false while the description explicitly says 'Idempotent — unsaving an unsaved model is a success.' This is a direct annotation contradiction, which per the rubric requires a score of 1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: it front-loads the core purpose, then covers use case, state transitions, idempotency, return shape, and important caveats. The structured Args/Returns sections make the information easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and an output schema, the description covers everything needed: behavior, model state after unsaving, listing visibility, hash addressability, re-saving, idempotency, pruning caveat, shared-model caveat, and return fields. The only downside is the annotation contradiction, but the description itself is contextually complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter model_hash has zero percent schema description coverage, but the description compensates fully by explaining it as 'Hash of the model whose slot to release.' This is precise and sufficient for an agent to populate the parameter correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Release a model's saved slot without deleting anything' and explicitly frames it as the inverse of save_model. It clearly separates unsave_model from delete_model and save_model, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: at the 20-saved-models cap, unsave instead of delete. It also states caveats for shared models and the auto-pruning of unsaved models, helping the agent decide whether this tool is appropriate in a given situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_recipe_draftUpdate Recipe DraftADestructiveIdempotent
Replace authoring state using the version from get_recipe_draft. Retain every unedited field, including original priors and source bytes. Stale changes fail; reload and reconcile explicitly. Identical retries return the current draft. Does not publish or launch.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| draft_id | Yes | ||
| snapshot | Yes | Lossless editor authoring state. Backend is authoritative; preserve unknown nested fields. Source bytes are copied into encrypted draft storage (10 MB source limit); metadata limit is 5 MB. No local filesystem paths or executable code. | |
| expected_version | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds real semantics beyond the annotations: merge/retention rules ("Retain every unedited field, including original priors and source bytes"), optimistic-concurrency failure ("Stale changes fail; reload and reconcile explicitly"), and scope ("Does not publish or launch"). The idempotent-retry sentence largely restates idempotentHint=true, only adding "returns the current draft."
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, front-loaded sentences that each carry a distinct point (source of version, retention rule, concurrency behavior, retry behavior, scope). No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex nested mutation with an output schema and full safety annotations, the description covers mutation semantics, concurrency, idempotency, and scope adequately. The remaining gap is the undocumented top-level identity inputs (draft_id, name), which is minor given rich nested schema documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Top-level schema coverage is 25% and the description never names draft_id, name, or snapshot directly, so those params rely on the nested schema. It does add meaning for expected_version (stale-fail/reload) and for source bytes/priors retention, which partially compensates for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Replace authoring state" gives a specific verb+resource and the phrase "Does not publish or launch" distinguishes it from publish_recipe_draft and create_recipe_draft. It is clear but slightly indirect about the object (recipe draft) since the title carries that load.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"using the version from get_recipe_draft" tells the agent when/how to sequence the call, and "Does not publish or launch" excludes the publish path. It stops short of naming the alternative tool (publish_recipe_draft) explicitly or stating full prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_runUpdate RunADestructive
Rename / annotate a saved optimizer or scenario run.
Runs are auto-named at creation (e.g. "$1.2M · 12mo · Jan 5"); renaming makes run history carry the analysis ("holiday cut -10%", "stretch 130%"). Renaming permanently flips the run's auto_named flag to false so future auto-naming never overwrites it. Only the fields you provide are changed.
Args: artifact: "optimizer" (run_id "opt_...") or "scenario" ("scn_..."). model_hash: Hash of the model the run belongs to. run_id: The run's stable id from run history. name: New display name (non-empty when given; capped at 255 chars). notes: Free-text annotation. Omit to leave untouched; pass "" to clear. tags: Replacement tag list (max 20 tags, 64 chars each).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| tags | No | ||
| notes | No | ||
| run_id | Yes | ||
| artifact | Yes | ||
| model_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing that renaming permanently flips the auto_named flag to false, that only provided fields are changed, and that notes can be omitted to leave untouched or passed as empty string to clear. This is valuable behavioral context that complements the destructiveHint and readOnlyHint annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a concise summary line followed by a clear Args list. Every sentence earns its place: the auto-naming example illustrates intent, the permanent flag note warns about irreversibility, and the parameter constraints are precise. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all six parameters, the partial-update behavior, the permanent auto_named flip, and edge-case semantics for notes. Since an output schema is present and annotations already convey the read/write/destructive profile, an agent has everything necessary to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the Args section fully compensates by documenting every parameter: artifact values with run_id prefixes, model_hash, run_id, name constraints (non-empty, capped at 255), notes semantics (omit vs empty), and tags limits (max 20 tags, 64 chars each). The description adds significant meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening line uses a specific verb-resource pair ('Rename / annotate a saved optimizer or scenario run'), which clearly identifies the tool's purpose and distinguishes it from sibling tools like set_run_pinned or update_recipe_draft. It also names the two artifact types (optimizer/scenario) and their run_id prefixes, leaving no ambiguity about what the tool operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: runs are auto-named at creation and renaming makes run history carry the analysis, implying this tool is for annotating/renaming existing runs. It doesn't explicitly reference alternatives or exclusion conditions, but the context is strong enough that an agent can infer the appropriate use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_studyUpdate StudyADestructive
Replace owner-controlled study settings. Send the full object: every field is assigned, so read the study first and pass its current values plus version (stale versions fail with 412). Omitting state or limits does not preserve them; omitting context keeps the stored context and an empty string clears it. State is active, paused or archived; paused blocks new reservations and does not cancel running work. Editing question or context triggers no action.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| state | No | Backend state: active, paused or archived. | active |
| context | No | Optional scope, data caveats and assumptions a reader needs to interpret results. Not rules; see question. Omit to leave unchanged on update; send an empty string to clear. | |
| version | Yes | ||
| question | Yes | What the study should find out, in one or two sentences. Stored and shown to humans; never executed or used as an acceptance rule. Do not put thresholds, validation rules or fit limits here: those belong in a quality policy, recipe revisions and max_attempts/max_concurrent. Exploratory and reliability questions are valid. | |
| study_id | Yes | ||
| max_attempts | Yes | ||
| max_concurrent | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as destructive and non-idempotent, but the description adds substantial behavioral detail: every field is assigned, stale versions fail with 412, omitting state or limits drops them, omitting context preserves it, empty string clears it, and paused does not cancel running work. This goes well beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries essential operational information, and the most important concept—full replacement—is front-loaded. The description is dense but efficient, with no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent update with 8 parameters and low schema coverage, the description covers the major gotchas: read-before-update, optimistic concurrency via version, omission semantics, and state behavior. The output schema exists, so return values need no description, but a bit more detail on max_attempts/max_concurrent and the exact meaning of 'triggers no action' would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 38%, but the description compensates for the most important parameters: version's concurrency role, state's allowed values and side effects, context's preserve/clear behavior, and question's 'no action' semantics. Parameters like max_attempts and max_concurrent are only addressed as 'limits', so not every parameter is fully elaborated, but the critical ones are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Replace owner-controlled study settings.' It immediately clarifies that this is a full-object replacement, not a partial update, and the resource 'study settings' distinguishes it from sibling tools like update_run or update_recipe_draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong process guidance: read the study first, pass current values plus version, and understand omission semantics for state, limits, and context. It does not explicitly name alternatives or when-not-to-use conditions, but the context is clear enough for an agent to use this tool correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_dataUpload DataA
Upload a CSV dataset to Simba for use in model building.
Provide EXACTLY ONE of csv_content (raw CSV text) or csv_path (a file path on the machine running this MCP server). Prefer csv_path for anything beyond trivial size — it avoids passing megabytes of CSV through the conversation.
The CSV should follow the canonical schema: one row per time period with date, KPI, multiplier, hierarchy, media activity/spend columns, and optional control variables.
IMPORTANT:
CSV only (not Excel). Maximum file size: 10 MB (API-enforced).
Row minimum: check get_data_schema -> x-simba-constraints.min_rows for the declared minimum; enforcement may be more permissive, and the upload response's
warningsfield is authoritative. More rows = tighter posteriors (104+ weekly rows recommended).Media columns must follow naming: {channel}_activity and {channel}_spend.
Use 0 for inactive periods, not blank or NA.
csv_path is only available when the server runs locally (stdio). On HTTP/SSE deployments it is disabled unless SIMBA_MCP_ALLOW_LOCAL_FILES=1.
Args: csv_content: The full CSV text content (not base64, just raw CSV text). csv_path: Path to a .csv file readable by the MCP server process. name: Optional dataset name for identification. Defaults to the file stem when csv_path is used. filename: Optional original filename to record alongside the dataset. roles: Optional column roles for get_data_report, stored with the dataset: {column: role} or {column: {"role": role, "channel": name}}. Roles are declared, never guessed; see get_data_report for the vocabulary. An unknown role or a column the CSV lacks is refused.
Returns the uploaded file ID (needed for create_model), row/column counts, and any validation warnings.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| roles | No | ||
| csv_path | No | ||
| filename | No | ||
| csv_content | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations, which only flag a non-readonly, non-idempotent, open-world write. The description adds hard operational constraints the agent cannot get elsewhere: CSV-only (not Excel), 10 MB API-enforced cap, row-minimum sourcing from get_data_schema with warnings as authoritative, naming conventions, and the stdio-vs-HTTP availability of csv_path. This is exactly the kind of context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose and mutual-exclusion rule, then groups constraints under an IMPORTANT block. Dense but every bullet (size cap, row minimum, naming, 0-fill, csv_path availability) is actionable. Slightly verbose, but the length is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite an output schema already existing, the description still summarizes the return (file ID for create_model, counts, warnings), and covers constraints, deployment caveats, and role semantics. For a 5-param upload tool with no schema descriptions, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden, and it does: it documents all five parameters in an Args block, including the mutual exclusion of csv_content/csv_path, defaulting behavior of name, and the shape and refusal semantics of roles. This adds real meaning beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (upload), resource (CSV dataset), destination (Simba), and immediate downstream purpose (model building). This clearly distinguishes it from sibling read tools like get_upload or list_uploads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit rule for choosing between the two mutually exclusive input modes: 'Provide EXACTLY ONE... Prefer csv_path for anything beyond trivial size.' It also explains the deployment condition under which csv_path is available. No alternative tool is named, but the routing guidance for its own arguments is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_study_recipeValidate Study RecipeARead-onlyIdempotent
Resolve and validate a recipe without creating a run. Returns effective settings, provenance limits and the same inspection block a saved revision would carry (authored/default settings, inert prior fields with their gate, engine state, lineage with the recorded dataset's availability and display line), so an agent can check for inert fields and an unavailable dataset before freezing.
| Name | Required | Description | Default |
|---|---|---|---|
| specification | Yes | Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, unknown config keys are rejected, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, so the safety profile is covered; the description adds genuinely new behavioral detail about what the inspection block carries (gate on inert prior fields, engine state, lineage with dataset availability). It also notes dataset availability is 'checked now and never inferred,' which is useful non-obvious behavior, though much of the return-shape detail is duplicated from the richer schema description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause and the rest enumerates what is returned. It is a single long sentence with a dense parenthetical, which costs a little readability, but every clause carries information and nothing is redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only validation tool with an output schema present, the definition covers purpose, decision context, and the shape of the inspection payload. Return values need not be spelled out given the output schema, so remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single nested specification parameter is documented in depth (api_mmm vs model_snapshot constraints, rejected keys, review-only limits). The description adds no parameter-level meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource pair ('Resolve and validate a recipe') and immediately scopes it against the obvious write-path siblings by stating 'without creating a run.' An agent can distinguish this from create_study_recipe, revise_study_recipe, and launch_study_run without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear decision context: use it to 'check for inert fields and an unavailable dataset before freezing,' which implies the pre-commit checkpoint role. It does not explicitly name the alternative tool to use once validation passes (e.g., create_study_recipe/launch_study_run), so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.17.0- Changed
create_incrementality_test3 fields changed- changed
Input schema / properties / record / descriptionPrevious value: -"One incrementality test as analysed in its own tool (Simba fits nothing here). type picks the design block: geo, owned_media_ab or platform_lift. Dates are YYYY-MM-DD; measured_through is the end of the carryover window (default end_date). result.lift_abs is the total lift in KPI units over the window; give its interval (two-sided, level e.g. 0.9) and/or sd. Null and negative lifts are valid. spend.incremental is the extra spend the test caused (signed)."New value: +"One incrementality test as analysed in its own tool (Simba fits nothing here). type picks the design block: geo, owned_media_ab or platform_lift. Dates are YYYY-MM-DD; measured_through is the end of the carryover window (default end_date). result.lift_abs is the total lift in KPI units over the window; give its interval (two-sided, level e.g. 0.9) and/or sd. Null and negative lifts are valid. spend.incremental is the extra spend the test caused (signed). The time_holdout type and block appear on records the server builds from a test-design calculation (save_incrementality_test_design); they are read and listed here but not accepted on create or import, and a status planned record carries no result." - added
Input schema / properties / record / properties / time_holdoutAdded value: +{ + "description": "Server-built from a test design: the outcome in the holdout window is contrasted with the model's forecast.", + "properties": { + "analysis_method": { + "type": "string" + }, + "carryover_periods": { + "minimum": 0, + "type": "integer" + }, + "counterfactual": { + "enum": [ + "model_forecast" + ], + "type": "string" + } + }, + "type": "object" +} - changed
Input schema / properties / record / properties / type / enumPrevious value: -[ - "geo", - "owned_media_ab", - "platform_lift" -]New value: +[ + "geo", + "owned_media_ab", + "platform_lift", + "time_holdout" +]
- Added
design_incrementality_test - Added
get_incrementality_test_design - Changed
list_incrementality_tests1 field changed- changed
Input schema / properties / type / anyOfPrevious value: -[ - { - "enum": [ - "geo", - "owned_media_ab", - "platform_lift" - ], - "type": "string" - }, - { - "type": "null" - } -]New value: +[ + { + "enum": [ + "geo", + "owned_media_ab", + "platform_lift", + "time_holdout" + ], + "type": "string" + }, + { + "type": "null" + } +]
- Added
save_incrementality_test_design
6 tool updates
v0.16.0- Added
get_campaign_marginal_returns - Added
recommend_campaign_budgets - Added
recommend_incrementality_tests - Added
show_decomposition - Added
show_optimizer_allocation - Added
show_response_curves
4 tool updates
v0.14.0- Added
get_campaign_incrementality - Added
get_campaign_report - Added
list_campaigns - Added
set_campaign_mapping
17 tool updates
v0.12.0- Added
create_incrementality_test - Changed
create_model1 field changed- added
Input schema / properties / calibrationAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "description": "Calibrate the fit with lift-test likelihood observations, in exactly one form: {tests: [{test_id, version?, channel?, confirm_kpi?}]} (1-50 recorded tests, each derived against this model's own data), or {units: 'revenue' | 'response', observations: [{channel, x, delta_x, delta_y, sigma, sigma_low?, sigma_high?}]} (1-200 rows you derived yourself). If any test can't calibrate this model the call fails with calibration_refused and a reason per test in `tests`; check a test first with get_incrementality_test(test_id, model_hash=...) on a saved model built on the same data.", + "properties": { + "observations": { + "items": { + "properties": { + "channel": { + "type": "string" + }, + "delta_x": { + "type": "number" + }, + "delta_y": { + "type": "number" + }, + "sigma": { + "exclusiveMinimum": 0, + "type": "number" + }, + "sigma_high": { + "exclusiveMinimum": 0, + "type": "number" + }, + "sigma_low": { + "exclusiveMinimum": 0, + "type": "number" + }, + "x": { + "type": "number" + } + }, + "required": [ + "channel", + "x", + "delta_x", + "delta_y", + "sigma" + ], + "type": "object" + }, + "maxItems": 200, + "minItems": 1, + "type": "array" + }, + "tests": { + "items": { + "additionalProperties": false, + "description": "A recorded test by reference; the row is derived against the model's own data.", + "properties": { + "channel": { + "description": "Model activity column to calibrate; default: the test's model_channel.", + "type": "string" + }, + "confirm_kpi": { + "description": "Assert the test's outcome is the model's KPI when their names differ (recorded in the steps).", + "type": "boolean" + }, + "test_id": { + "type": "string" + }, + "version": { + "description": "Default: the current version.", + "minimum": 1, + "type": "integer" + } + }, + "required": [ + "test_id" + ], + "type": "object" + }, + "maxItems": 50, + "minItems": 1, + "type": "array" + }, + "units": { + "enum": [ + "revenue", + "response" + ], + "type": "string" + } + }, + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Calibration" +}
- Changed
create_recipe_draft1 field changed- added
Input schema / properties / snapshot / properties / incrementality_testsAdded value: +{ + "description": "Optional recorded incrementality tests by reference ({test_id, version?, channel?, confirm_kpi?}); each is derived against every prepared brand's own data when the draft is published, and the revision records which test versions it used. A test that can't calibrate a brand fails publication with its reason.", + "items": { + "type": "object" + }, + "maxItems": 50, + "type": "array" +}
- Changed
create_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, unknown config keys are rejected, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."
- Added
get_data_report - Added
get_incrementality_test - Changed
get_model_results3 fields changed- added
Input schema / properties / endAdded value: +{ + "default": "", + "title": "End", + "type": "string" +} - added
Input schema / properties / granularityAdded value: +{ + "default": "", + "title": "Granularity", + "type": "string" +} - added
Input schema / properties / startAdded value: +{ + "default": "", + "title": "Start", + "type": "string" +}
- Added
get_pipeline_run - Added
get_workflow_guidance - Added
import_incrementality_tests - Added
list_incrementality_tests - Changed
revise_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, unknown config keys are rejected, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."
- Added
run_pipeline - Added
set_pipeline_schedule - Changed
update_recipe_draft1 field changed- added
Input schema / properties / snapshot / properties / incrementality_testsAdded value: +{ + "description": "Optional recorded incrementality tests by reference ({test_id, version?, channel?, confirm_kpi?}); each is derived against every prepared brand's own data when the draft is published, and the revision records which test versions it used. A test that can't calibrate a brand fails publication with its reason.", + "items": { + "type": "object" + }, + "maxItems": 50, + "type": "array" +}
- Changed
upload_data1 field changed- added
Input schema / properties / rolesAdded value: +{ + "anyOf": [ + { + "additionalProperties": true, + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Roles" +}
- Changed
validate_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, unknown config keys are rejected, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."
6 tool updates
v0.8.2- Changed
create_recipe_draft1 field changed- added
Input schema / properties / targetAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "description": "The recipe an edit session works on (jellyfish #881). recipe_id must belong to the draft's study; base_revision_id, when given, must belong to that recipe (404 otherwise) and is recorded as the new revision's source. Immutable after creation: a replay naming another target is refused (409).", + "properties": { + "base_revision_id": { + "type": [ + "string", + "null" + ] + }, + "recipe_id": { + "type": "string" + } + }, + "required": [ + "recipe_id" + ], + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Target" +}
- Changed
create_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."
- Added
diff_recipe_revisions - Changed
publish_recipe_draft1 field changed- added
Input schema / properties / targetAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "description": "Publish into an existing recipe as its next revision (jellyfish #881). expected_version is the recipe version you last read; 412 stale_version when it moved, with the draft and every edit kept.", + "properties": { + "expected_version": { + "minimum": 1, + "type": "integer" + }, + "recipe_id": { + "type": "string" + } + }, + "required": [ + "recipe_id", + "expected_version" + ], + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Target" +}
- Changed
revise_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."
- Changed
validate_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation. Every saved revision's read-time inspection carries lineage: the dataset origin recorded when the model was built, checked for availability now and never inferred after the fact."
3 tool updates
v0.7.3- Changed
create_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request; model_snapshot requires model_hash and is review-only. Unknown fields are forwarded for backend validation."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation."
- Changed
revise_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request; model_snapshot requires model_hash and is review-only. Unknown fields are forwarded for backend validation."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation."
- Changed
validate_study_recipe1 field changed- changed
Input schema / properties / specification / descriptionPrevious value: -"Backend recipe envelope. api_mmm requires request; model_snapshot requires model_hash and is review-only. Unknown fields are forwarded for backend validation."New value: +"Backend recipe envelope. api_mmm requires request (a create_model body; model_type must be mmm, config.auto_prior must be false, channel_map and var_model_hash are rejected); model_snapshot requires model_hash and is review-only (launch refused, lineage unknown). Smart priors and VAR recipes are available only through the authoring-draft tools. Unknown fields are forwarded for backend validation."
4 tool updates
v0.7.1- Changed
create_quality_policy4 fields changed- changed
Input schema / properties / checks / items / descriptionPrevious value: -"Built-in metric + maximum, or custom numeric gate with metric custom:<slug>, name, units, operator and applicable bounds. Boolean/manual checks use kind, operator equals and strict boolean expected. Manual expected must be true and sign-off is session-only. At least one required gate per policy."New value: +"Built-in metric + maximum (or operator gte/between with minimum), or custom numeric gate with metric custom:<slug>, name, units, operator and applicable bounds. Boolean/manual checks use kind, operator equals and strict boolean expected. Manual expected must be true and sign-off is session-only. Rules over saved artifacts need no external evidence: saved diagnostics (r_squared, mape, durbin_watson, pareto_k_pct, normality_p, loo_cv; bounds may be negative for loo_cv and r_squared), retained-sampling counts (retained_chains, retained_draws_per_chain, ess_bulk_min, ess_tail_min, divergences; non-negative) and provenance status (provenance:holdout, provenance:prior; operator equals, expected review_required — blocked fails, an absent record is not_collected). Every threshold is the author's; an absent artifact never passes. At least one required gate per policy; at most 20 checks in total." - changed
Input schema / properties / checks / items / properties / expected / typePrevious value: -"boolean"New value: +[ + "boolean", + "string" +] - changed
Input schema / properties / checks / items / properties / metric / examplesPrevious value: -[ - "r_hat_max", - "mae", - "rmse", - "wape", - "prediction_mae", - "prediction_rmse", - "prediction_wape", - "custom:benchmark_deviation" -]New value: +[ + "r_hat_max", + "mae", + "rmse", + "wape", + "prediction_mae", + "prediction_rmse", + "prediction_wape", + "r_squared", + "mape", + "durbin_watson", + "pareto_k_pct", + "normality_p", + "loo_cv", + "retained_chains", + "retained_draws_per_chain", + "ess_bulk_min", + "ess_tail_min", + "divergences", + "provenance:holdout", + "provenance:prior", + "custom:benchmark_deviation" +] - added
Input schema / properties / derived_from_policy_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Derived From Policy Id" +}
- Added
diff_quality_policies - Changed
evaluate_study_run3 fields changed- added
Input schema / properties / carry_forwardAdded value: +{ + "anyOf": [ + { + "additionalProperties": false, + "description": "Reuse an earlier assessment's external evidence for the named custom metrics. Allowed only when the preview listed the metric under carry_forward_available for that assessment: same run, same evidence basis, same policy rules. Manual sign-off is never carried.", + "properties": { + "from_evaluation_id": { + "type": "string" + }, + "metrics": { + "items": { + "pattern": "^custom:[a-z][a-z0-9_]{0,63}$", + "type": "string" + }, + "maxItems": 20, + "minItems": 1, + "type": "array" + } + }, + "required": [ + "from_evaluation_id", + "metrics" + ], + "type": "object" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Carry Forward" +} - changed
Input schema / properties / external_evidence / anyOfPrevious value: -[ - { - "items": { - "additionalProperties": false, - "description": "Externally calculated numeric or strict boolean evidence; server applies policy. Manual sign-off is session-only and cannot be submitted with an API key. Source/method/digest are submitter-reported, not independently verified. Never submit a pass/fail status.", - "properties": { - "method": { - "maxLength": 5000, - "type": "string" - }, - "metric": { - "pattern": "^custom:[a-z][a-z0-9_]{0,63}$", - "type": "string" - }, - "source_reference": { - "maxLength": 2000, - "type": "string" - }, - "source_sha256": { - "pattern": "^[a-f0-9]{64}$", - "type": "string" - }, - "value": { - "type": [ - "number", - "boolean" - ] - } - }, - "required": [ - "metric", - "value", - "method", - "source_reference" - ], - "type": "object" - }, - "type": "array" - }, - { - "type": "null" - } -]New value: +[ + { + "items": { + "additionalProperties": false, + "description": "Externally calculated numeric or strict boolean evidence; server applies policy. Manual sign-off is session-only and cannot be submitted with an API key. Source/method/digest are submitter-reported, not independently verified. Never submit a pass/fail status. To reuse an earlier submission on the same basis, use carry_forward instead of retyping.", + "properties": { + "method": { + "maxLength": 5000, + "type": "string" + }, + "metric": { + "pattern": "^custom:[a-z][a-z0-9_]{0,63}$", + "type": "string" + }, + "source_reference": { + "maxLength": 2000, + "type": "string" + }, + "source_sha256": { + "pattern": "^[a-f0-9]{64}$", + "type": "string" + }, + "value": { + "type": [ + "number", + "boolean" + ] + } + }, + "required": [ + "metric", + "value", + "method", + "source_reference" + ], + "type": "object" + }, + "type": "array" + }, + { + "type": "null" + } +] - added
Input schema / properties / previewAdded value: +{ + "default": false, + "title": "Preview", + "type": "boolean" +}
- Added
get_quality_policy
15 tool updates
v0.6.1- Changed
create_study3 fields changed- added
Input schema / properties / contextAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional scope, data caveats and assumptions a reader needs to interpret results. Not rules; see question. Omit to leave unchanged on update; send an empty string to clear.", + "title": "Context" +} - added
Input schema / properties / question / descriptionAdded value: +"What the study should find out, in one or two sentences. Stored and shown to humans; never executed or used as an acceptance rule. Do not put thresholds, validation rules or fit limits here: those belong in a quality policy, recipe revisions and max_attempts/max_concurrent. Exploratory and reliability questions are valid." - added
Input schema / properties / question / examplesAdded value: +[ + "How much do paid search and paid social contribute to weekly sales after price, promotions and seasonality?", + "Which carryover and saturation choices does the data support for TV, and how sensitive are contributions to them?", + "Can a weekly model across all channels produce estimates that hold up on a temporal holdout?" +]
- Added
get_launch_eligibility - Added
get_study_overview - Added
list_pipeline_versions - Added
list_pipelines - Changed
list_quality_policies2 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Changed
list_recipe_drafts2 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Changed
list_studies2 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Changed
list_study_decisions2 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Changed
list_study_evaluations3 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / expandAdded value: +{ + "anyOf": [ + { + "items": { + "const": "report", + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Expand" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Changed
list_study_recipes3 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / expandAdded value: +{ + "anyOf": [ + { + "items": { + "enum": [ + "effective", + "inspection", + "specification" + ], + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Expand" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Changed
list_study_runs2 fields changed- added
Input schema / properties / cursorAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Cursor" +} - added
Input schema / properties / limitAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Limit" +}
- Added
refreeze_recipe_revision - Added
retire_quality_policy - Changed
update_study3 fields changed- added
Input schema / properties / contextAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Optional scope, data caveats and assumptions a reader needs to interpret results. Not rules; see question. Omit to leave unchanged on update; send an empty string to clear.", + "title": "Context" +} - added
Input schema / properties / question / descriptionAdded value: +"What the study should find out, in one or two sentences. Stored and shown to humans; never executed or used as an acceptance rule. Do not put thresholds, validation rules or fit limits here: those belong in a quality policy, recipe revisions and max_attempts/max_concurrent. Exploratory and reliability questions are valid." - added
Input schema / properties / question / examplesAdded value: +[ + "How much do paid search and paid social contribute to weekly sales after price, promotions and seasonality?", + "Which carryover and saturation choices does the data support for TV, and how sensitive are contributions to them?", + "Can a weekly model across all channels produce estimates that hold up on a temporal holdout?" +]
63 tool updates
v0.5.0- Added
adopt_model_into_study - Added
assess_study_validation_pair - Added
cancel_study_run - Added
compare_study_runs - Changed
create_model5 fields changed- added
Input schema / properties / channels / items / descriptionAdded value: +"Media channel binding. Exact activity-column keys identify channels in result/optimizer calls." - added
Input schema / properties / channels / items / propertiesAdded value: +{ + "activity_column": { + "type": "string" + }, + "name": { + "type": "string" + }, + "spend_column": { + "type": "string" + } +} - added
Input schema / properties / channels / items / requiredAdded value: +[ + "name", + "activity_column", + "spend_column" +] - added
Input schema / properties / control_priorsAdded value: +{ + "anyOf": [ + { + "items": { + "additionalProperties": true, + "description": "Optional override for a selected control column. Values and priors use transformed units; discover backend support first.", + "properties": { + "control": { + "type": "string" + }, + "distribution": { + "examples": [ + "normal", + "inversegamma", + "truncatednormal", + "halfnormal" + ], + "type": "string" + }, + "lower": { + "type": "number" + }, + "mean": { + "type": "number" + }, + "sd": { + "type": "number" + }, + "transform": { + "description": "N unchanged; DM divide by variable mean; STA scale by sample SD without centering; DDM divide by variable mean within hierarchy; LOG log(x/mean(x)). Backend validates applicability.", + "examples": [ + "N", + "DM", + "STA", + "DDM", + "LOG" + ], + "type": "string" + }, + "upper": { + "type": "number" + } + }, + "required": [ + "control" + ], + "type": "object" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Control Priors" +} - changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "create_modelDictOutput", + "type": "object" +}
- Changed
create_project1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "create_projectDictOutput", + "type": "object" +}
- Added
create_quality_policy - Added
create_recipe_draft - Added
create_study - Added
create_study_recipe - Changed
create_var_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "create_var_modelDictOutput", + "type": "object" +}
- Added
declare_study_holdout_use - Changed
delete_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "delete_modelDictOutput", + "type": "object" +}
- Added
evaluate_study_run - Added
get_backend_capabilities - Changed
get_contribution_groups1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_contribution_groupsDictOutput", + "type": "object" +}
- Changed
get_data_schema1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_data_schemaDictOutput", + "type": "object" +}
- Changed
get_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_modelDictOutput", + "type": "object" +}
- Changed
get_model_results2 fields changed- added
Input schema / properties / max_response_bytesAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Max Response Bytes" +} - changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_model_resultsDictOutput", + "type": "object" +}
- Changed
get_model_status1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_model_statusDictOutput", + "type": "object" +}
- Changed
get_optimizer_results1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_optimizer_resultsDictOutput", + "type": "object" +}
- Added
get_recipe_draft - Added
get_recipe_draft_template - Added
get_recipe_revision - Added
get_recipe_revision_authoring - Changed
get_scenario_results1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_scenario_resultsDictOutput", + "type": "object" +}
- Changed
get_scenario_template1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_scenario_templateDictOutput", + "type": "object" +}
- Added
get_study - Added
get_study_champion - Added
get_study_prediction_access - Added
get_study_run - Added
get_study_validation_resolutions - Changed
get_upload1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "get_uploadDictOutput", + "type": "object" +}
- Added
launch_study_run - Changed
link_var_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "link_var_modelDictOutput", + "type": "object" +}
- Changed
list_models1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "list_modelsDictOutput", + "type": "object" +}
- Changed
list_projects1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "list_projectsDictOutput", + "type": "object" +}
- Added
list_quality_policies - Added
list_recipe_drafts - Changed
list_runs1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "list_runsDictOutput", + "type": "object" +}
- Added
list_studies - Added
list_study_decisions - Added
list_study_evaluations - Added
list_study_recipes - Added
list_study_runs - Changed
list_uploads1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "list_uploadsDictOutput", + "type": "object" +}
- Added
publish_recipe_draft - Added
recommend_study_run - Changed
rename_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "rename_modelDictOutput", + "type": "object" +}
- Changed
rename_project1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "rename_projectDictOutput", + "type": "object" +}
- Added
revise_study_recipe - Changed
run_optimizer1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "run_optimizerDictOutput", + "type": "object" +}
- Changed
run_scenario1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "run_scenarioDictOutput", + "type": "object" +}
- Changed
save_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "save_modelDictOutput", + "type": "object" +}
- Changed
set_contribution_groups1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "set_contribution_groupsDictOutput", + "type": "object" +}
- Changed
set_run_pinned1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "set_run_pinnedDictOutput", + "type": "object" +}
- Changed
unlink_var_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "unlink_var_modelDictOutput", + "type": "object" +}
- Changed
unsave_model1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "unsave_modelDictOutput", + "type": "object" +}
- Added
update_recipe_draft - Changed
update_run1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "update_runDictOutput", + "type": "object" +}
- Added
update_study - Changed
upload_data1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "title": "upload_dataDictOutput", + "type": "object" +}
- Added
validate_study_recipe
29 tool updates
v0.3.2- First observed
create_model - First observed
create_project - First observed
create_var_model - First observed
delete_model - First observed
get_contribution_groups - First observed
get_data_schema - First observed
get_model - First observed
get_model_results - First observed
get_model_status - First observed
get_optimizer_results - First observed
get_scenario_results - First observed
get_scenario_template - First observed
get_upload - First observed
link_var_model - First observed
list_models - First observed
list_projects - First observed
list_runs - First observed
list_uploads - First observed
rename_model - First observed
rename_project - First observed
run_optimizer - First observed
run_scenario - First observed
save_model - First observed
set_contribution_groups - First observed
set_run_pinned - First observed
unlink_var_model - First observed
unsave_model - First observed
update_run - First observed
upload_data
TDQS
Scored across 94 tools
With 94 tools spanning overlapping MMM, study, recipe, revision, incrementality, campaign, and pipeline domains, many tools share nouns and require long descriptions to distinguish. Boundaries between recipe drafts, study recipes, revisions, and incrementality design/save/create operations are not self-evident to an agent without careful reading. Several descriptions include warnings about confusing related tools, which further signals latent ambiguity.
Tool names consistently use snake_case with a predictable verb_noun pattern: list_*, get_*, create_*, update_*, run_*, set_*, and delete_*. Domain prefixes such as study_, recipe_, campaign_, pipeline_, and incrementality_ are applied consistently. Although many names are long, the convention is stable throughout.
94 tools far exceeds the typical well-scoped range of 3-15 and lands in the extreme-mismatch band. Even for a broad MMM platform, this surface is so large that tool selection, documentation, and maintenance become unwieldy. The count alone indicates over-exposure rather than a tight, purpose-built tool set.
The surface is extensive, covering data upload/reporting, model creation and results, optimizers, scenarios, studies, recipes, quality policies, evaluations, champion tracking, campaigns, pipelines, and incrementality tests. There are some intentional gaps, such as deletion only for failed models and no API deletion for projects or unreferenced policies, but agents can generally work around these. Overall lifecycle coverage is strong, though not perfectly complete.
Maintenance
Related MCP Connectors
Conversational access to advertising performance data, creative analysis, and campaign insights
Conversational access to advertising performance data, creative analysis, and campaign insights
AI marketing agent for Google Ads, Meta, GA4, TikTok, LinkedIn, Shopify, HubSpot and more.
Ask live marketing data anything to get verified answers, client-ready reports, and next steps.
Related MCP Servers
- -licenseNot gradedqualityBmaintenanceConnects AI assistants to marketing mix models, enabling natural language data upload, performance modeling, budget optimization, and scenario testing.-
- AlicenseNot gradedqualityCmaintenanceEnables AI assistants to create, analyze, and optimize ad campaigns across Google Ads, Meta Ads, TikTok Ads, LinkedIn Ads, Amazon Ads, and ChatGPT Ads through natural language using 400+ tools.98MIT
- FlicenseNot gradedqualityCmaintenanceEnables marketing optimization tasks such as copywriting, campaign analysis, social media planning, audience segmentation, and KPI tracking through natural language.99 npm-

PaidSync MCP Serverofficial
AlicenseNot gradedqualityFmaintenanceConnects Google Ads, Meta Ads, and LinkedIn Ads to AI assistants, enabling natural language ad campaign management, reporting, and optimization across platforms.MIT