Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

sdid

Read-only

Estimates treatment effects from panel data using synthetic difference-in-differences, validating parallel trends via control weighting and producing placebo/bootstrap confidence intervals.

Instructions

Synthetic Difference-in-Differences estimator (and SC / DID variants). Validation: certified parity evidence. Do NOT use when: there is no clean pre-treatment block for every unit -- the unit and time weights are fit on the pre-period grid. Cost: Placebo / bootstrap standard errors refit the full weighting problem n_reps times; the point estimate alone is cheap. Lower n_reps while iterating. Assumptions: Parallel trends in the absence of treatment, after the synthetic/DiD weighting; No anticipation and no interference between units (SUTVA); The control pool's outcome process is stable around the intervention. Pre-conditions: Panel with treated and control units and a clear treatment date; Pre-treatment periods available to assess comparability of trends. Failure modes: Weighted pre-treatment trends still diverge between treated and synthetic control -> Inspect the unit/time weights and pre-trend fit; consider event-study DiD with honest bounds. Alternatives: sp.synth, sp.augsynth, sp.callaway_santanna, sp.gardner_did. Typical minimum N: 15.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
yNoOutcome variable column name or outcome array.
seedNoRandom seed for reproducibility.
timeNoTime period column.
unitNoUnit identifier column.
alphaNoSignificance level for confidence intervals.
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
methodNo* ``'sdid'`` -- Synthetic DID (unit + time weights) * ``'sc'`` -- Synthetic Control (unit weights only) * ``'did'`` -- DID (uniform weights)sdid
n_repsNoReplications for placebo / bootstrap SE.
backendNo``'native'`` uses StatsPAI's Python implementation. ``'synthdid'``/``'r'`` delegates to the R ``synthdid`` package through ``Rscript`` and returns the reference package's point estimate and ``synthdid_se`` standard error. The R backend is mainly for exact cross-language parity claims; the dependency- light native implementation remains the default.native
outcomeNoOutcome variable column. Alias ``y=`` accepted for R-style calls.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
se_methodNoStandard-error method (see Notes).placebo
covariatesNoReserved for future covariate-adjusted extensions.
treat_timeNotreat_time parameter.
treat_unitNotreat_unit parameter.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
treated_unitNoTreated unit(s). Alias ``treat_unit=`` accepted.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.
treatment_timeNoFirst treatment period (inclusive). Alias ``treat_time=`` accepted.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/openWorldHint annotations, it discloses cost behavior (placebo/bootstrap SE refit the full weighting problem n_reps times; point estimate alone is cheap), assumptions (parallel trends, SUTVA, stable control process), and failure modes (divergent pre-trends -> inspect weights, consider event-study DiD). No contradiction with annotations; readOnlyHint is consistent with an estimator that only reads data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Each labeled section (validation, do-not-use, cost, assumptions, pre-conditions, failure modes, alternatives, min N) carries actionable information, and the first sentence states the purpose. Though dense, the length is justified for a complex estimator with 21 parameters and many sibling tools.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's causal-inference complexity, the description covers exclusions, cost, assumptions, preconditions, failure modes, alternatives, and sample-size guidance. The output schema exists to document return values, and the input schema covers all parameters, so nothing essential is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema already covers all 21 parameters at 100%, so the baseline is 3. The description adds practical meaning to n_reps by explaining the computational cost and advising 'Lower n_reps while iterating', and clarifies method variants (sdid/sc/did) and backend parity. This goes beyond the schema's mechanical descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with 'Synthetic Difference-in-Differences estimator (and SC / DID variants)', naming the exact method and scope. This distinguishes it from sibling estimators like synth, augsynth, callaway_santanna, and gardner_did, which are also listed as alternatives. Clear verb/resource with no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states 'Do NOT use when: there is no clean pre-treatment block for every unit' and gives pre-conditions such as a panel with treated/control units, a clear treatment date, and pre-treatment periods. It also names four concrete alternatives (sp.synth, sp.augsynth, sp.callaway_santanna, sp.gardner_did) so an agent can route correctly. This is explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools