Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

callaway_santanna

Read-only

Estimate staggered difference-in-differences effects with Callaway-Sant'Anna, delivering group-time average treatment effects robust to heterogeneous impacts where two-way fixed effects fail.

Instructions

Staggered DID (Callaway-Sant'Anna 2021): group-time ATT with doubly-robust, IPW, or regression-adjusted estimators. Robust to heterogeneous treatment effects where TWFE fails. Requires a cohort column g (first-treatment period; 0 = never-treated). Validation: certified evidence with scoped limitations. Known limitations: clustervars is not yet supported with bstrap=False; the analytic standard errors do not account for within-cluster dependence, so the multiplier bootstrap is required.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
gYesFirst-treatment-period cohort column (0 for never-treated).
iYesUnit ID column
tYesTime column
yYes
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathYesAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
estimatorNodr
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
control_groupNonevertreated
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the description's burden is reduced, but it adds value by disclosing specific limitations: clustervars not supported with bstrap=False, analytic standard errors ignoring cluster dependence, and the necessity of multiplier bootstrap. This goes beyond the annotations and helps the agent anticipate constraints. The vague 'Validation: certified evidence' line is a minor detractor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, compact sentence that front-loads the core method and purpose, then tacks on necessary limitations and requirements. It avoids redundancy with the schema for most parameters. The 'Validation: certified evidence' phrase is filler but not harmful. Overall, it's appropriately sized and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (12 parameters, DID estimation) and the presence of an output schema, the description covers the essential method, required inputs (g, i, t, y implied), and known limitations. It does not explain the detail parameter's role in payload depth, but that is in the schema. The description, combined with a rich schema and annotations, is adequate for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, which is below the 80% threshold, so the description should compensate for undocumented parameters. It does explain the g parameter's role (cohort column) and the estimator options (dr/ipw/reg), but it offers no additional meaning for other parameters like detail, as_handle, or control_group. The schema already covers most of g and estimator, so the description adds marginal value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a Callaway-Sant'Anna staggered DID estimator computing group-time ATTs with three estimator options (doubly-robust, IPW, regression-adjusted). It states the method's advantage over TWFE, which helps disambiguate from many DID siblings. However, it does not explicitly name any sibling (e.g., staggered_cs, stacked_did) to differentiate from, so it's not a perfect 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when TWFE fails or heterogeneous treatment effects are present ('Robust to heterogeneous treatment effects where TWFE fails'), but it does not explicitly state when not to use it or name alternative tools. The 'Requires a cohort column g' is a prerequisite but not a usage guideline. This falls under implied usage rather than explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools