Skip to main content
Glama
vikranthviki

Causal Decision Agent

by vikranthviki

parallel_trends_robustness

Read-only

Run a robustness pipeline for difference-in-differences findings: test parallel-trends assumptions, assess sensitivity to pre-trend violations, and determine if conclusions hold.

Instructions

Run the full parallel-trends robustness pipeline on a DiD result. Validation: validated evidence tier (known-truth, reference, external-parity, or Monte Carlo artifact).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
eNoRelative time whose effect the sensitivity analysis targets.
alphaNoSignificance level.
deltaNoHypothesised pre-trend violation passed to ``pretrends_power``.
detailNoPayload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip.agent
m_gridNoGrid of violation magnitudes. Default: the ``honest_did`` default, multiples of the standard error at ``e``.
resultYesA fitted DiD/event-study result carrying event-study estimates in ``result.model_info['event_study']`` (e.g. from ``sp.event_study``, ``sp.callaway_santanna``, ``sp.sun_abraham``). .. note:: If the result does not carry a pre-period covariance matrix in ``model_info['vcv_pre']``, the pre-trend test and power fall back to assuming the pre-period coefficients are mutually independent and warn loudly. ``sp.event_study`` computes the full cluster-robust covariance; pass ``expose_pre_vcov=True`` to it to have this pipeline use the correct covariance instead of the diagonal fallback.
familiesNoRestriction families. ``"SD"`` maps to ``honest_did``'s ``method='smoothness'`` (bounded second differences); ``"RM"`` maps to ``method='relative_magnitude'``.
as_handleNoIf true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running.
data_pathNoAbsolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://.
result_idNoOptional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns.
data_columnsNoOptional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads.
data_sample_nNoOptional uniform random subsample size (seed=0, deterministic) — useful on huge panels.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=true and openWorldHint=false, covering the safety profile (no mutation, closed-world). The description adds a cryptic validation tier note ('known-truth, reference, external-parity, or Monte Carlo artifact') which is informative about evidence status but not about runtime behavior, side effects, or output characteristics. It does not contradict annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no fluff, front-loading the main purpose. However, the second sentence about validation tier is somewhat opaque and may not earn its place for an agent deciding how to call the tool. Still, it is compact and clear enough.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 12-parameter tool with a rich schema (100% parameter coverage), readOnlyHint, and an output schema, the description need not explain return values or parameter details. It clearly identifies the domain ('DiD result') and the pipeline nature. It lacks explicit notes on common pitfalls or when to consider lighter-weight siblings, but the schema compensates for most needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with detailed parameter descriptions (e.g., 'result' explains the required event_study and vcv_pre structure, 'detail' explains payload depth). The description itself does not add parameter-level meaning, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Run') and resource ('full parallel-trends robustness pipeline') on a clear target ('a DiD result'). It is not merely a restatement of the name, and the 'full pipeline' wording hints at comprehensiveness, though it does not explicitly name or contrast with sibling tools like parallel_trends_plot or pretrends_test.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied rather than explicit. 'Full ... pipeline' suggests this is the comprehensive robustness check, but the description does not state when to choose this over alternative tools, nor does it give exclusions or prerequisites beyond what is buried in parameter docs. No direct when-to-use/when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools