rdhte
Estimate conditional average treatment effects in regression discontinuity designs, revealing how treatment impact varies with covariates at the cutoff.
Instructions
Estimate conditional average treatment effects (CATE) in RD designs. Assumptions: Conditional expectations of potential outcomes are continuous at the cutoff; Units cannot precisely manipulate the running variable around the cutoff (no sorting); For fuzzy designs: monotonicity of treatment take-up at the cutoff. Pre-conditions: A continuous running/forcing variable with a known cutoff that (sharply or fuzzily) assigns treatment; Enough observations in a neighbourhood of the cutoff to fit a local polynomial. Failure modes: Density of the running variable jumps at the cutoff (manipulation / sorting) -> Run a McCrary / density test (rdplotdensity); if manipulation is present the design is invalid near the cutoff; Estimate swings with the bandwidth -- results are not robust -> Report a bandwidth-sensitivity curve and use a data-driven MSE-optimal bandwidth. Alternatives: sp.rdrobust, sp.rdrandinf, sp.rdbwselect. Typical minimum N: 500.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| b | No | Bandwidth for bias correction. Defaults to h. | |
| c | No | RD cutoff value. | |
| h | No | Bandwidth for estimation. If None, MSE-optimal bandwidth is selected. | |
| p | No | Polynomial order for the running variable (1 = local linear). | |
| x | Yes | Running variable name. | |
| y | Yes | Outcome variable name. | |
| z | Yes | Covariate(s) for treatment effect heterogeneity. | |
| alpha | No | Significance level for confidence intervals. | |
| detail | No | Payload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip. | agent |
| kernel | No | Kernel function: 'triangular', 'uniform', or 'epanechnikov'. | triangular |
| n_eval | No | Number of evaluation points when eval_points is not provided. | |
| cluster | No | Cluster variable name for cluster-robust standard errors. | |
| bwselect | No | Bandwidth selection method: 'mserd' or 'msetwo'. | mserd |
| as_handle | No | If true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running. | |
| data_path | Yes | Absolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://. | |
| result_id | No | Optional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns. | |
| eval_points | No | Z values at which to evaluate CATE. Each row is a point in Z-space. If None, n_eval equally spaced quantiles (10th to 90th pctile) are used. | |
| data_columns | No | Optional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads. | |
| data_sample_n | No | Optional uniform random subsample size (seed=0, deterministic) — useful on huge panels. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||