rd_flex
Estimates regression discontinuity treatment effects using cross-fit machine-learning residualisation of covariates, reducing outcome variance at the cutoff for more precise causal inference.
Instructions
RD with flexible covariate adjustment via cross-fit ML residualisation (Noack-Olma-Rothe 2025). Reduces variance of tau at the cutoff by subtracting an ML estimate of E[Y|W] before running rdrobust; consistent under free-of-cutoff continuity of eta, asymptotically efficient when eta converges to E[Y|X=c, W]. Assumptions: Continuity-based RD identification at the cutoff: potential outcomes are continuous in the running variable except for the treatment jump; Cross-fit ML residualisation of the outcome (and treatment, when fuzzy) on covariates only removes outcome variance and does not bias the cutoff estimate, requiring honest K-fold cross-fitting; Covariates predict the outcome well enough to shorten CIs relative to plain rdrobust; covariates are pre-determined (not affected by treatment). Pre-conditions: data has continuous running variable with adequate mass on both sides of the cutoff; Covariates list valid pre-treatment columns (or is empty/None to fall back to rdrobust); n_folds>=2 for genuine cross-fitting. Failure modes: Sparse data near the cutoff makes the local fit and learner unstable -> Widen the bandwidth via bwselect or collect more mass around the cutoff; Covariates...
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| W | No | Covariates used by the flexible adjustment | |
| c | No | c parameter (float). | |
| x | Yes | Primary running variable, regressor, or feature input for this estimator. | |
| y | Yes | Outcome variable column name or outcome array. | |
| alpha | No | Significance level for confidence intervals and tests. | |
| fuzzy | No | fuzzy parameter (str). | |
| detail | No | Payload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip. | agent |
| kernel | No | Kernel function used for weighting or smoothing. | triangular |
| learner | No | Built-in learner | boost |
| n_folds | No | Cross-fit folds (1 disables CV) | |
| as_handle | No | If true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running. | |
| data_path | Yes | Absolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://. | |
| result_id | No | Optional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns. | |
| data_columns | No | Optional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads. | |
| data_sample_n | No | Optional uniform random subsample size (seed=0, deterministic) — useful on huge panels. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||