select_pci_proxies
Selects and ranks candidate proxy variables for proximal causal inference, identifying valid negative controls to adjust for unobserved confounding and improve causal estimates.
Instructions
Score and rank candidate proxies for PCI. Assumptions: The proxies are valid negative controls (relevant to the confounder, excluded from the causal channel); A bridge function exists (completeness conditions hold). Pre-conditions: Treatment-inducing and outcome-inducing proxy variables (negative controls) for the unobserved confounder. Failure modes: Proxies are weak or invalid -- the bridge function is poorly identified -> Test proxy relevance, select stronger proxies, or fall back to sensitivity analysis. Alternatives: sp.select_pci_proxies, sp.dml. Typical minimum N: 500.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| y | Yes | Outcome variable column name or outcome array. | |
| top_k | No | Number of top candidates to recommend per side. | |
| treat | Yes | Treatment indicator or first-treatment-period column. | |
| detail | No | Payload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip. | agent |
| as_handle | No | If true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running. | |
| data_path | Yes | Absolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://. | |
| result_id | No | Optional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns. | |
| candidates | Yes | All variables that could plausibly serve as proxies. | |
| covariates | No | Covariate matrix, DataFrame, or column names. | |
| data_columns | No | Optional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads. | |
| data_sample_n | No | Optional uniform random subsample size (seed=0, deterministic) — useful on huge panels. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||