direct_method
Estimate a target policy's expected reward by fitting an outcome regression model to logged actions, contexts, and rewards. Use when the outcome model is correctly specified and no unmeasured confounders exist.
Instructions
Direct outcome regression (plug-in Q-model) OPE. Assumptions: Plug-in outcome regression (Q-model) is correctly specified; No unmeasured confounding in the logged data. Pre-conditions: X (context), A (logged action), R (reward) are available to fit the outcome model. Failure modes: Model misspecification bias -- the Q-model extrapolates outside the logged action support -> Prefer the doubly-robust estimator, which is robust to Q-model misspecification. Alternatives: sp.doubly_robust, sp.ips, sp.snips. Typical minimum N: 500.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| A | Yes | A parameter (np.ndarray). | |
| R | Yes | R parameter (np.ndarray). | |
| X | Yes | Feature matrix or covariate DataFrame. | |
| alpha | No | Significance level for confidence intervals and tests. | |
| detail | No | Payload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip. | agent |
| as_handle | No | If true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running. | |
| data_path | No | Absolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://. | |
| n_actions | No | Number of actions. | |
| pi_target | Yes | pi_target parameter. | |
| result_id | No | Optional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns. | |
| data_columns | No | Optional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads. | |
| data_sample_n | No | Optional uniform random subsample size (seed=0, deterministic) — useful on huge panels. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||