xtabond
Run Arellano-Bond or Blundell-Bond GMM on panel data to correct for endogeneity and dynamic effects, including tests for instrument validity and serial correlation.
Instructions
Arellano-Bond / Blundell-Bond GMM for dynamic panels (standalone). Validation: certified parity evidence. Assumptions: No second-order serial correlation in differenced errors; Internal instruments are valid and not too numerous. Pre-conditions: Panel data include unit, time, outcome, and lagged dependent variable structure; Number of time periods is moderate relative to units. Failure modes: Instrument proliferation or AR(2) test rejects -> Collapse instruments, reduce lag depth, or compare with fixed-effects estimates. Alternatives: sp.panel, sp.feols. Typical minimum N: 100.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| h | No | xtabond2 h(): one-step error covariance; 3 is xtabond2's default, 2 zeroes the system cross quadrants (Stata xtdpdsys), 1 the identity. | |
| x | No | Exogenous regressors | |
| y | Yes | Dependent variable | |
| id | No | Unit identifier | id |
| lags | No | lags parameter (int). | |
| time | No | Time column | time |
| steps | No | Number of GMM steps, or 'iterated' / 'cue' | |
| detail | No | Payload depth: 'minimal' (~150 tokens) for sub-step calls where only the point estimate is needed; 'standard' (~1K tokens) for diagnostics + coefficient table; 'agent' (~2K tokens, default) adds violations / next_steps / suggested_functions so the LLM can plan its next call without another round-trip. | agent |
| method | No | difference (Arellano-Bond), system (Blundell-Bond) or ah (Anderson-Hsiao IV) | difference |
| cluster | No | Cluster SEs on a coarser unit than the panel id (must be constant within unit) | |
| twostep | No | twostep parameter (bool). | |
| collapse | No | Collapse instruments (Roodman 2009) to curb proliferation | |
| as_handle | No | If true, cache the fitted result on the server and return result_id + result_uri alongside the JSON payload so a subsequent tools/call can chain without re-running. | |
| data_path | Yes | Absolute path or URL to a data file. Supported: .csv / .tsv / .txt (delimited), .parquet / .pq, .feather / .arrow, .xlsx / .xls, .dta (Stata), .json / .jsonl. Schemes: file://, s3://, gs://, https://. | |
| result_id | No | Optional handle to a previously-fitted result (returned by an earlier call when as_handle=true). Tools that operate on a fitted object accept this in place of re-supplying data_path + columns. | |
| endogenous | No | Endogenous regressors (own lags 2+ as instruments) | |
| orthogonal | No | Forward orthogonal deviations instead of first differences | |
| iv_equation | No | System GMM: equation(s) the exogenous regressors instrument; None means 'both' (xtabond2), 'diff' is Stata xtdpdsys. | |
| data_columns | No | Optional column projection. Parquet/Feather/Stata loaders honour this for fast partial reads. | |
| time_dummies | No | Add period dummies as regressors and instruments | |
| ah_instrument | No | Anderson-Hsiao instrument for method='ah' | levels |
| data_sample_n | No | Optional uniform random subsample size (seed=0, deterministic) — useful on huge panels. | |
| predetermined | No | Predetermined regressors (own lags 1+ as instruments) |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||