Integrating Double/Debiased Machine Learning into Doubly Robust Estimators for Causal Inference
Saved in:
| Title: | Integrating Double/Debiased Machine Learning into Doubly Robust Estimators for Causal Inference |
|---|---|
| Language: | English |
| Authors: | Muwon Kwon, Peter M. Steiner, Society for Research on Educational Effectiveness (SREE) |
| Source: | Society for Research on Educational Effectiveness. 2025. |
| Availability: | Society for Research on Educational Effectiveness. 2040 Sheridan Road, Evanston, IL 60208. Tel: 202-495-0920; e-mail: contact@sree.org; Web site: https://www.sree.org/ |
| Peer Reviewed: | Y |
| Publication Date: | 2025 |
| Document Type: | Reports - Research |
| Descriptors: | Artificial Intelligence, Statistical Analysis, Computation, Inferences, Robustness (Statistics), Statistical Bias |
| Abstract: | Background: Double/debiased machine learning (DML) methods have been proposed to overcome the regularization bias from the naive approach of ML methods (Chernozhukov et al., 2018). DML methods use a partialling-out approach which removes the effect of confounders from both the treatment and outcome and then regresses the residualized outcome on the residualized treatment (Frisch & Waugh, 1933; Lovell, 1963, 2008). By using the partialling-out approach, DML estimators attain Neyman orthogonality under which the target parameter (e.g., causal effect) is locally insensitive to perturbations of the nuisance parameters (Chernozhukov et al., 2018; Neyman, 1979). Then Neyman orthogonality of the DML estimators ensures [square root]n-consistency of the estimators where n is the sample size; that is, the DML estimators overcome for sufficiently large samples the regularization bias. However, even though the entire covariates in the model satisfy unconfoundedness that is needed for causal inference, the subset of the covariates selected by the regularized ML methods may not meet unconfoundedness (i.e., incorrect covariate selection). For example, with the entire set of covariates that satisfy unconfoundedness, ML methods using regularization may select a collider variable without selecting variables that block the collider path, which results in collider bias; select an instrumental variable while failing to include all confounding variables, which leads to amplified omitted variable bias; or, include only a subset of confounding variables, leaving residual confounding bias unaddressed. To overcome the misspecification regarding covariate selection (i.e., the violation of unconfoundedness due to incorrect covariate selection) we integrate DML methods into doubly robust (DR) estimators as a potential solution. DR estimators are robust to the misspecification of either the propensity score (PS) model or the outcome model (Glynn & Quinn, 2010; Tsiatis, 2006). That is, DR estimators provide two opportunities to remove confounding bias, either by correctly specifying the PS model or the outcome model with regard to a covariate set that meets the unconfoundedness assumption (Robins & Rotnitzky, 2001). Thus, combining DML methods with DR estimators yield a consistent estimator even for high-dimensional datasets as long as the estimated PS model or the outcome model approximates the respective true model sufficiently close. Although whether unconfoundedness holds solely depends on the true DGP, previous studies on DML methods have generally assumed unconfoundedness without considering the causal structure of DGPs (Chernozhukov et al., 2024; Dukes & Vansteelandt, 2021; Knaus, 2022). As a result, little attention has been paid to settings where only one or multiple subsets of the entire covariate set meet unconfoundedness. Purpose of Study: This simulation study has two primary objectives. First, the study demonstrates that DML estimators can be biased with the DGP where unconfoundedness is met by the entire covariate set but violated by some subsets of covariates (imperfect covariate selection). Second, the study examines whether incorporating DML methods into DR estimators can mitigate this bias. The two objectives indicate that this simulation study is only concerned with covariate selection issues. Setting & Theoretical Derivations: This study considers a commonly used DML method that incorporates covariate selection: the double Lasso (Belloni et al., 2011; Chernozhukovet al., 2018; Tibshirani, 1996). For the DR estimator, the inverse probability weighted regression (IPWR) estimator is considered, as it is one of the most basic and frequently used DR estimators (Imbens & Wooldridge, 2009; Wooldridge, 2007). The IPWR estimator uses the inverse probability weight (IPW) to fit weighted regression models for treatment and control groups. Since the integration of DML methods into DR estimators has not been explored in previous research, they are combined in the most intuitive manner by applying the inverse probability weight (IPW) directly to the outcome regression, which is accommodated by the IPWR estimator. The demonstration of the potential violation of unconfoundedness due to incorrect covariate selection relies on the DGP shown by the causal graph in Figure 1 (Steiner & Lyu, 2024). Here, Z and Y are the binary treatment indicator and outcome, respectively. The parameter of interest is [tau], the treatment effect of Z on Y. Covariates U and V as well as the set of covariates X are independent exogenous variables, where X is a p-dimensional vector. M is a collider variable. In this simulation, high-dimensional settings are considered only in terms of the number of confounders p, as it plays a crucial role in determining whether unconfoundedness holds (Chernozhukov et al., 2018, 2024). In the causal graph in Figure 1, let S denote the set of all variables except for Z and Y: S = {X, U, V, M}. The set S meets the unconfoundedness assumption. Then, the double Lasso is integrated into the doubly robust IPWR estimator as follows:[equations omitted]. Although the covariate set S included in both the PS model and the outcome model meets unconfoundedness, the two selected covariate sets S[subscript X] and S[subscript Y] may not, due to incorrect covariate selection. The description above indicates that for the regression of the residualized Y on the residualized Z, which is the DML approach, the inverse probability weight (IPW) is applied. Evaluation and Comparison of Estimators: Using simulated data, we compare the performance of the proposed integrated approach with that of the double Lasso without integrating the IPWR estimator. The results of the simulation (Figure 2) indicate that when the double Lasso is integrated into the IPWR estimator, the bias is reduced. Although, in some conditions, the bias adjustment is larger than necessary, leading to negative bias, this bias is not substantial compared to the bias of the estimator without using the IPWR estimator. Conclusions: Incorporating DML into the DR estimators helps reduce the bias arising from failing to control for the entire bias of the model by the DML method alone (due to imperfect covariate selection). Although, the bias adjustment can be larger than needed resulting in negative bias, it is worth integrating DML into the DR estimators because the absolute value of the adjusted bias is not substantial compared to the double Lasso alone. Since one can never know whether approximate sparsity is satisfied in practice, taking the proposed integrated approach protects to some extent against remaining confounding bias and, thus, should be preferred to the double Lasso. |
| Abstractor: | As Provided |
| Entry Date: | 2026 |
| Access URL: | https://www.sree.org/2025-conference |
| Accession Number: | ED677778 |
| Database: | ERIC |
| Abstract: | Background: Double/debiased machine learning (DML) methods have been proposed to overcome the regularization bias from the naive approach of ML methods (Chernozhukov et al., 2018). DML methods use a partialling-out approach which removes the effect of confounders from both the treatment and outcome and then regresses the residualized outcome on the residualized treatment (Frisch & Waugh, 1933; Lovell, 1963, 2008). By using the partialling-out approach, DML estimators attain Neyman orthogonality under which the target parameter (e.g., causal effect) is locally insensitive to perturbations of the nuisance parameters (Chernozhukov et al., 2018; Neyman, 1979). Then Neyman orthogonality of the DML estimators ensures [square root]n-consistency of the estimators where n is the sample size; that is, the DML estimators overcome for sufficiently large samples the regularization bias. However, even though the entire covariates in the model satisfy unconfoundedness that is needed for causal inference, the subset of the covariates selected by the regularized ML methods may not meet unconfoundedness (i.e., incorrect covariate selection). For example, with the entire set of covariates that satisfy unconfoundedness, ML methods using regularization may select a collider variable without selecting variables that block the collider path, which results in collider bias; select an instrumental variable while failing to include all confounding variables, which leads to amplified omitted variable bias; or, include only a subset of confounding variables, leaving residual confounding bias unaddressed. To overcome the misspecification regarding covariate selection (i.e., the violation of unconfoundedness due to incorrect covariate selection) we integrate DML methods into doubly robust (DR) estimators as a potential solution. DR estimators are robust to the misspecification of either the propensity score (PS) model or the outcome model (Glynn & Quinn, 2010; Tsiatis, 2006). That is, DR estimators provide two opportunities to remove confounding bias, either by correctly specifying the PS model or the outcome model with regard to a covariate set that meets the unconfoundedness assumption (Robins & Rotnitzky, 2001). Thus, combining DML methods with DR estimators yield a consistent estimator even for high-dimensional datasets as long as the estimated PS model or the outcome model approximates the respective true model sufficiently close. Although whether unconfoundedness holds solely depends on the true DGP, previous studies on DML methods have generally assumed unconfoundedness without considering the causal structure of DGPs (Chernozhukov et al., 2024; Dukes & Vansteelandt, 2021; Knaus, 2022). As a result, little attention has been paid to settings where only one or multiple subsets of the entire covariate set meet unconfoundedness. Purpose of Study: This simulation study has two primary objectives. First, the study demonstrates that DML estimators can be biased with the DGP where unconfoundedness is met by the entire covariate set but violated by some subsets of covariates (imperfect covariate selection). Second, the study examines whether incorporating DML methods into DR estimators can mitigate this bias. The two objectives indicate that this simulation study is only concerned with covariate selection issues. Setting & Theoretical Derivations: This study considers a commonly used DML method that incorporates covariate selection: the double Lasso (Belloni et al., 2011; Chernozhukovet al., 2018; Tibshirani, 1996). For the DR estimator, the inverse probability weighted regression (IPWR) estimator is considered, as it is one of the most basic and frequently used DR estimators (Imbens & Wooldridge, 2009; Wooldridge, 2007). The IPWR estimator uses the inverse probability weight (IPW) to fit weighted regression models for treatment and control groups. Since the integration of DML methods into DR estimators has not been explored in previous research, they are combined in the most intuitive manner by applying the inverse probability weight (IPW) directly to the outcome regression, which is accommodated by the IPWR estimator. The demonstration of the potential violation of unconfoundedness due to incorrect covariate selection relies on the DGP shown by the causal graph in Figure 1 (Steiner & Lyu, 2024). Here, Z and Y are the binary treatment indicator and outcome, respectively. The parameter of interest is [tau], the treatment effect of Z on Y. Covariates U and V as well as the set of covariates X are independent exogenous variables, where X is a p-dimensional vector. M is a collider variable. In this simulation, high-dimensional settings are considered only in terms of the number of confounders p, as it plays a crucial role in determining whether unconfoundedness holds (Chernozhukov et al., 2018, 2024). In the causal graph in Figure 1, let S denote the set of all variables except for Z and Y: S = {X, U, V, M}. The set S meets the unconfoundedness assumption. Then, the double Lasso is integrated into the doubly robust IPWR estimator as follows:[equations omitted]. Although the covariate set S included in both the PS model and the outcome model meets unconfoundedness, the two selected covariate sets S[subscript X] and S[subscript Y] may not, due to incorrect covariate selection. The description above indicates that for the regression of the residualized Y on the residualized Z, which is the DML approach, the inverse probability weight (IPW) is applied. Evaluation and Comparison of Estimators: Using simulated data, we compare the performance of the proposed integrated approach with that of the double Lasso without integrating the IPWR estimator. The results of the simulation (Figure 2) indicate that when the double Lasso is integrated into the IPWR estimator, the bias is reduced. Although, in some conditions, the bias adjustment is larger than necessary, leading to negative bias, this bias is not substantial compared to the bias of the estimator without using the IPWR estimator. Conclusions: Incorporating DML into the DR estimators helps reduce the bias arising from failing to control for the entire bias of the model by the DML method alone (due to imperfect covariate selection). Although, the bias adjustment can be larger than needed resulting in negative bias, it is worth integrating DML into the DR estimators because the absolute value of the adjusted bias is not substantial compared to the double Lasso alone. Since one can never know whether approximate sparsity is satisfied in practice, taking the proposed integrated approach protects to some extent against remaining confounding bias and, thus, should be preferred to the double Lasso. |
|---|