2011ORCA Online Research @Cardiff (Cardiff University)Open access

Propensity score matching with missing covariatesvia iterated, sequential multiple imputation

Robin Mitra, Jerome P. Reiter

Open full text 1 citations

Abstract

In many observational studies, analysts estimate causal effects using propensity score matching. Estimation of propensity scores is complicated when covariate values intended for collection are in fact missing. To handle the missing data, one approach is to use multiple imputation to create completed datasets, and compute propensity scores from these datasets. However, inaccurate imputation models can result in ineffective matching, thereby limiting reductions in bias. We propose a multiple imputation approach based on chained equations in which the researcher gradually reduces the set of control units used to estimate the imputation models. This approach can reduce the influence of control records far from the treated units' region of the covariate space on the estimation of parameters in the imputation model, which can result in more plausible imputations and better balance in the true covariate distributions. This approach can be conveniently implemented with standard multiple imputation software for missing data. Using simulations, we find that the approach can improve estimation when imputation models are mis-specified; however, it can be ineffective when imputation models are correctly specified. This suggests using the approach as part of sensitivity analysis in causal inference. We apply the approach to an observational study of the effect of breast-feeding on the child?s educational outcomes later in life.

Open-access reader

About this research paper

What this paper is about

In many observational studies, analysts estimate causal effects using propensity score matching. Estimation of propensity scores is complicated when covariate values intended for collection are in fact missing. To handle the missing data, one approach is to use multiple imputation to create completed datasets, and compute propensity scores from these datasets. However, inaccurate imputation models can result in ineffective matching, thereby limiting reductions in bias. We propose a multiple imputation approach based on chained equations in which the researcher gradually reduces the set of control units used to estimate the imputation models. This approach can reduce the influence of control records far from the treated units' region of the covariate space on the estimation of parameters in the imputation model, which can result in more plausible imputations and better balance in the true covariate distributions. This approach can be conveniently implemented with standard multiple imputation software for missing data. Using simulations, we find that the approach can improve estimation when imputation models are mis-specified; however, it can be ineffective when imputation models are correctly specified. This suggests using the approach as part of sensitivity analysis in causal inference. We apply the approach to an observational study of the effect of breast-feeding on the child?s educational outcomes later in life.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

In many observational studies, analysts estimate causal effects using propensity score matching. Estimation of propensity scores is complicated when covariate values intended for collection are in fact missing. To handle the missing data, one approach is to use multiple imputation to create completed datasets, and compute propensity scores from these datasets. However, inaccurate imputation models can result in ineffective matching, thereby limiting reductions in bias. We propose a multiple imputation approach based on chained equations in which the researcher gradually reduces the set of control units used to estimate the imputation models. This approach can reduce the influence of control records far from the treated units' region of the covariate space on the estimation of parameters in the imputation model, which can result in more plausible imputations and better balance in the true covariate distributions. This approach can be conveniently implemented with standard multiple imputation software for missing data. Using simulations, we find that the approach can improve estimation when imputation models are mis-specified; however, it can be ineffective when imputation models are correctly specified. This suggests using the approach as part of sensitivity analysis in causal inference. We apply the approach to an observational study of the effect of breast-feeding on the child?s educational outcomes later in life.

Key concepts: Imputation (statistics), Covariate, Propensity score matching, Missing data, Computer science, Observational study, Statistics, Causal inference

Related papers

Back to paper searchBrowse research topicsOriginal source
Propensity score matching with missing covariatesvia iterated, sequential multiple imputation — Research Paper | ScholarLens