Propensity Score Analysis: Recent Debate and Discussion
Shenyang Guo, Mark W. Fräser, Qi Chen
Abstract
Open-access reader
Shenyang Guo, Mark W. Fräser, Qi Chen
Abstract
Open-access reader
Propensity score analysis is often used to address selection bias in program evaluation with observational data. However, a recent study suggested that propensity score matching may accomplish the opposite of its intended goal—increasing imbalance, inefficiency, model dependence, and bias. We assess common propensity score models and offer our responses to these criticisms. We used Monte Carlo methods to simulate two alternative settings of data creation—selection on observed variables versus selection on unobserved variables—and compared eight propensity score models on bias reduction and sample-size retention. Based on the simulations, no single propensity score method reduced bias across all scenarios. Optimal results depend on the fit between assumptions embedded in the analytic model and the process of data generation. Methodologic knowledge of model assumptions and substantive knowledge of causal mechanisms, including sources of selection bias, should inform the choice of analytic strategies involving propensity scores.
OpenAlex reports 155 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Propensity score analysis is often used to address selection bias in program evaluation with observational data. However, a recent study suggested that propensity score matching may accomplish the opposite of its intended goal—increasing imbalance, inefficiency, model dependence, and bias. We assess common propensity score models and offer our responses to these criticisms. We used Monte Carlo methods to simulate two alternative settings of data creation—selection on observed variables versus selection on unobserved variables—and compared eight propensity score models on bias reduction and sample-size retention. Based on the simulations, no single propensity score method reduced bias across all scenarios. Optimal results depend on the fit between assumptions embedded in the analytic model and the process of data generation. Methodologic knowledge of model assumptions and substantive knowledge of causal mechanisms, including sources of selection bias, should inform the choice of analytic strategies involving propensity scores.
Key concepts: Propensity score matching, Selection bias, Observational study, Selection (genetic algorithm), Econometrics, Matching (statistics), Inefficiency, Sample size determination