Predictive Performance of Bayesian Stacking in Multilevel Education Data
Mingya Huang, David Kaplan
Abstract
Open-access reader
Mingya Huang, David Kaplan
Abstract
Open-access reader
The issue of model uncertainty has been gaining interest in education and the social sciences community over the years, and the dominant methods for handling model uncertainty are based on Bayesian inference, and particularly, Bayesian model averaging. However, Bayesian model averaging assumes that the true data-generating model is within the candidate model space over which averaging is taking place. Unlike Bayesian model averaging, the method of Bayesian stacking can account for model uncertainty without assuming that a true model exists. An issue with Bayesian stacking, however, is that it is an optimization technique that uses predictor-independent model weights and is, therefore, not fully Bayesian. Bayesian hierarchical stacking, proposed by \citeA{yao2021bayesian}, further incorporates uncertainty by applying a hyperprior to the stacking weights. Considering the importance of multilevel models commonly applied in educational settings, this paper investigates via a simulation study and a real data example the predictive performance of original Bayesian stacking and Bayesian hierarchical stacking along with two other readily available weighting methods, pseudo-BMA and pseudo-BMA bootstrap (PBMA and PBMA+). Predictive performance is measured by the Kullback-Leibler divergence score. Although the differences in predictive performance among these four weighting methods in Bayesian stacking are small, we still find that Bayesian hierarchical stacking performs as well as conventional stacking, PBMA, and PBMA+ in settings where a true model is not assumed to exist.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The issue of model uncertainty has been gaining interest in education and the social sciences community over the years, and the dominant methods for handling model uncertainty are based on Bayesian inference, and particularly, Bayesian model averaging. However, Bayesian model averaging assumes that the true data-generating model is within the candidate model space over which averaging is taking place. Unlike Bayesian model averaging, the method of Bayesian stacking can account for model uncertainty without assuming that a true model exists. An issue with Bayesian stacking, however, is that it is an optimization technique that uses predictor-independent model weights and is, therefore, not fully Bayesian. Bayesian hierarchical stacking, proposed by \citeA{yao2021bayesian}, further incorporates uncertainty by applying a hyperprior to the stacking weights. Considering the importance of multilevel models commonly applied in educational settings, this paper investigates via a simulation study and a real data example the predictive performance of original Bayesian stacking and Bayesian hierarchical stacking along with two other readily available weighting methods, pseudo-BMA and pseudo-BMA bootstrap (PBMA and PBMA+). Predictive performance is measured by the Kullback-Leibler divergence score. Although the differences in predictive performance among these four weighting methods in Bayesian stacking are small, we still find that Bayesian hierarchical stacking performs as well as conventional stacking, PBMA, and PBMA+ in settings where a true model is not assumed to exist.
Key concepts: Bayesian average, Bayesian hierarchical modeling, Bayesian probability, Bayesian inference, Stacking, Weighting, Variable-order Bayesian network, Bayesian statistics