Skip to main content
eScholarship
Open Access Publications from the University of California

UC Merced

UC Merced Electronic Theses and Dissertations bannerUC Merced

Advanced Methods for Implementation of Bayesian Growth Mixture Modeling

Abstract

Growth mixture models (GMMs) constitute a versatile statistical framework for capturing population heterogeneity in longitudinal data by identifying latent classes of growth trajectories. Since their introduction in 1999, the models have been extended and disseminated through sustained methodological advancements, yet uncharted gaps remain in their implementation. This dissertation develops and evaluates advanced Bayesian methods to address two such gaps, predictor selection and missing data handling, across two studies. In Study 1, I focused on predictor selection in conditional GMMs. Beyond the usefulness of GMMs in describing heterogeneous growth patterns, predicting these patterns offers an enhanced understanding of substantive phenomena. However, the question of how best to select important predictors has remained unaddressed in the GMM literature. As a principled approach to this issue, I proposed Bayesian variable selection via shrinkage priors and evaluated seven priors (ridge, lasso, hyperlasso, elastic net, horseshoe, regularized horseshoe, and spike-and-slab) through a Monte Carlo simulation. Results showed that the horseshoe, regularized horseshoe, and spike-and-slab offered the best balance between detecting true nonzero predictors, shrinking small effects, and controlling false selections. In addition, the regularized horseshoe, spike-and-slab, and horseshoe achieved superior predictive accuracy, though the regularized horseshoe and spike-and-slab came at the cost of lower convergence rates. The ridge prior showed limited effectiveness in selecting important predictors and high false selection rates. Practical recommendations are provided for choosing shrinkage priors, with caveats and other implementation considerations. In Study 2, I addressed missing data in GMMs under attrition scenarios common in longitudinal research. Existing missing data techniques either discard incomplete cases, require the number of latent classes to be pre-specified as part of the missing data treatment, or ignore the population heterogeneity. To overcome these limitations, I developed a Bayesian nonparametric multiple imputation approach using the Chinese restaurant process (MI-CRP). This new approach retains existing values, imputes missing values, and accounts for population heterogeneity without requiring the number of classes to be predetermined. A Monte Carlo simulation compared MI-CRP with four existing methods, including full information maximum likelihood, Bayesian estimation, multiple imputation via joint modeling, and multiple imputation via fully conditional specification. Results showed that MI-CRP adequately recovered the growth factor means and avoided extreme covariance bias. MI-CRP also achieved higher convergence and valid replication rates at low class separation with smaller sample sizes, and higher classification accuracy with less variability under challenging conditions. Overall, the advanced Bayesian methods developed and evaluated in this dissertation fill unresolved methodological gaps in GMMs and equip applied and methodological researchers with principled tools for implementing these models.