
Muda and Yangxuan recommended a ridge-calibrated test statistic for controlling Type I error inflation in confirmatory factor analysis under non-normality. Although improving finite-sample inference is an important goal, their empirical evaluation is undermined by incorrectly implemented benchmark statistics, misattributed comparison methods, and a simulation design restricted to asymptotically robust conditions. Using corrected code, we reanalyze the original conditions and replace the misattributed benchmarks with the penalized eigenvalue procedures pEBA4RLS and pOLS2RLS. Under the original conditions, the best-performing procedures are the ridge-calibrated TCsFCr statistic and the penalized eigenvalue methods. However, under additional non-asymptotically robust conditions, including high-dimensional models, pEBA4RLS provides the most reliable Type I error control and outperforms the ridge-calibrated competitors.
As empirical applications and methodological research on second-order exploratory factor analysis have matured, important refinements to estimation procedures, interpretive frameworks, and theoretical expectations have emerged. These advances establish second-order EFA as a rigorous alternative to bi-factor EFA and to conventional EFA models characterized by substantial factor correlations. We further demonstrate that bi-factor EFA models constitute a reparameterization of second-order EFA with direct effects, clarifying the formal relationship between these approaches. The unification of bi-factor and second-order EFA models reveals a novel rotation approach for bi-factor solutions that outperforms bi-geomin in simulation studies and yields more interpretable factor solutions in empirical studies.
Structural equation modeling (SEM) is a popular methodological approach for representing and testing hypothesized relationships among observed and unobserved, or latent, variables. Researchers in the social and behavioral sciences commonly encounter three general situations for testing such structural equation models-strictly confirmatory, testing alternative or competing hypothesized models, and model generation. However, none of these involve the examination of potential impacts of omitted confounders on the conclusions generated from the model estimates, known as a sensitivity analysis. While SEM researchers are well-versed at specifying and testing models, methods for sensitivity analysis in SEM are still not widely available and typically require a priori specification of the path coefficients and pattern of interrelationships. The current study uses the metaheuristic Tabu Search algorithm for automating the process of the sensitivity analysis in SEM and examines its overall performance in accomplishing this task. We present sensitivity analysis results obtained from the evaluation of a published empirical study on emotional and behavioral problems among foster children. Finally, a discussion of the implications of the findings is provided.
Identifying observations that disproportionately influence model estimates is essential for evaluating the robustness of structural equation models (SEMs), yet existing case-level diagnostics remain computationally intensive and limited in interpretive depth. This paper introduces a scalable influence diagnostic framework for SEM by adapting a curvature-aware influence function approximation from the machine learning literature. The approach combines casewise score contributions with curvature information from the observed information matrix to produce a vector-valued influence estimate for each observation without requiring repeated model fitting. Unlike scalar diagnostics such as generalized Cook's distance, the resulting influence vectors preserve the joint directional structure of each case's effect across all model parameters simultaneously, enabling parameter-level interpretation and multivariate visualization via principal component analysis. A Monte Carlo simulation study demonstrated that the proposed method detected known influential observations with greater accuracy than the approximate diagnostic implemented in the semfindr package across all 12 experimental conditions, with recall advantages ranging from .044 to .137. An empirical application to the Holzinger-Swineford dataset illustrated how the framework identifies not only which observations are influential, but how and where in the parameter space their influence is concentrated. Confirmatory case deletion verified that the influence function correctly predicted the direction of parameter change for all nine factor loadings. The proposed approach offers a theoretically grounded, computationally efficient, and interpretively richer alternative to deletion-based diagnostics for applied SEM researchers.
A Monte Carlo simulation was used to evaluate the performance of latent growth curve (LGC) models with binary observed variables when the attrition pattern was missing not at random (MNAR) for five time points. Parameter and standard error biases for three estimation methods were compared: weighted least squares with mean and variance adjustment (WLSMV), categorical robust marginal maximum likelihood (categorical MLR), and Bayes. The results indicated that robust diagonal weighted least squares paired with multiple imputation (MI) performed best when values were missing at random. When data were missing not at random, Bayesian estimates performed best.
Model estimations given mixed data types proceed in two stages. Stage 1 obtains a covariance matrix of the continuous indicators and the variables underlying the categorical indicators. Stage 2 fits the researcher's model to data with weighted, diagonally weighted, or unweighted least squares. For a misspecified model, three problems have been overlooked about defining and estimating population standardized parameters beta. First, given the researcher's model and population covariance matrix, only one fit function can yield unique population unstandardized parameters theta. Second, standardizing theta into beta cannot remove the influence from data scales; beta is not unique if data scales change. Third, the prevailing standard error method for theta (known as the "robust SE") may have large bias even for slightly misspecified models, and thus is not suitable for finding SE(beta). Unless the researcher's model is perfect, the three problems above will persist. We propose solutions and conduct simulations to verify the solutions are effective.
We embed Bayesian graphical modeling within partially confirmatory factor analysis (PCFA) to estimate sparse latent factor correlations, thereby overcoming the limitations of the inverse-Wishart prior. We regularize the factor correlation matrix using graphical lasso, horseshoe, and spike-and-slab priors, yielding a unified framework that applies sparsity to loadings and factor correlations. For complex structures, we further introduce partial regularization for multitrait-multimethod and extended bifactor models. Across simulation studies, the spike-and-slab prior achieves superior false-positive control and high accuracy, mainly in large-sample, complex, or sparse settings. The horseshoe prior demonstrates stable performance across sample sizes but yields more false positives, whereas the lasso prior shows moderate and reliable performance. For model identification, significance-based criteria outperform threshold-based rules for regularized estimators, especially in complex designs and empirical analyses. A personality-inventory application shows that partial regularization improves interpretability and model fit in estimating factor correlations and loadings.
Matrix Decomposition Structural Equation Modeling (MDSEM) provides a data-matrix-based alternative to conventional SEM procedures. MDSEM is shown to be highly vulnerable to multicollinearity among latent variables, often resulting in biased and unstable estimates of parameters. This study proposes Regularized Matrix Decomposition Structural Equation Modeling (RMDSEM), an extension of MDSEM that introduces explicit L2 regularization. RMDSEM addresses the multicollinearity issue by regularizing the path coefficient matrix while preserving the matrix-decomposition framework. An iterative estimation algorithm and a K-fold cross-validation procedure for tuning multiple regularization parameters are derived. Simulation results show that MDSEM and maximum-likelihood SEM yield inflated bias and standard errors under moderate to strong multicollinearity, whereas RMDSEM achieves stable estimation with substantially reduced mean squared errors. Empirical applications further demonstrate that RMDSEM suppresses unreasonable estimates and yields interpretable structural relations.
The growing availability of large-scale survey data has increased interest in comparing relations among latent variables (often called structural relations) across many groups (e.g., countries or schools) using Structural Equation Modeling (SEM). Traditional multigroup or multilevel SEM become impractical in these settings, as they require pairwise comparisons to identify specific differences between groups. These pairwise comparisons become daunting as the number of groups increases. Mixture Multigroup SEM (MMG-SEM) addresses this challenge by clustering groups that share similar structural relations, while accounting for measurement invariance and non-invariance. Nevertheless, MMG-SEM has lacked an open software implementation until now. This paper introduces mmgsem, an open-source R package that provides a user-friendly implementation of MMG-SEM. We present a step-by-step tutorial covering key aspects of the method, including handling measurement invariance, selecting the number of clusters, and hypothesis testing of structural parameters. A simulated example illustrates the workflow, from data preparation to result interpretation.
Moderated Nonlinear Latent Factor Analysis (MNLFA) has been introduced as a flexible approach for testing measurement invariance among categorical and continuous covariates. Equipped with Bayesian shrinkage priors, MNLFA can handle large numbers of covariates and potentially invariant item parameters. The present study extends the capabilities of the Bayesian MNLFA to multilevel and longitudinal confirmatory factor analysis. We show how a Bayesian hierarchical MNLFA (BH-MNLFA) can be implemented and provide two simulation studies to demonstrate its functionality. Focusing on invariance explorations in experience sampling data as a potential use case in the context of longitudinal data analysis, we showcase the utility of BH-MNLFA with data from educational psychology, and test invariance of state self-concepts measures across time and school subjects.
This simulation study investigates the utility of the percentage of noninvariant groups and the R & sup2; statistic as indicators of approximate measurement invariance (AMI) within the alignment optimization framework, particularly when measurement noninvariance varies continuously across groups in large-scale assessments. The design manipulated the number of groups, the number of factors, and the degree of AMI. Results showed that, for a given level of AMI, the threshold percentage of noninvariant groups depended on the total number of groups, with higher thresholds observed as the number of groups increased. R & sup2; values were generally higher for intercepts than for factor loadings under equivalent conditions, and R & sup2; appeared to be a more informative indicator of invariance for factor loadings than for intercepts. The number of factors had negligible impact on the outcome metrics. Correlations between estimated and true factor means indicated that values below 0.95 were suggestive of noninvariance. Implications for applied researchers were discussed.
This study introduces a joint modeling framework that integrates structural equation model (SEM) with multi-state semi-competing risks model to investigate the complex causal mechanism among exposure, latent mediators, and multiple hierarchically correlated survival outcomes. The framework first employs a linear SEM to characterize latent mediators underlying highly correlated observable surrogates. These latent mediators are then incorporated into semi-parametric proportional hazards models with frailty to assess their mediating roles in the relationships between exposure and hierarchically correlated survival outcomes. Estimation and inference are conducted within an integrated Bayesian framework leveraging Monte Carlo methods, Gibbs sampling, and the Metropolis-Hastings algorithm. Extensive simulation studies demonstrate the robustness and accuracy of the proposed method. Application to the Framingham Heart Study uncovers causal pathways linking smoking and metabolic health factors and evaluates their combined impact on long-term cardiovascular outcomes, offering new insights into the complex interplay between lifestyle factors and health risks.
The performance of linear growth mixture models (GMMs) estimated in the multilevel modeling (MLM) framework and the structural equation modeling (SEM) framework in Mplus is examined. Empirical analyses conducted using data from the Early Childhood Longitudinal Study - Kindergarten Class of 1998-1999 (ECLS-K) led to different estimates of a 2-class linear GMM when estimated in the MLM versus the SEM frameworks via Mplus. A simulation study was conducted under varying conditions of sample size and class separation to better understand when differences in model fit and parameter estimates emerge. Results suggested the frameworks yielded comparable model fit information and parameter estimates when class separation was large; however, differences emerged as class separation decreased with the MLM estimation routine tending to outperform the SEM estimation routine. Overall, these findings highlight the importance of considering the estimation framework when estimating linear GMMs in Mplus.
This study focuses on Level 2 (L2) predictor-outcome bias in standardized coefficients as obtained through traditional multilevel models and multilevel structural equation models when sampling and measurement error is present in the outcome. While unreliability in the outcome variable does not bias unstandardized coefficients, the current study demonstrates how it can bias estimates in their more commonly reported standardized forms. Using simulations, we evaluate this issue under a variety of conditions in which sampling and measurement error are present in the outcome as a function of aggregation method. In doing so, we consider four methods for modeling L2 aggregate variables, including doubly manifest aggregation (MLM), doubly latent aggregation (MLSEM), and their two hybrids, in relation to the type of L2 variables researchers may encounter in practice and the extent to which a given methods can produce unbiased standardized coefficient values.
This Teacher's Corner article provides a comprehensive illustration of consistent partial least squares structural equation modeling (PLSc-SEM), a variant of the original PLS-SEM method, which corrects construct correlations for attenuation. The method thereby allows estimating common factor models within a composite-based SEM framework. Our descriptions draw on SmartPLS 4, the most frequently used software for conducting PLS-SEM analyses, illustrating how to set up and estimate the model and interpret the estimates. We further contrast our results with those from other composite-based SEM estimators and from covariance-based SEM using maximum likelihood estimation, as implemented in SmartPLS.
The present study employs penalized structural equation modeling (PSEM) to conduct 1-step analyses of latent class models with covariates. Through the utilization of simulation studies, it is demonstrated that the 1-step approach yields unbiased parameter estimates, which are comparable to those obtained through the application of multi-step bias-correction methods, including the three-step, and BCH procedures. In contradistinction to multi-step approaches, the PSEM 1-step method necessitates only a single estimation step to incorporate covariates into the latent class model, thereby obviating the necessity for additional procedures to adjust for classification error. Moreover, we illustrate the practical application of the PSEM 1-step approach through an empirical example and provide Mplus syntax as a reference for applied researchers.
Parceling is routinely recommended to SEM researchers with many items (high p) and modest N in order to "lower indicator-to-sample-size ratio." Here we show that this directive is a roundabout response to address chi-square distortion at high p and modest N, which in turn leads to egregiously inflated Type I error for the item-level model's chi-square test of absolute fit. We explain, demonstrate, and provide software to implement two more-direct strategies to address this chi-square distortion. These strategies involve fitting an item-level model and either (i) empirically correcting the chi-square critical value or (ii) employing a high-p and modest-N adjustment to the chi-square statistic itself (and other chi-square based fit indices). Both strategies are shown to control Type I error more consistently than parceling and, critically, avoid across-item-to-parcel-allocation variability in Type I error rates. Both item-level model strategies can also legitimately improve fit of a correct item-level model by addressing high-p, modest-N distortion of chi-square-without inducing parcel-allocation variability in model fit.
In modeling intensive longitudinal data with the dynamic structural equation modeling framework, this demonstration analyzes the impact of Gaussian process priors in the presence or absence of a significant trend component on the parameter estimates obtained through Bayesian analysis. The results show that even when not the focus of a research study, failing to account for the presence of trend may severely impact statistical inference and Gaussian processes present a viable option for accounting for them. Also, in the absence of trends, Gaussian processes do not impact the quality of parameter estimates when introduced in an over-specified model.