In psychological research, researchers may wish to examine whether two groups have similar means on a target variable, such as whether online and in-class learning are equally effective. However, researchers often rely on traditional null hypothesis testing (TNHST) to claim equivalence based on non-significant results, an approach that may lead to misleading conclusions. This article explains the limitations of this practice and introduces the concepts and methods of equivalence testing, also referred to as the negligible effect significance test. In brief, researchers specify a smallest effect size of interest (SESOI), below which an effect is considered practically negligible. They then test whether the observed statistic is significantly smaller than this threshold. If so, it can be concluded that the population effect is sufficiently small to be considered negligible. The article also provides examples of applications in psychological research and demonstrates how equivalence can be evaluated using outputs from commonly used statistical software or online tools.
Understanding latent interactions is an important need for the structural equation modeler. Plotting and probing latent interactions, however, has not been well defined. We describe methods for plotting and probing two- and three-way latent interactions fit with a variety of approaches (LMS/QML, residual centering, double mean centering). The methods are demonstrated through a small simulation and examples based on existing data.
To provide researchers with a means of assessing the fit of the structural component of structural equation models, structural fit indices- modifications of the composite fit indices, RMSEA, SRMR, and CFI- have recently been developed. We investigated the performance of four of these structural fit indices- RMSEA-P, RMSEA_S, SRMR_S, and CFI_S-, when paired with widely accepted cutoff values, in the service of detecting structural misspecification. In particular, by way of simulation study, for each of seven fit indices- 3 composite and 4 structural-, and the traditional chi-square test of perfect composite fit, we estimated the following rates: a) Type I error rate (i.e., the probability of (incorrect) rejection of a correctly specified structural component), under each of four degrees of misspecification in the measurement component; and b) Power (i.e., the probability of (correct) rejection of an incorrectly specified structural model), under each condition formed of the pairing of one of three degrees of structural misspecification with one of four degrees of measurement component misspecification. In addition to sample size, the impacts of two model features, incidental to model misspecification- number of manifest variables per latent variable and magnitude of factor loading- were investigated. The results suggested that, although the structural fit indices performed relatively better than the composite fit indices, none of the goodness-of-fit index with a fixed cutoff value pairings was capable of delivering an entirely satisfactory Type I error rate/Power balance, [RMSEA_S,.05] failing entirely in this regard. Of the remaining pairings; a) RMSEA-P and CFI_S suffered from a severely inflated Type I error rate; b) despite the fact that they were designed to pick up on structural features of candidate models, all pairings- and especially, RMSEA-P and CFI_S- manifested sensitivities to model features, incidental to structural misspecification; and c) although, in the main, behaving in a sensible fashion, SRMR_S was only sensitive to structural misspecification when it occurred at a relatively high degree.
We simulated Bayesian CFA models to investigate the power of PPP to detect model misspecification by manipulating sample size, strongly and weakly informative priors for nontarget parameters, degree of misspecification, and whether data were generated and analyzed as normal or ordinal. Rejection rates indicate that PPP lacks power to reject an inappropriate model unless priors are unrealistically restrictive (essentially equivalent to fixing nontarget parameters to zero) and both sample size and misspecification are quite large. We suggest researchers evaluate global fit without priors for nontarget parameters, then search for neglected parameters if PPP indicates poor fit.
A composite score is the sum of a set of components. For example, a total test score can be defined as the sum of the individual items. The reliability of composite scores is of interest in a wide variety of contexts due to their widespread use and applicability to many disciplines. The psychometric literature has devoted considerable time to discussing how to best estimate the population reliability value. However, all point estimates of a reliability coefficient fail to convey the uncertainty associated with the estimate as it estimates the population value. Correspondingly, a confidence interval is recommended to convey the uncertainty with which the population value of the reliability coefficient has been estimated. However, many confidence interval methods for bracketing the population reliability coefficient exist and it is not clear which method is most appropriate in general or in a variety of specific circumstances. We evaluate these confidence interval methods for 4 reliability coefficients (coefficient alpha, coefficient omega, hierarchical omega, and categorical omega) under a variety of conditions with 3 large-scale Monte Carlo simulation studies. Our findings lead us to generally recommend bootstrap confidence intervals for hierarchical omega for continuous items and categorical omega for categorical items. All of the methods we discuss are implemented in the freely available R language and environment via the MBESS package.
Planned missing data designs allow researchers to increase the amount and quality of data collected in a single study. Unfortunately, the effect of planned missing data designs on power is not straightforward. Under certain conditions using a planned missing design will increase power, whereas in other situations using a planned missing design will decrease power. Thus, when designing a study utilizing planned missing data researchers need to perform a power analysis. In this article, we describe methods for power analysis and sample size determination for planned missing data designs using Monte Carlo simulations. We also describe a new, more efficient method of Monte Carlo power analysis, software that can be used in these approaches, and several examples of popular planned missing data designs.
This article aims to show the mathematical reasoning behind all effect sizes used in the partialInvariance and partialInvarianceCat functions in semTools package. In the functions, the following statistics are compared across groups: factor loadings, item intercepts (for continuous items), item thresholds (for categorical items), measurement error variances, and factor means.
In many situations, researchers collect multilevel (clustered or nested) data yet analyze the data either ignoring the clustering (disaggregation) or averaging the micro-level units within each cluster and analyzing the aggregated data at the macro level (aggregation). In this study we investigate the effects of ignoring the nested nature of data in confirmatory factor analysis (CFA). The bias incurred by ignoring clustering is examined in terms of model fit and standardized parameter estimates, which are usually of interest to researchers who use CFA. We find that the disaggregation approach increases model misfit, especially when the intraclass correlation (ICC) is high, whereas the aggregation approach results in accurate detection of model misfit in the macro level. Standardized parameter estimates from the disaggregation and aggregation approaches are deviated toward the values of the macro- and micro-level standardized parameter estimates, respectively. The degree of deviation depends on ICC and cluster size, particularly for the aggregation method. The standard errors of standardized parameter estimates from the disaggregation approach depend on the macro-level item communalities. Those from the aggregation approach underestimate the standard errors in multilevel CFA (MCFA), especially when ICC is low. Thus, we conclude that MCFA or an alternative approach should be used if possible.
When planning to conduct a study, not only is it important to select a sample size that will ensure adequate statistical power, often it is important to select a sample size that results in accurate effect size estimates. In cluster-randomized designs (CRD), such planning presents special challenges. In CRD studies, instead of assigning individual objects to treatment conditions, objects are grouped in clusters, and these clusters are then assigned to different treatment conditions. Sample size in CRD studies is a function of 2 components: the number of clusters and the cluster size. Planning to conduct a CRD study is difficult because 2 distinct sample size combinations might be associated with similar costs but can result in dramatically different levels of statistical power and accuracy in effect size estimation. Thus, we present a method that assists researchers in finding the least expensive sample size combination that still results in adequate accuracy in effect size estimation. Alternatively, if researchers have a fixed budget, they can select the sample size combination that results in the most precise estimate of effect size. A free computer program that automates these procedures is available.
This paper proposes a Monte Carlo approach for nested model comparisons. This approach allows for test of approximate equivalency in fit between nested models and customizing cutoff criteria for difference in a fit index. Different methods to account for trivial misspecification in the Monte Carlo approach are also discussed. A simulation study is conducted to compare the Monte Carlo approach with different methods of imposing trivial misspecification to chi-square difference test and change in comparative fit index (CFI) with suggested cutoffs. The simulation study shows that the Monte Carlo approach is superior to the chi-square difference test by correctly retaining the nested model with trivial misspecification. It is also superior to the change in CFI by offering higher power to detect severe misspecification.
Salivary cortisol is often used as an index of physiological and psychological stress in exercise science and psychoneuroendocrine research. A primary concern when designing research studies examining cortisol stems from the high cost of analysis. Planned missing data designs involve intentionally omitting a random subset of observations from data collection, reducing both the cost of data collection and participant burden. These designs have the potential to result in more efficient, cost-effective analyses with minimal power loss. Using salivary cortisol data from a previous study (Hogue, Fry, Fry, & Pressman, 2013), this article examines statistical power and estimated costs of six different planned missing data designs using growth curve modeling. Results indicate that using a planned missing data design would have provided the same results at a lower cost relative to the traditional, complete data analysis of salivary cortisol.
Residual centering is a useful tool for orthogonalizing variables and latent constructs, yet it is underused in the literature. The purpose of this article is to encourage residual centering's use by highlighting instances where it can be helpful: modeling higher order latent variable interactions, removing collinearity from latent constructs, creating phantom indicators for multiple group models, and controlling for covariates prior to latent variable analysis. Residual centering is not without its limitations, however, and the authors also discuss caveats to be mindful of when implementing this technique. They discuss the perils of double orthogonalization (i.e., simultaneously orthogonalizing A relative to B and B relative to the original A), the unintended consequences of orthogonalization on model fit, the removal of a mean structure, and the effects of nonnormal data on residual centering.
The rankings of the Akaike information criterion (AIC) or Bayesian information criterion (BIC) are often used for nonnested model comparison. When the competing models do not include the population...
Residual centering is a useful tool for orthogonalizing variables and latent constructs, yet it is underused in the literature. The purpose of this article is to encourage residual centering’s use by highlighting instances where it can be helpful: modeling higher order latent variable interactions, removing collinearity from latent constructs, creating phantom indicators for multiple group models, and controlling for covariates prior to latent variable analysis. Residual centering is not without its limitations, however, and the authors also discuss caveats to be mindful of when implementing this technique. They discuss the perils of double orthogonalization (i.e., simultaneously orthogonalizing A relative to B and B relative to the original A), the unintended consequences of orthogonalization on model fit, the removal of a mean structure, and the effects of nonnormal data on residual centering.
Directional dependency is a method to determine the likely causal direction of effect between two variables. This article aims to critique and improve upon the use of directional dependency as a technique to infer causal associations. We comment on several issues raised by von Eye and DeShon (2012), including: encouraging the use of the signs of skewness and excessive kurtosis of both variables, discouraging the use of D'Agostino's K2, and encouraging the use of directional dependency to compare variables only within time points. We offer improved steps for determining directional dependency that fix the problems we note. Next, we discuss how to integrate directional dependency into longitudinal data analysis with two variables. We also examine the accuracy of directional dependency evaluations when several regression assumptions are violated. Directional dependency can suggest the direction of a relation if (a) the regression error in population is normal, (b) an unobserved explanatory variable correlates with any variables equal to or less than .2, (c) a curvilinear relation between both variables is not strong (standardized regression coefficient ≤ .2), (d) there are no bivariate outliers, and (e) both variables are continuous.