A discussion is provided of several issues related to behavioral measurement that arise from Pfadt et al. (2026, Psychometrika, 2026, 1-35). The note may be viewed in part as a complement to their developments regarding precision estimation for individual test scores.
A recent discussion of potentially serious drawbacks of the widely used scale sum scores is extended to nonnormal observed variables, and a didactic discussion is provided of their underpinnings. It is pointed out that the possibly marked bias and mean squared error as well as parameter estimator inconsistency, which result if using the sum scores as predictors of response measures, are not limited to normally distributed manifest variables. The asymptotically distribution-free estimation approach within the structural equation modeling framework is recommended to consider with large samples and nonnormal observed measures, in lieu of the common and traditional application of the scale sum scores as predictors of outcome variables. The approach is demonstrated in empirically relevant settings with substantially nonnormal data, where it is found to decidedly outperform that frequent application of the popular sum scores.
The present study aimed to examine (a) how fatigue severity changes during the course of chemotherapy in patients with hematologic cancer and (b) whether cytokines (IL-1 alpha, IL-1 beta, IL-6) are associated with fatigue change after controlling for demographic and clinical factors (e.g., hemoglobin/hematocrit, medications, comorbid conditions). This observational cohort study used data from 148 hematological cancer patients four times: prior to chemotherapy, on the last day of chemotherapy, 1 week after the chemotherapy completion, and 1 month after baseline assessment. Latent growth curve modeling was used to examine the longitudinal association of fatigue severity with cytokines and hemoglobin. A quadratic growth curve model fit the data well, indicating model tenability, and explained a large amount of variance in fatigue across measurement time points. Fatigue slightly increased toward the end of chemotherapy and decreased with time after chemotherapy completion. The influence of IL-6 on fatigue was significant at all time points except at the last assessment occasion (i.e., 1 month after the baseline assessment). The influence of IL-6 on fatigue was independent (unique) from the impact of hemoglobin level. Age and chemotherapy given for the first line of treatment significantly influenced the rate of fatigue change over time. Age also influenced the change pattern’s shape. Fatigue severity changes across the course of chemotherapy within the context of IL-6 activity, not the hemoglobin level. The influence of IL-6 may be limited during and shortly after chemotherapy. These findings inform the development of new symptom management strategies.
This article revisits the popular and widely used sum scores resulting from multiple-component measuring instruments, in the light of a recent resurgence of interest in them. We draw attention to a potentially serious disadvantage of sum scores that can be associated with (i) substantially larger bias and mean squared error than an alternative approach to parameter estimation, in addition to (ii) the inconsistency feature of an estimator frequently of special interest when utilizing the sum scores in behavioral and social studies. This drawback can be counteracted using a readily applicable structural equation modeling approach, which we outline and illustrate with data from empirically relevant settings. The article concludes with a discussion of extensions and limitations of the described procedure for examining the relationships between response variables and constructs evaluated with multi-component scales.
A procedure for interval estimation of the difference in the adjusted R-square index for nested linear models is discussed. The method yields as a byproduct confidence intervals for their standard R-square difference, as well as for the adjusted and standard R-squares associated with each model. The resulting interval estimate of the difference in adjusted R-square represents a useful and informative complement to the commonly used R-square change statistic and its significance test in model selection and contains substantially more information than that test. The outlined procedure is readily employed with popular software in empirical educational and psychological studies and is illustrated with numerical data.
The large-sample behavior of the asymptotically distribution-free estimator in covariance structure analysis is studied. It is shown that with increasing sample size and under suitable conditions, the estimator almost surely converges to the true parameter value. This strong convergence (i) represents numerical convergence with probability 1 of the resulting estimates of model parameters to their respective population values, as well as (ii) is stronger than the consistency and convergence in distribution of a parameter estimator, thus implying the latter two types of large-sample behavior. The demonstrated asymptotic convergence of the asymptotically distribution-free estimator for covariance structure models is illustrated on data.
A procedure for evaluation of the proportion explained component variance by the underlying trait in behavioral scales with second-order structure is outlined. The resulting index of accounted for variance over all scale components is a useful and informative complement to the conventional omega-hierarchical coefficient as well as the proportion of explained component correlation. A point and interval estimation method is described for the discussed index, which utilizes a confirmatory factor analysis approach within the latent variable modeling methodology. The procedure can be used with widely available software and is illustrated on data.
This note is concerned with the chance of the one-parameter logistic (1PL-) model or the Rasch model being true for a unidimensional multi-item measuring instrument. It is pointed out that if a single dimension underlies a scale consisting of dichotomous items, then the probability of either model being correct for that scale can be zero. The question is then addressed, what the consequences could be of removing items not following these models. Using a large number of simulated data sets, a pair of empirically relevant settings is presented where such item elimination can be problematic. Specifically, dropping items from a unidimensional instrument due to them not satisfying the 1PL-model, or the Rasch model, can yield potentially seriously misleading ability estimates with increased standard errors and prediction error with respect to the latent trait. Implications for educational and behavioral research are discussed.
An index extending the widely used omega-hierarchical coefficient is discussed, which can be used for evaluating the influence of a second-order factor on the interrelationships among the components of a hierarchical measuring instrument. The index represents a useful and informative complement to the traditional omega-hierarchical measure of explained overall scale score variance by that underlying construct. A point and interval estimation procedure is outlined for the described index, which is based on model reparameterization and is developed within the latent variable modeling framework. The method is readily applicable with popular software and is illustrated with examples.
This article is concerned with the assumption of linear temporal development that is often advanced in structural equation modeling-based longitudinal research. The linearity hypothesis is implemented in particular in the popular intercept-and-slope model as well as in more general models containing it as a component, such as longitudinal structural models with covariates, or models for the study of predictors and correlates of change. In empirical research applications, currently behavioral and social scientists typically evaluate only overall goodness of fit for a considered model. However, this omnibus fit assessment may miss violations of the underlying linearity assumption. To respond to this limitation, the present article discusses a testing procedure for examining the hypothesis of linear growth or decline separately from the widely used overall fit evaluation process. The method is readily utilized with popular latent variable modeling software and is illustrated using a numerical example.
An application of Bayesian factor analysis for evaluation of scale reliability is discussed, which is developed within the framework of latent variable modeling. The method permits direct point and interval estimation of the reliability coefficient of multiple-component measuring instruments using Bayesian inference. The approach allows also point and interval estimation of the population discrepancy between the popular coefficient alpha and instrument reliability. The procedure is readily applied in empirical measurement research employing widely available statistical software. The outlined method is illustrated using numerical data.
This note intends to complement the recent discussion in Hayes and Coutts (2020) by focusing on (i) the loading equality condition for the population identity of coefficient alpha and reliability of multiple-indicator measurement scales, as well as (ii) the potential utility of alpha when this condition is not satisfied. We show that the alpha and reliability coefficients can be very close at the population level in certain cases of loading inequality. In addition, we point out that in any studied population the identity of alpha and scale reliability (coefficient omega) is an improbable event. We discuss implications for communication and behavioral research with large samples, which are becoming increasingly widely used in large-scale and nationally representative studies. Findings of the article are then illustrated using numerical data. We conclude with proposed recommendations for the use of coefficients alpha and omega by communication and behavioral scientists concerned with evaluating scale reliability.
This note demonstrates that the widely used Bayesian Information Criterion (BIC) need not be generally viewed as a routinely dependable index for model selection when the bifactor and second-order factor models are examined as rival means for data description and explanation. To this end, we use an empirically relevant setting with multidimensional measuring instrument components, where the bifactor model is found consistently inferior to the second-order model in terms of the BIC even though the data on a large number of replications at different sample sizes were generated following the bifactor model. We therefore caution researchers that routine reliance on the BIC for the purpose of discriminating between these two widely used models may not always lead to correct decisions with respect to model choice.
A procedure is outlined for point and interval estimation of location parameters associated with polytomous items, or raters assessing studied subjects or cases, which follow the rating scale model. The method is developed within the framework of latent variable modeling, and is readily applied in empirical research using popular software. The approach permits testing the goodness of fit of this widely used model, which represents a rather parsimonious item response theory model as a means of description and explanation of an analyzed data set. The procedure allows examination of important aspects of the functioning of measuring instruments with polytomous ordinal items, which may also constitute person assessments furnished by teachers, counselors, judges, raters, or clinicians. The described method is illustrated using an empirical example.
A Bayesian statistics-based approach is discussed that can be used for direct evaluation of the popular Cronbach’s coefficient alpha as an internal consistency index for multiple-component measuring instruments, as well as for testing its identity to scale reliability. The method represents an application of confirmatory factor analysis within the Bayesian inference framework and is widely applicable in empirical measurement research using popular latent variable modeling software. The procedure readily furnishes posterior median point estimates and credible intervals of coefficient alpha. The approach also permits testing a necessary and sufficient condition for population equality of the alpha and scale reliability coefficients, and under its plausibility provides in addition a dependable means for estimation of instrument reliability. The outlined procedure is illustrated using numerical data.
A multiple-step procedure is outlined that can be used for examining the latent structure of behavior measurement instruments in complex empirical settings. The method permits one to study their latent structure after assessing the need to account for clustering effects and the necessity of its examination within individual levels of fixed factors, such as gender or group membership of substantive relevance. The approach is readily applicable with binary or binary-scored items using popular and widely available software. The described procedure is illustrated with empirical data from a student behavior screening instrument.