
Estimation of Latent class item response theory (LC-IRT) models is typically performed using full-information maximum likelihood, which may be computationally demanding, susceptible to convergence problems, and subject to interpretational confounding. We propose a two-step estimation approach for unidimensional and multidimensional LC-IRT models with covariates and differential item functioning (DIF). In the first step, the measurement model is estimated, including direct covariate effects on item responses when DIF is part of the final measurement specification. In the second step, covariate effects on latent class membership are estimated while treating the measurement parameters as fixed. We evaluate the two-step estimator through simulation studies and identify the conditions under which it yields low bias, particularly when measurement quality and entropy are adequate. We then apply the proposed approach to data on student performance in statistics.
This Monte Carlo study examined how covariate inclusion affects class enumeration in one-factor, two-class factor mixture models (FMMs). Data were generated under two scenarios: a continuous covariate predicted latent class membership only, or predicted both latent class membership and the latent factor. Conditions varied by class separation, sample size, mixing proportion, and covariate-effect magnitude. Unconditional and one-step conditional FMMs were evaluated using information criteria, likelihood-ratio-based tests, entropy, classification accuracy, and parameter coverage. Results showed that class separation was the strongest determinant of performance, with sample size providing secondary benefits. When the covariate predicted class membership only, conditional FMMs generally improved classification accuracy and became more competitive for class enumeration as the covariate effect increased. When the covariate also predicted the latent factor, these advantages weakened and became more criterion-dependent. Across conditions, the Bayesian information criterion and consistent Akaike’s information criterion were the most stable enumeration criteria, whereas entropy and likelihood-ratio-based tests were less reliable. The findings indicate that the value of covariate inclusion in FMM depends on the pathway through which the covariate affects the population model.
Researchers using latent profile and latent class analysis (LPA/LCA) commonly assign individuals to their modal class without evaluating whether this simplification distorts reported class sizes, profile means, or high-severity subgroups. Existing classification-quality diagnostics-entropy, average posterior probabilities, and Masyn's odds of correct classification and classification probability-assess how sharply a model separates its classes, not whether hard-assigned summaries diverge from probability-weighted ones. We address this through an empirical benchmark (four-class SCL-90 solution; N = 59,408), a fully crossed simulation (class separation, class balance, and indicator-class discrimination precision; 27 conditions), and a six-index diagnostic framework. Hard-assigned and probability-weighted summaries were interchangeable under favorable conditions but diverged under low class separation, severe imbalance, or low precision-most acutely for the smallest, most extreme class. The framework offers simulation-calibrated thresholds for documenting assignment adequacy in continuous-indicator LPA; extension to categorical LCA requires further validation. Hard assignment is most defensible when its adequacy is documented rather than assumed.
Differential item functioning (DIF) is an issue of a measure that can lead to differences in item scores between subpopulations (e.g., sex, race, age) while controlling for a latent trait or ability level. It is important to report DIF effect size measures to understand the practical significance of DIF results as well. We reviewed 15 different DIF effect size measures in the literature, and no previous study had examined whether the 15 DIF effect size measures agreed with one another, which has implications in the number and type of measures to report. The present study investigated the relationship among DIF effect size measures in 1PL, 2PL, and GRM models and provide clear recommendations to applied researchers for their reporting. Our simulation results showed that, for each model, three effect size measures can be reported to obtain the most unique information about uniform DIF and the least bias from sampling conditions (e.g., sample size and impact): Mantel-Haenszel Delta, McFadden's Pseudo R-squared, and a signed DIF measure. However, the optimal signed DIF measure varies by model, with Signed Area under the curve for the 1PL model, Standardized P-difference for the 2PL, and Signed Item Difference in the Sample/Normal Distribution for the GRM. The recommended approach to reporting DIF effect size improves the detection of DIF through practical significance and provides researchers with more tools to address bias.
Local item dependence (LID) is a persistent threat to dimensionality assessment in ordinal data, yet its effects on factor retention methods remain incompletely understood. Five factor retention methods (minimum average partial [MAP], parallel analysis [PA], Exploratory Graph Analysis [EGA]-Walk, EGA-Leiden, EGA-TMFG) combined with six correlation matrix estimators (Pearson, polychoric, and four fuzzy-based variants) were evaluated in a Monte Carlo simulation spanning 900 conditions and 450,000 datasets. First, pairwise and testlet LID structures produced asymmetric effects on factor retention: pair-LID induced systematic overfactoring in GLASSO-based EGA methods and PA, while testlet-LID produced near-ceiling MAP accuracy (.99) through mechanical error compensation rather than correct structural identification. Second, correlation matrix choice had negligible influence on retention performance. All four fuzzy variants were statistically indistinguishable from Pearson across all conditions, while polychoric showed a condition-dependent profile, with modest advantages under local independence but poorer performance under severe testlet-LID. Third, threshold distribution shape exerted no meaningful influence on retention outcomes, either as a main effect or in interaction with LID condition. EGA-TMFG showed the strongest resilience and is recommended when local dependence is plausible. LID structure and factor loading strength were the primary determinants of retention accuracy; correlation matrix preprocessing provided no practical corrective benefit.
We propose a flexible multimodal model for predicting all dichotomous and polytomous item parameters from text, images, and metadata by fusing representations from encoder Transformer vision and language models. This deep learning model accommodates heterogeneous item formats, including items with any number of components such as correct and incorrect options, stimuli, and images. Answer key indicators distinguish correct options from distractors, while an attention pooling technique weights the relative importance of these components. Based on the two-parameter logistic model and the generalized partial credit model, we predict all item parameters jointly using a masking strategy to ensure that only relevant parameters contribute to the loss. Item-level and item component-level metadata are also included. We evaluate the approach using 40,965 English language arts and mathematics items for Grades 3-11. A single model accommodated both exam subjects, all 11 item types, and all item parameters, eliminating the need for multiple specialized models. However, results indicate that the full model was unable to leverage all input data. Often, prediction accuracy was unchanged as features were removed. Images were good predictors on their own (item intercept R 2 = . 25 ) but did not consistently contribute unique information when combined with text and metadata. Of the strongest-performing models, the most parsimonious variant achieved R 2 values of .67, .45, .39, .34, .66, .33, and .74 for item intercept, discrimination, difficulty, and four polytomous step threshold parameters, respectively. Findings suggest that current training methods may limit the learnability of complex multimodal deep fusion models. Example Python code from this study is available on Github: https://github.com/hotakamaeda/multimodal_prediction.
Differential test functioning (DTF) evaluates whether a test exhibits group differences beyond those attributable to latent trait differences. Its magnitude is often most interpretable on the raw-score scale. However, most existing DTF effect size measures rely on item response theory (IRT) modeling. Building on the renewed practical value of sum scores, this study proposes a simple observed-score-based approach to estimate the DTF magnitude directly in raw-score points. The proposed Index S stratifies examinees by an anchor-based sum score composed of items not flagged for differential item functioning (DIF), summarizes the within-stratum mean differences in total test scores, and aggregates these conditional differences using weights based on the observed distribution of the matching variable. To stabilize the estimation when score strata are sparse, adjacent strata are merged to satisfy a minimum per-group sample size requirement, and Index S_std provides a standardized version. We evaluated the indices using two-parameter logistic (2PL) simulation studies by varying the sample size, test length, DIF type, DIF proportion, DIF direction, and focal group latent trait distribution. The utility was assessed in terms of estimation accuracy, including Bias and root mean square error (RMSE), and correlation with established IRT-based DTF indices. An empirical application to TIMSS 2023 Grade 8 mathematics (Japan vs. the United States) illustrates how the proposed indices provide an accessible raw-score interpretation of the DTF magnitude for both psychometric and non-psychometric stakeholders.
Agreement between raters does not, by itself, show how that agreement is produced. In binary judgement tasks, raters may classify cases similarly either because they rely on similar decision thresholds or because those thresholds operate under favourable marginal conditions. The present study examined whether convergence in raters’ decision thresholds must be imposed exogenously or can emerge endogenously through strategic adjustment. Within an equal-variance signal-detection framework, raters were modelled as choosing thresholds that balance expected classification accuracy against an incentive to align with the rest of the panel. Equilibrium thresholds were obtained by Nash best-response updating, and threshold variance, the Strategic Convergence Index (SCI), and mean pairwise Cohen’s κ were examined as functions of alignment pressure and prevalence. Increasing alignment pressure produced a sharp collapse in threshold dispersion and drove SCI towards 1 under the chosen normalisation, whereas κ increased only within a comparatively narrow range. When alignment pressure was held constant, SCI remained effectively unchanged across prevalence conditions, while κ varied substantially. These findings indicate that convergence in decision thresholds can arise as an equilibrium outcome of strategic interaction yet remains analytically distinct from observable agreement. They also clarify the status of Cohen’s κ in this setting: it tracks agreement at the level of realised classifications and therefore remains sensitive to marginal conditions, whereas SCI captures convergence in the latent decision structure from which those classifications arise. SCI is thus not proposed as a replacement for agreement coefficients but as a model-based quantity that makes explicit a form of alignment that κ only partially reflects.
Exploratory factor analysis (EFA) is widely used to identify latent structures measured by observed indicators. Nevertheless, determining the number of factors remains a methodological challenge, especially when the size of the EFA model is large (e.g., models with more than 10 indicators per factor and more than five factors). This study evaluated the performance of 15 factor-retention methods under large factor analysis models, including parallel analysis (PA) using the mean and 95th percentiles of threshold, exploratory graph analysis (EGA) with GLASSO and TMFG estimations, sequential χ2, fit indices (Comparative Fit Index [CFI], Tucker-Lewis Index [TLI], Root Mean Square Error of Approximation [RMSEA]), Kaiser Criterion (K1 rule), the very simple structure (VSS) method, the comparison data (CD) approach, the minimum average partial (MAP) test, the Hull method using the comparative fit index, Cattell's acceleration factor criteria, and Bayesian information criterion (BIC). Specifically, we manipulated the number of factors (5, 6, and 7), indicators per factor (10 and 15), sample sizes (100, 200, 500, 800, 1,000, 1,500, 2,000, and 4,000), inter-factor correlations (.30, .50, and .70), and factor loadings (.40 and .70). Results reveal that PA-M and EGA-TMFG are the most robust methods across varying conditions. PA-M excels under low-to-moderate inter-factor correlations, while EGA-TMFG is more accurate in high-correlation scenarios when sample sizes are sufficient. Notably, EGA-Glasso may fail to converge when sample size is insufficient in large factor models.
Computerized adaptive testing (CAT) aims to optimize measurement by tailoring item administration to individual examinees. The efficiency and precision of a CAT heavily depend on the choice of ability ( θ ) estimator and the termination criterion (stopping rule). Prior research suggests these components interact, but comprehensive evaluations across varying item bank characteristics remain limited. This simulation study investigated the interactive effects of four θ estimators (maximum likelihood [MLE], weighted likelihood [WLE], maximum a posteriori [MAP], and expected a posteriori [EAP]) and four termination criteria (fixed-length, standard error of measurement [SEM], minimum information [MI], and change-in-estimate [Δ θ ]) on measurement bias, precision (RMSE), and test length. These combinations were evaluated across low- (100-item) and high-information (500-item) item banks with both flat and peaked information distributions using the three-parameter logistic model. The results demonstrated that the optimal CAT configuration is contingent on item bank size and shape. Across all conditions, WLE emerged as the most robust estimator, effectively neutralizing the boundary estimation issues of MLE and the shrinkage bias characteristic of Bayesian estimators. In high-information banks, the SEM and fixed-length rules yielded the lowest conditional RMSE and bias regardless of bank shape. However, in low-information peaked banks, the strict SEM rule frequently failed to reach precision targets at the θ extremes, resulting in inefficient, maximum-length tests. Under these sparse conditions, the Δθ rule paired with WLE provided a superior balance of accuracy and efficiency by halting administration when precision gains stagnated. Conversely, the MI rule consistently exhibited the highest bias and RMSE. These findings underscore that optimal CAT design is not a one-size-fits-all solution. For high-quality banks, WLE paired with an SEM or fixed-length rule is recommended. For lower-quality banks, practitioners should adopt a Δ θ rule or a hybrid SEM approach to prevent inefficient test elongation.
Misreporting and other forms of aberrant responding can undermine the validity of survey-based inferences. Person-level evaluation of aberrant responses is rarely conducted because inspecting individual response patterns is time-intensive. This study proposes an integrated approach for identifying, classifying, and interpreting misfitting response patterns using nonparametric visualizations of person response functions combined with clustering of person response functions. The first step is to calibrate the survey items using an IRT model, such as the Rasch model, to establish an interpretable latent continuum with item-location ordering. Next, person-fit statistics, such as infit and outfit mean square error statistics, are examined, and a smaller subset of response patterns is flagged as misfitting. The third step is to use a nonparametric Hanning procedure to create person response functions, followed by clustering misfitting person response functions using Partitioning Around Medoids (PAM). The advantage of PAM over other clustering methods is that an observed response pattern is identified as a representative case for each cluster. Clusters can then be identified that correspond to an appropriate interpretation for the cluster, such as underreporting, inconsistent reporting, and overreporting patterns. Finally, decisions can be made about how to address aberrant person response patterns. The Household Food Security Survey Module from the U.S. Census is used as an illustration. These visualizations can support transparent data-quality evaluation with the potential for survey improvements.
Self-generated identification codes (SGICs) are commonly used in longitudinal research studies to anonymously link data collected from a participant across time points. However, the lack of research evaluating the methods of developing SGICs is concerning, given that low match rates of participants across time points are common, resulting in significant data loss over time and diminishing statistical power. The current study aims to fill this gap by evaluating the effectiveness and reliability of a newly developed SGIC. Participants (n = 135) completed an SGIC during two sessions, 8 weeks apart. Researchers informed the experimental group that the second survey included sensitive information and the control group that the study focused solely on the SGIC. Although no sensitive information was requested, this allowed analysis of how the perception that sensitive information was being collected might affect the consequent matching of the SGIC. On average, more than 91.9% (91.5% experimental group, 92.2% control group) of participants had at least 10 correct SGIC element matches. Upon examining perfect matches, only 40.0% of participants matched on all 12 elements across the two sessions. Multivariate logistic analyses revealed that the trust rating was a significant predictor of group membership with higher trust ratings being associated with control group membership (odds ratio [OR] = 0.541, 95% confidence interval [CI] = [0.32, 0.91], p = .021) and trustworthiness was a significant predictor of matching on all 12 SGIC elements (OR = 4.975, 95% CI = [1.70, 14.55], p = .003). This study demonstrates that SGICs can function effectively under the right conditions, but reliability depends on the structure of the code itself and on participants' perceptions of trustworthiness, anonymity, and engagement. As researchers continue to seek methods that balance participant privacy with data accuracy, the field must move toward empirically grounded standards.
Multivariate assessment profiles contain two conceptually distinct sources of variation: overall level and profile shape. Existing approaches recover some aspects of this structure, but none jointly establishes the replicability of latent pattern dimensions and provides a principled, population-referenced summary of person profile differentiation independent of overall level. We introduce the Aggregated Latent Profile Index (ALPI), a variance-weighted Euclidean distance that quantifies the degree to which an individual's profile departs from a flat population reference within a bootstrap-validated latent profile space. ALPI is derived from the Aggregated Latent Profile Space (ALPS), a framework that combines parallel analysis with bootstrap stability diagnostics - principal angles between subspaces and Tucker's congruence coefficients - to identify replicable latent pattern dimensions, then assembles the retained dimensions into a K-dimensional latent space through singular value decomposition, into which individuals and variables are jointly projected. In a simulation benchmark, ALPS recovered the known four-factor population structure as three replicable ipsatized pattern dimensions, confirming that the pipeline performs as intended when the true structure is known. In an application to normative WAIS-IV data (n = 900), ALPS identified a stable three-dimensional pattern space, and ALPI distinguished individuals with identical Full-Scale IQ - ranging from the 7th to the 85th percentile - on the basis of their ipsatized subtest configurations. ALPI provides assessment researchers and clinicians with a single, measurement-principled index of person profile differentiation that is grounded in replicable latent structure and independent of overall score level.
Anomalous survey responses, including random, careless, extreme, acquiescent, straightline, and alternating responding, threaten the validity of survey-based research. Machine learning (ML) algorithms offer flexible, model-agnostic alternatives to traditional detection methods, yet their relative effectiveness across anomaly types remains poorly understood. This study evaluated 11 unsupervised anomaly detection algorithms spanning four paradigms (distance-based, density-based, reconstruction-based, and tree/boundary-based) against six simulated anomaly types embedded in a realistic survey dataset (N = 3,000). Results revealed pronounced differential sensitivity: globally deviant patterns (random, extreme, alternating) were universally detectable, whereas careless and acquiescent responding required reconstruction- or boundary-based methods, and straightline responding resisted detection by all algorithms (maximum area under the receiver operating characteristic curve [AUC-ROC] < .70). No single algorithm dominated across all types. These findings argue for multimethod approaches combining ML algorithms with traditional response quality indicators, and provide a framework for selecting detection methods based on anticipated anomaly types.
In classical test theory, independence of measurement errors constitutes a central assumption when estimating the reliability of measures. Furthermore, it is well known that this assumption cannot be tested with standard methods that rely on second-order moments (variances, covariances). The present study, therefore, explores properties of non-Gaussian parallel and congeneric measures (i.e., when observed scores deviate from the Gaussian distribution) with respect to their capabilities of identifying violations of the error independence assumption. We show that, under non-Gaussianity and inequality of hidden confounding effects, third and fourth cumulant-based test statistics can be derived which enable researchers to detect non-independent error structures. We describe identifiability conditions under which the proposed test statistics can be expected to have adequate statistical power in parallel and congeneric measures and present results of Monte-Carlo simulation experiments. Results suggest that third-order tests adequately protect the nominal significance level. However, fourth-order tests can produce inflated Type I error rates, in particular, when error variances are unequal. In general, the power to detect non-independent errors increases with the sample size, the magnitude of non-Gaussianity, the degree of inequality of hidden confounding effects, and the degree of error non-independence. A real-world data example is presented for illustrative purposes. The current study presents important insights for developing statistical methods to detect error non-independence. However, these methods also rest on crucial assumptions, emphasizing the severity of the issue of statistically detecting non-independent measurement errors.
Simultaneous estimation of structural and measurement models in structural equation modeling (SEM) is not always tenable in small samples. In such cases, it may be necessary or advantageous to obtain scores. Current scoring recommendations draw predominantly from simulations with sample sizes greater than N = 200. This paper extends these recommendations to small N, directly comparing factor scores to sum scores. In addition, scores computed from an essentially tau-equivalent factor model are introduced as an alternative scoring option aimed at balancing the competing benefits of sum scores and factor scores. Findings largely suggest that factor scores from an essentially tau-equivalent factor model are advantageous when considering convergence and stability, even when not supported by the data, so long as departures from their assumptions are not substantial. They are obtainable when congeneric factor models fail to converge and have similar correlations with true scores compared with typical factor scores in samples at or less than N = 200.
Evidence of external validity based on individual score estimates is still relevant in many psychometric applications. From a model-based perspective, however, the topic appears to have been rather neglected in recent decades. Thus, in structural equation modelling (SEM), this evidence is sought to be obtained structurally, bypassing the scoring stage. And, in item response theory (IRT), the score interest mostly focuses on internal properties. Taking this state of affairs into account, this paper develops and proposes a model-based approach, intended for noncognitive measures, that combines SEM and IRT developments, and which allows a detailed assessment of the external validity of a class of score estimates to be carried out. The starting point is a general extended model that also includes the relevant external variables. From this general model, four well-known extended IRT models can be derived and fitted at the structural level. Next, on the basis of the structural results, a series of unconditional (population-dependent) and conditional (population-independent) indices that describe the model-implied relation between the score estimates and each external variable are developed and proposed. The practical relevance of the proposal is discussed mainly around three applications: assessing model appropriateness, obtaining point and interval prediction estimates at the individual level, and shortening a test while optimizing the external validity of the resulting version. The functioning of the proposal is illustrated using a real-data example.
Structural equation modeling (SEM) is widely used in educational and behavioral research, but applied SEM often involves simultaneous tests of many structural paths. When many coefficients are evaluated at nominal thresholds, the probability of false positives and the expected number of false discoveries can be substantial even when global fit indices indicate close fit, encouraging substantive interpretation of chance findings. Building on prior work on multiplicity control in SEM, this article presents a practical workflow for false discovery rate (FDR) adjustment of families of SEM parameter tests obtained from fitted lavaan model objects, including the dependence-robust Benjamini-Yekutieli (BY) procedure, and provides an R implementation to support routine use. In a Monte Carlo study (1,000 replications; N = 500) with nine latent factors, a correctly specified measurement model, and an overspecified structural model with 33 candidate regressions (8 non-zero), nominal p < .05 produced at least one false positive in 69.3% of samples and a mean of 1.182 false-positive paths. BY adjustment reduced the mean number of false positives to 0.073, while the mean number of detected true effects declined from 6.358 to 5.857. A sensitivity analysis across three dependency conditions indicated that BY-FDR was more robust to the direction and magnitude of parameter dependence, whereas BH's false-positive control weakened under negative dependence. These results suggest that dependence-robust FDR adjustment can be integrated into a standard SEM workflow with lavaan in R, and may substantially reduce false positives with a modest reduction in detected true effects.
Advances in large language models can provide opportunities to evaluate the characteristics of scales prior to data collection. In this study, we explore if item text can be used to predict a scale's psychometric properties. Specifically, we examine if clustering consensus (i.e., the frequency by which items are grouped with other items from the same underlying factor across multiple clustering algorithms), and a cosine similarity metric (i.e., the semantic similarity of items to other items from the same factor), can be used to predict exploratory factor analysis (EFA) factor loadings. Across six scales with varying sample sizes, number of factors/items, we found that both the cosine similarity and ensemble clustering consensus methods predicted factor loading values. While the methods share some conceptual and empirical overlap, and results vary by scale, the ensemble clustering approach explains incremental variance above and beyond cosine similarity in predicting factor loadings. Using both methods in conjunction can be a useful way to identify problematic items prior to data collection and help researchers develop more optimal scales from the onset, thereby potentially saving time, resources, and increasing the likelihood of developing sound measures.
Extreme response style (ERS), the tendency of participants to endorse the extreme categories of an item partially independent of item content, has repeatedly been found to decrease the validity of Likert-type scale results. For this reason, many IRT models have been developed that attempt to detect and correct for ERS. Despite the substantive literature on ERS and modeling of ERS, several important questions remain. To date, there is no clear estimate of how often ERS occurs in practice across a variety of scales and populations. In addition, there is little guidance on what item parameters for ERS models are commonly found in empirical data, while this information is crucial to inform future methodological studies utilizing ERS models. Finally, there is only limited information available on which ERS models tend to fit the data best. The current study sets out to address these three issues by analyzing data from the Programme for International Student Assessment using a generalized partial credit model, several multidimensional nominal response models, and several IRTree models. Results indicate an extremely high prevalence of ERS across scales, populations, and timepoints. Item parameters for future methodological studies are presented, and a general preference for IRTree models over MNRM models is found in many datasets. Implications for futures studies are discussed, and recommendations for practice are made.