
Differential item functioning (DIF) analysis is essential for evaluating measurement invariance in educational and psychological assessments. In cognitive diagnostic assessment, however, most existing methods require prespecified comparison subgroups and anchor items. When subgroup membership and anchor items are unavailable or mis-specified, DIF detection and parameter estimation may be biased. To overcome these limitations, this study puts forward a DIF detection method that incorporates an extended modelling framework and a two-stage estimation algorithm. The proposed modelling framework directly integrates DIF parameters into the measurement model and uses a structural model to characterize subgroup differences in attribute mastery distributions. A two-stage expectation maximization algorithm with an adaptive lasso penalty is developed to identify anchor items, classify respondents into latent subgroups and estimate model parameters. The performance of the proposed method was evaluated through a simulation study and an empirical data analysis. Simulation results indicated generally satisfactory DIF detection and parameter recovery, although subgroup-classification accuracy varied across conditions. When applied to the empirical data, the proposed method identified 10 of the 28 items as exhibiting DIF.
Sensitivity analysis for unmeasured confounders is essential for assessing the robustness of causal mediation conclusions. Most existing methods rely on parametric assumptions, which are ill-suited for machine learning-based estimators that are not tied to specific parametric models. This study develops a sensitivity analysis method for mediation analyses based on the debiased machine learning approach. The proposed method can accommodate categorical (e.g. binary) or continuous mediators and outcomes, quantify robustness to unmeasured pretreatment confounders through contour plots and robustness values defined by R 2 -type measures, and allow researchers to avoid parametric assumptions by incorporating data-adaptive machine learning methods. Simulation studies are conducted to evaluate the performance of the proposed method. An empirical example is provided to illustrate the application. We hope this study provides a novel method for quantifying the sensitivity to unmeasured confounding in causal evaluations of mediation effects.
The logistic positive exponent (LPE) and its reflection (RLPE) models accommodate asymmetric item characteristic curves in item response theory, offering greater flexibility than traditional symmetric specifications. While several asymmetric IRT models exist, the LPE and RLPE framework provides a compelling balance of theoretical foundation, interpretability, and computational tractability. Despite their appeal, key aspects of Bayesian estimation for these models remain understudied, including systematic comparison of prior specifications for the asymmetry parameter, model selection performance relative to symmetric alternatives, and application of modern convergence diagnostics. This study provides a comprehensive methodological investigation of Bayesian estimation for both LPE and RLPE models. We make four primary contributions: First, we conduct extensive simulation studies examining parameter recovery, computational efficiency, and sensitivity to alternative prior specifications for the asymmetry parameter. Second, we present the first systematic evaluation of model selection performance for discriminating between asymmetric and symmetric IRT models. Third, we provide theoretical comparisons with alternative asymmetric IRT approaches, clarifying when LPE/RLPE models offer advantages and when alternatives may be preferable. Fourth, an empirical application to mathematics assessment data demonstrates the practical utility of the framework, with the RLPE model clearly outperforming symmetric alternatives and revealing meaningful asymmetry patterns across items.
Traditional item response theory (IRT) models involve symmetric response probability functions, but a developing area of research has focused on asymmetric alternatives. To date, these studies have focused primarily on introducing new asymmetric models, detailing their mathematical formulations, explaining how to interpret their item parameters and demonstrating their distinctive psychometric properties through simulations and empirical data analysis. We add to this burgeoning research area by exploring model evaluation in the context of asymmetric IRT. We begin by introducing a new asymmetric IRT model based on the Aranda-Ordaz (Biometrika, 68, 1981, 357) link function and comparing it to two established models that parameterize asymmetry in distinct ways. Our comparison of these models then allows us to demonstrate three uncommon methods of IRT model evaluation: (1) characterizing asymmetry in the response probability and pseudo-information functions, (2) quantifying configural complexity through fitting propensity analysis and (3) exploring the pattern signatures of each model. Through this in-depth evaluation, we find that all three models possess subtly unique psychometric properties that will benefit applied research. More generally, the model evaluation methods that we consider in this work can offer unique insights into any IRT model.
Standardized mean differences (SMDs) are widely used to quantify treatment effects in cluster-randomized trials. However, covariate adjustment in hierarchical linear models reduces the residual variance components used for standardization, which artificially inflates effect size estimates and undermines comparability across studies. We propose a unified family of estimators that recover the unadjusted variance components by rescaling the covariate-adjusted variance components using pseudo- R 2 $$ {R}^2 $$ indices. This rescaling places effect size estimates on a common reference scale, thereby improving comparability across studies and model specifications under standard modeling assumptions. The framework accommodates three covariate adjustment scenarios including level-1, level-2, and simultaneous both-level adjustments. Furthermore, it introduces three estimator types spanning method of moments, maximum likelihood, and a t-statistic reformulation suitable for meta-analysis from published summaries, alongside delta-method variance approximations for each. An empirical example and a simulation study apply the proposed covariate-adjusted SMDs across these scenarios to illustrate their implementation and demonstrate the consequences of omitting the correction.
This article revisits the concepts of reliability and measurement precision across classical test theory (CTT), item response theory (IRT) and cognitive diagnosis models (CDMs), emphasizing their conceptual differences and proposing a unified framework. We show how continuous score estimators in CTT and IRT relate to mastery probabilities in CDMs through integration over proficiency estimator distributions relative to defined thresholds. Building on this link, we introduce reliability indices grounded in the coefficient of determination ( R 2 $$ {R}^2 $$ ) as a common measure of association between true scores and estimated true scores, applicable to both continuous proficiencies and discrete classifications. This unified perspective clarifies how reliability can be consistently estimated and reported across frameworks, promoting coherence in psychometric practice and supporting more transparent interpretation of test results. Ultimately, this study aims to clarify reliability concepts, facilitating a unified perspective and aiming to promote consistent reliability reporting across diverse psychometric applications.
When high-stakes decisions depend on test scores, it is natural to ask whether examinees' relative standing is determined by the response evidence or by the scoring rule. This paper characterizes when four IRT proficiency estimators-maximum likelihood, maximum a posteriori, expected a posteriori, and Warm's weighted likelihood estimator-can change the rank ordering of examinees under the 1PL (Rasch), 2PL, and 3PL models with fixed item parameters. We show that, under the 1PL and 2PL, and restricting attention to response patterns for which the relevant estimators are well-defined, these four estimators agree on every strict comparison implied by the model's evidence order. The reason is structural: with fixed item parameters, the person likelihood in both the 1PL and 2PL depends on the response pattern through a single evidence statistic, yielding a monotone likelihood-ratio order in that statistic. Under the stated conditions, MLE, EAP, MAP, and WLE all increase in this same statistic. By contrast, the 3PL does not retain this same single-statistic structure in the examinee latent trait ( θ $$ \theta $$ ), so global rank invariance is not guaranteed without additional restrictions. We identify suborders where invariance still holds and clarify where disagreements are possible, along with the implications for measurement specialists and policymakers.
We discuss two approaches to estimating the true score from classical test theory, each with a corresponding measure of uncertainty due to measurement error: the classical method with the standard error of measurement (SEM) and Kelley's method with the standard error of estimation (SEE). For both approaches, we examined the bias and sampling variability of the true-score, SEM and SEE estimators, as these properties were largely unknown. For each estimator, we derived analytic expressions for the bias and proposed approximations for practical bias assessment by omitting cumbersome terms. We also proposed a standard error estimator for each estimator. Simulations were used to assess the accuracy of the bias approximations and standard error estimators and to examine the impact of bias and sampling variability on both true-score estimation methods. Results indicate that both the bias approximations and standard error estimators are sufficiently accurate. Whereas the classical true-score estimator is unbiased, Kelley's estimator may exhibit substantial bias for extreme true scores. The bias of the estimated SEM and SEE may also be substantial when reliability is high and the reliability estimator is negatively biased. The impact of sampling variability is typically small.
Categorical structural equation models (cat-SEM) typically rely on tetrachoric/polychoric correlations under a latent multivariate normality assumption. This can be generalized to an elliptical latent trait, whose radial symmetry justifies treating a single correlation parameter as the target of estimation. This article makes two contributions. First, it introduces an elliptical sieve estimator that profiles the latent correlation over a non-parametric radial scale mixture. Simulations under a variety of latent elliptical densities show reduced pseudo-likelihood bias and stable performance across realistic threshold schemes (including skewed floor/ceiling patterns), category numbers (3-5) and sample sizes typical of cat-SEM. Second, an ordinal tail asymmetry test is proposed that uses checkerboard-copula tail contrasts and a randomization reference distribution to diagnose violations of latent radial symmetry. For items with 4-5 categories and symmetric thresholds, or thresholds skewed in the same direction, the test maintains near-nominal Type I error under elliptical densities, and achieves high power against non-elliptical copulas. However, with binary items and with three-category items whose thresholds are strongly asymmetric and oriented in opposite directions, Type I error inflates even under ellipticity. In such coarse, highly unbalanced designs, fully parametric latent-copula models remain preferable to semiparametric ordinal diagnostics.
A model for multiple-choice (MC) items based on signal detection theory (SDT), the MC-SDT model (DeCarlo, 2021a), follows from assumptions about perceptual and decision processes involved when examinees choose alternatives for MC items. The model can be expressed as a hierarchical model with an 'item-level', Level 1, and an 'examinee-level', Level 2. Here it is shown that cognitive diagnosis models (CDMs) can also be viewed as consisting of two levels, with the MC-SDT model serving as the first-level model, whereas the second-level model determines the type of CDM. Thus, the theory about how examinees make choices for MC items is unified across different CDMs. The resulting MC-SDT-CDM models are shown to be parsimonious sub-models of MC-CDMs. The models have straightforward interpretations (MC-SDT at Level 1), avoid estimation problems, and are useful for small sample sizes. The models are illustrated with an application to MC items from the TIMSS 2007 4th grade exam.
Latent variable item response and structural equation models are widely used to model constructs and acknowledge measurement error in research settings and operational assessments. Such work often proceeds in stages, where the results of an analysis from an earlier stage are fed into the analysis at a later stage. However, common practices in single-stage and multistage estimation have weaknesses, including when viewed from a Bayesian perspective. This work extends recent developments in 2-Stage Bayesian approaches in structural equation modelling (SEM) to advance a general multistage approach where models are viewed as comprised of fragments that can be assembled in a modular way. The proposed approach is more in line with Bayesian principles and offers advantages over existing approaches. The approach is illustrated by applications in several different modelling scenarios, including: SEM for a wider class of situations than existing approaches have considered; and calibration and scoring situations encountered in operational assessment using item response theory (IRT). R functions for executing these analyses by interfacing with Mplus are provided and documented, as is code for running the examples.
Unidimensional item response theory (IRT) models are widely used even in settings where assessment data exhibit subtle forms of multidimensionality. Recent empirical evidence suggests that when item difficulty is associated with dimensionality, asymmetric item characteristic curves (ICCs) emerge in the unidimensional approximation. Through theoretical derivation and extensive simulation, this paper develops a framework for understanding the emergence and the degree of ICC asymmetry in UIRT models when applied to multidimensional data with difficulty-dimensionality associations. An empirical analysis of the Virginia Language & Literacy Screener (VALLSS) confirms the predicted patterns. These results highlight ICC asymmetry as an anticipated consequence of unidimensional approximation and underscore its relevance for applications such as vertical scaling and test linking.
Although full-information maximum likelihood (FIML) estimation is widely used for diagnostic classification models (DCMs), its computational efficiency deteriorates sharply in high-dimensional settings. This scalability challenge is increasingly critical as DCMs are applied to large-scale assessments, psychological testing and longitudinal studies involving many attributes. We propose a composite marginal likelihood (CML) estimation approach via expectation-maximization (EM) algorithm (CML-EM) for higher-order DCMs (HO-DCMs) as an alternative. The central premise is that, because response probabilities depend only on the attributes specified by the Q-matrix and because of the conditional independence assumption of HO-DCMs, the full likelihood can be partitioned into low-dimensional subsets of items and attributes. This reduces both attribute and response spaces in the E-step to that of each subset, resulting in substantial computational gains that become more pronounced with larger sample sizes and numbers of attributes. We also introduce a subset-construction procedure that ensures both efficiency and feasibility of CML-EM and present two methods for attribute classification. Simulation results demonstrate that CML-EM is significantly faster than FIML while maintaining accurate parameter recovery and acceptable classification performance. The practical utility of the method is further illustrated through an empirical application to a high-dimensional personality assessment.
The social relations model (SRM) is commonly used in psychological research to analyse interdependent data from round-robin designs, where all members of a group rate each other. Based on the recently suggested social relations confirmatory factor analysis (SR-CFA), we present general formulas for determining the reliability of composites of round-robin judgments and also derive simpler variants when specific restrictions are applied to the parameters of the SR-CFA model. In the unidimensional case, this results in an omega-type reliability measure, which can be converted into an alpha-type reliability measure through further restrictions. We also discuss how standard errors of the reliability coefficients can be obtained, illustrate the suggested methods using an empirical example, and we examine the suitability of the estimation approach in a small simulation study. Finally, we discuss questions for future methodological research and how the person-level composites are related to other SRM effect estimates proposed in the SRM literature.
Sinharay (Psychometrika, 2016, 81, 992) suggested the asymptotically correct standardized version of a class of person-fit statistics for mixed-format tests. This paper provides an alternative and arguably simpler derivation of the standardization. The derivation leads to several benefits including a simpler formula of the (asymptotically correct) standardized person-fit statistics and a theoretical explanation of simulation results reported in literature on person-fit statistics. This paper promises to make several person-fit statistics for mixed-format tests more accessible to researchers and practitioners.
Study design and subsequent data analysis is ideally a collaborative endeavour between applied researchers and statistical experts (e.g. methodologists, data scientists). Applied researchers know what research questions need to be answered and statistical experts, in conjunction with the applied researchers, plan study details (e.g. sample size, measurement instruments) and the analytic approach to best answer the research questions. Research questions regarding change processes are especially challenging as there are more study design decisions and more analytic options to consider. Moreover, the optimal analytic approach may not be feasible given the available longitudinal data. Additionally, many studies of change utilize secondary data, where decisions regarding the measures, the number and timing of assessments and sample size were not planned for the current research questions. In these cases, longitudinal model development is often a compromise between what is ideal and what is feasible given the available data. In this paper, we discuss two empirical projects and how collaboration between applied researchers and developmental methodologists informed the analytic models.
Asymmetric item response theory (asymIRT) has emerged as an important extension of classical IRT, motivated by empirical evidence and theoretical arguments that symmetric item response functions (IRFs) often inadequately describe real response processes. Despite rapid model development, there remains ambiguity regarding what constitutes asymmetry, how different models relate to one another, and how asymmetry should be quantified. This paper provides a unified framework for defining, interpreting, and measuring asymmetry in IRT models. Refining Samejima's notion of point symmetry, we propose general definitions of IRF symmetry based on properties of the first derivative of the IRF. These definitions clarify the status of various models, including the 3PL, unipolar models, and recently proposed asymmetric functions. We further introduce quantile-based measures of skewness as convenient indices of the magnitude and direction of item asymmetry and demonstrate how these measures behave across several asymmetric models. Through analytic results and numerical illustrations, we show that asymmetry has meaningful consequences for latent trait estimation, particularly in how items penalize or reward responses at different trait levels. This work positions asymmetry as a fundamental item characteristic, alongside difficulty and discrimination, and provides practical tools for comparing asymmetric IRT models and understanding their substantive implications.
In the Bayesian Graphical Modeling framework, priors on network structure encode theoretical assumptions and uncertainty about the topology of psychological constructs under study. For instance, the Bernoulli prior specifies the probability of each pairwise interaction, the Beta–Bernoulli prior governs expected network density, and the Stochastic Block prior models clustering. In practice, however, specifying informed hyperparameters is challenging: theoretical guidance is limited, and default choices can be overly simplistic or restrictive. To address this, we introduce an LLM-based prior elicitation framework in which a large language model provides inclusion judgments for each variable pair. These judgments are converted into edge-specific prior probabilities for the Bernoulli prior and used to derive hyperparameters for the Beta–Bernoulli and Stochastic Block priors. To make the approach accessible, we provide an R package, bgmElicit, with a Shiny app implementing the methodology. We illustrate the framework in two examples. First, a validation on a subset of a PTSD network from a meta-analysis compares OpenAI GPT models across several conditions. Second, an empirical analysis of 17 PTSD symptoms shows that elicited priors can modestly strengthen evidence regarding edge presence and absence. Taken together, this work is a proof of concept, complementary to expert judgment and prior sensitivity checks.
In a recent review, Liu et al. (Psychological Methods, 2025b) classified reliability coefficients into two types: classical test theory (CTT) reliability and proportional reduction in mean squared error (PRMSE). This article focuses on quantifying the sampling variability of these coefficients under item response theory (IRT) models. While some existing standard error (SE) formulas are accurate when variability arises only from item parameter estimation, the reliability estimators considered in our work involve additional variability from substituting population moments with sample moments. We propose a general strategy to derive SEs that incorporates both sources of sampling error simultaneously, enabling the estimation of model-based reliability coefficients and their SEs in such settings. We then apply our general theory to derive SEs for two specific estimators under the graded response model: (1) CTT reliability for the expected a posteriori score of the latent variable and (2) PRMSE for the latent variable. Simulation results show that the derived SEs accurately capture the sampling variability across various test lengths in moderate to large samples. We conclude with an empirical illustration and directions for future research.