
Detecting and interpreting differential item functioning (DIF) is critical for ensuring the fairness and validity of measurement instruments. Combinations of psychometric models with data-driven machine learning (ML) approaches have been introduced to facilitate comprehensive DIF analysis across diverse sets of covariates. We review two ML methods for detecting uniform DIF effects (i.e. variation of item difficulty parameters). Rasch trees explore DIF effects with nonparametric decision trees and divide the dataset into subgroups. Regularized moderated Rasch models specify a parametric model for predicting DIF effects. A simulation study illustrates that the optimization strategy and functional form assumptions can result in different conclusions. Finally, we demonstrate the implementation of the methods in practice, using data ( N = 5,193) from the National Educational Panel Study, and provide documented analysis code.
Nonignorable item nonresponses commonly occur in educational and psychological measurement and pose challenges for statistical inference in item response theory (IRT) models. In multidimensional IRT (MIRT), a key issue is identifying the relationships between multiple latent abilities and test items, known as latent variable selection. However, the latent variable selection in MIRT models with nonresponses remains largely unexplored. A common strategy for modeling omissions introduces a latent propensity variable to capture individuals’ tendency to omit items, leading to a joint MIRT model that combines a Rasch model for missingness and a multidimensional two-parameter logistic model for responses. Existing latent variable selection methods are developed mainly for fully observed response data and are not directly applicable when omissions are incorporated through this joint MIRT model. In this article, we develop an efficient expectation model selection (EMS) algorithm for MIRT models with omitted items, termed EMS-OI. Simulation studies show that EMS-OI performs competitively in both latent variable selection and parameter estimation, compared with existing EMS- and expectation maximization-based alternatives. An application to the PISA 2012 dataset further illustrates the proposed EMS-OI algorithm.
In most psychological tests using Likert-type scales, graded response model (GRM) is frequently used to describe the latent traits of subjects. However, many psychological constructs, such as dysthymia disorder, are typically non-normally distributed in a general population. Thus, GRM incorporated with Ramsay curves – graded response model (RC-GRM) allowing for flexible latent trait distributions is proposed in this article. In addition, an estimation procedure using metropolis–hastings (MH) sampling method for the RC-GRM model is given to estimate the item parameters in GRM and the shape parameter in the RC simultaneously. The algorithm is easy to apply and not restricted to the close form of parameter estimate. Four simulation studies reveal that the recovery results of this MH algorithm are accurate and the proposed algorithm is not sensitive to the chosen prior density distributions and specifications of knots and degree of B-spline functions. Moreover, RC-GRM has an obvious advantage to the traditional GRM when the true latent trait is non-normal according to estimation accuracy. The application to a real data set from the Programme for International Student Assessment (PISA) 2015 tests demonstrates that the proposed model and algorithm are effective in some real circumstances.
This study explores the integration of multiple-choice, multiple-attempt test items within the computerized adaptive testing (CAT) framework, named as MM-CAT. Using the sequential item response theory model for multiple-choice, multiple-attempt items, a simulation study was conducted to investigate the effectiveness of a MM-CAT design in improving ability estimation accuracy compared to traditional CAT, which relies on single-attempt, dichotomously scored items. Results show that MM-CAT substantially reduces the standard error of measurement, bias and root mean square error, particularly for examinees with lower ability levels. Furthermore, we examine the impact of item exposure control procedures and find that while both the Sympson-and-Hetter method (SH) and the Randomesque method are useful, the SH method is particularly effective in exposure control when paired with MM-CAT, minimizing the severity of over-exposed items without sacrificing the measurement precision. Taken together, these findings suggest that MM-CAT is a promising approach for enhancing the precision and fairness of adaptive testing, especially in educational contexts where multiple attempts may support both assessment and learning.
Recently, a new Bayesian model assessment criterion () has been proposed to separately assess the contributions of different sources of data in the joint model. In order to evaluate the performance of with a more complex response time model by jointly modeling with accuracy, we develop efficient computational algorithms to calculate the assessment criterion based on the decomposition of the deviance information criterion and the logarithm of the pseudo-marginal likelihood under the joint IRT and generalized odds-rate hazards model for accuracy and response time data. Extensive simulation studies are conducted to examine the empirical performance of the proposed methodology, and a detailed analysis of empirical data is carried out to demonstrate the usefulness of the assessment criterion.
Sensitivity analyses can inform evidence-based policy by quantifying the hypothetical conditions necessary to change an inference. Perhaps the most prevalent index used for sensitivity analyses is Oster’s Coefficient of Proportionality (COP) which expresses how strong selection on unobserved covariates would have to be relative to selection on observed covariates to nullify an estimated effect. But Oster has been critiqued based on its two-stage conceptualization of the COP and its estimation of the COP based on coefficient stability across estimated models. In this article, we reconceptualize the COP as a function of unobserved covariates’ correlations with the focal predictor (e.g., treatment) and with the outcome. Our correlation-based approach addresses the critiques of Oster while preserving the comparison of selection on unobserved covariates to selection on observed covariates. As importantly, our expressions do not depend on analysts’ subjective choices of covariates to include in a baseline model, are adapted to a threshold for inference based on statistical significance, and can be directly calculated from conventionally reported quantities (e.g., estimated effect, standard error) through the Konfound packages in R or Stata or the R-shiny app https://konfound-project.shinyapps.io/konfound-it/ . Thus, for most published studies in the social sciences our correlation-based COP index can be easily applied and intuitively interpreted.
Item response theory models have been extended by including guessing parameters to represent the response tendencies of individuals with low trait levels. Particular cases of these models arise when the guessing parameters are fixed at boundary values. For example, the two-parameter logistic model is derived from the three-parameter logistic model by fixing the guessing parameter to zero. However, the likelihood ratio test statistic for nested pairs of models asymptotically fails to follow a chi-square distribution when some parameters take boundary values, resulting in an unreliable p-value. This article proposes two solutions to this problem: a chi-square distribution with estimated degrees of freedom and a score test to evaluate fit at the item level. Guessing parameters are incorporated into the nominal categories model, and models with and without guessing are compared using the proposed chi-square statistic. The special case of the comparison between the three-parameter logistic and two-parameter logistic models is also considered. A simulation study was conducted to evaluate the performance of the proposed chi-square statistic with dichotomous and polytomous data. The results indicate that it is a viable alternative for detecting significant guessing effects. The solution is further illustrated with an empirical example to demonstrate its practical application.
Halo effects refer to a persistent rater error posing a significant threat to the validity and fairness of assessments involving human raters. This research introduces a novel approach to analyzing halo effects by considering rating times recorded automatically in technology-based, online, or onscreen assessments as an additional data source. For this purpose, we propose the mixture Rasch facets model for halo with rating time. Utilizing Bayesian parameter estimation methods, we applied the model to a real dataset from a large-scale Chinese writing assessment. We found that rating time predicted illusory halo, such that longer rating times were associated with a greater likelihood of observing halo effects. Compared to traditional models, considering rating time as an additional variable resulted in better data-model fit and preserved the integrity of the latent scale, maintaining the informative value each rating criterion conveyed regarding the examinees' performances. A simulation study confirmed the model's superior parameter recovery under different levels of impact rating time had on the likelihood of halo error. Findings suggest that integrating rating time into measurement models can enhance the detection of illusory halo effects, enrich research on the psychological mechanisms underlying these effects, and improve the overall psychometric quality of assessments. The discussion focuses on implications for model development and future research into rater effects.
Cognitive diagnostic models (CDMs) are widely applied in educational and psychological measurement, and variational Bayes (VB) algorithms have recently gained attention for parameter estimation. In this paper, we propose a novel VB-based algorithm, termed Variational Bayes via P & oacute;lya-Gamma (VB-PG), which leverages the P & oacute;lya-Gamma data augmentation technique. The method is developed for the reparameterized deterministic inputs, noisy "and" gate (DINA) model, providing an efficient and accurate framework for parameter estimation. By introducing P & oacute;lya-Gamma data augmentation method, VB-PG algorithm achieves a streamlined implementation with fully closed-form updates for all unknown variables. Simulation studies show that the VB-PG algorithm outperforms both existing VB methods, the expectation-maximization (EM) algorithm, and the Markov chain Monte Carlo (MCMC) algorithm, yielding higher estimation accuracy, particularly for slope parameters under small samples, and superior computational efficiency. An empirical example analysis further confirms its robustness and practical advantages, demonstrating that the VB-PG algorithm provides a simpler, faster, and more accurate solution for parameter estimation in the DINA model.
Social network relationships are often not directly observable or easily measured. Reliance on a single survey item to capture such relationships is likely to introduce construct measurement error. Although latent constructs and latent variable measurement models are widely used in the social sciences, they are rarely applied to the measurement of social network ties and, to our knowledge, have not been used to account for construct measurement error in network selection models. To address this gap, we propose an item response theory-based latent space model (IRT-LSM) that employs a multi-item scale measurement model to more accurately measure social relationships, which are then predicted using a latent space model. In addition to introducing the model, we present a simulation study demonstrating its capabilities and improvements over existing approaches. Finally, an empirical data analysis illustrates how substantive conclusions may vary depending on the modeling approach used.
Nonlinear random effects models (NREMs) are particularly useful for modeling longitudinal data that follow intrinsically nonlinear trends. However, NREMs assume both random effects and random errors to be normally distributed, which is likely violated when the outcome variable does not appear normal. No existing software under the frequentist framework provides researchers the flexibility to specify non-normal random effects to date. To address this significant gap in the application and research of NREMs, we developed an R routine NEfit.R that allows the specification of non-normal random effects and random errors using Gauss-Hermite Quadrature approximation and Newton-Raphson iteration. This study includes a motivating real-data application and a comprehensive Monte Carlo simulation study that supports the robustness of the performance of NEfit.R.
Assessment based on fine-grained latent traits can provide more detailed information about the subjects. Cognitive diagnostic assessment (CDA) is a framework of education and psychological measurement that is grounded in the assessment of fine-grained latent traits. The Q -matrix, defining the relationship between items and attributes, is the basis for CDA. Data-driven Q -matrix estimation has become a research hotspot due to its high efficiency and objectivity. However, the existing method for Q -matrix estimation primarily focuses on scenarios with low or middle correlations between attributes (latent traits), and they point out that the accuracy of Q -matrix estimation significantly declines in situations with high attribute correlations. To address the limitations of existing methods in scenarios with high attribute correlation. This paper proposes a majority-class symmetric undersampling (MCSU) method tailored for CDA. To evaluate its performance, two simulation studies are conducted. The simulation results under a wide variety of conditions show that the MCSU can improve the estimation accuracy of attribute number and Q -matrix in highly correlated scenarios. Finally, a real dataset of the Examination for the Certificate of Proficiency in English is analyzed using the proposed method.
Various Bayesian regularization methods have been incorporated into cognitive diagnosis models to enhance Q-matrix inference. The purpose of implementing regularization is to identify significant parameters while penalizing insignificant ones toward zero, thereby reducing model complexity and promoting a sparser solution. In this study, we aim to: (a) formulate the partially confirmatory cognitive diagnosis modeling (PCCDM) framework under different link functions (probit and logit); (b) compare Bayesian regularization methods for Q-matrix inference within the PCCDM framework; and (c) evaluate the effectiveness of the cut-off-based approach and the Wald-test-based approach for determining the significance of q-entries. The investigation of penalty included spike-and-slab and Least Absolute Shrinkage and Selection Operator. Implemented by Markov chain Monte Carlo estimation using Gibbs sampler, the results were demonstrated from both simulated and real data.
Often in educational and behavioral randomized trials, researchers discover that subgroups of individuals respond differently to the same intervention due to underlying individual differences. Previous studies have investigated such heterogeneous treatment effects (HTE) across different latent classes for both cross-sectional and longitudinal (repeated measures) outcomes, but not in the context of longitudinal mediation analysis. This study develops Bayesian (nonlinear) growth mixture mediation models, B(N)GMMMs, to assess the HTE of the intervention variable X on the longitudinal dependent variable Y, via the longitudinal mediator M. We consider Y following linear or nonlinear trajectories and incorporate class predictive and growth predictive covariates. Monte Carlo simulations were conducted to evaluate model performance, and an empirical example demonstrates the model's application.
Evaluating and comparing models with respect to their predictive performance is a cornerstone of Bayesian statistics. Two related and important techniques are leave-one-out cross-validation and stacking. Both quantify and set in relation the ability to predict unseen observational units from the same data-generating process for a set of models. Recent advancements in software development-in particular, the Stan modeling framework-have made it possible to apply these techniques easily to a wide range of models. However, in more complex models, such as the widely applied classes of hierarchical models and mixture models, the choice of observational unit is not trivial and can result in the need for numerical integration, in particular in non-normal models. We present a case study of Bayesian mixture item response models for cross-classified multirater data, where the most parsimonious choice of observational unit required two-dimensional integration. We show that implementing a numerical quadrature scheme directly within the Stan model code, which is available as Supplemental Data, allows for efficient and accurate estimation of predictive performance.
While regression discontinuity designs (RDDs) offer credible identification of treatment effects near cutoff points, their external validity remains inherently limited. To address this, a conditional independence assumption may be required. However, in multiple-score RDDs, this requires joint independence between all running variables and both potential outcomes and treatment assignments-a restriction often too strong to hold in practice. We relax this by assuming only mean independence between potential outcomes and the interaction of running variables, conditional on covariates. This weaker assumption allows us to link treatment effects at boundary and non-boundary points. We propose a non-parametric testing procedure and illustrate our method using simulation experiments and an empirical application in Colombian education policy, respectively.
When studying latent outcomes in causal inference, causal effects can occur not only on the mean of the latent variable but also on all parts of the measurement model: the mean and residual variance of the latent variable, as well as the item parameters. This article proposes causal parameter moderation, the application of moderated nonlinear factor analysis (MNLFA) to causal inference with latent outcomes. The proposed approach can handle an arbitrary number of continuous or categorical covariates, making it well-suited to handling heterogenous treatment effects. Causal parameter moderation is compared to item-level heterogenous treatment effect (IL-HTE) models. The models are best suited to handle different types of differential item functioning (DIF), and thus the circumstances under which each is appropriate differ: IL-HTE models assume a normal distribution of DIF across all items, so they do not substantially correct bias introduced by a subset of items being extra-sensitive to treatment. The proposed MNLFA approach is shown to handle this scenario more effectively. The approach is demonstrated in a reanalysis of data from a randomized controlled trial of a reading intervention's effects on science vocabulary, where there is substantial evidence of both a main effect on the latent variable and additional effects on a subset of items.
In educational assessments, testlet-based designs are widely used but often violate the local independence assumption of item response theory (IRT). This work proposes a flexible Bayesian approach for modeling local item dependence (LID) in testlet data, using a multivariate two-parameter probit IRT model with a structured correlation matrix defined via antedependence models. By incorporating Toeplitz structures, the model captures nuanced within-testlet dependencies. We implement an efficient posterior sampling scheme using the No-U-Turn Sampler via Stan package. A simulation study shows accurate parameter recovery and reduced bias in the presence of LID. In addition, we provide applications to real educational datasets, including large-scale assessments in Brazil, in which we show that our approach yields more interpretable and stable parameter estimates compared to standard IRT models.
Too often young researchers choose topics to work on for the topic's convenience rather than its importance. In earlier work we have laid out in broad strokes some areas that would reward progress, but in this essay, we provide guidance toward topic choice that is much more specific, focusing on a fundamental irony of item response theory (IRT). IRT is the name given to a family of models that provide a stochastic description of what happens when a person meets a test item. In general, models are most useful when we can lean on them in areas where data are sparse. Strong models are required when data are very limited, weaker models are possible when there are more data. The recommended size of data samples required to fit IRT models has historically seemed excessively large, yielding the apparent irony that IRT can only be used when data samples are so large that no model is needed. Bayesian work over the past decade or so has revealed that with modern estimation methods, IRT models can be successfully fit with surprisingly modest samples. One practical consequence of expanding the use of IRT to such samples is that it makes possible its use on classroom tests. This can have profound positive implications. In this essay, we describe the apparent irony of IRT and then outline a two-part research project that may resolve matters, thus making the potential provided by IRT models available to a much wider range of applications.
Sometimes a treatment, such as receiving a high school diploma, is assigned to students if their scores on two inputs (e.g., math and English test scores) are above established cutoffs. This forms a multidimensional regression discontinuity design (RDD) where there are two running variables instead of one. Present methods for estimating such designs either collapse the two running variables into a single running variable, estimate two separate one-dimensional RDDs, or jointly model the entire response surface. The first two approaches may lose valuable information, while the third approach can be very sensitive to model misspecification. We examine an alternative approach, developed in the context of geographic RDDs, which uses Gaussian processes to flexibly model the response surfaces and estimate the impact of treatment along the full range of students who were on the margin of receiving treatment. We also discuss parametric and nonparametric surface response methods in general, which have been under explored in multidimensional RDDs for education. We demonstrate theoretically, in simulation, and in an applied example, that the Gaussian process approach has several advantages over current approaches, including other surface response methods. In particular, using Gaussian process regression in two-dimensional RDDs shows strong coverage and standard error estimation and allows for easy examination of treatment effect variation for students with different patterns of running variables and outcomes. As nonparametric approaches are new in education-specific RDDs, we also provide an R package for users to estimate treatment effects using these methods.