
In this article, we develop A-optimal designs for a longitudinal Rasch Poisson Gamma counts model (RPGCM). The model is particularly well suited for repeated testing designs in educational and psychological research when the outcomes are count data. Within the RPGCM framework, test scores depend on two types of parameters: respondents' abilities, which are modeled as gamma-distributed random effects, and item difficulties, which are treated as fixed parameters. To account for dependencies among repeated measurements, we incorporate a compound symmetry structure into the RPGCM. Parameter estimation is performed using maximum quasi-likelihood methods. Based on this framework, we derive A-optimal designs that allow for efficient estimation of mean abilities and commonly used functions thereof. In analogy to classical ANOVA models, no constraints are imposed on the mean structure. The resulting A-optimal designs are substantially less restrictive than commonly used D-optimal designs. Moreover, the derived designs are also valid for a more general class of longitudinal Poisson Gamma models with repeated measurements. These models can be applied to analyze growth curves for count data in other disciplines, including epidemiology, medicine, and criminology.
In the context of classical test theory, we study under which conditions addition of an item to a unifactorial test increases the reliability of the test score. We distinguish item weights used to construct a sum score and item factor loadings. We show that the reliability of a weighted sum score increases if the scoring weight of the new item is set sufficiently low. We conclude that any negative contribution to the reliability is not caused by inclusion of the new item per se, but by inadequate weighting of the new item. We also show that if the unweighted sum score is being used with items that have equal error variances, inclusion of a new item increases reliability if the factor loading of the new item is greater than about half of the mean loading of the other items. For standardized items, with loadings based on correlations, we derive lower bounds for the loading of a new item based on the mean and variance of the existing loadings. Instead of removing items when their factor loading is below 0.30, we recommend lowering their scoring weight if they reduce reliability.
Recent studies on cognitive diagnostic models (CDMs) have extended the framework to longitudinal data. Various methods combined CDMs and hidden Markov models (HMMs) to assess changes in attributes over time and to evaluate the effects of interventions and covariates. A requirement for model fitting, inference, and interpretation is that models are identifiable. In this article, we derive identifiability conditions for experimental design used in education research. Specifically, we consider three designs: pretest/posttest single-group design, counterbalancing, and multiple-group longitudinal design. Drawing on existing HMM research, we examine the extent to which these designs satisfy common identifiability assumptions and propose new constraints for the counterbalancing and multiple-group longitudinal designs. These two setups are recommended in situations where item parameters differ over time or when there is heterogeneity in transition patterns across groups. We introduce a general HMM model for the multiple-group longitudinal design and a Gibbs sampling algorithm to estimate the parameters. We assess parameter recovery through a Monte Carlo simulation study and apply the model to a dataset from a study which evaluates the effectiveness of two interventions relative to a control condition. The results demonstrate the flexibility of the model and its potential to offer new insights into learning processes.
This study investigates the relationship between daily interpersonal stress (binary, time-varying) and suicidal behavior (binary, time-varying) using 90 days of daily diary data from 106 adolescents assessed immediately after discharge from acute psychiatric treatment. It addresses two key complexities: the rarity of suicidal events and non-monotone, non-ignorable missingness in both the outcome and the predictor. Because existing methods often fail to accommodate these complexities, leading to biased estimates, a Bayesian selection model is specified. The model integrates a mixed-effects complementary log-log regression for rare events with a missingness model that accounts for non-monotone, non-ignorable missingness in the outcome. A probit mixed-effects model is used for the time-varying predictor, along with a corresponding missingness model for its non-monotone, non-ignorable missingness. Empirical results support the applicability of the specified model to longitudinal studies involving rare events and complex missing-data structures. Furthermore, a simulation study demonstrates parameter recovery and highlights bias in focal parameters when sensitivity parameters in the outcome and missingness models are ignored.
Over 45 years ago, William Revelle proposed a reliability measure based on the worst split-half of a test or scale, commonly known as Revelle's beta, to assess the general factor saturation. However, to this day, there is no reliable method for computing this measure, as existing approaches are either computationally infeasible or insufficiently accurate in identifying the worst split-half. This difficulty arises because the number of candidate splits increases exponentially with the number of items. In this article, we show that computing Revelle's beta is conceptually equivalent to divisive ("top-down") hierarchical clustering. This insight allows us to reduce the number of candidate splits to a quadratic problem, making the computation feasible. We specify theoretical conditions under which this approach is guaranteed to recover the worst split-half. To validate the efficiency of our approach, we conduct simulation studies and analyze real-world data. Code implementations accompanying this work are available online, together with Supplementary Material.
Traditional perceptual models are ill-equipped for the high-dimensional data, such as text embeddings, central to modern psychology and AI. We introduce the double machine learning lens model, a framework that utilizes machine learning to handle such data. We applied this model to analyze how a modern AI and human perceivers judge social class from 9,513 aspirational essays written by 11-year-olds in 1969. A systematic comparison of 45 analytical approaches revealed that regularized linear models using dimensionality-reduced language embeddings significantly outperformed traditional dictionary-based methods and more complex non-linear models. Our top model accurately predicted human $(R<^>{2}_{CV} =0.61)$ and AI $(R<^>{2}_{CV} =0.56)$ social class perceptions, capturing over 85% of the total accuracy. These results suggest that "unmodeled knowledge" in perception may be an artifact of insufficient measurement tools rather than an unmeasurable intuitive process. We find that both AI and humans use many of the same textual cues (e.g., grammar, occupations, and cultural activities), only a subset of which are valid. Both appear to amplify subtle, real-world patterns into powerful, yet potentially discriminatory heuristics, where a small difference in actual social class creates a large difference in perception.
Classical symmetric association measures, such as correlation and chi-square indices, are widely used in applied psychology. However, these indices have limitations in identifying asymmetric implicative relationships. Standard regression analysis of Y on X, frequently interpreted as evidence of a directed dependence X -> Y $X o Y$ upper X right arrow upper Y , does not preclude the reverse direction ( Y -> X $Y o X$ upper Y right arrow upper X ). While various proposals in the literature have sought to provide non-symmetric association measures between binary events, most have overlooked the potential information in the contrapositive ( B & strns; -> A & strns; $\bar {B} o \bar {A}$ upper B overbar right arrow upper A overbar ), in addition to the main assertion ( A -> B $A o B$ upper A right arrow upper B ). When multiple variables are involved, asymmetric dependence is frequently represented as intricate dependency networks, which can be challenging to summarize and interpret in terms of higher-order clusters or latent dimensions. This article introduces a novel statistical implication index designed to address both limitations. The efficacy of this asymmetric index is demonstrated through its ability to detect one-way implication relationships, using both positive and contrapositive evidence. It also facilitates dimensional reduction by establishing aligned sets of nodes in a graph representation, under the condition that a Rasch model holds on these nodes, thus filling the gap between graphical and dimensional models. The efficacy of this index is substantiated through both simulated and real-world data illustrations.
A discussion is provided of several issues related to behavioral measurement that arise from Pfadt et al. (2026, Psychometrika, 2026, 1-35). The note may be viewed in part as a complement to their developments regarding precision estimation for individual test scores.
In this commentary on Pfadt et al. (2026, Psychometrika , 1-35), I first make the case for implementing psychometric methods, such as the conditional standard error of measurement (CSEM), in software that is user-friendly from a practitioner's perspective. Furthermore, I argue that bias and variance in CSEM estimates are still poorly understood and I report a small simulation study comparing the coverage rates of the CSEM estimate recommended by Pfadt et al. with those of the estimated (unconditional) standard error of measurement. The results point to possible directions for future research on CSEM estimation.
Pfadt et al. present accessible methods for estimating conditional standard errors of measurement (CSEMs) and implement them in the open-source software JASP. Their emphasis on individual-level precision represents an important contribution to applied measurement practice. This commentary discusses several conceptual issues that clarify and extend the authors' treatment of dimensionality, error definition, and score interpretation. The aim is to strengthen alignment between CSEM estimation and the interpretive purposes for which test scores are used.
Not-reached (dropout) and omitted (intermittent missingness) behaviors are often inevitable in timed computerized tests. These missingness behaviors may be related to the subject's latent traits, the difficulty of the item, or even the unobserved item response itself. In order to better understand the underlying test-taking behaviors, a Bayesian hierarchical framework is adopted to jointly model the item response, the item response time, the not-reached, and the omitted behaviors. For missing data, a sequential multinomial model is developed for the not-reached behavior, and a conditional model is proposed for the omitted behavior conditioning on the not-reached behavior. A decomposed logarithm of the pseudo marginal likelihood (LPML) is then developed to assess the fit of the missing data models. It can be further used to quantify the importance of modeling item response and response time jointly versus individually in identifying the missing data mechanism. The empirical performance of the proposed models and the model assessment criterion are examined through extensive simulations. The proposed methodology is further applied to analyze the Program for International Student Assessment (PISA) 2018 Test.
Recent research shows that amortized variational inference (AVI) can be used to efficiently estimate high-dimensional latent variable models on large datasets. However, its use has remained limited to item response theory (IRT), and generalizing the approach to discrete latent variable models is not straightforward. We propose two ways to deal with this problem. In an initial simulation, we verify that these approaches can be used to estimate simple discrete latent variable models, such as latent class analysis and the generalized deterministic inputs, noisy and gate model. In these cases, AVI provides accurate parameter estimates, although the computational advantage over marginal maximum likelihood (MML) and standard variational inference (VI) is limited. We then apply the same approach to estimate mixture IRT models. In this case, AVI is computationally faster than MML estimation and standard VI. To demonstrate the practical applicability of our AVI approach, we use it to fit a seven-dimensional mixture IRT model to a narcissism inventory. Whereas quadrature-based methods cannot feasibly estimate models of this dimensionality, the efficient AVI approach even allows for computation of bootstrapped standard errors. We provide our code, along with an easy-to-use tool for fitting these models to new datasets.
Test speededness, caused by time constraints, can impact examinees' performance, leading to decreased response accuracy, particularly toward the end of the test. Most existing methods for detecting test speededness rely on specific distributional assumptions for response times (RTs), such as the lognormal distribution, which may lead to incorrect statistical inference if the true data distribution deviates from these assumptions. This article proposes a novel Bootstrap-CUSUM method for detecting test speededness, which is robust to non-normality in log-RTs. By constructing a cumulative sum (CUSUM) person-fit statistic for log-RTs and using the multiplier bootstrap to estimate its empirical distribution, our method facilitates individual-level detection and changepoint estimation. We prove the theoretical consistency of the method under both null and alternative hypotheses. Simulation studies show that the Bootstrap-CUSUM method outperforms the likelihood ratio test, Wald test, and score test in terms of correct classification rate, true detection rate, and false positive rate, demonstrating superior robustness and adaptability across different data distributions. The real data analysis further demonstrates the practical utility of the proposed method for detecting test speededness.