
facomplex R package implements procedures for estimating complexity and simplicity derived from factor-analytic results of multidimensional scales. In this type of analysis, factor loadings are estimated to represent the relationship between the items and their latent factors or constructs, as well as with the other factors in the model where cross-factor loadings are estimated. An evaluation of these cross-factorial loadings helps determine whether the result is factorially simple or complex. There are indicators for assessing the complexity and simplicity of factorial results, and facomplex calculates all known indicators from various approaches focused on the direct estimation of factorial simplicity and complexity at the item, dimension, and total scale levels: the loading simplicity index, the factor simplicity index, Hofmann's complexity coefficient (for items and factors), the Bentler Simplicity Index, and the proportion of explained variance derived from the squared sum of target versus non-target factor loadings.
projectCAT is an open-source R Shiny application providing a complete, code-free computerized adaptive testing (CAT) simulation environment. It integrates IRT calibration (1PL, 2PL, 3PL via marginal maximum likelihood), three adaptive item selection methods, namely, Maximum Fisher Information (MFI), Mean Expected Information (MEI), and Minimum Expected Posterior Variance (MEPV), three ability estimation methods (EAP, ML, and MAP), Sympson-Hetter (SH) exposure control, and content balancing with user-specified target proportions across multiple domains. Both SH and content balancing operate identically for single-examinee and full multi-examinee simulations. Applied to the built-in 300-item 2PL demo bank ( n = 500 , 20 items, MFI + EAP), RMSE = . 36 , correlation = . 94 , and item overexposure was reduced to 11.3% under SH ( r max = . 20 ) . Content proportions matched a 50/30/20 user-specified split exactly. projectCAT is freely available at https://hdmeasurement.shinyapps.io/projectCAT/ .
Projective Item Response Theory (PIRT) modeling is a measurement framework in which multidimensional item response models are expressed in terms of lower-dimensional (e.g., unidimensional) proxy IRT models. Existing PIRT methods and their associated estimators currently rely on logistic function approximations that are limited to a narrow class of ordered, monotonic response functions. Though unexplored to date, these methods also require computationally intensive procedures to obtain sampling variability estimates for their resulting PIRT estimates. These and other limitations restrict the practical utility and broad application of PIRT models, particularly in empirical settings where only moderate sample sizes are available. To address these challenges, we introduce a general maximum marginal likelihood PIRT (MML-PIRT) approach that leverages components of the expectation-maximization (EM) algorithm commonly used in MML estimation. The proposed method uses expected count information generated during the EM-MML to fit proxy response functions for any focal trait of interest, and for a much broader class of multidimensional IRT models. In addition, MML-PIRT provides accurate and efficient large-sample variability estimates of the resulting PIRT model, thereby enhancing both the flexibility and statistical efficacy of PIRT model applications.
Generalizability Theory (G-Theory) extends classical reliability estimation by simultaneously de-composing score variance into multiple sources of measurement error such as raters, occasions, and items within a single analytical framework. Despite its methodological advantages, G-Theory analyses typically require specialized programming expertise, limiting accessibility for applied researchers. This paper introduces projectGST , an open-source R Shiny application that provides a unified, code-free interface for conducting G-Theory analyses. The application integrates preliminary statistical analyses including descriptive statistics, normality tests, inter-rater correlation matrices, Cronbach’s alpha, and intraclass correlation coefficients with G-Study variance component estimation and D-Study reliability optimization. projectgst supports both crossed and nested measurement designs and is illustrated using simulated datasets representative of nonformal education assessment contexts, including tutor evaluations at community learning centers and portfolio-based assessments in equivalency education programs. The application is openly accessible via a hosted Shiny interface and GitHub repository.
The rapid adoption of Large Language Models (LLMs) in educational assessment has reshaped scoring practices, yet evaluation remains tethered to aggregate reliability metrics like Quadratic Weighted Kappa, which obscure discrimination and rater effects. This study applies Signal Detection Theory to evaluate eight state-of-the-art LLMs (including Claude 3.5 Haiku, DeepSeek-V3, Gemini 3 Flash, GPT-4o, and Grok 4.1) against expert human raters across 1,726 essays. By decoupling discrimination from response criteria, I provide a diagnostic analysis of AI scoring behavior. Results indicate that human raters exhibit significantly superior evaluative precision, with average discrimination estimates approximately double those of the AI models. Furthermore, LLMs are prone to pronounced centrality effects and score compression, systematically failing to award the highest rubric tiers. These findings demonstrate that low human-machine agreement stems from both a deficit in discriminative accuracy and systematic shifts in response criteria. Ultimately, this research provides a robust framework for calibrating and selecting AI scoring systems based on specific pedagogical goals and fairness requirements.
Linear regression and analysis of variance are widely used in applied psychological measurement to estimate group, condition, and covariate effects, yet statistical efficiency and conventional inference can be compromised when outcome variance changes across groups or covariate levels. This article introduces varGuid for R and varguid for Python, open-source implementations of variance-guided regression for linear models. The method estimates a covariate-dependent mean-variance relationship and uses it to iteratively reweight the original mean model. Ordinary analyses use iteratively reweighted least squares, whereas sparse analyses use an iteratively reweighted lasso. Because the original design matrix and outcome scale are retained, regression coefficients and ANOVA contrasts remain directly comparable with conventional effect estimates. Robustness here refers to relaxation of the homoscedasticity assumption rather than resistance to outliers. Under the conditions established for variance-guided regression, the estimator matches the homoscedastic baseline in population predictive quasi-risk when variance is constant and improves on that baseline when variance depends on covariates. The packages accept general linear-model design matrices, including ANOVA-style encodings, and provide baseline and variance-guided predictions, example data, and heteroscedasticity-consistent summaries for non-lasso fits. The Python implementation also supports NumPy and pandas inputs, Patsy formulas, model summaries, and a scikit-learn-compatible estimator. Both implementations are operating-system independent and require no unusual hardware. Source code, documentation, examples, and installable files are available through CRAN, PyPI, GitHub, and Zenodo. The packages provide accessible tools for applying variance-guided regression as a primary or companion analysis when homogeneity of variance is uncertain in routine measurement research and related quantitative applications.
The traditional thresholds for product-moment correlation ( r ) usually condensed in terms of “very small,” “small,” “medium,” “large,” “very large,” and “huge” are derived on the scale of d . Properly transforming an estimate of r to the scale of Cohen’s d is important for qualitatively evaluating the magnitude of the r effect size. A proper formula exists for a transformation between r and d in the binary and dichotomous settings, but not in polytomous or continuous settings. The traditional formula for polytomous settings assumes an equal number of cases in the subpopulations and can lead to a radically misleading transformation if the group sizes are severely imbalanced. Two general formulae are derived for transformations applicable to dichotomous, polytomous, and continuous cases, regardless of the imbalance in the group sizes. Real-life examples are given which illustrate the use of the general forms and difference between the general and the traditional forms.
scindex is an R package for analysing inter-rater reliability in binary classification tasks. The package computes Cohen's κ and Fleiss' κ and, when ground-truth labels are available, estimates signal detection theory parameters, including sensitivity, specificity, and decision thresholds. It also implements the Strategic Convergence Index (SCI), a measure of convergence in raters' response criteria.
babebi is an R package for analysing complete two-time, two-rater pre-post rating designs. The package estimates pre-post effects using a linear model with a rater indicator as covariate and provides adjusted estimates of change, posterior summaries, and BIC-based Bayes factor approximations. It also includes Monte Carlo validation routines calibrated from the observed design to evaluate inferential performance under study-specific conditions.
Choosing suitable estimation methods for cognitive diagnostic models (CDMs) is critical. However, practitioners often face issues like non-convergence, boundary estimates, extreme values, and unstable suboptimal solutions, which affect the accuracy and reliability of parameter estimates. In this study, we compared expectation-maximization (EM), Bayesian modal estimation (BM), their monotonic constraint variants (EMM and BMM), and variational Bayes (VB) methods. A simulation study was conducted, manipulating factors such as sample size (50, 200, 1000), test length (15, 30), item quality (high, low), and attribute distribution (uniform and multivariate normal). The performance was assessed based on the empirical frequency of each issue, the recovery accuracy of the parameters, and the sensitivity to algorithm initialization. The results, analyzed using the generalized deterministic inputs noisy "and" gate model, reveal three main findings. First, an insufficient sample size was identified as a key factor in problems related to parameter estimation. Second, methods that incorporate prior information (BM and VB) exhibited fewer cases of non-convergence and extreme estimates than EM. Third, the sensitivity analysis showed that the stability of solutions was affected by the choice of initial values, emphasizing the need for proper initialization to reduce the risk of becoming trapped in local suboptimal solutions. This systematic comparison demonstrates that no single estimator is universally superior, and the choice depends on practical constraints. Our findings offer evidence-based guidance for selecting context-sensitive methods, thereby improving the validity of CDMs in real-world applications.
Self-report questionnaires are widely used in research and practice. In most applications, the vulnerability of these questionnaires to response biases like faking is ignored. However, especially in high-stakes situations such as personnel selection, measurement can be severely biased when test-takers engage in faking to present themselves more favorably. To separate faking-related variance from substantive trait variance, the Multidimensional Nominal Response Model (MNRM) has been used to reduce systematic bias in trait estimation by allowing for item-specific relations between response categories and social desirability. A critical but untested assumption of this approach is that perceptions of social desirability are homogeneous across test-takers. However, individuals may differ considerably in how they perceive the desirability of the item content. Here, we conducted simulation studies to investigate how violations of this assumption affect the MNRM’s ability to recover substantive trait person parameters. We implemented three distinct manipulations of heterogeneous desirability perceptions and examined their impact on person parameter recovery. Results showed that the MNRM is robust against violations of homogeneous social desirability perceptions as long as test-takers’ faking behavior is aligned with their perceived desirability of the item content. In contrast, when test-takers fake responses in ways that are inconsistent with item-wise desirability perceptions, parameter recovery seems to decline. Implications for practice and possible model extensions are discussed.
This study aims to examine the performance of adaptive quadrature (AQ) estimation method for ordinal confirmatory factor analysis (CFA). Specifically, we compared four link functions (complimentary log-log [CLL], logit, log-log, and probit) of the AQ estimation method across varying factor structures, sample sizes, distributional shape of latent trait, and number of quadrature points. The study is conducted via a simulation study and using empirical data. The results demonstrate that the probit link function exhibits superiority across the vast majority of conditions, consistently yielding the highest proper convergence rates and, among successfully converged solutions, the lowest parameter recovery errors, and the best relative fit, whereas the logit generally showed the weakest performance. Additionally, a critical divergence was discovered regarding asymmetric link functions: while the probit link generally provided the best model fit across the positively skewed simulation conditions, the log-log link yielded the best relative model fit for the positively skewed empirical data. Furthermore, the study reveals the complex role of quadrature points in multidimensional spaces. Although using eight quadrature points may be necessary in more complex simulated models, it frequently causes severe estimation failures when applied to sparse real-world data.
In Automatic Item Generation (AIG), item incidentals refer to surface characteristics of an item that are assumed not to influence item parameters (e.g., item difficulty), whereas item radicals refer to attributes that are presumed to affect these parameters. Within the empirical validation process of the item generator, subjects and incidentals may either be sampled independently so that every subject sees every incidental (cross-classified sampling) for a radical, or incidentals may be sampled within each subject so that every subject only sees a specific set of incidentals (two-level sampling) for a radical. We present an approach for scrutinizing the effect of item incidentals relying on two classical test theory models that adhere to the stochastic sampling space of cross-classified and two-level sampling, respectively. We show how these may be used in combination to enable a more optimized investigation of incidental-induced variance within the item generator. We illustrate the approach with the figural short-term memory item-generator "figumem." Results show that incidentals have little effect on item difficulty in the cross-classified model/sample and that the model parameters generalize to a larger set of incidentals in the two-level model/sample. Implications, limitations, and future research are discussed.
Differential item functioning (DIF) detection is an important yet understudied problem in computerized adaptive testing (CAT). In this article, we proposed a two-level logistic model to improve DIF detection in CAT by explicitly accounting for nuisance effects arising from CAT-induced structural dependency. First, we conceptualized that adaptive item selection induces systematic dependencies among examinees and items through provisional ability estimates, whereas traditional single-level DIF methods assume independent observations and may yield misleading results in CAT settings. Then, using a numeric example and Monte Carlo simulations, we compared our proposed two-level model with competing single-level models under various CAT conditions, manipulating test length, exposure control, ability estimator, DIF type, and DIF prevalence. Item-level Type-I error and statistical power conditional on joint model convergence were reported for each model. We showed that the proposed two-level model has improved control of spurious DIF and competitive power relative to single-level models, particularly with shorter tests and smaller exposure rates. However, we observed that the model convergence varied systematically across simulated conditions, highlighting that inferential accuracy and convergence reliability are intertwined in complex CAT DIF settings. Through this study, we underscored both the promise of multilevel DIF modeling in CAT and the need for future research to jointly evaluate convergence and inferential performance when assessing DIF models.
Researchers understand that conducting numerous pairwise comparisons between group means increases the Type I error rate, prompting the use of planned contrasts like orthogonal contrast sets. Implicit to orthogonal contrast sets is the principal assumption that groups are balanced in size. Further, when dealing with complex variables like latent constructs, specialized modeling is necessary. Understanding how violating the assumptions of orthogonal contrasts, specifically under conditions of sample imbalance, can help identify variability in parameter recovery. This study examines the effect of sample size imbalance and modeling approach on the accuracy of latent group mean difference estimates when using orthogonal contrasts. Monte Carlo simulations compared the Multiple Indicators Multiple Causes (MIMIC) and re-parameterized multigroup confirmatory factor analysis models while manipulating sample sizes, group proportions, and effect size. Results suggest declining parameter recovery as group imbalance increased, particularly in small samples, with some estimates falling below acceptable thresholds for power, Type I error, and bias. The MIMIC model consistently produced more accurate estimates, though is replete with implicit measurement assumptions that are seldom tested. These findings suggest that researchers using orthogonal contrasts when comparing groups on a latent variable continuum must (a) be aware of examined group's sample size proportions and the impact of group size inequalities on estimate accuracy, and (b) carefully consider the costs and benefits of the latent variable modeling approach, including how the model addresses measurement non-invariance.
The study compared the effectiveness of four methods for detecting differential item functioning (DIF) in polytomous multidimensional data with a simple structure: the item response theory likelihood ratio test (IRT-LR), two ordinal logistic regression approaches (using raw scores vs. latent trait estimates as the matching variable), and the multidimensional MIMIC-interaction method. Data were generated under a two-dimensional graded response model with 28 five-category items. Simulation conditions manipulated DIF type (uniform, nonuniform), DIF magnitude (0, 0.3, 0.6), group size ratio (1:1, 3:1), latent trait correlation (ρ = 0, 0.5), and the presence of group impact, yielding 40 conditions with 100 replications each. Across conditions, IRT-LR and both logistic regression approaches generally maintained Type I error within acceptable limits, whereas the MIMIC-interaction model showed inflated Type I error in the presence of impact. All methods demonstrated high power for moderate uniform DIF, but detection rates declined substantially for low DIF and for nonuniform DIF. Logistic regression with latent trait estimates showed the most stable overall performance, combining adequate Type I error control with comparatively high power across conditions. Logistic regression with raw scores demonstrated relatively stronger performance for moderate nonuniform DIF. In contrast, IRT-LR exhibited lower power despite conservative Type I error control. Results suggest that regression-based approaches, particularly logistic regression using latent trait estimates, provide robust performance for DIF detection in multidimensional polytomous assessments under simple structure.
Latent structure analysis methods, including latent profile analysis (LPA), latent class analysis (LCA), item response theory (IRT), exploratory factor analysis (EFA), and confirmatory factor analysis (CFA), are widely used in psychological and educational research to model unobserved constructs and identify heterogeneity across individuals. However, applying these methods often requires advanced statistical expertise and the use of multiple specialized software packages with different workflows, which can limit accessibility and increase analytical complexity. This paper introduces projectLSA, a Shiny-based application designed to provide an integrated and user-friendly platform for conducting latent structure analyses. The application enables users to upload data, specify models, estimate parameters, compare model fit using standard indices, and visualize results within a single interface. By integrating several established R packages, projectLSA supports a unified analytical workflow without requiring users to write code. The practical utility of the application is illustrated using built-in simulated datasets that support multiple analytical procedures, including LPA, LCA, IRT, EFA, and CFA. These examples demonstrate how users can estimate, compare, and interpret models efficiently within a consistent workflow. Overall, projectLSA enhances accessibility, consistency, and efficiency in latent structure analysis by reducing technical barriers and supporting interactive and reproducible data analysis.