
Abstract Cut scores are typically applied to observed or estimated scores that contain measurement error. Even if the criterion cut score itself were error free, error in examinee scores leads to false‐negative and false‐positive classifications. The relative costs of these classification errors may be asymmetric and may be further affected by retake policies that increase the probability of passing across repeated administrations. We propose a decision‐theoretic procedure for selecting an operational cut score that minimizes a weighted expected misclassification loss while respecting the committee‐designated criterion cut score. The method is formulated for item response theory (IRT) scales and accommodates one‐, two‐, or three‐parameter models. Rather than relying on a single reliability‐based standard error, we represent measurement error through the conditional sampling distribution of the proficiency estimator given true proficiency, approximated via parametric bootstrap. Expected false‐negative and false‐positive rates are obtained by Monte Carlo integration over a specified target proficiency distribution, and the optimal operational cut is found by numerical minimization over a feasible cut‐score set. The key feature is that misclassification probabilities are computed from an approximation to the conditional sampling distribution of the proficiency estimator, rather than from a reliability‐based approximation. An empirical illustration using a three‐parameter logistic item response theory model fit to response data from a 36‐item examination shows that the optimal operational cut increases as false positives are weighted more heavily and as the number of permitted administrations increases. The proposed approach provides a practical, model‐based tool for aligning cut‐score policy with explicit misclassification costs and retake rules.
Abstract Achievement tests are often designed to measure student growth over time using vertical scales while also providing diagnostic information on mastery of specific skill areas. However, traditional assessments rarely integrate psychometric models appropriate for both purposes. This study examines an approach combining vertical scaling with Diagnostic Classification Models (DCMs) to provide both overall achievement scores and subtest‐level diagnostic classifications to provide growth and diagnostic information simultaneously. We compared four model specifications against traditional unidimensional vertical scaling using empirical data from the Virginia Language and Literacy Screening System with 17,302 kindergarten and 18,339 first‐grade students. We evaluated model fit, score comparability, and classification reliability, and added Monte Carlo simulations to assess parameter recovery. Results revealed important tradeoffs between approaches. While no single approach provided an ideal solution, second‐order DCMs with appropriate constraints offer practical viability for operational assessment contexts that wish to measure student growth while also providing diagnostic information.
This manuscript introduces a culturally specific assessment framework for STEM leadership, developed through qualitative analysis of leadership narratives from presidents, provosts, deans, department chairs, and professors at Historically Black Colleges and Universities (HBCUs). While grounded in the HBCU context, the framework is designed to guide the future development of a rubric-based assessment tool, with broader implications for leadership assessment in other culturally specific educational environments. A formal instrument development process was used to ensure that a leadership assessment framework reflected the voices and experiences of HBCU leaders. Through collaborative in-depth qualitative analyses, six dimensions of STEM leadership were identified that represent the epistemologies, values, and practices embedded in these institutions. While the resulting framework is tailored to the HBCU context, the development process provides a rigorous and transferable approach to assessment design. The methodology repositions validity not as a post hoc adjustment, but as a foundational commitment to epistemic alignment between what is being measured and who is being measured. In doing so, assessment development for culturally specific environments is presented as both methodologically rigorous and conceptually sound.
Longitudinal data from repeated measurements are commonly used in social and behavioral sciences to study students' growth. When scores are assigned by raters, they become subject to rater effects such as variability in rater stringency. Therefore, in longitudinal assessments with rater-assigned scores, valid inference on growth requires accounting for both within-student dependencies and rater effects. To address the need, we propose a longitudinal cross-classified item response theory (L-CCIRT) model that integrates the longitudinal IRT model with the CCIRT model to incorporate timepoint-specific latent variables on the student side while accounting for rater influences. We also describe how latent change parameterization can be derived from the latent state parameterization to draw inferences about incremental changes at each timepoint. Our simulation studies indicated successful recovery of both item and group parameters of the L-CCIRT model under correct model specification and underscored that group parameter estimates can be affected by unmodeled student-rater interactions. Finally, we applied the model to an empirical dataset from three yearly administrations of a scenario-based assessment to demonstrate the utility of the proposed L-CCIRT. We concluded this paper by discussing ways the L-CCIRT model can be utilized in practice and pointing out future research directions.
This study examines the intellectual structure and historical development of the Journal of Educational Measurement over its six-decade history (1964-2025). Using computational text mining and network analysis, we analyzed 1,678 articles based on a deterministic, expert-defined taxonomy of 34 psychometric categories to quantify research trends. Results identify item response theory (IRT) as the dominant paradigm and the most connected area in the publication network. Community detection reveals a tripartite structure dividing the discipline into psychometric modeling, applied test operations, and classical measurement foundations. While IRT constitutes the primary theoretical focus, parameter estimation and validity emerge as intermediary concepts linking statistical modeling to practical applications. The analysis also highlights a historical transition from classical test theory to latent trait modeling, with computerized adaptive testing and differential item functioning identified as major contemporary topics. Furthermore, artificial intelligence appears as an emerging area, showing limited activity for many years followed by rapid growth beginning in 2022, primarily driven by automated scoring research. Overall, these findings provide a data-driven framework for monitoring how the field continues to evolve in response to technological developments.
This study applies backward construct mapping, beginning from empirical Rasch modeling and working toward empirical construct definition, to develop and validate an assessment of inflectional morphological processing. This project was a part of the UCSF Dyslexia Center's Dyslexia Phenotype Project, which is a comprehensive multidisciplinary investigation examining cognitive function across multiple domains (language, visuospatial processing, social cognition, executive function, attention, and emotional processing). We administered 64 items in a 2 & times; 2 & times; 2 factorial design (regularity, frequency, word class) to 131 participants (84 with developmental dyslexia, 47 typical readers; ages 7-16 years) using an oral sentence completion task. Multidimensional Rasch analysis validated unidimensionality (latent correlation = .84) and revealed strong measurement properties (person separation reliability = .92). Regularity (beta = 3.63) and frequency (beta = .99) accounted for 77.5% of item difficulty variance, confirming dual-route theory. Wright Map analysis revealed five empirically derived waypoints representing natural boundaries in processing demands: from regular high-frequency forms to irregular low-frequency forms.Despite a meaningful ability difference between groups (1.0 logit; similar to 25 percentage point difference at average item difficulty), both groups demonstrated similar item difficulty hierarchies, supporting measurement invariance. Situated within Wilson's (2005, 2023) BEAR Assessment System, this study demonstrates how beginning from an existing dataset and Rasch measurement modeling (Building Block 4) and working toward empirical construct definition (Building Block 1) can transform ordinal rankings into genuine educational measurement, providing an initial, replicable framework for assessment development and actionable insights for differential instruction in students with developmental dyslexia.
This study develops and tests a Large Language Model-based assessment framework that uses a multi-agent system to analyze students' written responses, generate scoring rationales, identify uncertainty levels, and assign final scores to support learning. The framework was tested using chemistry responses from 834 middle school students scored with a dichotomous analytic rubric. Through prompt engineering and aggregation methods, the multi-agent system-enhanced by rubric revisions and human scoring insights informed by Artificial Intelligence (AI)-generated rationales-achieved 94% scoring accuracy, representing a 16% improvement over the 78% accuracy of a single-agent model using the original rubric. The system's uncertainty detection closely aligned with areas where human raters also indicated uncertainty. Results indicate a strong relationship between AI confidence and scoring accuracy: when the AI was confident, its scores were largely correct, and low-confidence cases were often inaccurate. These findings demonstrate the value of a multi-agent system with human-AI collaboration, in which the AI identifies unclear cases and teachers review and refine uncertain scores. This collaboration approach shows the framework's potential to enhance classroom assessment by providing timely, reliable scores and feedback to inform teaching and learning. The future work may improve accuracy further through retrieval-augmented generation, human-in-the-loop, and fine-tuning with synthetic data.
This study presents a secondary analysis of word recognition patterns of 650 second-grade students on an untimed oral reading fluency assessment, using explanatory item-response models (EIRMs). The analytic sample included 50 passages containing 1,267 word-tokens. Rasch model calibration showed most words clustered around the lower third of student ability distribution, capturing variability among struggling readers. Subsequent EIRM analyses indicated that several factors had significant effects on word recognition difficulty: Lexical features including word frequency, age of acquisition, and first-syllable vowel patterns (especially r-controlled vowels, diphthongs, and long vowels) were key predictors. Content words (nouns and verbs) proved easier than function words, while more concrete words were unexpectedly more difficult. Word position within passages significantly predicted difficulty. When examined simultaneously, first-syllable vowel patterns emerged as the most robust decoding-related predictor, explaining variance previously attributed to broader decoding measures. The final model explained 40% of item variance. By identifying specific item features that challenge developing readers in oral reading, this research provides a foundation for more precise instructional approaches that target particular sources of word recognition difficulty rather than relying on conventional aggregate reading indicators.
Multidimensional Computerized Adaptive Testing (MCAT) can improve testing efficiency, but it can also lead to test issues such as item overexposure and unbalanced use of items in the item bank. In MCAT, item selection strategies based on response time typically ignore item exposure. Therefore, this study combined four response time-based item selection strategies (DT-inc, DT, AT-inc, and AT) with two different exposure control methods (RT and RPG), designing eight new item selection strategies to examine the performance of item selection strategies under conditions with and without exposure control approaches. Results of simulation and empirical studies indicated that eight new item selection strategies effectively improved test item pool utilization and item exposure uniformity, significantly reduced test-taking time for examinees, and did not significantly affect estimation accuracy.
Engagement has been debated for years across various perspectives and educational contexts. However, few studies clearly define or distinguish School Engagement and Child Engagement, often conflating them with related constructs. This literature review aims to address three central questions: How are these constructs defined? What are their properties? What kind of information do they provide? To explore these questions, 356 articles were analyzed, 287 focused on School Engagement and 69 on Child Engagement. For each study, the objective, main variables, and research findings were examined. The articles were further categorized based on the engagement model or dimensions used to operationalize Child or School Engagement. The review reveals both differences and commonalities between the constructs. While the operationalization and research topics differ between Child Engagement and School Engagement, the studies generally acknowledge the dynamic nature of engagement. This review emphasizes that measurement is not a neutral act but a designed and intentional process. How engagement is defined and operationalized shapes the type of information produced, reflecting both theoretical choices and educational priorities.
This study presents the development and validation of two construct maps representing confidence in the adoption of practices associated with an anti-racism (AR) value stance among in-service teachers using and exploratory item response modeling (IRM) approach. The Anti-Racism Value Stance Survey (ARVSS) was developed to assess teachers' confidence in understanding critical AR concepts and engaging in AR-aligned actions following completion of a state-sponsored professional development. The AR-Concepts (AR-C) construct reflects three hierarchically related stages starting at foundational critiques of racism to systemic, intersectional understandings, while the AR-Actions (AR-A) construct outlines four hierarchically related stages of action starting at classroom-level interventions to systemic advocacy. We provide evidence for the ARVSS's reliability and provide validity based on content and internal structure. Findings suggest that confidence in AR conceptual understanding and AR action-readiness progress concurrently. Our constructs hold implications for curriculum design, teacher preparation, and future assessment development in multicultural and anti-racist education.
The operational scaling in large-scale assessments (LSAs) comprises two stages: item response theory and latent regression modeling. Principal component analyses (PCA) are routinely performed before latent regression modeling for dimension reduction, but this approach retains too many principal components (PCs), threatening numerical stability. This study proposes a new approach: adding process variables to the usual contextual variables and replacing PCA with variable selection for latent regression modeling. We found that using Lasso, random forests, and ultimately gradient boosting for variable selection led to measurement precision similar to or higher than the PCA approach, but with considerably fewer covariates; including process variables into latent regression models consistently improved measurement precision. Integrating process data and variable selection yielded the highest measurement precision while achieving parsimony: The latent regression models with the 150 most important variables, including process variables, outperformed those with 330 PCs in France and 270 PCs in the Republic of Korea. The findings suggest that our proposed approach can effectively solve the overparameterization problem in LSA scaling while preserving measurement precision.
This article proposes the Ability-Attenuated 3PL (AA-3PL), an extension of the three-parameter logistic model designed to represent systematic discrepancies between latent ability and expressed performance (the ability-expression gap). The model retains the conventional interpretation of 3PL item parameters while adding a parsimonious mechanism to capture structured performance attenuation that can arise from non-ability influences in longitudinal and applied testing contexts. Evidence is established through a Bayesian simulation program spanning multiple study modules, including baseline multistage conditions, shape-misspecification stress tests, identification-constraint checks, a growth-related extension, a one-stage boundary case, and prior-sensitivity analyses. Across attenuation conditions, AA-3PL consistently reduces difficulty misattribution and improves overall model performance, while showing negligible differences from 3PL when no gap is present; robustness depends on aligning the attenuation shape with the underlying data pattern. An empirical two-wave WJ-PV illustration provides convergent support: AA-family models better capture follow-up systematic deviations than 3PL, with the shift-augmented extension offering the most coherent improvements when cross-wave shifts are present.
International assessments, such as the Programme for International Student Assessment (PISA), play a critical role in monitoring educational progress, evaluating teaching effectiveness, and supporting institutional accountability. The validity of these interpretations rests on the assumption that students exert sufficient effort and perform to the best of their abilities, condition necessary for tests to reflect students' true proficiency. However, researchers have documented concerns about low effort in low-stakes test contexts where results carry no direct consequences for examinees. Despite growing awareness of this issue, research examining low effort correlates in large-scale assessments like PISA remains limited. Moreover, existing approaches often overlook the correlation between low effort and ability, which can bias estimates of both constructs.The current work used an explanatory IRTree model that jointly estimated science ability and rapid-guessing propensity to examine item-level predictors of disengagement in PISA 2022 science, including Item Position, item interactivity, cognitive demand, and assessment sequencing. Interactive items decreased disengagement which rose with Item Position. Cognitive demand was the strongest predictor, with high-demand items sharply increasing disengagement. Also, science items were less likely to be rapidly guessed when students started with adaptive math or reading rather than science, even after controlling for position. As expected, disengagement was strongly negatively related to science ability, underscoring the value of jointly modeling proficiency and effort in low-stakes contexts. We discuss these results and explore avenues for future research.
The importance of the construct to the fields of measurement and cognitive science cannot be overstated. Without an understanding of the construct, including clarity and scope of the substantive and structural aspects, one cannot ensure valid use and interpretation of a constructed measure. In this sense, measurement and cognitive science are fundamentally connected through the construct. One important construct that lacks clarification in definition and structure is a sense of belonging in science. This is especially true for adolescent age students. Unfortunately, this attribute is often vaguely conceptualized and largely missing from the literature for school-age students, and therefore, in need of a rigorous theoretical foundation supported by empirical evidence to ensure meaningful use of any assessment or interpretation of results. Using the BEAR Assessment System's approach to constructing measures, which includes using Rasch family modeling, data were collected from high school students through eight exploratory interviews, 12 think-aloud interviews, focus groups comprised of 9-12 participants per group (45 students in total), and 537 surveys; and from five experts to develop and validate a 12-item self-report survey of high school students' sense of belonging in science. Sources of evidence for reliability, validity (content, internal structures, response processes, and relation to other variables), and fairness were collected and analyzed to produce a construct-centered survey. Overall, findings produced strong evidence for the valid, reliable, and fair interpretation and use of the Sense of Belonging in Science Survey (BiSS). This research highlights the importance of the construct as the driver of the design and quality assurance of a measure.
Unbiasedness for proficiency estimates is important for autoscoring engines since the outcome might be used for future learning or placement. Imbalanced training data may lead to certain biases and lower the prediction accuracy for classification algorithms. In this article, we investigated several data augmentation methods to lower the negative effect of imbalanced data in measurement settings. Four approaches were examined: (1) Resampling methods, either oversampling or undersampling; (2) Active resampling methods, where the resampling weight is based on representativeness in the training set; (3) Data expansion methods using synonym Replacement, slightly changing the meaning or semantics of the original answers; and (4) Content recreation method using Generative AI (e.g., ChatGPT) to create responses for less populated scores. We compared the performance (e.g., Accuracy, QWK, F1) as well as the distance metric for different combinations of the methods. Two datasets with different imbalanced distributions were used. Results show that all four methods can help to mitigate the bias issue and the efficacy was influenced by the imbalance level, representativeness of the original data and the level of increment in the variety of the response (i.e., lexical diversity). In general, resampling and GenAI with active resampling showed the best overall performance.
Cognitive Diagnostic Models (CDMs) provide fine-grained diagnostic feedback, but their central component-the Q-matrix-remains costly and labor-intensive to construct. This study explores the automated generation of Q-matrices using general-purpose AI, including ChatGPT-4o, Gemini-2.5-pro, and Claude-sonnet-4. We evaluated two prompting strategies (all-at-once and one-by-one) across TIMSS 2007, TIMSS 2011, and PISA 2012 mathematics assessments. Results show that AI-generated Q-matrices approximate human baselines with competitive model fitting performance (AIC, BIC, log-likelihood, and SRMSR) and acceptable classification discrepancies. While AI predictions for larger and more complicated assessments (TIMSS 07 and 11) were generally sparser than human-generated Q-matrices, they still achieved equal or better fit statistics under most CDMs. In contrast, for the smaller and less complicated PISA 2012 assessment, AI-generated Q-matrices matched human density and fitting quality. Importantly, chatbot-human matching accuracy remained high across models, with Gemini benefiting from all-at-once prompting, ChatGPT-4o maintaining stable performance under both strategies, and Claude showing sensitivity to prompt structure. These findings highlight both the promise and current limitations of automated Q-matrix generation, underscoring opportunities for integrating LLMs into scalable diagnostic assessment practices.
Recently the "Hexagon Measurement Framework" has been proposed as a unifying interpretation of measurement across the physical and human sciences. While the steps for using this Framework to understand and to perform a measurement are extensively described and discussed in previous publications of ours, what is discussed only briefly there is how the Framework can also be exploited as a blueprint to help underpin the process of framing a concept of the property to be measured and then developing a measurement for it: this is the object of the present article.First, we outline how the Hexagon Framework allows us to interpret the development of physical measurement, taking Chang's account of the historical advancements of temperature measurement as an illustration. Second, we show how the Framework is mirrored in a contemporary measurement development approach used in the human sciences: the BEAR Assessment System (BAS). We describe how the BAS is overlaid on the Hexagon Framework as the measurement development proceeds and may well iterate several times in that process. Third, we use this underlay of the Framework to illustrate and exemplify how the BAS has been used to design and develop a measuring instrument in the human sciences, specifically to measure an educational achievement construct intended for use in schools: Scientific Argumentation. Finally, we make comparisons of the above account with the so-called "norm-referenced" approach to measurement and also discuss some caveats and limitations of the account here.
In this study, we propose a new fixed person parameter calibration (FPC) strategy that incorporates measurement error in examinee ability estimates. Specifically, the proposed FPC method is an application of the mixed-effect structural equation model of Junker et al. (2012) to the small-sample item calibration context and relies on a Bayesian iterative sampling procedure for parameter estimation. We evaluated the proposed FPC method using simulated data sets that varied in terms of sample size, item composition, and examinee ability distribution. The parameter recovery performance of the proposed method was compared to those from alternative small-sample calibration methods two other FPC methods and the state-of-the-art fixed item parameter calibration (FIC) method. The results from the simulation study showed that the proposed method consistently outperformed the compared FPC and FIC methods. The encouraging performance of the proposed method demonstrates the impact of properly accounting for measurement error and provides a justification for its use as a competent small-sample item calibration method.