
We advocate that the educational measurement community take a proactive stance to provide tools that improve the fairness of the use of standardized tests in college admissions, above and beyond removing biased items. We believe this can be achieved without resorting to a complete abandonment of comparability. Drawing on the evolution of high-stakes testing in ancient China, we examine recurring tensions between standardization and equity that resonate with contemporary U.S. admissions debates. We then describe elements thought to influence fairness in admissions. We argue that an alternative statistical formalism can clearly state what additional information would be necessary to reduce or perhaps eliminate the distributional differences in test scores across subgroups that lie at the core of the claim that standardized testing itself is unfair. We believe that more data, rather than less, when rigorously measured and contextualized, can do more to promote equitable access to higher education.
Fairness issues in admissions testing date from the very inception of our field over 100 years ago. This article selectively recounts that history through the time of this writing. As other recountings have noted, conceptions of fairness have evolved considerably, sometimes jarringly so. In the current recounting, the conceptual evolution of the idea is constructed as initially centering on fairness for individuals, engineered through standardization. That idea was later complemented by an awareness of the need to conceptualize fairness for groups. Last is an emerging return to fairness at the individual level but, rather than achieved through standardization, engineered through personalization.
Multidimensional item response theory (MIRT) model and multiple group item response theory model (MGM) are widely applied for measuring students' learning and change. Recently, there are several estimation methods for multidimensional item response theory (MIRT) model and multiple group item response theory model (MGM), but previous studies have not applied these methods to longitudinal data. This study investigated the performance of MIRT and MGM with two popular estimation methods for the longitudinal data, including the Metropolis-Hastings Robbins-Monro (MH-RM) algorithm and the second-order Laplace (Lap2) approximation. The results suggest that Lap2 outperforms both MH-RM and the expectation maximization (EM) algorithm under the MIRT model and MGM in longitudinal data. The MIRT model with Lap2 produces more accuracy of ability estimates than the MGM with both EM and Lap2, especially when the correlation between abilities is 0.8. When the number of time points is 2 and 3, the MIRT model and MGM yield highly similar results in terms of RMSE of ability estimates, ability means, and item parameters, regardless of the estimation method.
This paper examines the use of generative large language models (LLM) to produce selected-response (SR) test items. In study 1, we compared 201 prompt-engineered LLM-generated items to 52 items developed by subject matter experts (SME). We found that 60% of LLM-generated items met psychometric criteria, compared to 74% of SME items. LLM-generated items had an average difficulty (0.73) and discrimination (0.16) similar to SME items (0.78 and 0.18). However, LLM-generated items exhibited a higher frequency of psychometric flags (e.g., high difficulty) and higher initial rejection rates. In study 2, a fine-tuned LLM generated 154 SR items across two batches. Fine-tuning improved item retention rates by 16% when compared to prompt engineering and by 10% between batches. Batch performance showed stable difficulty (0.69) and improved discrimination (0.24 to .27). While LLM cannot yet replace SME-driven item development, it may offer efficiency gains when integrated into the item development model.
This study advances hierarchical rater modeling by relaxing the common assumption of equal discrimination across latent classes in constructed-response (CR) scoring. Building on univariate signal detection approaches, we propose a hierarchical model with ordered perceptual distributions, allowing rater discrimination to vary across score categories within an item. Through simulation studies, we evaluate parameter recovery, compare performance with equal-discrimination models, and assess how fit indices can identify the correct model. An empirical application to a financial assessment dataset illustrates practical utility. Diagnostic outputs, including tables and visualizations, demonstrate how the model can inform rubric refinement and support targeted rater training in multi-item CR assessments.
Test-taking disengagement in international large-scale assessments (ILSAs) is a widely explored topic, but little is known about students' disengagement in ILSA surveys or how it compares internationally. In this study, we used data from PIRLS 2021 to compare six measures of survey disengagement across 25 countries. We also tested how the relationships between these disengagement measures, the survey scores, and reading achievement vary internationally. Countries accounted for a significant variation in the disengagement measures, especially in the percentage of slow responses (eta(2) = .11). Fast responses predicted experiencing bullying and poorer reading achievement (beta = 0.26, -0.37, respectively). However, there were substantial variations between countries in the relationships between the disengagement measures and students' scores. Our results suggest that there are meaningful differences between countries in survey disengagement and in the associations between disengagement and the target scores, such that uniform criteria for identifying disengaged survey-takers might not be appropriate.
This paper presents the automated assembly of multidimensional forced-choice questionnaires (MFCQs) using mixed integer programming (MIP). The unique features of using MIP to assemble MFCQs, distinct from those of traditional cognitive tests and Likert scales, are emphasized. The assemblies are demonstrated in three use cases: one form, two parallel forms, and one multistage test (one Stage 1 routing form and three Stage 2 forms). The objective functions are to maximize the expected test information. The MIP models are implemented using the R package, eatATA, an MIP test assembly tool. Simulation results provided empirical evidence that the test assembly objectives were achieved using MIP. The benefits and limitations of MIP are discussed.
Multiple-choice (MC) items have been used in educational testing for almost a century. The shift from paper-and-pencil to online administration has dramatically increased the popularity of technology-enhanced (TE) items relative to MC items. However, little research has examined how to use item statistics to develop high-quality TE items. The present study extended previous research on using item analysis for MC item development to TE item development. The proposed data visualization has been used in actual data review sessions. Feedback from subject matter experts (SMEs), along with item performance after the suggested revisions, demonstrates that the data visualization introduced in this paper is effective in helping SMEs identify problematic components of TE items. This, in turn, leads to more efficient and targeted revisions of TE items.
Effective score report design requires alignment between reporting practices, users, and the interpretations and uses that the assessment is intended to support. Although parents frequently request specific, actionable information about their child's academic progress, large-scale assessment reports rarely provide that level of diagnostic detail. Assessments scored with diagnostic classification modeling (DCM) can provide mastery information at the skill or attribute level, better meeting parents' information needs. However, parents' perceptions of mastery-based score reports have not yet been studied. To evaluate parents' perceptions of the design, interpretability, and usability of mastery-based score reports, we conducted feedback sessions with 21 parents and guardians in three states. Parents generally understood mastery-based reporting and found the reports useful for understanding their child's knowledge in the subject and for communicating with teachers and supporting learning at home. The findings expand current evidence on stakeholder interpretation of diagnostic score reports in operational contexts.
Careless responding (CR) threatens validity evidence, yet its potential to yield spurious factors when testing unidimensionality remains underexplored. We simulated polytomous data for 500 participants, manipulating CR type (fixed, midpoint, random), prevalence (10%, 20%, 30%), and severity (25%, 50%, 75%) in a fully crossed 27-condition design with 500 replications each. Unidimensionality was assessed using parallel analysis (PA; principal axis factoring with polychoric correlations) and item factor analysis (IFA). For IFA, one- and two-factor models were fit and compared using descriptive indices and inferential tests. The outcome was Type I error, defined as falsely concluding multidimensionality in truly unidimensional data. PA falsely detected additional factors in 52% of conditions. Within IFA, BIC was most robust (21% Type I error). Two of four three-way interactions were significant: Method x Type x Prevalence and Type x Prevalence x Severity, explaining 23% and 20% of the variance in Type I error. Results underscore cleaning low-stakes data before assessing unidimensionality.
Logistic regression is one of the common methods for differential item functioning (DIF) detection and is typically estimated by maximum likelihood estimation (ML), despite the fact that ML may produce biased estimates in situations of rare event data and small sample sizes. Firth's penalization method (PML) may address this issue. In the current study, we compared PML with ML estimation in logistic regression for polytomous DIF detection, focusing on rare event, small, and unbalanced data. The manipulated factors included item average difficulties, sample sizes, group impact, and DIF types and magnitudes. PML demonstrated higher power than ML in unbalanced samples, especially in non-uniform DIF conditions, but also demonstrated a trade-off of having higher Type I error rates than ML. We provide practitioners with DIF estimation suggestions across various sample sizes and item difficulties with respect to concerns of power and Type I error rates.
Multiple-choice is the most common test format in educational assessment. It offers many advantages but has one shortcoming: examinees might earn extra points through strategic guessing. Item-writing guidelines on response option placement have been formulated to prevent examinees from strategically using the answer keys' distribution as clues. However, several guidelines co-exist in literature, untested by mathematical methods. In this study, the answer keys' distributions that may emerge when implementing the most common guidelines on answer keys' placement (balancing and randomization) were modeled, and the effectiveness of using well-known option-position-based response strategies ("when in doubt, choose C"; balanced guessing) was analyzed through a simulation study. Gains obtained in the different scenarios were compared for tests with different lengths and examinees with different achievement levels. Results show that randomization is the only guideline not associated with potential test results validity or equity issues. A set of recommendations for test developers is provided.
Mathematics is a core domain in large-scale assessments (LSA), yet item development remains resource-intensive, limiting scalability and innovation. Automatic Item Generation (AIG) offers a promising solution, but empirical validations remain rare. This study investigates the psychometric functioning and fairness of 48 cognitive item models designed to generate language-reduced, image-based math items for Grades 1, 3, and 5. Treating these models as proto-theories, we generated 612 item instances varying in cognitive demands and contextual features. Using data from Luxembourg's school monitoring (N = 35,058), we found that item difficulty was mainly driven by predefined cognitive factors, with stronger contextual influences in early grades. We introduce Differential Radical Functioning to evaluate whether AIG-based items permit comparable score interpretations across subgroups. Results reveal meaningful differences by cultural background, regardless of language proficiency. These findings highlight the importance of contextual embedding and demonstrate the potential of cognitive modeling in AIG for scalable, valid, and equitable assessments.
Teacher efficacy measurements predominantly rely on self-report Likert scales, raising concerns about inter-individual comparability due to variations in response patterns. This study explored the anchoring vignette (AV) method as a complementary approach by examining how domain-general and domain-specific AV corrections influence self-reported teacher efficacy across three domains: instructional strategies, classroom management, and innovative learning support. Findings from a sample of 2,803 Korean elementary teachers revealed that while both AV correction methods maintained the original factor structure, they differentially impacted score distributions and interpretations. Domain-general AV correction increased correlations between subdomains, while domain-specific correction enhanced discriminant validity by decreasing subdomain correlations. Both methods improved scale reliability and measurement precision, though they showed reduced covariance with external variables, potentially addressing common method variance. The findings demonstrate how different AV approaches can enhance the validity of teacher efficacy measurements by accounting for individual differences in capability conceptualization.
The authors investigated using a large language model (LLM) for writing test questions for a real estate licensing exam. In Study 1 items were generated by GPT-4 and rated by subject matter experts (SMEs). These items were on-topic,relevant, and generally appropriate. Item difficulty manipulation was ineffective. Cognitive level matching was harder as cognitive level increased. Study 2 compared human and LLM items using SME and content developer ratings. Human and LLM items were similar in blueprint alignment, relevance, factual errors, and key quality. LLM items had better stem quality and cognitive level matching. Human distractors had an edge in quality. In Study 3 investigated content overlap and breadth of coverage. Similar prompts frequently generated overlapping content. The range of content represented in large sets of generated items did not cover the breadth of the generating content areas. Results suggest LLMs are as good as SMEs at generating first-draft items.
Think-aloud methods provide a window into test-takers' real-time cognitive processes, offering valuable evidence for validity based on response processes. Yet, their use in complex verbal assessments remains limited, particularly concerning the choice between concurrent and retrospective protocols. This multiple-method study directly compared these two approaches in the context of a high-stakes verbal selection test, involving 10 biomedical students equally divided between conditions. Using a combination of deductive coding (based on the assessment blueprint) and inductive thematic analysis, we explored the richness and nature of participants' explanations. Quantitative analysis (interpreted cautiously due to the small sample) further supported the conclusions from the qualitative findings. Participants in the concurrent condition verbalized more frequently, often reading items aloud and articulating their reasoning with greater clarity and depth. In contrast, retrospective participants tended to offer shorter, more fragmented responses, with less transparency in their thought processes. Quantitative results revealed higher assessment scores and significantly more verbal engagement in the concurrent group. Moreover, participants reported the concurrent think-aloud as easier and more natural to perform. These findings may challenge the prevailing assumption that concurrent think-aloud is unsuitable for verbal tasks due to cognitive load. Instead, our results suggest that it can elicit richer, more authentic process data, even in complex verbal assessment contexts, although the preliminary nature of the current study must be emphasized. This study offers a practical contribution to the collection of response process evidence, with clear implications for both researchers and practitioners aiming to enhance the validity and design of educational assessments.