
In educational testing, differential item functioning (DIF) examines whether item parameters differ across examinee characteristics. In this study, DIF is viewed not as bias but as item parameter heterogeneity reflecting differential engagement with related subskills. Using recursive partitioning based on parametric logistic item response theory (PL IRTree) models, this study investigated item functioning heterogeneity in a high-stakes reading comprehension test regarding vocabulary, grammar, and gender based on responses of 14,936 examinees. The comparison of Rasch tree, 2PL, 3PL, and 4PL IRTrees showed that the Rasch tree could better capture the structure of the data. The Educational Testing Service (ETS) classification scheme using the Mantel-Haenszel log odds ratio was then used to quantify effect sizes. The Rasch tree yielded 11 nodes, indicating subgroup differences in item difficulty, with seven items showing moderate parameter variation across five nodes. Except for one vocabulary-focused item showing systematic sensitivity to vocabulary knowledge, item parameter variation associated with lexico-grammatical knowledge was small, suggesting limited heterogeneity attributable to these subskills. Gender-related differences in item functioning appeared only within a subgroup of examinees with relatively high grammar-vocabulary scores. PL IRTree models can reveal item-level heterogeneity and the role of correlated subskills in shaping reading comprehension item functioning.
Complex syntax in psychological test items may hinder some respondents' comprehension, yet the prevalence of specific syntactic complexity features and their implications for item responding and validity are not well understood. This two-part study provides new empirical insight into these relationships. Study 1 examines whether different types of dependent clauses are associated with item responses and whether these associations vary according to respondents' vocabulary familiarity. Study 2 uses a corpus-based register analysis to determine how frequently dependent clause types occur in psychological test items, providing context for the practical significance of the Study 1 findings. Of the five clause types examined in Study 1, only nonfinite complement clauses showed a significant association with item responses, and this association was moderated by vocabulary familiarity. Study 2 further showed that nonfinite complement clauses are the most prevalent dependent clause type in a corpus of psychological test items. Together, these findings suggest that dependent clause types should not be treated as uniformly difficult to process. This study demonstrates how integrating psychometric modeling with corpus-based linguistic analysis can build validity evidence and inform fairer assessment practices for linguistically diverse populations.
This study investigates the cross-national measurement invariance (MI) of the Information and Communication Technologies (ICT) Efficacy Scale in PISA 2022 across 26 OECD countries. ICT efficacy, defined as students' confidence in using digital tools for learning, is central to understanding educational outcomes in technology-rich environments, yet cross-country comparability cannot be assumed. Using the alignment method, the study identifies substantial non-invariance in both factor loadings and intercepts, indicating that direct comparisons of observed ICT efficacy scores may be misleading. More than 25% of model parameters exceeded recommended non-invariance thresholds in the sample of 176,621 students. Monte Carlo simulations confirmed the robustness of the alignment estimates, supporting valid comparisons of latent factor means. These findings highlight the importance of diagnostic evaluation of cross-national instruments and provide evidence-based guidance for policymakers and researchers seeking accurate and equitable assessments of students' ICT-related competencies.
This systematic review maps and synthesizes psychometric evidence from the past decade on digital and online tools assessing executive functions (EF) in healthy adults. It examines six contemporary domains of validity: content validity, structural validity, external validity, response processes, consequential validity, and reliability. Searches across PsycNet, PubMed, Embase, Web of Science, and the Virtual Health Library followed PRISMA guidelines. Thirty-one studies met inclusion criteria, encompassing 11,246 participants. Most tools assessed core EF domains-working memory, inhibition, and cognitive flexibility-through performance-based tasks or online questionnaires. Reliability was reported in 23 studies, though often via single indices. Content validity appeared in 26 studies but frequently lacked methodological rigor. Structural and external validity were reported in 8 and 17 studies, respectively. Evidence regarding response processes (k = 22) and consequential validity (k = 31) was commonly cited but rarely examined in depth. No study comprehensively addressed all six domains. Risk of bias was low for administration but high for participant selection, and applicability concerns included unrepresentative samples and limited construct alignment. Overall, although digital EF assessment is expanding, psychometric validation remains uneven across key domains.
A set of items that includes a common stimulus such as a reading passage, picture, graphic, or scenario is called testlet and has been used frequently in educational assessments. However, due to the presence of a common stimulus in testlet, it is possible that the local independence assumption may be violated. The violation of local item independence due to testlets may also affect the performance of computerized adaptive testing (CAT). This study examines how different factors impact the performance of testlet-based CAT in a simulation study. The manipulated conditions included adaptation level (testlets vs. items), testlet/item selection approaches, the magnitude of testlet dependencies, and test length, testlet size, and number of testlets. One of the important findings of the study is that adaptation not only between testlets but also within testlets improved the performance of CAT, resulting in less biased latent trait estimation and a smaller standard error of measurement.
Nonparametric methods adapted from Mokken Scale Analysis (MSA) offer an exploratory approach to evaluating the quality of rater judgments in performance assessments that is grounded in invariance requirements. However, typical applications of MSA-based methods require complete data, where all raters score all examinees. We use a real data illustration and simulation study to demonstrate how invariant ordering analyses adapted from MSA can be used with incomplete rating designs. We focus on anchor rating designs, where all raters score a common group of examinee performances, but all other performances are scored by two randomly selected raters. Our real data demonstration uses data from a rater-mediated music performance assessment, and our simulation study includes conditions that vary with respect to sample size, anchor set characteristics, and the presence and type of rater effects. Our results support the use of the adapted MSA techniques to evaluate rating quality in performance assessment contexts with anchor rating designs. We discuss the implications of our results for research and practice.
International large-scale assessments (ILSAs) like TIMSS, PIRLS, and PISA use planned missingness completely at random (MCAR) but suffer from non-response bias with missing data not at random (MNAR), especially in home questionnaires where parental omissions may correlate with item content, compromising cross-country comparisons. This study evaluated whether PIRLS 2021's Home Socioeconomic Status (HSES) could be reliably imputed using Multivariate Imputation by Chained Equations (MICE) based on student and school data. Six methods-Predictive Mean Matching (PMM), Weighted PMM, CART, Random Forests, Bayesian Linear Regression, and Lasso-were evaluated on 5,500 students with 10-60% simulated missingness. Imputation accuracy declined from 36% explained variance (correlation similar to 0.6) at 10% missingness to 16% (correlation <0.4) at 60% missingness. PMM performed best but exhibited shrinkage, overestimating low HSES and underestimating high HSES. Results underscore the critical need to address MNAR in ILSAs to ensure equitable and reliable international educational comparisons.
This study investigates the accuracy and reliability of large language models (LLMs) in scoring writing tasks from the Advanced Placement (AP) Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and compares score consistency (reliability) and alignment with human raters (accuracy) across two types of AP Chinese free-response writing tasks: story narration and email response. These essays were independently scored by two trained human raters and seven artificial intelligence (AI) raters. Each essay received four scores: one holistic score and three analytic scores corresponding to the domains of task completion, delivery, and language use. Results indicate that although human raters produced more reliable scores overall, LLMs demonstrated reasonable consistency under certain conditions, particularly for story narration tasks. Composite scoring that incorporates both human and AI raters enhanced score dependability and human-AI agreement, which supports that hybrid scoring models improve both reliability and accuracy in large-scale writing assessments.
This study examined how acquiescent responding influences vocational-interest items and evaluated the equivalence of a new Pictorial Inventory of Vocational Interests (PIVI) and a parallel verbal form (VIVI). A sample of 179 participants completed the instruments and a brief set of "bogus" preference items indexing acquiescence. Using SEM, we fit baseline CFA models and MIMIC models with the acquiescence index predicting item responses. Internal-structure analyses showed acceptable fit consistent with the RIASEC model in both formats, with adequate factor loadings. We then tested invariance of item parameters across formats (configural, metric, scalar): model fit changed minimally when factor loadings and thresholds were constrained equal (Delta CFI < .01; Delta RMSEA < .01), supporting metric and scalar invariance. Acquiescence exerted direct effects on multiple items in both formats, with the pictorial version slightly less vulnerable. At the scale level, acquiescence correlated more strongly with profile elevation than with differentiation. Overall, pictorial assessments appear to be a viable alternative for measuring vocational interests, and the results underscore the importance of modeling response bias when measuring vocational interests.
Low test-taking effort remains a persistent problem in low-stakes assessments as it introduces construct-irrelevant variance into the scores which can compromise the validity interpretations made with these scores. This study applies the IRTree Model for Disengagement to jointly estimate disengagement and science ability in the U.S.A. sample of the 2022 Programme for International Student Assessment. We examined how disengagement relates to student-level characteristics, including gender, grade level, whether the language spoken at home is different from that of the test, socioeconomic status, growth mindset, sense of belonging, family support, and experiences of bullying. Results indicated that male students and those in higher grade levels were significantly more likely to disengage. Language background also exhibited a credible effect, with students speaking English at home showing higher disengagement. Socioeconomic status, sense of belonging, and family support showed directional tendencies that provided moderate evidence of association with disengagement, whereas growth mindset, bullying, and educational expectations exhibited no significant effects. A strong negative latent correlation between science ability and disengagement confirmed that lower-performing students were substantially more likely to disengage. The study underscores the importance of accounting for individual differences in test-taking motivation interventions and highlights the utility of IRTree models for producing more valid inferences from low-stakes assessment data.
Large language models (LLMs) have transformed the field of natural language processing (NLP), demonstrating remarkable performance across a wide range of NLP tasks. This article provides an overview of the fine-tuning of LLMs, with a particular focus on their application in psychological and educational assessment contexts. Fine-tuning is introduced as a method for adapting pretrained LLMs to domain-specific tasks, allowing them to capture the nuanced complexities inherent in psychological and educational constructs. We present a tutorial on two different approaches to fine-tuning LLMs: open-source and closed-source (also known as proprietary) models. Using two illustrative examples, we guide readers through the step-by-step process of fine-tuning LLMs for real-world assessment scenarios, highlighting the challenges and considerations involved. This paper aims to serve as a practical resource for researchers, providing the knowledge and tools needed to effectively leverage LLMs in applied assessment contexts.
Psychological testing is an important component of professional practice, but its use in Latin America remains insufficiently documented. To address this, we surveyed 2,319 psychologists from nine countries using a modified European Federation of Psychologists' Associations (EFPA) questionnaire. The study examined perceptions of test misuse and training, professional attitudes toward testing, and the need for regulatory frameworks. Measurement invariance was assessed, and age and gender effects were analyzed with linear mixed-effects models. Results indicate generally positive attitudes toward testing, together with persistent shortcomings in psychometric training and regulation. Age was associated with differences in attitudes and perceived need for regulation, whereas gender showed no significant effects. The findings underscore the importance of strengthening training and regulatory systems and fostering collaboration among national and international institutions to advance standardized testing practices across Latin America.
Careless and inattentive responding pose a significant challenge to the validity of psychological tests. However, existing indicators often fail to fully utilize the available data and lack sensitivity to partial degradation. This study introduces a machine learning-based item prediction approach to assess the extent to which responses deviate from expected patterns by predicting each item's response based on the responses to the remaining items. Data were collected from 941 participants who completed a 195-item scale measuring clinical symptoms and personality traits (Millon Clinical Multiaxial Inventory-IV). The dataset was divided into a training set, a test set, and a partially randomized set, and the effectiveness was examined using indicators such as the item prediction approach, Mahalanobis distance, and person-fit statistics. The results showed that the item prediction approach showed better performance in differentiating partially randomized datasets compared to other methods. Moreover, the item prediction approach showed enhanced performance within the ranges where the proportion of randomized data was relatively small, or the training dataset was sufficiently large. This study suggests that utilizing extensive item information leveraged through machine learning can effectively detect careless response patterns.
Effort tends to be lower and more variable in low-stakes versus high-stakes testing contexts. To understand individual differences in effort, we examined students' perceived normativity of giving effort on low-stakes assessments. During a low-stakes testing session, 794 first-year college students indicated they personally believed that students should expend high effort on these tests (high personal normative beliefs; PNB) while also believing that only half of students expend high effort (low empirical expectations; EE) and that other students believe minimal effort is acceptable (low normative expectations; NE). After controlling for personality traits and gender, we found that EE and PNB had significant direct and indirect associations with effort and had significant indirect associations with test performance via effort. In short, perceived norms related to individual differences in how students behaved and performed in low-stakes testing contexts, which encourage future study of perceived norms about expending effort during low-stakes testing.
Examinations of the internal structure of the Depression, Anxiety, and Stress Scale-21 (DASS-21) have yielded inconsistent conclusions within and across cultural contexts. This study examined the dimensionality and reliability of the DASS-21 across three theoretically plausible factor structures (i.e., unidimensional, oblique three-factor, and bifactor) as well as measurement equivalence/invariance of the DASS-21 using two different approaches (i.e., multigroup confirmatory factor analysis and the alignment approach) with a large, diverse sample of 2,920 young adult college student participants from nine countries/regions (i.e., Australia, Brazil, Germany, Hong Kong, Lithuania, Taiwan, T & uuml;rkiye, United Arab Emirates, and the United States). Results showed an excellent fit of the bifactor model in all countries/regions except the UAE and the US in which the model did not converge. Regarding parameter equivalence, we found configural, threshold, and loading invariance for the oblique three-factor model (across the nine studied countries/regions) and for the bifactor model (across seven countries/regions). Results indicate that DASS-21 scores measure a general psychological distress factor with more validity and reliability than depression, anxiety, or stress constructs independently. Findings supported the bifactor structure of DASS-21 and demonstrated that cross-cultural comparisons using this scale should be conducted using proper procedures, such as the alignment approach.
This study conducted a comprehensive comparison of Item Response Theory (IRT) linking methods applied to a bifactor model, examining their performance on both multiple choice (MC) and mixed format tests within the common item nonequivalent group design framework. Four distinct multidimensional IRT linking approaches were explored, consisting of two methods that incorporated linking coefficients, namely extensions of Haebara and Stocking-Lord, and two methods that did not involve linking coefficients, specifically concurrent calibration and fixed item parameter calibration (FIPC). The study involved an evaluation of the linked item parameters, with a focus on their bias, standard error of estimate (SEE), and root mean squared error (RMSE). The findings revealed that both linking methods, those with and without linking coefficients, demonstrated proficient recovery of item parameters. Notably, the methods lacking linking coefficients exhibited superior performance compared to their counterparts with coefficients. Remarkably, the FIPC linking method emerged as particularly adept at recovering item parameters, especially with regard to difficulty parameters, within the context of the bifactor model.
This article describes the 2022 ITC/ATP Guidelines for Technology-Based Assessment (TBA), a collaborative effort by the International Test Commission (ITC) and the Association of Test Publishers (ATP) to address digital assessment challenges. Developed by over 100 global experts, these Guidelines emphasize fairness, accessibility, security, and psychometric quality. Tracing their evolution from prior standards, the article highlights the Guidelines' relevance to current technological and ethical issues in educational and psychological assessments. The Guidelines serve as a framework for equitable, high-quality TBA practices worldwide, supporting accessible and valid assessments and adapting to technological advancements in many contexts, including under-resourced settings.
Acquiescence is the tendency of participants to shift their responses to agreement. Lechner et al. (2019) introduced the following mechanisms of acquiescence: social deference and cognitive processing. We added their interaction into a theoretical framework. The sample consists of 557 participants. We found significant medium strong relationship between acquiescence and social deference and significant but negligible correlation between acquiescence and selective attention, perception of task difficulty, and cognitive reflection, and non-significant relationships between acquiescence, verbal cognitive reflection, and general cognitive factor. We did not find significant interactions between social deference and factors of cognitive processing. It is possible that the selected cognitive factors do not play a role in explaining acquiescence. Future studies should focus on finding another way of measuring task difficulty.
We explored the practicality of relatively small item pools in the context of low-stakes Computer-Adaptive Testing (CAT), such as CAT procedures that might be used for quick diagnostic or screening exams. We used a basic CAT algorithm without content balancing and exposure control restrictions to reflect low stakes testing scenarios. We examined the effects of small item pools under various testing conditions using a series of Monte Carlo simulations. We examined the effects of these conditions on the accuracy, precision, and stability of examinee achievement estimates. Our results showed that the effects of item pool size are strongest when there is less-precise targeting between item and person location parameters. Our findings suggest that small item pools can effectively support satisfactory performance in CAT, particularly when examinee targeting is adequate and a variable-length test with a standard error stopping rule is implemented. We consider implications for research and practice.