Small-scale (e.g., classroom) assessment represents the most common and needed scenario for cognitive diagnostic testing. In such settings, polytomously scored items (e.g., constructed-response tasks) are widely used, as they provide more fine-grained measurement of students' skills and cognitive processes. However, a significant gap remains between the current methods and pressing practical needs. On one hand, parametric cognitive diagnosis models capable of handling polytomous response data require large samples for stable estimation, making them unsuitable for small-scale classroom use. On the other hand, existing nonparametric classification methods, while robust in small samples, are largely confined to dichotomous (0/1) response data. There is a lack of dedicated nonparametric methods for polytomous responses, creating a disconnect between practical testing and diagnostic tools. To address this real-world necessity, this study proposes the seq-GNPED method. It extends the generalized nonparametric classification framework to polytomous data by introducing weighted ideal category response and a collapsed class iterative algorithm. Simulations and empirical applications confirm that seq-GNPED achieves robust and accurate diagnosis under small sample conditions where parametric models falter, effectively leveraging the informational richness of polytomous items. This work bridges a critical gap by providing a practical, nonparametric tool tailored for fine-grained, classroom-ready cognitive diagnosis.
With the growing recognition of the advantages of cognitive diagnosis(CD)in psychological and educational measurement,applying the CD framework to test development has become an important research direction in the field of psychology.In the development of cognitive diagnostic assessments,detecting differential item functioning(DIF)remains a crucial quality control procedure to ensure test fairness and validity.However,existing CD-based DIF detection methods typically focus on a single covariate at a time.While these approaches are effective for identifying main effect DIF induced by a single covariate,they are limited in detecting interactive DIF caused by the interaction among multiple covariates.Such limitations may compromise the fairness and interpretability of assessment outcomes.To address this issue,the present study integrates CD modeling with recursive partitioning techniques by proposing a novel DIF detection method,namely the Item-based Sequential Recursive Partitioning Method(ISRPM).Building on the core principles of recursive partitioning,the ISRPM allows the simultaneous consideration of multiple covariates within a single DIF detection procedure and facilitates the identification of both main effect DIF and interactive DIF in cognitive diagnostic assessments. To evaluate the performance of the proposed method,a series of Monte Carlo simulation studies were conducted focusing on two key objectives:(1)examining how factors such as sample size per group,DIF magnitude,DIF type,item quality,correlations among attributes,and the influence of demographic covariates on attribute mastery distribution affect the performance of ISRPM;and(2)comparing ISRPM with several existing DIF detection methods across varied experimental conditions.In addition,to illustrate its practical utility,ISRPM was applied to a cognitive diagnostic version of the Schizotypal Personality Questionnaire(DC-SPQ)and compared with five established DIF detection methods. The results showed that(1)sample size,DIF magnitude,and item quality substantially influenced the performance of all methods;and(2)when items exhibited interactive DIF,ISRPM achieved higher detection accuracy than the Wald,LR,FS-Wald,FS-LR,and Mantel-Haenszel(MH)approaches.When only the main effect DIF was present,the overall performance of ISRPM was comparable to that of the existing methods. These findings suggest that ISRPM provides a flexible and effective framework for identifying both main effect DIF and interactive DIF in cognitive diagnostic assessments,thereby contributing to methodological advancements in fairness evaluation and the broader application of CD-based measurement in psychological and educational measurement.
With the rise of online social activities, sexually objectifying content has become increasingly prevalent, contributing to women’s self-objectification. However, the Online Sexual Objectification Experiences Scale (OSOES) had not yet been validated in a Chinese or non-Western cultural context. This study evaluated the applicability of the OSOES within a Chinese cultural context using Classical Test Theory (CTT) and Item Response Theory (IRT). Participants were 644 Chinese female college students. CTT analyses assessed construct validity, criterion-related validity, and internal consistency. IRT examined item parameters, item and test information, and reliability. Results showed that (1) the three-factor structure of the Chinese version of OSOES (C-OSOES) was supported, with sound internal consistency and criterion-related validity; (2) all 18 items demonstrated appropriate discrimination and location parameters and high precision across most latent trait levels, though category characteristic curves revealed some issues. These findings offer insights for refining the C-OSOES and, more importantly, provide initial psychometric evidence for the applicability of the OSOES three-factor structure among Chinese female college students.
Abstract Background Suicide among university students in China constitutes a critical public health challenge. Although numerous studies have examined risk factors, most have relied on cross-sectional designs. Evidence from longitudinal studies, particularly regarding protective factors, remains limited. This systematic review and meta-analysis synthesizes longitudinal evidence to identify risk and protective factors for suicidal ideation and suicide attempts among Chinese university students. Methods Six electronic databases (three English-language and three Chinese-language databases) were systematically searched for longitudinal studies published up to October 2024. Eligible studies investigated predictors of suicidal ideation or suicide attempts among Chinese university students. Random-effects meta-analyses were conducted to estimate pooled effect sizes, expressed as odds ratios (ORs). Meta-regression analyses were used to examine temporal moderators, and heterogeneity and publication bias were assessed. Results Twenty-two longitudinal studies comprising 953,651 participants were included. The strongest predictors of subsequent suicidal ideation were prior suicidal ideation, anxiety, and early morning awakening. Suicide attempts were most strongly associated with prior non-suicidal self-injury, insomnia symptoms, and academic pressure. Higher levels of social support were associated with a reduced risk of suicidal ideation, whereas other examined protective factors were not statistically significant. Meta-regression analyses indicated that the predictive strength of depressive symptoms for suicidal ideation attenuated with longer follow-up periods. Conclusions Sleep disturbances, academic pressure, and prior suicidal ideation may represent important targets for early identification among Chinese university students. The findings highlight the importance of addressing both psychological vulnerabilities and contextual stressors in suicide prevention efforts. Future longitudinal research should adopt standardized measurements to enhance comparability and inform culturally tailored population-level prevention strategies.
Individuals typically employ multiple cognitive strategies rather than relying on a single approach in decision-making scenarios or problem-solving tasks. With recent advancements in measurement technology, the collection of process data has become increasingly common, with response times (RTs) and eye movement fixation counts (FCs) emerging as critical indicators of cognitive processing. Analysis of RTs and FCs can reveal problem-solving strategies that may not be discernible from response patterns alone. To enhance diagnostic accuracy and provide deeper insights into the cognitive processes underlying strategy selection, this study developed a multi-strategy cognitive diagnosis modeling framework that integrates individual RTs and FCs into a unified framework to define strategy selection (MS-CDM-RTFC). The empirical study utilized data from Raven’s Advanced Progressive Matrices (APM), a widely used measure of nonverbal reasoning and fluid intelligence, to evaluate the practical applicability of the MS-CDM-RTFC model. Simulation results based on the empirical analysis indicate that the MS-CDM-RTFC achieves higher parameter recovery and attribute classification accuracy, demonstrating significantly better performance than traditional multi-strategy models.
In standardized tests, examinees are likely to engage in either one or more following test behaviors: solution behavior, rapid guessing behavior, cheating behavior, nonresponse behavior, etc. Examinees do not always response all items with solution behavior due to various reasons (such as time constraint or low motivation). Aside from solution behavior, rapid guessing, cheating or nonresponse behavior can result in aberrant responses and inaccurate estimates of examinees' ability or trait, as well as item parameters, thus undermining the validity and fairness of the test. To address this issue, this paper aims to propose an IRTree model to that simultaneously considers rapid guessing, cheating and nonresponse behaviors in order to model the various behaviors exhibited by examinees. The proposed model offers a notable improvement over previous studies, as it provides additional classifications for examinee behaviors at both item and examinee levels. Furthermore, it is the first model to separate and simultaneously model guessing and cheating. Two real data sets are utilized to demonstrate the reasonableness and superiority of the proposed model. Subsequently, two simulation studies based on these real data sets are conducted to validate, revealing that it provide more precise estimates of person and item parameters compared to existing models, and explored the boundary condition of model application.
Recent advances in process data collection have made it possible to efficiently collect multimodal behavioral indicators, such as response times and eye-tracking measures. These multimodal data have been widely applied in cognitive and achievement assessments, where they have improved the accuracy of latent construct estimation. However, the use of informative multimodal process data in noncognitive assessments, such as personality measures widely used in organizational research, has received considerably less attention. To address this gap, we integrate response time and eye-tracking data into a conventional item response model to capture respondents' response processes, thereby improving differentiation across trait levels and enhancing noncognitive assessment. Simulation studies were conducted to evaluate the performance of the proposed model and compare it with a conventional IRT model. Results indicate that model parameters can be accurately recovered and that incorporating multimodal data significantly improves the accuracy of person latent trait estimates. Finally, an empirical analysis was conducted to demonstrate the applicability and advantages of the proposed model in personality assessment.
Although forced-choice (FC) assessments with social desirability matching reduce faking compared to Likert scales, the desirability of items may shift when matched in blocks and vary across contexts. Consequently. faking remains a persistent issue in FC assessments, compromising measurement accuracy and fairness. To address this, we propose a statistical model for detecting and mitigating faking in FC assessments, integrating B & ouml;ckenholt's (2014) model of RES faking theory with the Thurstone Item Response Theory (TIRT) model (Brown & Maydeu-Olivares, 2011). Our approach aims to minimize the adverse effects of faking and enhance the robustness of FC measures. Two simulation studies were conducted to evaluate the proposed RES-TIRT model. Simulation Study 1 examined model performance under varying conditions (sample size, FC scale format, item direction, trait correlation, and dimensionality). Results indicated optimal estimation accuracy when using 3-item blocks, 3 dimensions, a correlation of 0 between dimensions, and a mix of positive and negative item descriptions. Simulation Study 2 compared trait estimation accuracy between TIRT and RES-TIRT models under increasing faking prevalence. While the TIRT model performed better in faking-free conditions, its accuracy declined more sharply than the RES-TIRT model as faking increased particularly for item parameters-demonstrating the RES-TIRT model's superior resistance to faking. An empirical analysis further validated the model's applicability in real-world settings. Comparing honest responses with faked responses (simulating lawyer job applications), we found that applicants strategically inflated traits like openness, agreeableness, and extraversion to meet job requirements. The RES-TIRT model effectively detected these distortions, showing significant discrepancies in these dimensions compared to the TIRT model. Additionally, the RES-TIRT model effectively captured response distortion tendencies, as evidenced by significantly elevated latent trait values theta(E)(j) under faking conditions compared to honest responses. This indicates that, the faking behavior of the applicants can be successfully captureti by the RES-TIRT model. Moreover, the difficulty parameter beta(E)(im) triggering fake answers can be observed to determine whether a FC block is prone to faking. These empirically derived parameters enable targeted refinements in FC measure development, allowing test constructors to strategically modify or eliminate items with low faking thresholds, thereby enhancing the scale's overall resistance to response biases. In conclusion, both sinfulation afid effipirical studies irave demonstrated that the ERES-TIRT model is a viable alternative to the TIRT model. It can be employed to address the issue of faking in FC scales, particularly in high-stakes situations such as talent selection.
In testing programs or survey data, it is common to observe careless or inattentive responding due to time constraints, low motivation, or other factors. Such behavior can significantly jeopardize the validity of the test and bias the parameter estimates of examinees. Therefore, detecting such response behavior is of utmost importance. One of the most commonly observed random behaviors is back random responding (BRR). However, existing detection methods for BRR in the framework of cognitive diagnostic assessment (CDA) have shown limited power. Change point analysis (CPA) is a well-established statistical method that can be applied to detect whether aberrant response behaviors exist in a sequence of response data. Although existing CPA methods are mostly used in the item response theory framework and have shown encouraging performance in detecting aberrant response behaviors, it is not yet clear whether and how they can be applied to CDA framework, and what their performance would be in that context. To address these issues, we modify and improve the conventional CPA statistics based on the Bayesian framework and apply them to CDA. We then evaluate and compare their performances of CPA statistics through simulation study. Our results show that the proposed CPA methods have encouraging performances in detecting BRR with higher power while generating a well-controlled Type-I error rate. Finally, we demonstrate the utility of all CPA statistics by applying them to two real datasets.
Measurement and evaluation play a crucial role in psychology and pedagogy, with testing serving as the primary tool for assessment. Researchers and administrators consistently seek methods to accurately assess a subject's traits based on test results. Traditional person traits estimation methods heavily rely on the authenticity of responses and suppose all respondents honestly and normally respond to all items. However, when aberrant responses occur, biased results can arise with traditional methods, thereby diminishing the precision of person trait estimation. Robust estimation method is believed as an effectively method to mitigate the impact of aberrant responses on estimation accuracy. Nevertheless, extant robust estimation approaches, while reducing estimation bias for aberrant test-takers, also impede estimation precision for normal test-takers. To address this issue, we proposed an innovative robust estimation method that can balance the mitigation of aberrant behavior's impact on accuracy with the assurance of precision in normal test-taker estimation. Simulation findings reveal the newly proposed method consistently maintains exceptional estimation accuracy, demonstrating precise estimates even in the absence of anomalous behavior. The empirical study further clarifies the applicability and advantages of our method within psychological and educational assessments.
The ability to rapidly provide examinees with detailed and effective diagnostic information is a critical topic in psychology. Knowing what diagnostic criteria the examinees have met enables the practitioner to seek the solution to help them in a timely manner, and this can be achieved by cognitive diagnostic computerized adaptive testing (CD-CAT). However, the pervasive challenge of replenishing items in the CD-CAT item bank limits its practical application. Online calibration is a means to address item replenishment, but in CD-CAT, most existing online calibration methods that jointly calibrate the Q-matrix and item parameters of the new items are developed only for dichotomous responses and are time-consuming. Notably, previous studies pay no attention to polytomously scored items that are frequently observed in testing, even though they can offer additional evidence for the examinees’ diagnosis. To fill this gap, we propose a SCAD-based method (SCAD-EM) to calibrate the Q-matrix and item parameters of the new items with polytomous response data in order to promote the application of CD-CAT in practice. The performance of the SCAD-EM was investigated in two comprehensive simulation studies and compared against the revised single-item estimation method (SIE-BIC). Results indicated that the SCAD-EM produces a higher calibration accuracy for the category-level Q-matrix and is computationally more efficient across all conditions, but it produces a lower calibration accuracy for the item-level Q-matrix. An empirical study further demonstrated the utility of the SCAD-EM and the SIE-BIC methods in calibrating new items with a real dataset. The advantages of the proposed method, its limitations, and possible future research directions are offered at the end.
The diagnostic classification model (DCM) has been widely utilized in non-cognitive tests, offering diagnostic information on latent attributes. However, the model's reliance on single-stimulus (SS) items may lead to response biases (e.g., social desirability), jeopardizing the psychometric properties. As an alternative to SS scales, forced choice questionnaires (FCQ) can effectively control response biases. The combination of FCQs and the DCM not only circumvents response bias but also yields fine-grained diagnostic information on latent attributes. To the best of our knowledge, only one study (Huang, Educ. Psychol. Meas., 83, 2022, 146) has explored this topic and developed a DCM for forced choice (FC) items. However, the existing model has limitations in terms of its modelling assumption, the FC format and the number of attributes measured by statement. To address these limitations, this study proposes a ranking FC-DCM that (1) adopts a generalized assumption, (2) covers all FC formats and (3) eases the limitation on the number of attributes measured by each statement. The simulation study demonstrated that the proposed model exhibited satisfactory person and item parameter recovery under all conditions. This study provided an illustrative example by developing an FC version questionnaire to further explore the applications and advantages of the proposed model in real-world settings.
Continuous bounded responses in psychometrics usually come from the visual analog scale (VAS). The VAS is a rating scale measurement tool that requires respondents to report their agreement with items by tracing a mark somewhere on a fixed-length continuous horizontal segment with ends that are generally labeled “0% disagreement” to “100% agreement” (or other possible labeling) using continuous data. In recent years, the VAS has gradually appeared in medical, educational, and psychological research, such as research on pain, worry, rumination, anxiety, risk perception, and even personality trait measurement. However, there are very few cognitive diagnosis models (CDMs) in cognitive diagnostic assessment that can analyze such continuous bounded data from VAS-type scale. In this study, we propose a family of CDMs for the continuous bounded data in VAS-type scale and provide model selection methods for practice. Three simulation studies were used to examine parameter recovery, the impact of model misspecification on parameter recovery, and the effectiveness of the model selection method. Moreover, real data are used as an illustration to demonstrate the application and effectiveness of the proposed models.
Traditional IRT and IRTree models are not appropriate for analyzing the item that simultaneously consists of multiple-choice (MC) task and constructed-response (CR) task in one item. To address this issue, this study proposed an item response tree model (called as IRTree-MR) to accommodate items that contain different response types at different steps and multiple different cognitive processes behind each score to effectively investigate the cognitive process and achieve a more accurate evaluation of examinees. The proposed model employs appropriate processing function for each task and allows multiple paths to an observed outcome. The simulation studies were conducted to evaluate the performance of the proposed IRTree-MR, and results show the proposed model outperforms the traditional IRT model in terms of parameters recovery and model-fit. Moreover, an empirical study was carried out to verify the advantages of the proposed model.
For various reasons, respondents to forced-choice assessments (typically used for noncognitive psychological constructs) may respond randomly to individual items due to indecision or globally due to disengagement. Thus, random responding is a complex source of measurement bias and threatens the reliability of forced-choice assessments, which are essential in high-stakes organizational testing scenarios, such as hiring decisions. The traditional measurement models rely heavily on nonrandom, construct-relevant responses to yield accurate parameter estimates. When survey data contain many random responses, fitting traditional models may deliver biased results, which could attenuate measurement reliability. This study presents a new forced-choice measure-based mixture item response theory model (called M-TCIR) for simultaneously modeling normal and random responses (distinguishing completely and incompletely random). The feasibility of the M-TCIR was investigated via two Monte Carlo simulation studies. In addition, one empirical dataset was analyzed to illustrate the applicability of the M-TCIR in practice. The results revealed that most model parameters were adequately recovered, and the M-TCIR was a viable alternative to model both aberrant and normal responses with high efficiency.
Response times (RTs) facilitate the quantification of underlying cognitive processes in problem-solving behavior. To provide more comprehensive diagnostic feedback on strategy selection and attribute profiles with multistrategy cognitive diagnosis model (CDM) and utilize additional information for item RTs, this study develops a multistrategy cognitive diagnosis modeling framework combined with RTs. The proposed model integrates individual response accuracy and RT into a unified framework to define strategy selection and make it closer to the individual's strategy selection process. Simulation studies demonstrated that the proposed model had reasonable parameter recovery and attribute classification accuracy and outperformed the existing multistrategy CDMs and single-strategy CDMs in terms of performance. Empirical results further illustrated the practical application and the advantages of the proposed model.
Multidimensional three-parameter logistic model (M3PLM) often faces poor recovery of parameters due to overestimation or underestimation of guess parameters. In order to more reasonably consider the participant’s guesses, this study proposed a new multidimensional IRT modeling based on 2PLE, which equal to the guessing parameter with a function that integrates item characteristic according to the ability level of an examinee to quantify guessing behavior. Our new model, named as multidimensional two-parameter logistic model with ability-item-based guessing (M2PL-AIG) model, not only the number of parameters is smaller than the M3PL model, but also considers the participant’s guesses just like M3PL model. Two simulation studies were conducted to fully comparing the performance of the M2PL-AIG model with the M3PL model. In addition, real data from TIMSS 2011 is used as an application illustration and to compare with the M3PL model. The results show that the M2PL-AIG model outperforms the M3PL model in parameter recovery and model fit, which indicates that the M2PL-AIG is a rather interesting proposal, surely worth some further study and the consideration by the IRT community.
In recent decades, multidimensional forced-choice (MFC) tests have gained widespread popularity in organizational settings due to their effectiveness in reducing response biases. Detecting differential item functioning (DIF) is crucial in developing MFC tests, as it relates to test fairness and validity. However, existing methods appear insufficient for detecting DIF induced by the interaction between multiple covariates. Furthermore, for multi-category, ordered or continuous covariates, existing approaches often dichotomize them using a-priori cutoffs, commonly using the median of the covariates. This may lead to information loss and reduced power in detecting MFC DIF. To address these limitations, we propose a method to identify both main effect DIF and interactive DIF. This method can automatically search for the optimal cutoffs for ordered or continuous covariates without pre-defined cutoffs. We introduce the rationale behind the proposed method and evaluate its performance through three Monte Carlo simulation studies. Results demonstrate that the proposed method effectively identifies various DIF forms in MFC tests, thereby increasing detection power. Finally, we provide an empirical application to illustrate the practical applicability of the proposed method.
Cognitive diagnostic computerized adaptive testing(CD-CAT)provides a detailed diagnosis of an examinee's strengths and weaknesses in the content measured in a timely and accurate manner,which can be used as a reference for further study or remediation planning,thus meeting the practical need for efficient and detailed test results.The successful implementation of CD-CAT is based on an item bank,but its maintenance is a very challenging task.A psychometrically popular choice for maintaining an item bank is online calibration.Currently,the research on online calibration methods in the CD-CAT that can calibrate Q-matrix and item parameters simultaneously is very weak.The existing methods are basically developed based on the deterministic input,noisy and gate(DINA)model.Compared with the DINA model,the generalized DINA(G-DINA)model has been more widely applied because it is less restrictive and can meet the requirements of a large number of test data in psychological and educational assessment.Therefore,if the online calibration method that jointly calibrates the Q-matrix and item parameters can be developed for models with few constraints such as G-DINA,its meaning is understood without explanation. In current study,a new online calibration method,SCADOCM,was proposed,which was suitable for the G-DINA model.The construction of SCADOCM was based on the smoothly clipped absolute deviation penalty(SCAD)and marginalized maximum likelihood estimation(MMLE/EM)algorithm.For the new item j,the log-likelihood function with SCAD can be formulated based on the examinees'responses in this item and the examinees'attribute marginal mastery probability,and the q-vector of the new item can be estimated by the q-vector estimator based on SCAD.Then,the EM algorithm was used to estimate the item parameter of the new item j based on the posterior distributions of examinees'attribute patterns,the examinees'responses to new item j and the estimated q-vector. To examine the performance of the proposed SCADOCM and compare it with the SIE method,two simulation studies(Study 1 and Study 2)are conducted.Study 1 is based on a simulated item bank while Study 2 is based on the real item bank(Internet addiction item bank;Shi,2017).In these simulation studies,four factors were manipulated:the calibration sample size(nj=50 vs.100 vs.500 vs.1000 vs.2000),the distribution of the attribute pattern(uniform distribution vs.high-order distribution vs.normal distribution),the item quality(U(0.05,0.15)vs.U(0.1,0.3)),and the online calibration methods(SCADOCM vs.SIE).The results showed that(1)SCADOCM has satisfactory calibration accuracy and calibration efficiency,and is superior to the SIE method.In addition,the traditional SIE method is not applicable for the G-DINA model,and its Q-matrix estimation accuracy rate is low under all experimental conditions.(2)The item calibration accuracy of SCADOCM and SIE increases with the increase of calibration sample and item quality under most conditions,and its item calibration accuracy in the uniform distribution/higher-order distribution is greater than that in the normal distribution.(3)The calibration efficiency of SCADOCM decreases with the increase of calibration samples,but it is less affected by the item quality and the attribute pattern distribution;the calibration efficiency of SIE decreases with the increase of calibration samples,but it is less affected by the item quality.Moreover,the calibration efficiency of the SIE method in the normal distribution is slightly slower than that of uniform distribution/high-order distribution. To sum up the results,this study demonstrated that the SCADOCM has higher item calibration accuracy and calibration efficiency,and outperforms the SIE method;meanwhile,the traditional SIE method is not suitable for G-DINA model.All in all,this study provides an efficient and accurate method for item calibration in CD-CAT,and provides important support for further promoting the application of CD-CAT in practice.
As one of the three broad types of test cheating, item preknowledge has always been a severe threat to test validity and is widespread in various testing programs, especially within some high-volume certification testing programs. While many response-based methods have been proposed to identify and handle item preknowledge, most require the assumption that the compromised items are known. Moreover, few studies have considered that both examinee and item characteristics might affect this cheating behavior, and even fewer have considered the relationship between ability and such behavior when estimating ability. Therefore, this article proposes a mixture model with less strict assumptions on the compromised items. By modeling cheating behavior with a latent response approach, the model takes into account the effect of both characteristics at the item-by-examinee level and assesses how a person’s ability relates to such behavior. Two simulation studies demonstrate that the parameters of the proposed model can be effectively recovered, and when data contains item preknowledge, the model generally produces more accurate ability estimates than existing models. Finally, an empirical example based on a licensure test dataset illustrates the applicability of the new model.