
Developmental dyscalculia affects 3–7% of school-aged children, yet no validated diagnostic instrument exists for Arabic-speaking North African populations. This study developed a domain-specific battery for screening developmental dyscalculia in Moroccan fourth-grade children and assembled initial validity evidence for its score interpretations and intended use as a screening component within a broader multifactorial evaluation, rather than as a stand-alone diagnostic tool. The 79-item battery, assessing nine domains, was administered to 128 typically developing children and 32 children with an independent DSM-5 diagnosis of dyscalculia. Validity evidence was examined across three sources: internal structure, relations with nonverbal cognitive ability, and consequences for the intended use, together with test-retest stability. The battery showed good internal consistency (α = .76 to .79), a provisional two-factor structure, large group differences (Cohen’s d = 1.60) robust to age, and excellent diagnostic accuracy (AUC = .927). These findings provide initial support for its use for early identification in Arabic-speaking contexts.
Motivational experience shifts within academic task episodes, yet widely used measures assess stable beliefs, strategy use, or retrospective summaries. We developed and evaluated the Task-Episode Appraisal Scale (TEAS), a six-item semantic differential for low-burden, repeated deployment in routine coursework, capturing appraisals at task entry and pre-submission. University students ( N = 477) from four courses completed TEAS at task entry and pre-submission for one focal assignment via an embedded learning management system prompt, with administrations separated by at least 24 hours. We report initial validity evidence from test content, response processes (focus groups), and internal structure. Indicators shared substantial variance at both administrations; omega hierarchical was .827 at task entry and .879 at pre-submission. Single items function as interpretable indicators rather than subscales. Observed entry-to-pre-submission differences were small while individual differences in change were substantial. TEAS provides workflow-moment snapshots of task-episode appraisals in authentic coursework; replication across tasks and institutions is needed.
Background: Socially evaluative communication contexts can significantly hinder student engagement. Existing emotion regulation scales are primarily trait-based and lack state-based, scenario-specific utility. This study aimed to culturally adapt and psychometrically validate the Communication Anxiety Regulation Scale (CARS) in Turkish. Method: A sample of 203 university students completed the CARS-TR and the Behavioral Emotion Regulation Questionnaire (BERQ). Construct validity was examined using a competing models approach in confirmatory factor analysis (CFA). Reliability was assessed via internal consistency and 3-week test-retest stability. Results: The first-order, four-factor model (suppression, reappraisal, avoidance, venting) demonstrated acceptable fit (χ 2 /df = 2.10, GFI = .92, CFI = .93, RMSEA = 0.07). The CARS-TR exhibited strong internal consistency (α = .89–.91), satisfactory temporal stability (single-measure ICCs = .68–.82), and theoretically consistent convergent validity with BERQ subscales. Conclusion: The CARS-TR demonstrates acceptable preliminary psychometric properties as a state-based measure among university students. It provides educational practitioners with a tool to better understand communication anxiety regulation patterns and guide supportive psychoeducational strategies in high-stakes learning environments.
Morphological awareness contributes to literacy development through multiple components, yet few assessments with established evidence of internal construct structure are widely available and suitable for whole-class administration for efficiently characterizing elementary students' suffix-based morphological knowledge in written sentence contexts. This study reports on the development and validation of ROAR Morphology, a brief classroom-based assessment of suffix-based morphological knowledge in written sentence contexts for students in grades 2-5, administered in under 10 minutes to whole classrooms with automatic scoring. Items were designed to capture a learning progression of suffix-based morphological knowledge, varying suffix type (inflectional/derivational) and suffix commonality (common/less common), with careful attention to cognitive processing demands including number of derivational distractors. Calibration of response data from 735 students using Rasch modeling yielded high reliability (alpha = .91; with fit indices ranging from .79 to 1.22). Item difficulty analyses confirmed that derivational morphology was more challenging than inflectional morphology, and less common suffixes were more difficult than common suffixes. Cognitive processing demands, specifically the number of competing derivational distractors, contributed additional variance in item difficulty beyond linguistic features. Based on item difficulty modeling, we established four empirically derived learning progression waypoints reflecting proficiency with suffix-based morphological structures of increasing complexity, from foundational inflectional morphology with common suffixes to more complex derivational morphology with less common suffixes. Notably, base word characteristics did not drive item difficulty, confirming that the waypoints capture genuine differences in suffix-based morphological knowledge development. ROAR Morphology uniquely predicted literacy achievement beyond word reading and sentence reading measures (Delta R-2 = 7.2%, p < .001), supporting its discriminant validity. These findings demonstrate that suffix type, suffix commonality, and cognitive processing demands systematically influence suffix-based morphological knowledge development in written sentence contexts, and that empirically validated waypoints may inform instructional planning during the critical grades 2-5 window when this knowledge is rapidly developing.
Understanding and managing perfectionistic behaviours in students is a critical challenge for teachers. Consequently, this qualitative study sought to assess the strategies that elementary and secondary teachers in Ontario, Canada use to respond to perfectionism in their classrooms. Open-ended survey responses from 197 teachers (83.25% female, M age = 40.21, SD = 9.00; 66.50% elementary teachers) and semi-structured interviews with a subsample of 26 teachers (84.62% female, M age = 41.19, SD = 10.95; 65.38% elementary teachers) were analyzed using thematic analysis based on descriptive phenomenology. Findings indicated that teachers employ a range of strategies, such as encouraging a growth mindset, offering individualized support, fostering strong teacher-student relationships, and cultivating self-care among their students. Despite their efforts to build resilience and strengthen students’ sense of self, some teachers highlighted the persistent nature of perfectionism and expressed frustration with their limited capacity to effectively support perfectionistic students. These findings highlight the importance of recognizing perfectionism as a significant educational issue and suggest that insights from teachers’ observations and classroom strategies could inform psychoeducational assessment and targeted professional development.
Grounded in Self-Determination Theory (SDT), this study proposes the Teachers' Motivational Styles (TMS) questionnaire, a new tool designed to assess how teachers support or frustrate students' basic psychological needs. Items were developed based on a recent SDT-based taxonomy of need-supportive and need-thwarting teaching behaviors. The questionnaire was administered to 2454 secondary school students in Italy and the United Kingdom. Exploratory Structural Equation Modeling supported a bifactor structure with two general factors-Need Support and Need Frustration-and three specific factors for autonomy, competence, and relatedness. Measurement invariance confirmed the tool's robustness across countries. Need-supportive teaching predicted higher Positive Affect, lower Negative Affect, and lower Intention to Drop out of school. In contrast, need-frustrating teaching styles predicted increased negative affect and greater risk of school disengagement. The TMS is a psychometrically sound, theory-driven instrument. Its cross-cultural validation supports its use in international contexts, with important implications for research, teacher training, and interventions to promote student well-being.
Understanding the dimensionality of social-emotional learning is critical for valid assessment in youth. This study examined the psychometric properties of the Social-Emotional Learning Scale (SELS), and its associations with subjective well-being in a sample of 711 Portuguese students (50.6% girls; M age = 11.01, sd = .81). Structural equation modeling indicated that a bifactor structure (one general and two specific factors: Peer Relationships and Self-regulation) provided the best fit ( CFI robust = .966, RMSEA robust = .031, SRMR = .032). The bifactor indices supported unidimensionality ( ECV = .896, PUC = 0.78, Omega G = 0.88), but not the interpretation of subscale scores. Metric and scalar invariance across gender was supported. An Item Response Theory analysis indicated the SELS is most informative at lower levels of the latent trait. Higher SEL was related to greater positive affect, lower negative affect, and greater life satisfaction. Our findings support unidimensionality and the usefulness of the SELS for identifying youth with lower social-emotional competence.
While assessment literacy has long guided teacher’ assessment practice, the rise of digital assessment presents new challenges in educational measurement, necessitating a renewed focus on Teacher Measurement Literacy (TML). This study reconceptualizes TML and develops a validated framework to support teachers in this evolving context. Using the Delphi method, a three-dimensional, 14-element TML framework was refined through two rounds of expert consultation, demonstrating increased consensus. A corresponding instrument was then administered to 306 pre-service teachers. Confirmatory factor analysis supported the proposed three-dimensional structure. Furthermore, likelihood ratio tests comparing nested Rasch models indicated that a three-dimensional model provided a significantly better fit to the data than a unidimensional alternative. Cross-validation with a sample of 297 in-service teachers further supported the robustness of the multidimensional framework. These findings validate the proposed three-dimensional conception of TML and offer empirical grounding to strengthen teachers’ capacity in the context of evolving measurement practices.
Chronic absenteeism among high school students poses a significant threat to academic success. As schools address the root causes of absenteeism, assessments such as the Washington Assessment of Risks and Needs of Students (WARNS) help identify students’ risk factors and support needs. However, assessment scores depend on the quality and authenticity of student responses. When students are disengaged (e.g., rushing through items, exerting minimal effort), the data may misrepresent their actual needs, undermining assessment validity. We examined disengagement patterns among high school students completing the WARNS, focusing on response time as a behavioral indicator of engagement. A small percentage (<5%) of students displayed disengagement, which differed between males and females and assessment context. Differences in risk classification patterns were observed for students identified as rapid responders. Results highlighted the importance of incorporating response process data into assessment interpretation and suggested practical strategies for improving the accuracy of these types of assessments.
Student populations in higher education have diversified internationally. Enhancing students' sense of belonging is linked to educational success; yet, the applicability of existing measures in the diversified context is questioned. We developed and validated a sense of belonging measure in a population of students with and without a migration background. Study 1 involved creating the measure and testing its factor structure, reliability, and measurement invariance with 374 students from a large urban university. Study 2 confirmed the factor structure and examined construct validity with 151 students from a technical university. The "University Belonging - Acceptance, Recognition, Commonality and Support" (UB-ARCS) scale comprises 16 items across four dimensions. Metric invariance was found between students with and without a migration background. Strong correlations with "belongingness" and "academic efficacy" demonstrated convergent validity, while divergent validity was demonstrated with "conscientiousness." Therefore, we found preliminary evidence of the validity and reliability of the UB-ARCS scale. Additional research is necessary for refining the suitability of the UB-ARCS for minoritized student groups and revealing the stories behind the numbers, offering deeper insights into the complex nature of belonging and marginalization in HE.
The purpose of this study was to examine the generalizability and dependability of scores produced by Sentence Order Fluency, a novel approach to progress monitoring of reading comprehension. Analyses were conducted to evaluate the performance of three alternative scoring methods-Absolute Correct, Pairs Correct, and Levenshtein Similarity-as well as test length and levels of aggregation (the numbers of passages or paragraphs used to calculate scores). Absolute and Pairs Correct scores performed similarly and appeared to show greater generalizability than Levenshtein Similarity scores. Students contributed more variance than probes in most models. Minimally sufficient levels of reliability for progress monitoring decisions could be possible using scores based on administration of 2 passages or 6 paragraphs using Pairs Correct scores. Levenshtein Similarity appeared to require a greater number of probes to obtain comparable reliability, suggesting limited practical value relative to other scoring procedures.
We document results of validity generalization meta-analyses of the teacher, student, and parent forms of the BASC-3 Behavioral and Emotional Screening System (BESS). Effect sizes (reported here as absolute values, based on 16 studies and 428 correlation coefficients) ranged from 0.040 to 0.350 for constructs related to academics, 0.399 to 0.690 for executive functioning, 0.350 to 0.627 for externalizing problems, 0.287 to 0.815 for internalizing problems, 0.344 to 0.660 for prosocial functioning, and 0.478 to 0.875 for global risk indicators. Extracted coefficients were of the expected direction and magnitude with theoretically aligned constructs, although scores on the BESS scales showed relatively weak relationships with academic variables and results from the parent form were less consistent. Results indicate the broad utility of the BESS in identifying students for further assessment as part of the screening process and support the use of the broadband BESS Behavior and Emotional Risk Index (BERI) particularly for this purpose.
This study validated the Mandarin-Chinese Teacher Well-Being Scale (TWBS) and developed a context-specific extension (CS-TWBS) to capture culturally grounded dimensions of teacher well-being in China. Two studies were conducted with samples of in-service and former Chinese teachers. Study 1 evaluated the psychometric properties of the translated TWBS using exploratory and confirmatory factor analyses, reliability testing, and measurement invariance analyses. A bifactor model provided the best fit, supporting a predominantly unidimensional structure reflecting general teacher well-being. Study 2 developed and validated the CS-TWBS through cognitive interviews and psychometric testing. The CS-TWBS showed strong internal consistency, clear factorial structure, and expected associations with flourishing, burnout, and job stress. Together, the Mandarin-Chinese TWBS and CS-TWBS provide psychometrically sound and culturally appropriate instruments for assessing teacher well-being in Chinese educational settings and for informing targeted interventions.
Critical thinking is increasingly recognized as a key skill in higher education, but its systematic development and assessment remain limited in Eastern Europe. This study aimed to adapt and psychometrically validate the Slovak version of the Critical Thinking Disposition Scale (CTDS). Data were collected from 545 Slovak university students (M age = 21.94; SD age = 2.19). Confirmatory factor analysis compared one- and two-factor models and tested measurement invariance across gender. The two-factor model, including Critical Openness and Reflective Skepticism, showed a better fit (CFI = 0.971; RMSEA = 0.040) and was partially invariant across genders. The scale scores showed acceptable reliability, stability of Critical Openness over time, and convergent validity through positive correlations with selected constructs from the Motivated Strategies for Learning Questionnaire. The findings support the Slovak CTDS score as reliable and valid for assessing critical thinking dispositions in higher education, allowing for gender comparisons and contributing to cross-cultural research.
Motivational self-regulation is a key component of self-regulated learning. Research has revealed the variety of strategies students use to reach their learning goals, and several instruments, built on one another, have been developed. This study describes the development of the Motivational Regulation Strategies Inventory (MRSI), a French instrument that expands previous tools by measuring a broader range of strategies, including seeking support and emotion regulation. Two studies were conducted to assess its validity: one with 305 middle school students and another with 653 college students. Exploratory factor analysis (Study 1) and confirmatory factor analysis (Study 2) identified 10 strategies. Path analysis examined the nomological network of these strategies, which included intrinsic motivation, self-efficacy beliefs, and procrastination as sources, and academic perseverance as an outcome. The findings provide substantial evidence for the MRSI's validity.
With the growing use of touchscreen devices in cognitive and academic testing, understanding the impact of tap latency is important for ensuring test fairness, particularly as related to speed tasks. The present study aims to understand how tap latency influences participant test-taking behaviors and performance. Using a counterbalanced within-subjects design, 203 participants aged 5-19 completed three WJ V speeded subtests on both low-latency normative administration devices (i.e., iPads) and Android tablets with an experimentally imposed, noticeable 340-ms tap latency. While the scores achieved across the two different devices were generally consistent, the actual Android scores were significantly higher than scores predicted based solely on latency-related time loss across all tasks, suggesting behavioral compensation from the tap delay. While an argument can be made that scores are comparable and thus acceptable, given tap latency's behavioral effects and the absence of validated post-hoc score correction models, it is recommended that WJ V speeded tests be conducted on devices with minimal and consistent latency and devices with unknown, variable, or consistently higher than 340-ms tap latency should be used with caution for speeded testing.
ChatGPT has shown considerable potential for Automated Item Generation, but the quality of ChatGPT-generated items in language assessment remains insufficiently substantiated. This research recruited 121 participants to systematically compare the psychometric properties of the test items and the linguistic features of the reading passages in ChatGPT-generated and official CET-4 reading comprehension materials, using Item Response Theory and Coh-Metrix. Key findings are as follows: (1) generated items fell short in higher-order reading skills; (2) the generated items were less difficult than official ones, showing weaker discrimination and providing measurement information mainly for lower-performing students; (3) only 22.9% distractors functioned effectively, indicating insufficient distractor performance; and (4) ChatGPT-generated passages were characterized by irregular lexical distribution, higher lexical complexity, weaker cohesion but simpler sentences than CET-4 passages. Although ChatGPT-generated passages were less readable than CET-4 passages, the corresponding items were easier and showed lower discrimination. This discrepancy can be attributed to inadequate distractor functioning that facilitates option elimination without complete passage comprehension, as well as to the underrepresentation of higher-order reading skills. The findings corroborate the conclusion that ChatGPT may function effectively as a supplementary tool in low-stakes assessment; however, substantial refinements in item quality are imperative before its application in high-stakes testing.
Digital technologies have reshaped how individuals engage in creative activities, highlighting the need for updated and valid instruments to assess digital creative behavior. This study developed and examined the Digital Creative Behavior Instrument (DCBI) using two samples of university students in Taiwan (N = 300 and N = 412). Exploratory and confirmatory factor analyses generally supported a three-factor structure: (1) digital media creativity and engagement, (2) digital video editing and production, and (3) digital artwork and design. Overall, the pattern of fit indices (CFI, RMSEA, and SRMR) suggested that the model fit was within a reasonable range for interpretation, although the TLI fell slightly below conventional benchmarks. Evidence of reliability and initial construct validity was observed. Openness to experience was positively associated with digital creative behavior, although the magnitude of the associations was small, providing preliminary support for criterion-related validity. The findings offer initial psychometric support for the DCBI and suggest directions for further refinement and validation in future research. This study contributes to creativity research and offers a basis for educators and researchers to assess digital creative engagement in contemporary contexts.
The current study extended previous work by further examining the psychometric properties of the Mistake Rumination Scale and its associations with depression and evaluative fears, as well as the cognitive experience of perfectionism and procrastination. Most notably, this study also uniquely examined a possible link between mistake rumination and a perfectionistic self-presentational style in line with our view that needing to outwardly seem perfect reflects internal insecurities and ruminative brooding about mistakes. The Mistake Rumination Scale is a seven-item inventory measuring the tendency to ruminate about a past personal mistake. In a sample of 132 university students, the Mistake Rumination Scale had good psychometric properties, including acceptable internal consistency and concurrent validity in terms of its links with perfectionistic self-presentation and ruminative thoughts related to being perfect and procrastination. Mistake rumination was positively associated with all facets of perfectionistic self-presentation. The measures of mistake rumination and automatic thoughts related to perfectionism and procrastination were all positively linked with depression and social anxiety. Regression analyses showed that mistake rumination was the only significant predictor of depression, while mistake rumination and perfectionistic cognitions were both significant predictors of fear of negative evaluation. Our findings attest to the further use of the Mistake Rumination Scale, highlighting the need for interventions that promote a more positive orientation toward making mistakes and an explicit emphasis on reducing the tendency to ruminate about mistakes.
This study developed an Academic Self-Management Skills Scale to assess the self-management skills of middle school students. In the first phase, focus group interviews and short-answer forms were administered to students and teachers. Responses were analyzed using open and axial coding, along with thematic analysis, and a draft scale was constructed based on the analysis results. In the second phase, the scale's validity and reliability were evaluated through a pilot study. The final version was administered to a different student sample, and a cross-validation analysis was conducted. The findings indicate that the scale is a valid and reliable tool for measuring middle school students' self-management skills.