
In this paper, we review the notion of aberrant responses in educational and psychological assessments. We examine the literature on aberrant responding, from the early work of Lee Cronbach, to contemporary interpretations of aberrant response processes in digital assessments. We contrast approaches that portray aberrant responses as a de-facto threat to test integrity and validity (i.e. something to be detected and removed), with an alternative paradigm, informed by the work of Cronbach, which interprets anomalous and unexpected responses as a source of formative insight, with potential to improve understanding of test content and design, to illuminate the diverse ways in which tested populations engage with tests and questionnaires. We argue that the latter approach provides stronger foundations for psychometrics and for AI-based models of item responding.
Effective instructional decision-making in mathematics requires timely and accurate assessment of student learning. This study provides validity and reliability evidence for an adaptive mathematics assessment tool designed to support teachers in identifying areas where students need instructional support. A total of 4,875 Chilean students from fifth to eighth grade participated in the calibration of an item bank. Validity evidence was collected from four sources: test content, response processes, internal structure, and relationships with other variables. Reliability was assessed using a test-retest approach. The results provide validity evidence and reliability estimates that support the use of the tool for formative purposes. The assessment is well-suited to the Chilean educational context, offering a practical solution for evaluating mathematical learning and informing pedagogical interventions. Furthermore, the tool's framework serves as a model for developing adaptive mathematics assessments that can be adapted for use in other Spanish-speaking regions or countries seeking data-driven instructional practices.
Student perceptions of teaching quality allow teachers to obtain feedback on their lessons. However, the reliability and validity of student perceptions of teaching quality are subject of scientific debate. In this study, data of 717 students, all aged 14 or 15 years, taught by 26 different mathematic teachers in The Netherlands were analysed to assess the construct validity of student perceptions, using a combined Item Response Theory and Generalizability Theory model. Moreover, the global and local reliability of the scores were investigated. Results support the construct validity of the data. It was also found that the student scores are reliable measures of teaching quality. Reliability is achieved, using a D-study, with a minimum of 5 students conducting 3 separate lesson ratings at different time points. The findings support the use of these perceptions for the rating and teacher development. Limitations of the study and suggestions for future research are presented.
This article provides an overview of Canada's decentralised education systems and describes recent changes in provincial, national, and international large-scale assessment programmes. We focus our analysis on the last 17 years to describe notable developments in policy and practice across Canada's 10 autonomous provincial jurisdictions and 3 territories. The discussion examines the continued tensions that exist between formative and summative purposes of assessment, including large-scale assessment, along with the changing nature of educational accountability and importance of assessment fairness. The authors situate assessment policy reforms within the broader global context that has emphasised neoliberal reforms in education, with implications for international work in assessment.
This study examines second-year law students' experiences of evaluating administrative decisions across two successive course implementations. In 2023 (N=250), students were randomly assigned to either an Independent Scoring Format (ISF), where they rated individual decisions, or Comparative Judgement (CJ), where they compared pairs of decisions. Post-activity survey data, including open-ended responses, were analysed for both groups. In 2024 (N=270), Adaptive Comparative Judgement (ACJ) was introduced with a larger pool of decisions and structured written justifications. Ratings of perceived usefulness were positive, with no statistically significant differences between CJ and ISF. The qualitative analysis suggests that comparative formats supported learning through contrast and prioritisation, while ISF highlighted difficulties in applying overlapping criteria consistently to individual texts. ACJ made students' grounds for preference more explicit, but also increased cognitive demand, underscoring the need for scaffolding and calibration. The findings inform the design of formative assessment for fostering evaluative judgement in legal education.
Assessment of collaborative problem solving (CPS) competence has recently gained much research interest in education, given the importance of CPS in equipping students to succeed in the evolving demands of modern society. Specifically, there is an increased interest in computer-simulated, scenario-based tasks to provide environments for students to solve problems together and assess their CPS competence. Given the lack of systematic and critical reviews on the topic, this review provides an analysis of relevant, and systematically selected, empirical articles (n = 26) assessing students' CPS competence using computer-simulated, scenario-based tasks. To move the research on the CPS competence assessment forward, observing students in authentic situations, e.g. small-scale research designs, is suggested. In this way, evidence for validity aspects that is currently lacking could be evaluated. Other possibilities for the future of CPS research afforded by new assessment methods and analytical approaches are also discussed.
To prepare for high-stakes testing, students often use rereading and practice testing strategies, two of the most widely used approaches to test preparation. This study investigated how students combine these strategies to maximise the effectiveness of their preparation. Participants were 152 third-year students at a teacher training college in China who reviewed 16 educational psychology concepts via an online platform. For each concept, students could choose either testing or rereading as many times as they wished and rated difficulty and confidence before and after each attempt. Open-ended responses captured their perceptions of both strategies. Results indicated that students preferred one-time testing but rarely reread repeatedly. Strategy choice was predicted by their confidence and the type of question used for testing. Students in this study could recognise both the direct effect of testing on enhancing retention and its indirect role as a tool for monitoring and diagnosing their learning progress.
Implementing an inclusive school means that teachers should use exam accommodations to foster the participation of students with Special Educational Needs (SEN). However, due to the emphasis on merit in most school systems, this practice can create a dilemma between equality and equity that can notably influence teachers' perceived fairness of such accommodations. Three studies conducted in the French context with teachers, students and members of the public (i.e. convenience samples, N = 1,525) examined this dilemma. Results indicate a differentiated perception of accommodations in terms of fairness, but also comparability. Moreover, the more teachers believed assessment should support student learning, the more they perceived exam accommodations as facilitating comparisons of students' performance, and the fairer they perceived accommodations. These findings highlight a recurrent equality - equity tension in assessment and suggest that accommodations intended to promote equity may be viewed as less fair when they are perceived to undermine comparability.
Theoretical benefits attributable to formative assessment of students' academic achievement are widely recognised. Empirical evidence has, however, shown that this positive relationship is not straightforward, prompting investigations into factors that may influence such a relationship. In the current study, we examine the mediating role of student engagement in the relationship between formative assessment practices and students' academic achievement in mathematics. The survey sample consisted of 379 tenth-grade students from mainland China. Structural equation modelling revealed that formative assessment practices did not directly impact students' academic achievement. Rather, the analysis indicated that student engagement mediated formative assessment practices' indirect and positive impact on academic achievement. The findings suggest that student engagement plays an important mediating role in the relationship between formative assessment practices and students' academic achievement.
Each year, hundreds of thousands of 10-11-year-olds in England sit the Statutory Assessment Tests (SATs). Alongside overall mathematics scores, schools receive results for eight mathematics sub-domains intended to highlight specific areas of strength and weakness. However, little empirical evidence exists regarding the reliability or practical value of these sub-domain scores. This paper investigates this issue using data from England's Key Stage 2 mathematics assessment. Drawing on descriptive analyses, exploratory factor analysis and multidimensional item response theory, we demonstrate that the school-level sub-domain scores derived from the test lack the reliability, utility and precision necessary for meaningful diagnostic use. We argue that the current practice of reporting Key Stage 2 mathematics sub-domain scores to schools should either be discontinued or supported by a redesigned assessment that can validly fulfil this diagnostic purpose.