
Research on effective mathematics instruction informs teacher preparation, professional learning, student assessment, instructional materials, and classroom observation instruments. Observation protocols are designed to capture instructional quality and the extent to which research-based practices are embedded in classroom instruction. However, although such instruments are typically developed and validated within specific educational systems, they are increasingly used across national and cultural contexts, on the assumption that they will function similarly across settings. Relatively little research has examined how raters interpret and negotiate instructional quality when applying shared observation frameworks across contexts. This study investigates classroom observation as an interpretive process shaped by the interaction among protocol design, instructional context, and rater reasoning. Specifically, the study examines how researchers with different professional backgrounds used three validated observation instruments developed in the United States to analyze one mathematics lesson from the United States, Norway, and Finland. Using a mixed-methods design, the study combined observation scores with recorded collaborative scoring discussions to examine how raters justified and negotiated evaluative judgments. Through discourse-oriented qualitative analysis, the study identified recurring forms of interpretive variation related to operational ambiguity, contextual interpretation, construct overlap, and differing evidentiary thresholds across raters and instruments. The findings contribute to methodological discussions concerning interpretive validity, observer reasoning, and the transportability of classroom observation instruments across educational contexts. The study highlights the importance of examining not only observation scores, but also the interpretive processes through which judgments of instructional quality are produced.
This paper presents a case study of how an undetected wording change in the OECD’s Programme for International Student Assessment (PISA) may have led to erroneous conclusions. We focus on a survey question intended to measure student truancy across countries and over time. While the international documentation suggests that the same question was used across PISA cycles, national documentation reveals that the wording changed in some countries in 2015. We show how this change could easily be missed and how it may affect inferences about trends in truancy, particularly when comparing rates before and after the COVID-19 pandemic. We also examine whether artificial intelligence and large language models could help identify the problem. Instead, these tools provided overconfident but incorrect advice.
Providing effective feedback for student essays has been a longstanding challenge for EFL teachers. Although research has discussed the potentials of generative AI tools in rejuvenating automated writing feedback, few empirical studies have been conducted in authentic educational settings to demonstrate how teachers may collaborate with them in essay evaluation. The present study, adopting the approach of practical action research, examines a teaching assistant’s (TA) experience of employing ChatGPT to enhance teacher feedback in a mainstream college English course in China. Working with an instructor in two action-reflection cycles over a semester, the TA planned, implemented, observed and reflected on AI-assisted teacher feedback for two intermediate-proficiency classes. Data were collected from ChatGPT records, reflective notes, written feedback, student survey and semi-structured interviews to reveal the process and effect of AI-assisted teacher feedback. The first cycle showed that ChatGPT facilitated the TA’s evaluation while introducing new problems and leaving some student needs unaddressed. In the second cycle, a modified use of ChatGPT led to more efficient and satisfactory teacher feedback, yet individualized feedback still remained a longer-term goal. The study proposes a practical framework for integrating generative AI into teacher feedback and underscores the importance of iterative teacher-AI collaboration in advancing sustainable AI-empowered feedback practices.
Accurate self-assessment, or calibration, is a key component of students’ metacognitive monitoring and self-regulated learning. This study examined calibration and prediction confidence across two high-stakes course exams (initial and make-up) using both continuous and group-based analyses. Undergraduate students in early childhood education completed Exam 1 (n = 128), predicting their grade (0–10) and reporting confidence (1–5). Students who failed (n = 45) took a make-up exam one month later and repeated their self-assessment. In Exam 1, calibration followed a pattern consistent with the Dunning–Kruger Effect (DKE): low-performing students overestimated their performance, mid-range students were most accurate, and high-performing students slightly underestimated their results. Among students who took the make-up exam, absolute calibration accuracy improved and confidence increased, and these changes remained robust after controlling for regression-to-the-mean effects. Together, the findings show that performance-level calibration patterns consistent with the DKE persist under authentic assessment conditions, while absolute miscalibration and confidence levels can shift following feedback and repeated testing. These results highlight adaptive shifts in metacognitive monitoring accuracy across repeated high-stakes assessments and provide practical implications for supporting self-assessment in higher education.
Understanding learner engagement with peer feedback in socially-embedded contexts has been wellnoted, particularly regarding the importance of contextual factors including interpersonal factors like peer familiarity. However, the specific role of peer familiarity remains underexplored from an interpersonal perspective in authentic pedagogical settings. To get a more in-depth comprehension, this study conducted a quasi-experiment in a real online teaching environment. 154 university students participated in a peer feedback activity set in a 10-week Micro-video Production course, and the students were then divided into familiar and unfamiliar groups. A combination of statistical analysis, epistemic network analysis (ENA), and cluster analysis was used to handle both qualitative and quantitative data. The results suggested that the familiar groups usually showed higher behavioral engagement, more critical cognitive interactions, more dynamic affective exchanges, and more positive learner profiles. Moreover, ENA further revealed stronger interpersonal connections (e.g., personal opinion-suggestion and personal opinion-others) in familiar groups. The findings indicate that peer familiarity is notably associated with learners’ behavioral, cognitive and affective engagement with peer feedback. The study underscores the importance of fostering interpersonal familiarity and considering interpersonal factors in designing peer feedback activities to enhance learning engagement and outcomes.
Science education is widely recognized as a cornerstone for enhancing national competitiveness and cultivating innovative talent. The practical challenges and evolving demands of science education in schools render boundary-spanning collaboration increasingly imperative. In China, a series of policy documents have underscored the necessity of interdepartmental cooperation and multi-level accountability to establish science education as a collective responsibility of schools, government, and society. However, policy effectiveness inevitably attenuates during transmission through hierarchical administrative structures. Against this backdrop, this research investigates District P in Shanghai, drawing on 20 in-depth interviews with education officials and school administrators, complemented by an analysis of relevant public documents. Informed by these data and a theoretical framework grounded in hierarchical governance and networks, the study reveals a situational interaction between hierarchical management and boundary-spanning networks within the district’s science education governance system. The top-down hierarchical system is instrumental in fostering an overall social atmosphere for science education and mobilizing social resources, yet it also exhibits structural fractures, including misaligned demands, insufficient incentives, and ineffective accountability. Concurrently, schools have developed three types of boundary-spanning science education networks: symbolical-linked networks, transaction-oriented reciprocal networks and value-driven networks, though only the latter two serve a substantive bridging role. This study provides important insights for innovating regional governance mechanisms and shared responsibility frameworks in basic science education.
University teacher education equips prospective teachers with the competencies required for successful entry into the profession. However, evidence for this education’s effectiveness has long been a desideratum for researchers. Applying a multidimensional competence framework—encompassing knowledge and skills, professional beliefs, and affective-motivational characteristics—this study compares final-year pre‐service master’s-level teachers with students at bachelor’s level, accounting for program-specific effects and controlling for individual learning prerequisites. At one of Germany’s largest teacher education universities, 690 participants (234 master’s and 456 bachelor’s students) completed a comprehensive battery of standardized performance tests and self-report measures. Multiple regression analyses revealed that master’s students significantly outperformed bachelor’s students in most knowledge domains and held more favorable pedagogical beliefs. However, they exhibited slightly less favorable affect and motivation (e.g., lower self-efficacy). Significant main effects of the teacher education program type (e.g., higher pedagogical knowledge among prospective special education teachers) underscore university-based teacher education’s effectiveness. We discuss the findings’ implications for structuring and enhancing university teacher education.
The degree to which students are “engaged” when taking a test affects their performance. Despite this knowledge, little research has been done on test-taker engagement with respect to students with disabilities. In this study we investigate whether there are differences in test-taker engagement across students with and without disabilities and whether access to a text-to-speech affects such engagement. Using linear mixed modelling and generalized linear mixed modelling to analyze data from a large-scale test for students in grades K-4, we found special education students had lower test-taker engagement and text-to-speech has the potential to increase engagement for all students. The results support considering test-taker engagement when determining universal supports and accessibility features to support the needs of all students.
Student participation in assessment is essential for realising inclusion in culturally and linguistically diverse classrooms. This paper explores how building whole-school assessment practices and a shared assessment culture enable such student involvement and participation. Drawing on case study data from interviews with school leaders, teachers and students from five schools, we create a narrative account of the schools’ practices to examine the policies, guidelines and practices that guide teachers’ classroom assessment practices. In addition, the study examines how teachers and school leaders facilitate and cater for student interaction and participation in assessment, and how the different actors in the schools perceive and describe these practices. The five schools succeed to different degrees in facilitating student involvement by 1) facilitating a whole-school approach and building an assessment culture with shared practices understood, and enacted, throughout the school, and 2) involving teachers in building shared guidelines for assessments and facilitating teacher cooperation.
This explanatory sequential mixed-methods research examined Generative Artificial Intelligence (GenAI)- and student-authored cause and effect essays to reveal reliability, consistency, rater agreement between different scoring methods and to document decision-making strategies for identifying authorship. The participants (n = 107) in a teacher education program scored cause and effect essays through AI-generated rubrics, a standardized IELTS rubric, and a holistic general impression, and identified authorship of AI-generated and student-written essays. Data were analyzed through logistic regression, Generalizability Theory, and thematic analysis to examine detection patterns and pedagogical decision-making and reasoning processes. The results showed variation of scoring across rubric types which also interacted with rater type, rater expertise, and detection of authorship in complex ways. The instructor-pre-service teacher agreement was highest when using holistic scoring; AI-pre-service teacher agreement was highest for the IELTS rubric, and AI-instructor agreement was relatively high for the AI-generated rubric. The results showed that rubric type, rater expertise, and detection of authorship interact in complex ways. We revealed misdetections and a critical quality attribution paradox where sophisticated student writing was associated with AI, while polished AI-generated essay was identified as student-written. We documented multi-dimensional pedagogical reasoning strategies that moved from initial intuitive impressions to textual analysis that included indicators of originality and authorial stance. The results indicated that reliability of using GenAI for AES and teacher scoring may be at risk when authorship cannot be correctly identified. This study contributed to arguments about challenges to consistency and the need for control mechanisms and specialized training.
Relating students’ achievement goals to their engagement in academic dishonesty has been a fruitful research avenue for understanding the motivational underpinnings of cheating behavior in academia. However, most study results are grounded in self-reported measures of academic dishonesty, and heterogeneous findings – particularly regarding the links between performance goals and academic dishonesty – warrant an experimental investigation into the conditions influencing these effects. Through three experiments, we examined the effects and interactions of performance goals (appearance-approach) and performance evaluation standards (result-based vs. process-based) on cheating behaviors in an academic aptitude test. In Study 1 (pilot, N = 146), we tested the paradigm and developed suitable task types. In preregistered Study 2 (N = 238), university students completed a supposed academic aptitude test in the lab using a 2 (appearance goal induction vs. no induction) × 2 (result-based vs. process-based evaluation) design, wherein cheating was measured by indicating having solved unsolvable tasks and using illicit means to solve knowledge questions. In preregistered Study 3 (N = 253), we conducted a conceptual replication of Study 2 in an online setting. Taken together, the effects of goals and evaluation standards were inconsistent and differed between the type of cheating behavior and context. Methodological implications are discussed with respect to measures of cheating in both laboratory and online settings, while the practical implications focus on strategies to deter cheating through assessment modalities.
Grades are the currency of education, yet research on how teachers assign them remains limited. In this study, we examine the micro-decisions teachers make when grading classroom tests—small grading judgments that combine specific feedback with a partial point—focusing on both scoring types (addition, subtraction, imposing a maximum) and content priorities (e.g., conceptual understanding, procedural skills). Thirty-six secondary mathematics teachers graded 30 student tests on linear equations without predefined grading guidelines. Findings reveal four distinct clusters based on scoring types, such as teachers who consistently apply subtraction-based grading. Similarly, four clusters emerge based on content emphasis, including general praisers, who frequently provide praise with limited attention to mathematical content. While these findings reveal substantial variation in how teachers assign partial points and arrive at a total score, the analysis shows that only 4
Bibliometric reviews have proliferated across academic disciplines in recent years. This unprecedented publication trajectory has prompted concerns about quality and the contributions of these reviews to knowledge. This systematic review used descriptive statistics and content analysis to review 1,873 bibliometric reviews of educational research published in Scopus-indexed journals through 2024. The authors analyzed the research landscape, topical trends, and methodological quality of these education reviews. Remarkably, 87
Teacher performance management in many Western countries has been identified as jeopardising teacher professionalism rather than enhancing it. It has been argued that the teaching profession faces issues of mistrust, external reviews that demand unrealistic transparency, and constraints on professional judgement. To address these concerns, intelligent accountability has been proposed. Although there is some theoretical research on intelligent accountability, significant gaps remain in practical applications. This paper aims to theorise intelligent accountability in education policy by identifying its key aspects: trust (cooperation among stakeholders), collaboration with an emphasis on the important role of school leaders, and demystifying the relationship between big data and transparency. Additionally, this paper provides a specific example of how the principles of intelligent accountability can be enacted in practice. This study examines two education policies in New Zealand. In 2020, the New Zealand government introduced the Teacher Professional Growth Cycle (PGC) to replace the Teacher Performance Appraisal (TPA). This paper argues that this transition demonstrates a practical orientation towards intelligent accountability. By analysing the implementation of the professional growth cycle for teachers in New Zealand, the article shows that this initiative aligns perfectly with the principles of intelligent accountability. This paper contributes to the exploration of intelligent accountability within a specific socio-political context.
Standardized measures of achievement and teacher-assigned grades (henceforth grades) are both supposed to be measures of achievement. However, there might be concerns that grades are more subjective than standardized tests. The purpose of this study was to examine the relationship between standardized tests and grades in primary students. We collected data from 218 Grade 1 students and 124 were re-tested in Grade 2. Students completed subtests from the Woodcock-Johnson III Tests of Achievement (henceforth test scores) and grades were collected. Separate factor analyses were conducted on the test scores and grades, both yielding a 3-factor model with separate language, reading, and mathematics domains. We then examined the relationship between test scores and grades across domains, grade-level, genders, and schools. Structural equation models found that test scores explained half of the variance in grades across domains, except Grade 2 language. Bayesian paired t-test found that minimal discrepancies between test scores and grades existed across genders and schools. When a difference did exist, grades tended to be higher than test scores; discrepancies typically occurred in verbal subjects but attenuated by Grade 2; there was a female advantage for grades and male advantage for test scores; and, school differences depended on grade and domain. Understanding the degree of alignment between measures could bolster the use of teacher assessment as a valid index of performance in young children.
This article examines the linkage of grading criteria and students’ epistemic agency—students’ capability to use knowledge to shape their lives and society. It has been argued that assessment can both foster and limit students’ epistemic agency. Moreover, it is known that grading criteria are particularly powerful in signalling what knowledge is worth knowing and how students are expected to engage with it. Our study focused on Finnish basic education in which teachers assign students’ final grades based on national criteria. We employed Anderson and Krathwohl’s taxonomy to analyse how the notion of epistemic agency was manifested in the criteria. This led to two key findings. First, the criteria offered various manifestations of epistemic agency, which we organised into seven distinct types. The breadth of manifestations suggested that teacher-led criteria-based assessment, as opposed to high-stakes examinations, enables a more comprehensive reflection on educational objectives, including epistemic agency. Second, we found that manifestations of epistemic agency were concentrated in the higher-level grade descriptions, whereas the lowest acceptable level scarcely entailed them. Although it is natural to expect deeper engagement with knowledge at higher levels of expertise, the gap between grade levels appeared unnecessarily large. We conclude that grading criteria can provide a variety of affordances for students’ epistemic agency. Therefore, to promote equity, it is essential that criteria across all grade levels are consistent with educational aims, including epistemic agency.
The recent advances in technology, such as the emergence of generative artificial intelligence (GenAI) tools, warrant careful integration into education. In particular, exploring feedback and scores generated by both human raters and GenAI tools is crucial for assessing feedback alignment and score validity in L2 writing assessment. Moreover, L2 writing teachers’ agency in collaborating with these tools is a notable area of research. Given the importance of the topic, this mixed-methods research design aims to address three research questions: The alignment of GenAI and human scores and feedback on the same writing task responses; the justifications for scoring and feedback; and teachers’ agency in negotiating their roles in GenAI-supported assessment contexts. For that purpose, fifty essays (an IELTS retired task for Academic Writing Task 2) were rated by a human rater and ChatGPT-5 using the IELTS Task 2 criteria. The results displayed a strong correlation between human and ChatGPT-5 scores, confirming the scoring validity. Then, the rater was asked, and ChatGPT-5 was prompted to investigate the justifications for their scoring decisions. The findings yielded a contrast between the human rater and ChatGPT-5. These findings were also carefully interpreted following Kane’s argument-based approach to validity. Lastly, the thematic analysis of the semi-structured interview to navigate teachers’ agency in GenAI-mediated writing assessment was in accord with Priestley’s ecological model of agency. Overall, the findings illustrate the need for a hybrid model since blending GenAI-led surface-level evaluation with human-led cognitive, critical, and contextual evaluation is essential for a comprehensive and valid writing assessment.
A critical issue within the educational system is the availability of assessment tools capable of accurately describing the skills mastered by a student and those that are not. Competence-based Knowledge Space Theory (CbKST) is a mathematical theory for the formative assessment of skills that aims to precisely identify the set of skills an individual possesses based on the responses to test items. The current landscape in educational measurement includes lots of good tests designed under theories other than CbKST, which, after appropriate analysis, could be used for skill assessment. Identifying the skills underlying the responses to the items of these tests could enable educational stakeholders to obtain precise feedback about students' mastery of each skill. This work describes retrofitting of existing tests in the framework of CbKST. In particular, it provides a step-by-step guide useful to identify the skills measured by the test and the way they are associated with the items. Retrofitting is illustrated in practice through applications on empirical data collected using two educational tests. Finally, the information about the test, its items, and students' skills that can be gained from retrofitting is explored and discussed.