As governments react to international rankings, the Organisation for Economic Co-operation and Development's Programme for International Student Assessment (PISA) of 15-year-old students has arguably influenced many governments' mathematics education policies. In this paper, we look at the relationship between the latest cycle of PISA, policymaking and media in the economies where English is an official or substantial working language (Anglophone countries). We analysed three sources of media responses to PISA 2022 results, from governments, mainstream media, and social media, focusing on the number of the reactions, the sentiment of the reactions, and some relevant themes pertaining to PISA 2022 mathematics coverage. The results show that the PISA results have been used and discussed differently by different governments and that media responses seem to relate to pertinent policy objectives. Our contribution sheds light on how the PISA 2022 results interact with policymaking for mathematics education in Anglophone countries.
Conceptions of creativity differ systematically across cultures, shaping how they are expressed and valued. The introduction of a Creative Thinking (CT) domain in PISA 2022 marks the first large-scale, internationally comparable assessment of this competence among 15-year-olds. However, the availability of this rich cross-national data does not resolve the fundamental validity question of whether the creative thinking tasks function equivalently across diverse cultural contexts at the item level. This study investigates the extent, patterns, and mechanisms of cultural differences in the PISA 2022 CT assessment. We employ a sequential explanatory mixed-methods approach, pairing an IRT-based Differential Functioning of Items and Tests (DFIT) framework with qualitative item diagnosis. DFIT analyses of 30 PISA 2022 CT items across 10 Global Leadership and Organizational Behavior Effectiveness (GLOBE) cultural clusters reveal that while cultural non-invariance is pervasive (flagging 83.3% of items), the severity of bias is uneven. At the cluster level, bias does not suggest an "East-West" divide but distinct regional assessment profiles. To explain these quantitative patterns, a subsequent qualitative diagnosis of the most severely biased items was conducted. This analysis located a four-dimensional bias architecture that may contribute to the observed disadvantage: (1) context universality and familiarity, (2) the cultural match of scenario-activated cognitive prototypes, (3) implicit value assumptions, and (4) response-mode alignment. These findings call for a shift in the field from reactive, post-hoc bias screening to a proactive approach of embedding fairness considerations into the design process, aimed at creating culturally inclusive assessments.
Pursuing replicability - independent evidence for previous claims - is important for creating generalizable knowledge(1,2). Here we attempted replications of 274 claims of positive results from 164 quantitative papers published from 2009 to 2018 in 54 journals in the social and behavioural sciences. Replications were high powered on average to detect the original effect size (median of 99.6%), used original materials when relevant and available, and were peer reviewed in advance through a standardized internal protocol. Replications showed statistically significant results in the original pattern for 151 of 274 claims (55.1% (95% confidence interval (CI) 49.2-60.9%)) and for 80.8 of 164 papers (49.3% (95% CI 43.8-54.7%)), weighed for replicating multiple claims per paper. We observed modest variation in replication rates across disciplines (42.5-63.1%), although some estimates had high uncertainty. The median Pearson's r effect size was 0.25 (95% CI 0.21-0.27) for original studies and 0.10 (95% CI 0.09-0.13) for replication studies, an 82.4% (95% CI 67.8-88.2%) reduction in shared variance. Thirteen methods for evaluating replication success provided estimates ranging from 28.6% to 74.8% (median of 49.3%). Some decline in effect size and significance is expected based on power to detect original effects and regression to the mean because we replicated only positive results. We observe that challenges for replicability extend across social-behavioural sciences, illustrating the importance of identifying conditions that promote or inhibit replicability(3,4).
In recent years there have been more and more calls for the preparation of teachers to be increasingly evidence-informed. The assumption is that evidence-informed teaching practice improves the quality of teaching and subsequently student achievement will improve. Preparation for teaching using technology for student learning is well-established in pre-service teacher education programmes. However, the use of technologies to interrogate evidence by pre-service teachers is less well-researched. The work presented here offers insights as to how using technology can support and enhanced pre-service teachers in developing practices that are evidence-informed. Using a theoretical framework with elements of evidence-use mechanisms and Engeström’s expansive learning cycle, a programme for training pre-service mathematics teachers in one European country was redesigned, with the programme leaning strongly on technology use to query evidence. Using a design-based research approach, this article describes the design iterations, and presents a case study with five in-depth interviews with pre-service teachers on evidence-use. From these cases, it is concluded that the use of technology can aid and support evidence-use. The paper concludes by presenting a new redesign of the programme with the use of generative AI, pointing towards future ways in which technology can support the development of teachers’ evidence-use.
Online mathematics environments often provide feedback to support students’ mathematical learning. The extent to which students make use of feedback, so-called ‘help-seeking’, depends on numerous instructional variables, including the type of online platform, task difficulty and students’ precision. However, most student behaviour in such platforms occurs outside the view of the teacher, and engagements with such environments are not independent events: the order in which tasks are completed matters. This paper reports on two education phases where we explore the interplay of task difficulty, precision and feedback in two online mathematics environments. We use log file data from both platforms and analyse them with multivariate regression and sequence analysis. Study 1 uses student data from English students in grades 3 to 5 ( N=839 ), totalling 490,426 records. Study 2 uses data from undergraduate mathematics students ( N=232 ), totalling 138,632 records. The results show that task difficulty, precision and feedback interact, with help-seeking not necessarily relating to higher task difficulty or precision. This again confirms that simply recommending ‘more feedback’ is not the correct strategy for students engaging with online mathematics platforms.
Verbal Probability Expressions (VPEs), such as “likely” or “rarely,” are often used to communicate scientific information that is uncertain and are often preferred over numerical expressions despite a higher risk of misinterpretation (mode preference paradox). Prior research has focused on adult interpretations of VPEs, with only few studies focusing on how young people understand them. This study investigates how secondary school students in England interpret 29 common VPEs using a slider-based scale (0–100
Prior research indicates that spatial skills, such as Mental Rotation Skills (MRS), are a strong predictor for mathematics achievement. Other studies have shown that MRS can be improved through training. This paper explores whether a well-known puzzle-oriented tool for building houses with 3D cubes is effective in improving performance in a standardized MRS measure that recorded accuracy and speed. The field experiment took place with 85 year seven (11-12 yr olds) pupils from an independent secondary school in the south of England. We used two conditions in the experiment, with the puzzle-oriented training tool being the intervention condition. The findings show there was a significant effect for accuracy but not for speed. Contrary to prior research our findings did not show any gender effects. The findings and implications are discussed in light of the existing literature around spatial skills.
Game-based assessments (GBA) have attracted significant research attention because of their potential to predict dynamic models of students' ability states and to support personalised learning and flow. Using the Evidence-centered design (ECD) model as a theoretical framework, this systematic literature review synthesizes the current state of research in GBA. The findings reveal that GBA primarily focus on STEM domains and 21st-century skills assessment. Puzzle games remain the predominant choice which can embed assessment tasks more efficiently, although technologies such as Virtual Reality (VR) are emerging. Although the GBA study captured processual data from students' gameplaying, most data analysis models still haven't been able to dynamically analyse these processes. Mobile platforms are the most common delivery method. This study also discusses how the future direction of GBA can better exploit the strengths in promoting student enjoyment and improving learning outcomes and psychometric quality.
Social network analysis is useful for obtaining a better understanding of antecedents and mechanisms of relationship formation and interactions between individuals in educational and psychological contexts. Research utilising descriptive and cross-sectional applications of network analysis is regularly reported, but longitudinal analyses of networks have received less scrutiny. In this methodological article, we compare three commonly applied approaches for analysing longitudinal social network data: Multiple Regression Quadratic Assignment Procedure (MRQAP), Separable Temporal Exponential Random Graph Models (STERGM), and Stochastic Actor Oriented Modelling (SAOM) with research questions about correlations, social structures and mechanisms, respectively. We highlight advantages and disadvantages of the methods and illustrate differences between these methods by analysing longitudinal peer-communication network data of pre-service teachers. The key considerations by the researcher are summarised as 'FACTS' (Focus, Assumptions, Conceptualisation, Time points, and Size) as an aid to researchers in selecting the most appropriate method for the analysis of longitudinal social network data.
In recent decades, the different manifestations of mathematical problem posing (MPP) have become important objects of research in mathematics education. However, according to some researchers, these different manifestations pose the risk that the term “problem posing” might become “so diffuse as to undermine its analytic power and reduce it to an ephemeral sign”. This is because the activities of problem posing are formulated in different ways, using different mathematical idioms. This article studies the language of problem posing in three English-speaking countries: England, the USA, and Singapore. I analyzed the secondary school curriculum texts in the three countries. What language is used to denote MPP activities? What are the differences in language between England, the USA, and Singapore? The work will give insights into the way MPP plays a role in England, the USA, and Singapore and will conclude with the implications of the findings.
Families play a pivotal role in fostering children's science literacy, interests, and identities through everyday interactions and informal learning contexts, with parents as main facilitators. An essential, yet often underexplored, aspect of this process is the role of emotions in shaping science learning experiences. Emotions serve as powerful mediators of engagement, influencing key learning outcomes such as interest, motivation, achievement, and persistence. Despite the recognized importance of family engagement in science learning and the emotional dimensions associated with it, there is a significant gap in research specifically examining how families engage with science at home and the role emotions play in these settings. In this case study, we employed a mixed methods approach consisting of electro-dermal activity (physiological) and recorded observations (behavioral) to identify the emotional expressions of a mother as she engaged in five science activities with her children (ages 13, 11, 7, and 4) at home. All five activities were analyzed utilizing the following procedures: 1. Peak analysis, 2. Structural breaks, and 3. Microanalysis. We complemented our interpretation of the data with reflective notes and a reflective interview (self-reports) with the participant. The study reveals that mediated activities elicit more positive emotional expressions; the interrelationship between emotions and cognitive, social, and cultural domains needs to be accounted for while analyzing emotions, and highlights the methodological challenges of measuring emotions. By focusing on how a parent guides home science activities, it fills critical gaps in understanding family-based science engagement and sheds light on the affective dimensions of informal science learning. Employing a mixed methods approach provides a comprehensive understanding of emotional expressions during home science activities, which enhances the validity of the findings and captures the dynamic nature of emotions, offering a robust approach for analyzing the interplay between physiological, behavioral, and interpretive emotional expressions in real-world contexts.
Inspectors are tasked with judging the quality of provision based on visits to schools. They conduct these inspections sequentially, completing one before moving on to the next. However, empirical research in a range of settings outside education suggest that prior judgements in a sequence can influence subsequent judgements, despite being logically irrelevant. We investigate whether school inspectors in England display such sequential bias by testing whether they judge similar schools differently, depending on the judgements they reached in prior inspections. We find only limited evidence of sequential bias in primary school inspections. In particular, an inspector reaching an ‘Inadequate’ judgement in their previous inspection is associated with a 42% reduction in the odds of reaching another ‘Inadequate’ judgement in their next inspection. Only around 5% of inspection judgements result in an ‘Inadequate’ and we do not find consistent evidence of sequential bias at other grades, meaning this bias only affects a small minority of judgements. We also do not find the same results for secondary schools, albeit in a much smaller sample.
In England, a substantial proportion of school inspections are conducted by current school leaders. This may lead to concerns that this gives their school (about 2% of schools) an advantage in the inspection process when it is their turn to be inspected. Yet scant evidence exists on this issue. This paper thus presents the first evidence on this matter, using data obtained via a freedom of information request and linking this with other publicly available information about England's schools. We find that schools where a member of staff also works for Ofsted receive better inspection outcomes than schools without an inspector on their payroll. Our findings nevertheless suggest that other schools may benefit from having access to the training material and professional development opportunities Ofsted provides to its inspectors.
School inspections are a common feature of education systems across the world. In these inspections, trained professionals visit schools and reach a high-stakes judgement about the quality of education they provide. School inspections rely upon professional judgement, and are meant to reflect the quality of the provision of a school. There is currently little academic evidence investigating these reports at scale, including how they vary over time. We make use of two cut-off moments in the inspection process in the last two decades: (1) a document dispelling myths about what inspectors in England are looking for and (2) the introduction of the 2019 Education Inspection Framework (EIF). We present new empirical evidence on this matter, drawing upon data from more than 60,000 inspection reports for primary and secondary schools in England between 1997 and 2022. Several computational techniques show that themes in the reports did not change much after the myths document, while after the framework change reports put more emphasis on leadership, subject specialism and the curriculum.
ABSTRACT School inspections are a key component of the accountability system in many education systems, including England. The judgments and reports produced through these inspections are widely used by parents when they are choosing a school for their children. But should they be? This paper presents new evidence on this issue. We illustrate how parents selecting secondary schools using Ofsted judgments will often be basing their decision on dated information. Indeed, half the time, this will be based on a period in which the school had a different headteacher. We find there are almost no differences in future academic, behavioral, school leadership and parental satisfaction outcomes between schools rated as good, requiring improvement and inadequate in the inspection data available to parents at the point of school selection. That is, parents who choose a “good” secondary school for their child will not leave with appreciably better outcomes than a parent who selects an “inadequate” school. The one exception to this is an Outstanding judgment, which does predict future academic outcomes – though only if the inspection was conducted within the last five years. We thus advise parents that – besides choices involving Outstanding schools – Ofsted judgments are of limited use to them in selecting a school.
This is the final report of the Nuffield funded project Inspecting the Inspectorate.
Hugh C Davis合作论文数University of Southampton7