
Generative AI (GenAI) has shown promise in programming education, yet empirical evidence remains inconsistent due to varied pedagogical and technical contexts. This research addressed this inconsistency through a three-level Bayesian meta-analysis of 35 empirical studies with 131 effect sizes published between 2022 and 2025. Learning outcomes were classified into AI-assisted programming outcomes (AIPO), independent programming outcomes (IPO), higher-order skills (HOS), and motivational-emotional outcomes (MEO). Results demonstrated credible small-to-medium positive effects of GenAI on all four outcome categories, and no credible difference among categories was detected. The moderator analyses did not detect credible moderation by educational contexts, instructional designs, or GenAI system designs in any outcome category. A few exploratory pairwise differences were credible for MEO under specific educational contexts and strategy-training conditions. Within the current evidence, GenAI shows consistent positive average effects, whereas the conditions that strengthen or weaken these effects remain undetermined. Because the corpus is small and unevenly distributed across moderator levels, the absence of detected moderation may reflect limited statistical information rather than equivalence across conditions. Larger and more balanced studies with more complete reporting of implementation details are needed to identify these conditions.
Generative learning theory holds that learning depends on the learner’s own construction of relations between prior knowledge and new material, and its evidential practice has long inferred that construction from what the learner produces. External provision of such work is not new, but generative AI extends it across every generative strategy, supplies it on demand, and leaves little trace of its origin, so the locus of generative activity becomes instructionally and empirically ambiguous. Determining who performed the cognitive work that learning requires thus becomes a design and measurement problem. This paper examines that problem in mathematics learning. Through an integrative conceptual analysis of Fiorella and Mayer’s eight generative learning strategies, I identify four cross-cutting themes: the production-to-evaluation shift, metacognitive displacement, the scaffolding–substitution continuum, and reconfigured generative activity. Synthesizing these themes, I propose the AI-Enhanced Generative Learning in Mathematics (AI-GLM) framework, comprising (a) a typology of three AI roles, in which Generation Partner and Generation Substitute occupy the poles of a scaffolding–substitution continuum while Generation Catalyst is distinguished by a functional criterion rather than by a position on it; (b) a role determination mechanism specifying how tool design, task design, learner orientation, and teacher mediation combine to determine the role AI occupies in a given episode; and (c) five design principles. The framework addresses the representational, symbolic, and justificatory demands especially prominent in mathematics and yields propositions stated with the operationalizations and falsification conditions needed to test them.
Academic help-seeking (AHS) is widely recognized as an important component of self-regulated learning for preservice teachers, enabling them to address the challenges of their training. However, studies in this field show mixed results, reflecting the complexity of the phenomenon. To provide a comprehensive synthesis of the existing evidence, this review draws on 32 empirical studies published up to 2025 that met the established inclusion criteria and examines how AHS has been measured, which variables are associated with it, and which sources are accessed during training. Findings indicate substantial heterogeneity in measurement approaches, and AHS was frequently associated with academic achievement, self-regulation, and self-efficacy, although these associations were not consistent across studies. Additionally, preservice teachers tend to rely more on informal help sources. The results further suggest that contextual and social conditions play an important role in AHS, highlighting the importance of supportive learning environments and accessible formal guidance. Taken together, the evidence suggests that AHS may be understood as a contextually and motivationally influenced regulatory process rather than a stable individual disposition. In this sense, the review contributes to the literature by offering a more integrated conceptualization of AHS as a socially embedded regulatory process within self-regulated learning, highlighting the relational nature of AHS in teacher education and suggesting its potential implications for professional identity development. These findings have implications for teacher education programs, suggesting the value of fostering supportive learning environments and strengthening the relational accessibility of formal support structures.
Intensive longitudinal methods—such as the experience sampling method, ambulatory assessments, or daily diaries—have increased in popularity in recent years. These new methods offer valuable real-time insights into individuals’ daily lives by capturing phenomena as they occur. Their introduction, however, comes with challenges and uncertainties regarding study design, implementation, and reporting. This systematic literature review aims to provide an overview of the common practices in intensive longitudinal methods in school research and evaluates the transparency and consistency in the reporting of methods and rationales for methodological choices. Our review reveals substantial heterogeneity in the use of intensive longitudinal methods but a predominant focus on students and their experiences during school lessons. Methods and rationales for methodological choices are often insufficiently reported, and while we found an improvement in recent years, transparency still remains an issue. We also offer a comprehensive checklist of reporting criteria for intensive longitudinal methods. We emphasize the importance of reporting standards for intensive longitudinal studies to promote transparency and comparability.
Self-regulation (SR) has emerged as a pivotal construct in educational research, yet its diverse theoretical accounts often lack comprehensive integration. This paper addresses this by proposing an integrative conceptual framework that consolidates four core research traditions on SR in educational settings, each emphasizing distinct underlying mechanisms: (1) Stable Dispositions (e.g., personality traits; Matthews et al., 2009; Roberts et al., 2007; Song et al., 2020), (2) Limited Resources (e.g., working memory capacity and executive functioning; Friedman Miyake, 2017; Paas et al., 2003; Sweller, 1994), (3) Driving Forces (e.g., motivation, interest, and affect; Eccles Wigfield, 2002; Hidi Renninger, 2006; Trautwein et al., 2019), and (4) Learning Activities (e.g., self-regulated learning, cognitive and metacognitive strategies; Winne Hadwin, 1998; 2008; Zimmerman, 2000). By synthesizing these perspectives, our framework positions SR as a dynamic, interdependent system, emphasizing how interactions and compensatory mechanisms explain individual differences and influence learning outcomes. We further argue for an integrative methodological toolkit, leveraging advanced machine learning and computational models, to capture the multi-level, temporal, and social dynamics inherent in SR. This holistic approach is crucial for advancing a comprehensive and ecologically valid understanding of SR in educational contexts.
Academic resilience refers to students’ capacity to maintain or regain adaptive academic functioning when confronted with academic stressors, setbacks, failure, or threat. Although the construct is often considered central to educational psychology, its psychometric foundations remain fragmented. This systematic review and meta-analysis synthesized evidence from 46 studies evaluating academic resilience and closely related academic buoyancy measures. Evidence on dimensionality was mixed. Unidimensional models showed higher pass rates for CFI, TLI, and SRMR, whereas multidimensional models showed a higher RMSEA pass rate; however, exploratory Welch comparisons showed significant differences for CFI and TLI, but not for RMSEA or SRMR, indicating that the available global-fit evidence does not clearly favor either model family. Pooled internal consistency was high, α/ ω = 0.89, 95
Writing is a cornerstone of academic success, yet English language learners (ELL) in the United States must learn to write while also developing English proficiency. Although descriptive reviews suggest that writing instruction can benefit these learners, no prior meta-analysis has synthesized its effects quantitatively. This study addressed that gap by synthesizing 39 comparisons in 20 experimental and quasi-experimental investigations of writing interventions in K-12 classrooms with ELL students in the United States. From 122 effect sizes, the overall impact of writing instruction was statistically significant and positive (g = 0.33). Writing quality was the most frequently assessed outcome (in 33 comparisons across 16 studies) and showed an effect of 0.33. With regard to the most often implemented interventions, strategy instruction (g = 0.36) was statistically significant, while process writing (g = 0.51) yielded a meaningful but not statistically significant result. Moderator analyses found no significant variation by study quality indicators, participant characteristics, or intervention characteristics. Findings underscore the promise of writing interventions for English language learners in the United States and suggest directions for future research, including the importance of examining linguistic supports.
Vital to long-term academic success, student engagement is a key intervention target, especially during the early years of schooling when initial signs of disengagement may appear. The primary goal of this systematic review was to identify interventions promoting elementary student engagement and determine the best intervention approaches. We systematically searched Education Resources Information Center, PsycINFO, and Web of Science (July 2024) to identify intervention studies examining behavioral, cognitive, or affective engagement among elementary students, as compared to standard instruction, through experimental or quasi-experimental designs. Methodological bias was assessed using the Mixed Methods Appraisal Tool (Hong et al., 2019). Data were synthesized through meta-analyses and subgroup analyses using a three-level model with robust variance estimation. Forty-six interventions (38,225 students) met inclusion criteria, reporting data for behavioral (k = 41; n = 36,122), affective (k = 14; n = 5,065), and cognitive engagement (k = 5; n = 1,907). Three main types of intervention were identified and compared for behavioral engagement: teacher support (i.e., professional development; g = .503, 95
The Big-Fish–Little-Pond Effect (BFLPE) is the negative association between class-average achievement and students’ academic self-concept after controlling individual achievement. Over four decades, the BFLPE has become one of educational psychology’s most robust cross-national findings. However, psychological, economic, and contextual-judgment traditions propose different representations of social comparison: adaptation-level models emphasize contextual means, ordinal-standing models emphasize relative rank, and distributional-scaling models emphasize broader properties of local achievement distributions. Using TIMSS 2019 Grade 8 mathematics data from 46 countries (11,684 classes; 262,988 students), this study provides the first large-scale cross-national test of these formulations within a unified modeling framework. Consistent with prior BFLPE research, higher individual achievement positively predicted mathematics self-concept, whereas higher classroom mean achievement negatively predicted it. Ordinal standing also positively predicted self-concept in simpler models. However, strong ordinal-standing formulations predict that the negative contextual-mean effect should substantially attenuate or disappear when ordinal position is fixed at the top of the classroom distribution. Inconsistent with this prediction, the BFLPE remained evident even for the highest-ranked student in each class. Models integrating all three approaches further indicated that much of the apparent ordinal-rank effect reflected broader forms of distributionally scaled relative achievement embedded within classroom distributions. The findings support an integrative interpretation in which ordinal standing is relevant, not sufficient: contextual means, ordinal standing, and distributional structure jointly shape comparative self-evaluation, while contextual means remained the most robust comparative indicator. More broadly, the study highlights the Statistical–Mechanism Gap Fallacy: interpreting statistical contextual indicators as psychologically transparent representations of comparison mechanisms.
Traditional verbal theories in educational psychology often remain underspecified at the mechanistic level. While they offer rich descriptive constructs and conceptual insights, they provide limited accounts of how learning processes unfold dynamically and causally. Here, “verbal” denotes theories stated in prose and qualitative relations rather than as formal, computational process models. This lack of mechanistic precision constrains rigorous theory testing, limits integration across levels of analysis, and reduces the potential to design interventions grounded in explanatory understanding of learning processes. In this Review, we synthesize an emerging paradigm that treats artificial intelligence (AI) systems not merely as predictive tools or instructional technologies, but as cognitive models of learners, explicit, runnable instantiations of theoretical assumptions about cognition and learning. Our central question is not whether AI systems can serve as cognitive models, but when they should be allowed to count as such. We therefore organize the Review around explicit validity criteria, theoretical grounding, construct validity, mechanistic transparency, alignment with human learning trajectories, error-signature matching, causal-intervention tests, ecological validity, and instructional usefulness, that an AI system must satisfy before its cognitive-model status is granted rather than assumed. We examine how major families of AI models, including neural networks, reinforcement learning agents, cognitive architectures, and large language models, have been used to operationalize core educational constructs such as memory, strategy use, motivation, self-regulation, and social learning. Across these approaches, we highlight how mechanistic transparency, interpretability, and alignment with human learning trajectories and error patterns are essential for explanatory validity. We further discuss methodological tools, such as representation analysis, ablation, and trajectory-level comparison, that enable causal inference about learning mechanisms within models. Finally, we outline key challenges and future directions, including construct validity, ecological realism, individual differences, and ethical accountability. By positioning AI as a theoretical instrument rather than solely an engineering solution, this Review argues that AI-based cognitive models can, when they satisfy these criteria, help transform abstract learning theories into precise, testable, and educationally actionable accounts of how students learn.
The surge in Generative AI (GenAI) research highlights its potential for learning. A growing body of empirical studies has examined the application of GenAI across diverse contexts. However, there remains an underexplored research field concerning how GenAI can be integrated into learning processes and what recurring patterns are observed across studies reporting positive outcomes. Addressing this gap, this systematic review synthesized a corpus of 315 eligible empirical studies with three aims: (a) synthesizing reported outcome directions associated with GenAI-mediated learning across diverse stakeholders and disciplinary contexts; (b) identifying the instructional stages in which GenAI is integrated and the functions it performs; and (c) linking these recurring patterns to learning-outcome evidence reported in the included studies. Findings revealed that empirical research remains concentrated in higher education and language-related disciplines, with predominantly positive reported outcome directions for affective and behavioral outcomes. Pattern analysis identified a guided generative feedback configuration that integrates the functional triad of content generation, teaching support, and feedback across the instructional implementation stages. Nonetheless, more variable cognitive and metacognitive outcomes point to possible boundary conditions, suggesting that positive reported outcome directions may depend on whether learners remain active agents in the learning process. This review provides an evidence-informed mapping of GenAI integration patterns in the reviewed corpus to inform future theory-driven research and evidence-based instructional practice in educational psychology.
By the time they enter primary school, students have already developed a range of intuitive conceptions, heuristics, and other automatisms (Borst, 2020; Houdé, 2014). In certain contexts, these automatisms can lead to inappropriate responses or reasoning errors, interfering with students’ understanding of targeted learning topics (Jiang et al., 2019; Roëll et al., 2019). Research has highlighted the role of inhibitory control in overcoming such automatisms (Houdé Borst, 2015; Mason Zaccoletti, 2021; Medrano Prather, 2023). However, because this research is dispersed across various fields, it remains difficult to grasp the full range of learning topics whose acquisition may involve inhibitory control. This scoping review aims to provide a structured overview of primary school learning topics for which there is empirical evidence consistent with the involvement of inhibitory control in overcoming interfering automatisms. A total of 27 articles reporting 54 experiments with children aged 6 to 12 were analyzed. The review identified learning topics primarily in mathematics, but also in science and language, suggesting that inhibitory control may be involved across multiple areas of the curriculum. Moreover, recurring characteristics emerged across the competing automatisms involved in these learning topics, including strong perceptual salience, reliance on the “more A–more B” intuitive rule, susceptibility to whole-number bias, and frequent reinforcement through instruction and experience. These four characteristics could help identify other learning topics that may require inhibitory control but have not yet been explicitly studied. Lastly, the review explores variations in negative priming effects across learning topics, suggesting that some interfering automatisms may be more difficult to inhibit than others.
Primary-secondary school transitions are critical periods in children’s educational and emotional development, which pose heightened risk for poor attainment, attendance, social adjustment, and mental health, particularly for vulnerable children. Yet, the multiple challenges presented during this time remain difficult for English schools to manage without clear statutory guidance. Government reports consistently identify transitions as a systemic weakness, while practitioners report uncertainty about what to prioritise, resulting in highly variable, locally interpreted provision. It is important that best practice guidance is developed, drawing on evidence from research, policy, and practice, to inform an evidence-informed transitions strategy in England. To narrow this gap, the present research takes an exploratory sequential mixed-methods design, triangulating multiple stakeholder perspectives and multidisciplinary evidence to develop an evidence-based framework with policy recommendations. A systematic literature review was first conducted to explore key priority areas identified within existing research. Building on these insights, 10 round-table discussions were conducted to aggregate multi-disciplinary perspectives of 52 experts within educational practice, research, and policy, nationwide. Data were analysed using Thematic Framework Analysis, to identify barriers and facilitators in implementing high-quality primary-secondary school transitions provision, and what key priority areas should be included in national strategy. Meta-inferences were then drawn to develop an evidence-based framework outlining five policy recommendations, with implementation proposals of how each might be enacted in practice. This research offers a conceptual foundation to inform policy and could have practical utility for informing high-quality research and practice. Practical and contextual challenges associated with introducing a statutory transitions strategy in England are considered.
Understanding the relationship between language attitudes and proficiency is essential for improving language proficiency and shaping effective educational and policy strategies. This correlational meta-analysis synthesizes 132 effect sizes nested within 46 independent samples from 33 research reports (N = 20,505) to examine this relationship and how it is moderated by different attitude objects and contextual and methodological variables. The overall analysis revealed a significant small to medium association (r = .34, 95
Artificial intelligence has drawn on cognitive and educational psychology since its origins, and contemporary training pipelines for large language models and related systems (curated data curricula, staged prompting supports, feedback from human judges, and benchmark-driven evaluation) make this connection unusually concrete. Drawing on Messick’s (1995) unified theory of validity, we use educational psychology to derive a set of validity checks that distinguish robust competence from support dependence, proxy optimization, and rater-specific compliance. We organize the discussion around five domains where the analogy is especially revealing: curriculum design, scaffolding and instruction, social learning via human feedback, assessment and validity, and ethical alignment. For each domain, we map a familiar educational construct onto concrete training interventions such as data selection and sequencing, prompting and tool support, preference learning and other feedback loops, benchmark design, and reward modelling. For each construct, we identify what would count as evidence that an apparent training gain is, or is not, what it appears to be. We conclude with implications for model development practice, including documenting curriculum and scaffolding decisions, treating benchmarks and reward models as high-stakes assessments that can invite proxy optimization and “teaching to the test,” and making explicit how assessment targets are revised as models and deployment contexts change.
Self-regulation competence is known to be a critical driver of academic achievement, mental health, and lifelong success. However, research on education and educational practice in K-12 remains fragmented in that groups of researchers focus on conceptualizations from different traditions, including self-regulated learning (SRL), executive functions (EFs) and personality psychology (e.g., conscientiousness), and there is limited cross-talk between these groups. In this article, we propose an integrated multicomponent model that unifies these perspectives and conceptualizes self-regulation competence as a developmental “superpower” that is both predictive of a broad range of outcomes and amenable to educational support. We examine three core propositions: (1) Self-regulation competence—broadly defined—predicts important cognitive, social, and emotional outcomes in school and across the life span; (2) SRL, EFs, and conscientiousness show considerable conceptual and empirical overlap, supporting their integration into a comprehensive framework; and (3) self-regulation competence can be effectively fostered through targeted and embedded educational interventions. To fully develop and exploit the superpower qualities of self-regulation competence, promoting self-regulation competence should become a guiding principle of education systems. High-quality teaching—characterized by effective classroom management, cognitive activation, and student support—and whole-school approaches are needed to orchestrate the promotion of self-regulation competence in real-world educational contexts.
Education in Responsible Conduct of Research (RCR) aims to foster good research practices. Little, however, is known about how educators measure the impact of their courses, which outcome measures are used and if and how they are validated. In this review, we systematically identify measures used to evaluate RCR education and assess if and how they have been validated. We also analyse a selected subset of these measures. Specifically, those that include supporting evidence of validity and focus on research practices, in order to identify which Responsible Conduct of Research (RCR) concepts they address and which learning outcomes they are designed to assess. In total, 162 articles were included, of which 46 articles contained information about a measure described as ‘validated’. Measures from these articles (and four additional measures suggested by experts) were assessed on accessibility, evidence of validation, and content related to research practice, resulting in 15 unique outcome measures with supporting validation evidence. All contained items covering topics which we categorised as pertaining to ‘research integrity’, whereas a smaller selection contained items covering topics which we categorised as pertaining to ‘research ethics’. Whilst all 15 reported at least some form of reliability or validity evidence, only 3 met what we describe as ‘strong’ validation evidence. Our review revealed that the majority of RCR education evaluations in the published literature have not used measures related to research conduct with reported validation evidence. The comprehensive overview we provide can offer guidance for educators and researchers to select, or further develop measures aligned with core RCR competencies and targeted learning outcomes.
Interventions designed to improve student learning and achievement are essential to the mission of educational psychology and education in general. However, determining which interventions have strong empirical evidence and which do not can be challenging, especially for the general public (Cleary Robinson, 2025). In the present Reflection, we compared and contrasted the evidence base of two interventions: Growth Mindset Training (GMT) and retrieval practice. Whereas retrieval practice has considerable empirical support, GMT has a mixed and debated evidence base, with average effects that appear small. The former is typically studied using intervention research methods whereas the latter usually involves observational methods. Whereas retrieval practice takes long-term sustained effort and tends to be perceived by learners in the moment as less effective than re-reading material (and less comfortable due to retrieval difficulties), GMT is a relatively short, quick, intuitive activity. This contrast may lead strategies like retrieval practice to be less readily adopted than GMT. We propose that future research should examine whether easily adopted strategies like GMT could be beneficially bundled with harder-to-embrace effective learning strategies like retrieval practice to enhance the usage and impact of both.
Some students drop out of education. Recent work suggests that mental health issues – such as burnout – may be prominent reasons why they do so. To explore this idea further, the aim of the present study was to provide the first systematic review and meta-analysis of the relationship between student burnout and educational dropout. In doing so, we focused on both direct (dropout behavior) and indirect measures (dropout intentions) of dropout. Following pre-registration, we conducted a search of PsycINFO, PsycArticles, MEDLINE, Education Abstracts, Educational Administration Abstracts, and ProQuest Dissertations and Theses up to January 2026. We followed PRISMA guidelines and provided a narrative synthesis and multilevel meta-analysis. Our search found 31 studies (with 34 samples; N = 39,454) with most studies focusing on dropout intentions. In the three studies that had examined direct measures, burnout was shown to be a significant predictor of dropout behavior (including in one study examining enrollment status). In the remaining studies, burnout was consistently associated with dropout intentions both cross-sectionally, and, in several studies, over time. This was supported by the findings of the multilevel meta-analysis where burnout was positively associated with dropout intentions (r+ = .46, 95
Reading and mathematics are core components of children’s academic development and are linked to later educational, occupational, and psychosocial outcomes. Yet evidence on their predictors remains dispersed across biological, cognitive, behavioral, contextual, and neuroimaging traditions, and few syntheses have examined these literatures within a common developmental framework. In this systematic review, we synthesize 165 eligible studies to examine multilevel predictors of reading and mathematics achievement in children and adolescents and to distinguish shared from domain-specific patterns of prediction. We organize the evidence within a multilevel framework spanning perinatal and early-life conditions, cognitive skills, behavioral, emotional, and motivational processes, family and school contexts, and neuroimaging indicators. Across levels, the literature supports a developmental account in which early biological and contextual conditions are best understood as distal starting conditions, cognitive skills as the most proximal capacities for academic learning, and behavioral-emotional processes as regulatory pathways through which those capacities are expressed. Reading is more consistently associated with language-related and sound-symbol skills, whereas mathematics is more consistently associated with numerical concepts, symbolic processing, spatial resources, and math anxiety; executive function and general cognitive ability appear to provide a shared scaffold across both domains. We also include a worked multimodal neuroimaging illustration, explicitly framed as an illustrative application rather than an evidential extension, to show how the neural layer may be incorporated as a source of child-proximal markers. Overall, the review advances an evidence-calibrated multilevel framework for organizing shared and domain-specific predictors of reading and mathematics achievement and clarifies the interpretive limits of translating predictive evidence into intervention claims.