AI has become a partner in how people learn about, do, and engage with science, and the partnership takes three forms: a scientist works with a co-scientist whose output must be checked; a member of the public looks something up to decide whether a diet works or whether to fit solar panels; and a student takes up an inquiry with AI in a science class. Across all three, one thing decides whether the partnership helps or harms: whether the human evaluates what the AI returns or takes it on trust. I argue that this evaluation -- epistemic vigilance calibrated to how far a fallible source can be trusted -- is, given adequate prior knowledge, the binding constraint on productive augmentation. You can hand the AI a great deal precisely because you stay vigilant; vigilance makes generative partnership safe, so it licenses augmentation rather than restricting it. Vigilance is already invoked in science education but under-specified for the AI case; I specify its components, the mechanism tying it to learning, and a way to measure it without soliciting the evaluation it is meant to detect. What is distinctive is that the machine's fluent, confident prose reads as trustworthy whether or not it is, so its surface works against the human evaluating it. The argument bears hardest on education: the integrated conceptual knowledge instruction aims to foster forms only under deep processing, and vigilance sets how deeply a claim is processed, so it is the precondition for learning with AI. The design factors the field reports matter through whether they engage the learner's evaluation; none works around it. Untested is vigilance as a measured disposition, above all where the AI is confidently wrong. Because it is unevenly distributed, integrating AI uniformly is likely to widen the gap between better- and less-prepared students. I close on how it might be built by fading support as the learner takes over.
The Physics Lab Inventory of Critical Thinking (PLIC) measures three components of students' critical thinking in physics labs: evaluating data, evaluating methods, and proposing next steps. Prior work has analyzed these components in isolation or as a composite score. In this study, we apply latent profile analysis (LPA) to the three PLIC scales using a large, multi-institutional dataset of 5,513 matched pre/post student records to identify characteristic response patterns across the three components simultaneously. At both pre- and post-instruction, a two-profile solution best fit the data. Profile composition shifted substantially over instruction, with 48.4% of students in the lower-performing profile at pre-test transitioning to the higher-performing profile at post-test, while 43.6% of students moved in the opposite direction. Course type was statistically associated with profile membership at both timepoints, though the effect was small (Cramér's V ≈ 0.10). To examine the relationship between profile transitions and students' affective development, we estimated cross-lagged panel models (CLPMs) linking profile membership to belonging, recognition, self-efficacy, and agency. Belonging emerged as the principal upstream predictor, prospectively predicting recognition, self-efficacy, agency, and higher-knowledge profile membership. Agency and self-efficacy formed a reciprocal but asymmetric loop, with the path from agency to later self-efficacy being stronger. Recognition functioned primarily as a downstream construct over this timescale. These results provide the first person-centered, multidimensional characterization of PLIC performance and demonstrate that epistemic and identity-related constructs are interlinked in physics lab learning.
University physics lab courses should aim to develop students' experimental skills. These experimental skills have been-among others-conceptualized as critical thinking, i.e. as the ways in which one uses data and evidence to make decisions about what to trust and what to do in the lab. The Physics Lab Inventory of Critical Thinking (PLIC) was developed to measure students' critical thinking skills in the lab. PLIC measures three components: the ability to evaluate data, the ability to evaluate experimental methods and the ability to propose next steps for experimentation. However, little is known about the relationship between these components and how they develop. Thus, this study focuses on describing students' critical thinking skill profiles amongst physics students based on these three components. We also look for connections between the profiles and students' gender, self-efficacy, and view of experimental work. The data comes from a Finnish physics department, where students have answered to PLIC for 3 years in the beginning of a first-year lab course (N = 174). Latent profile analysis was used to classify the critical thinking profiles. Logistic regression was used to test to what extent students' gender, self-efficacy, and attitude towards experimental work was related to profile membership. The results show three distinct critical thinking profiles named 'average' (n = 122), 'method evaluators' (n = 29) and 'data evaluators' (n = 22). Profile membership was independent of gender, self-efficacy, and view of experimental work. Future research should focus on replicating these profiles with data from different contexts and to longitudinally follow the development of critical thinking of students from different profiles.
The same dataset can be analysed in different justifiable ways to answer the same research question, potentially challenging the robustness of empirical science1-3. In this crowd initiative, we investigated the degree to which research findings in the social and behavioural sciences are contingent on analysts' choices. We examined a stratified random sample of 100 studies published between 2009 and 2018, in which, for one claim per study, at least five reanalysts independently reanalysed the original data. The statistical appropriateness of the reanalyses was assessed in peer evaluations, and the robustness indicators were inspected along a range of research characteristics and study designs. We found that 34% of the independent reanalyses yielded the same result (within a tolerance region of ±0.05 Cohen's d) as the original report; with a four times broader tolerance region, this indicator increased to 57%. Of the reanalyses conducted, 74% reached the same conclusion as the original investigation, 24% yielded no effects or inconclusive results and 2% reported the opposite effect. This exploratory study indicates that the common single-path analyses in social and behavioural research should not be simply assumed to be robust to alternative analyses4. Therefore, we recommend the development and use of practices to explore and communicate this neglected source of uncertainty.
Students' diverse levels of knowledge and competence-shaped by individual interests and educational debts, including structural, systemic, and institutional barriers-create substantial cognitive heterogeneity in instructional settings. Adequately addressing this heterogeneity is challenging. Emerging studies applying artificial intelligence (AI) in education claim that advanced AI techniques like machine learning (ML) can mitigate educational debts by providing adaptive support. However, previous research offers limited clarity on how learning outcomes vary following AI-based adaptive instruction and which students improve their learning outcomes. To address these issues, this article quantitatively examines the extent to which ML-based adaptivity influences students' learning outcomes over time and identifies which students, considering their intersectional identities, benefit most from this adaptive support. Specifically, we illustrate a semester-long study conducted within an undergraduate organic chemistry course, where an ML model adaptively supported 266 students across four interventions on mechanistic reasoning. We identified five learning trajectories throughout these adaptive interventions. Our findings show that students with higher prior knowledge made greater progress than their peers. Additionally, men without an underrepresented minority (URM) status majoring in chemistry benefited more than URM women who are not chemistry majors. This indicates that the adaptive support maintained and partly exacerbated educational debts. Our study contributes to the literature by analyzing how ML-based adaptivity affects educational debts in undergraduate organic chemistry. In doing so, it adopts a theoretical framework-the enhanced educational debt framework-to assess when different aims of adaptive support are most appropriate in undergraduate education, informing an equity-centered design of adaptive support.
Recent research underscores the importance of inquiry learning for effective science education. Inquiry learning involves self-regulated learning (SRL), for example when students conduct investigations. Teachers face challenges in orchestrating and tracking student learning in such instruction; making it hard to adequately support students. Using AI methods such as machine learning (ML), the data that is generated when students interact in technology-enhanced classrooms can be used to track their learning and subsequently to inform teachers so that they can better support student learning. This study implemented digital workbooks in an inquiry-based physics unit, collecting cognitive, metacognitive, and affective data from 214 students. Using ML methods, an early warning system was developed to predict students’ learning outcomes. Explainable ML methods were used to unpack these predictions and analyses were conducted for potential biases. Results indicate that an integration of cognitive, metacognitive, and affective data can predict students’ productivity with an accuracy ranging from 60 to 100% as the unit progresses. Initially, affective and metacognitive variables dominate predictions, with cognitive variables becoming more significant later. Using only affective and metacognitive data, predictive accuracies ranged from 60 to 80% throughout. Bias was found to be highly dependent on the ML methods being used. The study highlights the potential of digital student workbooks to support SRL in inquiry-based science education, guiding future research and development to enhance instructional feedback and teacher insights into student engagement. Further, the study sheds new light on the data needed and the methodological challenges when using ML methods to investigate SRL processes in classrooms.
Abstract This chapter introduces a case study of how to apply unsupervised ML to science education-related language data. We will start with dimensionality reduction and hierarchical-agglomerative clustering. Later on, we will extend these analyses using LLMs and more involved clustering techniques.
Problem solving is considered an essential ability for becoming an expert in physics, and individualized feedback on the structure of problem-solving processes is a key component to support students in developing this ability. Problem-solving processes consist of multiple elements whose order forms the sequential structure of these processes. Specific sequential structures can be expected to better reflect expert problem solving and more likely lead to successful solutions. However, this sequential structure often receives limited attention in assessments, thereby neglecting possibly valuable diagnostic information that could be used for individualized feedback. Consequently, a deeper understanding of the sequential structure of students' written physics problem-solving approaches could leverage novel potentials for physics instruction and feedback provision. This study therefore aimed to examine how the sequential structure of written problem-solving approaches differs between high- and low-performing problem solvers as well as to what extent specific sequential elements are predictive of problem-solving performance. To achieve this, we employed methods from process mining and sequence analysis research. Our findings revealed that low-performing problem solvers often lack structure in their problem-solving approaches, contrasting with notably more systematic approaches of the high-performing problem solvers. Additionally, the order in which assumptions and conceptual aspects are addressed in a problem-solving approach seems to be an indicator of problem-solving performance. The findings of this study enhance our understanding of physics problem-solving processes and highlight opportunities for improving instruction and feedback for physics problem solving by considering the sequential structure of students' physics problem-solving approaches.
Abstract In this chapter, we explain why ML can be valuable for analyzing complex data in science (education) and show what types of data you might encounter in your research project. We single out important characteristics of these types of data that can become particularly important for the performance of ML algorithms.
Abstract In this chapter we will apply ML with the purpose of building a reliable classifier for either classifying students into groups, or predicting test scores of students.
Science education aims to foster knowledge-in-use, which is supported by the integration of scientific ideas. To study knowledge integration effectively, network analysis provides a valuable tool for visualizing and understanding how ideas are connected. Successful knowledge integration requires following a learning progression that leads to increasingly sophisticated connections between ideas. However, traditional learning progression models have limitations, as they often fail to account for the nonlinear and individualized nature of learning. This study explores the potential of digital learning environments and AI techniques to address these limitations by enabling frequent, high-resolution data collection and analysis in order to uncover individual students’ learning trajectories at a high resolution. We analyze a case study of middle school students’ learning about energy to investigate patterns and variations in their learning trajectories. Additionally, we explore how different learning trajectories influence the development of knowledge-in-use, leading to either productive or unproductive learning outcomes. Our findings aim to guide instruction for teachers and instructional designers, providing insights on how to develop more effectively adaptive learning environments that support diverse student learning trajectories.
Energy is one important concept in physics, but science education research has repeatedly shown that students struggle to develop a full understanding of energy. Especially challenging for students is the notion of potential energy. Overwhelmed by the sheer number of potential energy forms, students struggle to make connections between them. Students often struggle to develop a conceptual understanding of potential energy, resulting in difficulties in learning about energy in general and their continued learning about energy. To address this issue, scholars have proposed incorporating fields into energy instruction. Through fields, the various forms of potential energy can be connected and synthesized into two simple underlying principles: (1) fields mediate interaction-at-a-distance and (2) the energy is stored in a field with the amount of energy depending on the configuration of the objects. Recent studies suggest that incorporating fields in middle school energy instruction is feasible and effective; however, little is known about whether and how middle school students connect energy and fields ideas to benefit their learning. In response to this research gap, we developed a unit on energy with fields and a comparable unit without fields and compared students' learning on energy in these two units. In a mixed-methods approach, we examined students' learning on energy during an introductory and a continued learning unit on energy with N = 67 students from grade 7. Our findings suggest that students who learned about energy with fields outperformed students who learned about energy without fields. Furthermore, fields-based energy instruction seemed to support students in developing better-connected knowledge networks that reflect deeper conceptual understanding of energy. Our findings suggest that incorporating fields into energy instruction could help students to better understand energy and to better continue learning about energy.
This report summarizes the outcomes of a two-day international scoping workshop on the role of artificial intelligence (AI) in science education research. As AI rapidly reshapes scientific practice, classroom learning, and research methods, the field faces both new opportunities and significant challenges. The report clarifies key AI concepts to reduce ambiguity and reviews evidence of how AI influences scientific work, teaching practices, and disciplinary learning. It identifies how AI intersects with major areas of science education research, including curriculum development, assessment, epistemic cognition, inclusion, and teacher professional development, highlighting cases where AI can support human reasoning and cases where it may introduce risks to equity or validity. The report also examines how AI is transforming methodological approaches across quantitative, qualitative, ethnographic, and design-based traditions, giving rise to hybrid forms of analysis that combine human and computational strengths. To guide responsible integration, a systems-thinking heuristic is introduced that helps researchers consider stakeholder needs, potential risks, and ethical constraints. The report concludes with actionable recommendations for training, infrastructure, and standards, along with guidance for funders, policymakers, professional organizations, and academic departments. The goal is to support principled and methodologically sound use of AI in science education research.
Abstract In this chapter we revisit supervised ML and apply it to text data. We particularly utilize a LLM in the fine-tuning paradigm to showcase how these models can be used in science education research projects.