
This study analyzed the relation between self-efficacy (SE) and sense of belonging (BL) in introductory physics courses by measuring these variables at two time points, early in the semester and midsemester. The study was performed in the introductory, calculus-based mechanics class at a large university in the United States and includes 1373 complete records. An extensive set of Observed Prior Causes (OPCs), including Force and Motion Conceptual Evaluation pretest scores, college and high school preparation variables, high school physics and math variables, personality variables, and demographics, were found to be predictive of SE or BL. Early in the semester, SE and BL were highly correlated (r=0.56); after controlling for OPCs, r=0.37 of this correlation remained. SE fully explained the relation of BL to midsemester test average, while OPCs fully explained the relation of SE to test average. Changes to SE and BL in response to course feedback in the form of test averages were also explored. Test average explained little of the relation between early and midsemester SE or BL. The analysis supported a model where test average predicted midsemester SE, which in turn predicted midsemester BL. Mediation analysis found both directed and correlational paths were important in explaining the complex relation of these variables. Some moderation effects were detected, but were generally small.
This study examines the pedagogical quality and instructional adequacy of analogies included in high school physics textbooks (grades 9–12) in Türkiye. The analogies extracted systematically from 15 physics textbooks revealed a significant imbalance in their use across grade levels. The vast majority of the analogies were concentrated in the 10th grade textbooks, while their occurrence in other grades remained quite limited. A similar imbalance was also observed in subject focus: most of the analogies were allocated to the conceptually challenging and highly abstract topic of electricity, while other fundamental physics topics were largely neglected. In terms of pedagogical quality, all of the analogies were classified as either moderate or good according to Glynn’s teaching-with-analogies model. However, most of the analogies reviewed lacked a step for identifying their limitations, thereby increasing the risk of misalignment and the development of misconceptions. The analogies were also predominantly of the concrete-abstract type, which helps students connect abstract physics concepts to concrete examples from daily life, and were largely based on functional similarity, meaning they focus on how the concept works rather than its structural features. Furthermore, the increased emphasis on the relevance of analogies to learners’ backgrounds reflects a positive trend toward enhancing the contextual richness of analogies. To address the quantitative insufficiency of analogies and the imbalance across subject areas, it is essential that curriculum developers position analogical thinking as a required pedagogical tool for textbook authors.
Physics demonstrations are celebrated for enlivening lectures and anchoring understanding, yet remarkably little is known about how the choice of demonstration materials, namely, familiar household objects versus specialized laboratory equipment, may shape students’ affective evaluations. A total of 288 secondary and college students observed two resonance demonstrations—one using wine glasses as everyday focal objects and one using tuning forks as laboratory focal objects—and completed a survey requesting an overall preference (household/lab/no preference) and Likert ratings on seven perception items for each demonstration. Demonstration order was counterbalanced across class sessions to guard against sequence effects. Overall, students far more often preferred the wine-glass demonstration (58.3%) over the tuning-fork demonstration (14.9%), with 26.7% reporting no strong preference; this preference pattern was consistent across both secondary and college students and did not differ significantly between male and female students. Preference distributions were also independent of presentation order. Paired comparisons showed that the household demonstration was rated significantly higher in interestingness, novelty, relevance to daily life, and its ability to stimulate interest in science. Importantly, clarity of the phenomenon and self-reported support for understanding resonance did not differ significantly between the two demonstrations, supporting the interpretation that the preference for the household condition is not attributable to the phenomenon being easier to observe or the concept easier to grasp. The results are theoretically grounded in cognitive load and situated cognition frameworks to explain how material-based contextualization may enhance student engagement without increasing perceived cognitive burden. These findings suggest that a seemingly simple instructional design choice, i.e., using a recognizable household object as the focal material in a physics demonstration, can meaningfully enhance student engagement and perceived relevance without sacrificing conceptual clarity.
This paper presents the first iteration of development and classroom testing of Stars and Elements, a web-based learning resource for Norwegian upper secondary physics students focused on how heavy elements are produced in stars and neutron star mergers. Combining the model of educational reconstruction (MER) with the pedagogical link-making framework and design-based research (DBR), we identified core ideas to be communicated, reviewed the physics curriculum and research on student understanding and motivation, and formulated design hypotheses that guided development of a 90-min learning resource with animations, videos, interactive tasks, peer discussion, and teacher consolidation. Six classes (N = 120) participated; data comprised students’ written responses to three open tasks, brief evaluations, and classroom observations. Most students could describe central concepts and processes concerning red giant stars and neutron star mergers as sites for heavy element production, notably the slow and rapid neutron capture processes, and many benefited from animations and varied task formats. Some displayed uncertain grasp of concepts such as fusion and beta decay. Nature of science (NoS) aspects concerning how physics knowledge is developed were found to be insufficiently represented. From these findings, we propose design principles for research-based learning resources on similar topics, including emphasizing explicit link-making across scales and prior knowledge, scaffolding for challenging nuclear concepts, multimodal visualizations, and explicit NoS focus. We argue that combining MER, link-making, and DBR is a promising approach for developing learning resources that reflect the frontiers of science while having a solid grounding in the classroom.
In this paper, the focus has been on students’ application of special relativity theory after learning it. In order to determine students’ understanding and conceptualization of the theory of special relativity, paradoxes, which are contradictory problems, have been used in the study. How students solved paradoxes related to the relativity of simultaneity, length contraction, and time dilation has been investigated, and an answer has been sought to the question “What is the state of undergraduate students who have learnt special relativity theory in applying the theory to paradoxes?” The students who participated in the study had previously attended Physics I, General Physics 2, and General Physics 3 courses in which classical physics topics (force and motion, electricity and magnetism, and optics and waves) were taught. The data were analyzed using a phenomenographic approach to identify the categories emerging from students’ responses. Most students experience difficulty in understanding the nature of special relativity and applying the theory to other situations. For this reason, they did not use the concept of the reference frame in solving the paradoxes. From the students’ responses, it is understood that they mostly think within a single reference frame and experience difficulty in understanding the contradictions that arise when viewed from different reference frames.
The theoretical framework knowledge in pieces (KiP), proposed by Andrea diSessa, holds that intuitive knowledge in physics is organized through small cognitive structures called p-prims (phenomenological primitives). One of the heuristic principles on which the KiP framework is built—the principle of redescription—suggests that the descriptive framework of p-prims can be fine-tuned. However, the theoretical development of p-prim descriptions has remained largely limited to informal linguistic expressions drawn from students’ explanations, without a clearly defined syntactic or semantic structure. This lack of systematicity has led to variations in the naming and conceptualization of p-prims, making their consistent identification and analytical use difficult. In response to this limitation, this paper proposes a formal schematization of p-prims based on Fillmore’s Frame Semantics, supported by tools such as FrameNet. Our proposal aims to make explicit the core and peripheral components of each p-prim and to clarify the conditions under which they are evoked. By providing p-prims with a well-defined linguistic structure, the KiP framework is strengthened, opening new possibilities for the automated analysis of students’ reasoning using artificial intelligence (AI), as well as for the design of visual representations to support the development of instructional materials. Theoretical, methodological, and technological implications of this proposal for future research in physics education research (PER) are also discussed.
Physics students’ use of ChatGPT or other generative AI tools is a topic broadly discussed in physics faculties around the globe, with discussion often outpacing evidence. What remains unclear is how physics students truly engage with these tools and by what perspectives their choices are shaped. This study develops a data-driven typology of how students perceive the tool’s usefulness and its risks. Through a qualitative content analysis of 1189 survey responses, followed by latent class analysis (LCA), we identified two distinct user profiles. The majority (70%) are “Pragmatic Users” who, while aware of the tool’s inaccuracies, use it for scaffolding tasks like conceptual clarification and finding starting points for problem-set questions. The minority (30%) are “Skeptical Non-Users,” who avoid the tool over concerns about overreliance and its potential to hinder independent problem-solving. These profiles show students making a calculated trade-off between usefulness and risk, suggesting one-size-fits-all policies on AI are inadequate. Thus, rather than promoting uniform adoption, a differentiated pedagogy could create space for students to make informed, self-determined choices—supporting reflective AI use for those who opt in, while equally validating and enabling the learning practices of those who choose not to use generative AI.
We report results from a mixed-methods survey study (N=81) examining how undergraduate students in introductory physics courses use, trust, and prefer to integrate artificial intelligence (AI) tools into their learning. Among respondents, 91% (95% confidence interval (CI): [83%, 96%]) reported using AI for physics coursework, yet only 41% (95% CI: [30%, 52%]) trusted AI-generated physics explanations—a 50-percentage-point trust-utility gap consistent with domain-calibrated skepticism rather than uncritical adoption. Thematic analysis of open-ended responses (n=67; Cohen’s κ=0.78) identified eight themes; the most distinctly physics-specific finding was that 40% of qualitative respondents spontaneously articulated whereAI fails in physics—visual-spatial reasoning, circuit analysis, and abstract physical reasoning—aligning with physics education research on student difficulties with multirepresentational tasks. This skepticism is also consistent with Kortemeyer’s benchmarking results, which show that AI performance is weakest on precisely the multirepresentational task types that students identified as failure-prone, suggesting that student trust calibration may track underlying AI competence boundaries in physics. A majority (65%) preferred optional over mandatory AI integration. Because high-performing students were overrepresented by 13.9 percentage points, the estimated preference for optional integration (and/or resistance to mandatory) may be inflated in this sample. These findings highlight a critical unresolved priority—the verification gap between self-reported and actual verification behavior—and inform recommendations for AI-integrated physics pedagogy.
The ability of ChatGPT, especially its latest versions, attracts significant attention from educators and researchers in fields such as physics. Several studies on this topic explore the performance of ChatGPT in answering physics questions across different versions of ChatGPT. In this study, we evaluate the development of ChatGPT in answering physics questions using the Mechanical Waves Conceptual Survey (MWCS) to examine whether its capabilities improve across different versions. A total of 75 responses per item are generated using ChatGPT versions 4, 4.o, and 5.2, resulting in 4950 responses for comprehensive analysis. Based on the comparison between the three versions of ChatGPT, we find that some difficult types of questions that are answered incorrectly by previous versions, as reported in earlier studies, can now be answered correctly, although there are still minor mistakes, especially when examining the reasoning. A generalized linear mixed model (GLMM) analysis, together with post hoc comparisons, indicates that ChatGPT 5.2 significantly outperforms both ChatGPT 4 and 4.o, suggesting that the capabilities of the latest version included in this study significantly improve. These findings highlight a future direction for this type of research, as our study shows that the latest version of ChatGPT can correctly answer more types of questions in MWCS.
Learning progressions (LPs) are widely used to describe the long-term development of students’ understanding of core scientific ideas. However, as increasing attention has been paid to the context-dependent nature of student learning, the limitations of traditional LP models in capturing the complexity and dynamics of cognitive development have become more apparent. This study integrates the partial credit model (PCM) and model analysis to systematically investigate students’ cognitive developmental characteristics within the Force and Motion Learning Progression (FM-LP). Using Force Concept Inventory data from 1619 Chinese students, we repeated the study by Fulmer et al. on U.S. students to examine the applicability of the established FM-LP across Chinese student populations. The results indicate that the established FM-LP does not exhibit satisfactory model fit in the Chinese sample, providing further evidence that science learning is highly context- and population-dependent. To characterize students’ intermediate cognitive states, this study introduces a model-state combination analysis framework. The findings show that most students do not stably occupy a single LP level; instead, they display thinking patterns spanning multiple levels, with clear differences observed between middle school and high school students. Further analyses based on PCM reveal that, as students’ ability increases, their model states shift from strongly mixed states to less mixed states and eventually toward the relatively consistent use of the scientific conception. The model-state-based approach offers a new perspective on capturing cognitive complexity in learning progressions and provides important implications for progression-oriented assessment.
This paper is the second in a two-part series that investigates students’ reasoning on astronomical phenomena. In this work, we examined how (if at all) students rely on spatial scale assumptions when reasoning about Moon phases. To that end, we conducted N=25 semistructured interviews with last-year high school students. A personalized, physical scale model of the Sun-Earth-Moon system was constructed for every participant individually and placed on the interview table. Students were encouraged to clarify their thoughts by interacting with the scale model or by making a drawing. The scale estimates underpinning the scale model on the table were assessed through an interactive online survey, which was administered several days to weeks prior to the interviews. Only students with severely underestimated values for the relevant sizes and distances were selected for an interview. This selection procedure was motivated by both pragmatic and empirical considerations. We observed a wide variety of explanations for Moon phases, some of which were not found in the existing literature. Several students used their (inaccurate) scale ideas to reason about Moon phases. However, occasionally an interviewed student expressed arguments inconsistent with the scale model on the table. In these cases, scale assumptions seemed to be not taken into account when reasoning about Moon phases. The implications of these findings are discussed.
This paper is the first in a two-part series that investigates students’ reasoning on astronomical phenomena. We interviewed N=25 last-year high school students to uncover their reasoning on three observable phenomena, two of which will be discussed in this paper. The first phenomenon was the apparent sizes of the Moon and the Sun (as seen from Earth) and potential causes for their variation. The second phenomenon regarded the event of a solar eclipse and the difference between total and annular eclipses. The interviews were held in light of the students’ understanding of spatial scales in the Earth-Moon-Sun system. Based on their answers on a prior online survey assessing students’ estimates of these scales, a personal to-scale model of the Earth-Moon-Sun system was provided for every student. The students selected for interviews were those who showed a substantial underestimation of these scales. Students were repetitively encouraged to illustrate their reasoning on both phenomena by using the scale model or by providing a drawing. For both discussed astronomical phenomena, we encountered several alternative explanations that—as far as we know—were not previously documented in the research literature. Many of these relied on an inaccurate comprehension of the involved scales. Although for some students these explanations were in line with the scale model on the interview table representing their substantial underestimations, others did not seem to take spatial scales into account when making suggestions on the causes for these astronomical phenomena. The implications of these findings are discussed.
The present study describes the development and field testing of the nature of heat inventory (NOHI) that examines high school and undergraduate students’ understanding of heat transfer processes in the terms of kinetic molecular theory. NOHI comprises 10 items. Based on preliminary study and the history of science, the 30 distractors of NOHI were designed to reflect the following three categories of possible misconceptions regarding heat: (a) materialistic views, (b) hot atoms, and (c) expanding atoms. Field testing, carried out on 81 students for science teaching, showed the materialistic misconceptions of heat still to be the most popular. NOHI has been proven to be a valid and reliable tool that could be used to determine students’ thinking regarding classical heat processes.
Students interpret grades as signals of their strengths, and grades inform students’ decisions about majors and courses of study. Differences in grades that are not related to learning can impact this judgment and have real-world impact on course-taking and careers. Existing work has examined how an overemphasis on high-stakes exams can create equity gaps where female students and Black, Hispanic, and Native American students earn lower grades than male students and Asian and white students, respectively. Yet, minimal work has examined how the weighting of individual midterm exam scores can also contribute to equity gaps. In this study, we examine how three midterm exam score aggregation methods for final grades affect equity gaps. We collected midterm exam data from approximately 6000 students in an introductory physics course over 6 years at a large, research-intensive university. Using this dataset, we applied common midterm exam score aggregation methods to determine their impact on aggregated midterm exam grades: dropping the lowest midterm exam score, replacing the lowest midterm exam score with the final exam score if higher, and counting the highest midterm exam score more in the final grade calculation than the lowest midterm exam score. We find that dropping the lowest midterm exam score resulted in the largest increase in final grades, with students with lower grades benefiting the most. However, we find no evidence that these alternative midterm exam aggregation methods close equity gaps. Instead, we find some evidence that such practices could increase equity gaps between male and female students, though the effect is very small in relation to the overall grade increase. Implementing the alternative midterm exam aggregation methods examined here may be useful for instructors wanting to raise course grades or give lower-scoring students a boost. However, they do not appear to be effective in reducing equity gaps.
Accurate diagnosis of student misconceptions is crucial for effective science instruction. Traditional diagnostic tools, however, often rely on unreliable self-reported confidence levels. This study developed and validated a two-dimensional diagnostic model that integrates objective eye tracking data with test responses to more robustly identify misconceptions in Newtonian mechanics. The study involved three phases: developing a four-tier diagnostic test, conducting an eye tracking experiment to analyze visual attention patterns (fixation duration, revisits, scan paths) among high school students, and constructing the final model by correlating eye movements with confidence judgments. Consistent with previous research, results showed that students with misconceptions had shorter fixation durations on correct options and displayed inefficient scan paths. Conversely, those with scientific conceptions showed focused attention. As a novel contribution of this study, a “reversed U-shaped” pattern was observed between confidence and revisits, indicating that revisits can serve as an objective measure of cognitive uncertainty. Based on these findings, we propose a two-dimensional “eye tracking-test” model. This model offers a more precise method for diagnosing student understanding, providing a foundation for targeted instructional interventions and fostering a dynamic “teach-learn-assess” feedback loop for personalized physics education.
Peer discussion (PD) is widely used as an active-learning strategy in physics education, yet most prior research has focused on conceptual multiple-choice questions. Much less is known about how peer discussion functions when applied to mathematically structured problems typical of calculus-based mechanics courses. To address this gap, we compared the effects of PD and instructor explanation (IE) during short in-class exercise sessions in a calculus-based introductory mechanics course (N=87). Using a cluster-randomized crossover design with two course sections, instruction alternated between PD and IE across weekly exercise topics. After each lesson, students completed a 16-item post-test; 12 of these items were also administered on the final exam. Across the 16 post-test items, we found no statistically significant evidence that PD outperformed IE on any single item. In unadjusted item-level analyses, four items that required greater abstraction and multistep mathematical reasoning showed a trend favoring IE, although none of these differences remained statistically significant after Holm correction for multiple comparisons. On the final exam, no statistically significant differences were observed between conditions. These results suggest that, within the instructional conditions of this study, PD did not produce consistent performance gains relative to IE, and any item-level differences may depend on task characteristics such as mathematical abstraction, novelty, and multistep reasoning demands.
Generative artificial intelligence (GenAI) has shown growing potential in supporting teacher education, yet little is known about how preservice teachers collaborate with AI during instructional design (ID). This study examined the effects of DeepSeek, a Chinese large language model, on preservice physics teachers’ ID performance and cognitive processes. A quasiexperimental design was conducted with 220 participants completing a 45-min ID on Newton’s third law, either with or without DeepSeek assistance. Data from IDs, interaction logs, and interviews were analyzed through statistical tests, epistemic network analysis (ENA), and thematic coding. Results indicated that the DeepSeek-assisted group outperformed the control group in analytical and structural aspects of ID, such as teaching content analysis and learner analysis. ENA revealed that high-performing participants exhibited proactive “problem representation–AI interaction–monitoring” patterns, while low-performing participants relied on reactive prompting. Interviews further confirmed that AI was helpful in contextual introductions and experiment design but also revealed issues such as outdated curriculum references and overly absolute expressions that lacked sensitivity to teaching contexts. The findings suggest that GenAI can serve as a cognitive scaffold when used strategically and highlight the need to integrate AI literacy and metacognitive training into teacher education.
Student experiences in physics beyond the classroom and their role in supporting student development have been the subject of increasing research attention in recent years. Results, typically from small studies at single institutions, have illustrated that facilitating informal physics experiences for nonscientists can enhance student disciplinary identity, learning, sense of belonging, and career skills. However, it is essential to examine whether these impacts are the sole provenance of institutions with well-developed outreach programs or if they may be shared by institutions anywhere. This work reports on the analysis and findings of responses to three open-ended questions presented to students who indicated they had engaged in facilitating outreach programs as part of a national survey distributed through the Society of Physics Students network in spring 2023. Employing a network analysis with Girvan-Newman clusters revealed six core themes of student experiences: community participation, resilience, transformation, audience dialogue, disciplinary development, and disciplinary connectedness. The first four of these clusters were observed to be highly interconnected, providing evidence that the impacts and experiences within them are interrelated with other clusters, particularly interactions with the audience, which is a central feature of informal physics programs. In particular, student experiences highlighted that facilitating informal physics programs enhanced their resilience and belonging, grew their physics identity, provided opportunities to develop essential career skills, and cultivated a growth mindset.
We investigate how faculty share aspects of their identities to foster connections with students, finding significant nuances in what, how and when they decide to share. From interviews with 19 faculty, we identify and develop four personas: Brooke, the Trust Builder, prioritizes creating an environment of trust by openly discussing their identity, aiming to foster student openness. Wray adopts a walled-off approach, separating personal and professional life due to past negative experiences or a belief in the importance of that division. Casey, the Cautious Sharer, expresses concerns about potential alienation or backlash and approaches personal sharing with caution. Finally, Nour, the Identity Navigator, shares personal experiences to assist others in navigating their own identities, acknowledging the challenges of the college years. Communication of shared or adjacent identities is a key step toward developing an empathetic understanding that can facilitate taking appropriate supporting actions, and the nuances of how these identities are communicated add insight into how faculty can work within their comfort zones to better support students.
Models are commonly used in astronomy education to help students understand and discuss difficult or abstract concepts. From a physics education research perspective, metaphors and analogies are often used to support reasoning about abstract phenomena. This study investigates how students draw on classroom experiments as physical analogies (a multimodal instance of conceptual metaphor) when reasoning about exoplanetary atmospheres. We conducted two lessons: one about cloud formation on Earth and exoplanets and another about lightning on Earth and exoplanets, both including a hands-on experiment related to the topic. The students were then split into peer-led focus groups and tasked with discussing the experiment and its relation to weather in planetary atmospheres. We audio-recorded these group discussions and extracted relevant excerpts. We analyzed these excerpts sequentially using thematic analysis, network analysis, and conceptual blend analysis. We identified four input spaces used by students (relevant physics, the experiment, Earth conditions, and exoplanet conditions) and three ways in which students use analogies (mapping between domains, identifying limitations, and extending the analogy). Across these analyses, we found that students integrate these input spaces to construct blend spaces that mediate their reasoning about processes and concepts in Earth and exoplanetary weather phenomena. We also discuss implications for future research and teaching.