Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics provide limited insight into how grading performance varies across student responses of differing grading difficulty. We introduce an evaluation framework for LLM-based ASAG based on item response theory (IRT), which models grading correctness as a function of latent grader ability and response grading difficulty. This formulation enables response-level analysis of where LLM graders succeed or fail and reveals robustness differences that are not visible from aggregate scores alone. We apply the framework to 17 open-weight LLMs on the SciEntsBank and Beetle benchmarks. The results show that even models with similar overall performance differ substantially in how sharply their grading accuracy declines as response difficulty increases. In addition, confusion patterns show that errors on difficult responses concentrate disproportionately on the label, indicating a tendency toward intermediate-label collapse under ambiguity. To characterize difficult responses, we further analyze semantic and linguistic correlates of estimated difficulty. Across both datasets, higher difficulty is associated with weaker semantic alignment to the reference answer, stronger contradiction signals, and greater semantic isolation in embedding space. Overall, these results show that item response theory offers a useful framework for evaluating LLM-based ASAG beyond aggregate performance measures.
Artificial intelligence (AI) and digital technology provide extensive possibilities for education. But focusing only on their implementation does not address the challenges associated with them and they may even have a negative impact on learners. So far, the disciplines of psychology and computer science have not provided a theoretical framework for the competence needed to develop and use AI and digital technology in education in order to prepare learners for successful participation in modern societies. The main aim of this paper is to theoretically specify competent use of AI and digital technology in education and to provide a standardized instrument to evaluate courses that teach this competence. Therefore, we (a) combine theoretical contributions from both scientific disciplines to formulate a theoretical framework with four levels for the “Competence to Use Artificial Intelligence and Digital Technology in Educational Processes” (AIEDTEC competence), (b) introduce a questionnaire to evaluate courses that teach AIEDTEC competence, and (c) present results regarding its psychometric properties (N=240). The questionnaire showed good to very good psychometric properties and the assumed factor structure was supported by confirmatory factor analyses. The paper connects research on systems thinking and learning and instruction with recent developments regarding AI and digital technology and thereby provides an essential base for creating effective, modern, and safe learning environments in the future as well as a psychometric evaluation instrument.
Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabling the generation of synthetic feedback for educational assessment. Despite these advances, LLM-based grading remains imperfect, making reliable confidence estimates essential for safe and effective human-AI collaboration in educational decision-making. In this work, we investigate confidence estimation for ASAG with LLMs by jointly considering model-based confidence signals and dataset-derived uncertainty. We systematically compare three model-based confidence estimation strategies, namely verbalizing, latent, and consistency-based confidence estimation, and show that model-based confidence alone is insufficient to reliably capture uncertainty in ASAG. To address this limitation, we propose a hybrid confidence framework that integrates model-based confidence signals with an explicit estimate of dataset-derived aleatoric uncertainty. Aleatoric uncertainty is operationalized by clustering semantically embedded student responses and quantifying within-cluster heterogeneity. Our results demonstrate that the proposed hybrid confidence measure yields more reliable confidence estimates and improves selective grading performance compared to single-source approaches. Overall, this work advances confidence-aware LLM-based grading for human-in-the-loop assessment, supporting more trustworthy AI-assisted educational assessment systems.
Recent research underscores the importance of inquiry learning for effective science education. Inquiry learning involves self-regulated learning (SRL), for example when students conduct investigations. Teachers face challenges in orchestrating and tracking student learning in such instruction; making it hard to adequately support students. Using AI methods such as machine learning (ML), the data that is generated when students interact in technology-enhanced classrooms can be used to track their learning and subsequently to inform teachers so that they can better support student learning. This study implemented digital workbooks in an inquiry-based physics unit, collecting cognitive, metacognitive, and affective data from 214 students. Using ML methods, an early warning system was developed to predict students’ learning outcomes. Explainable ML methods were used to unpack these predictions and analyses were conducted for potential biases. Results indicate that an integration of cognitive, metacognitive, and affective data can predict students’ productivity with an accuracy ranging from 60 to 100% as the unit progresses. Initially, affective and metacognitive variables dominate predictions, with cognitive variables becoming more significant later. Using only affective and metacognitive data, predictive accuracies ranged from 60 to 80% throughout. Bias was found to be highly dependent on the ML methods being used. The study highlights the potential of digital student workbooks to support SRL in inquiry-based science education, guiding future research and development to enhance instructional feedback and teacher insights into student engagement. Further, the study sheds new light on the data needed and the methodological challenges when using ML methods to investigate SRL processes in classrooms.
Intelligent Tutoring Systems (ITS) for psychomotor skills provide an accessible, scalable, and efficient solution, compared to tutors. Despite successes in the past, the research progress seems stagnating. Part of the reasons can be due to developing ITS for a very specific skill or task. In consequence, skills and the applications of those in their entirety are not presented. Therefore, we conducted a systematic literature review, and examined existing ITS using Harrow’s taxonomy, a skill continua framework, and dimensions of psychomotor skill learning. We observed a lack in consideration of offering different tasks to promote skill proficiency. Skills supported by ITS are majorly fine, closed, internally paced, discrete, individual, and simple skills. ITS focus majorly on technical, thus coordination aspects of motor abilities. Feedback and repetition are key methods to promote psychomotor skill learning. There is potential considering other physical activities to promote skill proficiency. Similarly, it might be worth exploring ITS for skills that, for example, are gross, and open. Integrating tasks that target motor abilities, such as strength or flexibility, can be part of it. The integration of theories in ITS from related research fields, such as training periodization, can be investigated.
BackgroundLearning analytics dashboards (LAD) have been developed as feedback tools to help students self-regulate their learning (SRL) by using the large amounts of data generated by online learning platforms. Despite extensive research on LAD design, there remains a gap in understanding how learners make sense of information visualised on LADs and how they self-reflect using these tools.ObjectivesWe address this gap through an experimental study where a LAD delivered personalised SRL feedback based on interactions and progress to a treatment group, and minimal feedback based on the average scores of the lecture to a control group.MethodsAfter receiving feedback, students were asked to write down how they planned to adjust their study habits. These reflection texts are the target of this study. Three human coders analysed 1251 self-reflection texts from 417 students at three different times, using a coding system that categorised learning strategies, metacognitive strategies and learning materials.Results and ConclusionsOur results show that learners who received personalised feedback intend to focus on different aspects of their learning in comparison to the learners who received minimal feedback and that the content of the LAD influences how students formulate their self-reflection texts. Furthermore, the extent to which students incorporated suggested behavioural changes into their reflections was predicted by state measures like perceived helpfulness of the feedback. Our findings outline areas where support is needed to improve learners' sense-making of feedback on LADs and self-reflection.
Science education aims to foster knowledge-in-use, which is supported by the integration of scientific ideas. To study knowledge integration effectively, network analysis provides a valuable tool for visualizing and understanding how ideas are connected. Successful knowledge integration requires following a learning progression that leads to increasingly sophisticated connections between ideas. However, traditional learning progression models have limitations, as they often fail to account for the nonlinear and individualized nature of learning. This study explores the potential of digital learning environments and AI techniques to address these limitations by enabling frequent, high-resolution data collection and analysis in order to uncover individual students’ learning trajectories at a high resolution. We analyze a case study of middle school students’ learning about energy to investigate patterns and variations in their learning trajectories. Additionally, we explore how different learning trajectories influence the development of knowledge-in-use, leading to either productive or unproductive learning outcomes. Our findings aim to guide instruction for teachers and instructional designers, providing insights on how to develop more effectively adaptive learning environments that support diverse student learning trajectories.
Formative assessment is a key element of effective education. Recently, a growing body of evidence has highlighted the potential of learning analytics to support formative assessment. However, this evidence is sparse, and concerns have been raised about the connection between learning analytics results and formative assessment models. When learning analytics results do not align well with formative assessment models, teachers may be hesitant to place trust in, interpret, and act upon the insights provided by learning analytics for formative assessment. In this study, we address this concern by adopting a well-established formative assessment model and conducting a systematic review to explore how learning analytics can support formative assessment. A total of 93 relevant articles, published between 2011 and 2023 and retrieved from Web of Science (WOS), Scopus, IEEE Xplore, and ACM, were analyzed using a coding scheme grounded in a formative assessment model. Drawing on the model of formative assessment, we report how learning analytics can support formative assessment by connecting the three key actors, teachers, students as peers, and students as self-assessors with the three critical stages of the process: (a) where the learner is going, (b) where the learner is now, and (c) how the learner is getting there. This study extends our understanding of the alignment between learning analytics results and theoretical concepts and further contributes to the advancement of the application of learning analytics for formative assessment.
IntroductionThe purpose of this exploratory study was to examine the potential effects of virtual reality (VR) mental training, based on cognitive-behavioral (CB) techniques, on race preparation among long-distance recreational runners within a sports coaching context. Although VR interventions have shown promises for enhancing athletic performance, their integration with CB-based imagery and self-talk remains limited.MethodsUsing a single-subject A-B-A design, six recreational runners completed two races: a first race without mental training and a second race after a series of VR mental training sessions conducted alongside their usual physical training. Each participant used a VR headset equipped with an application that delivered strategy guidance (including pacing and drafting) while also targeting motivation through CB-based imagery and self-talk. Navigation occurred entirely in a virtual environment, with no physical movement. Background audio featured participant-generated self-talk statements. Performance data from VR sessions were recorded through log-file, and emotional responses were assessed with the Emotional Stress Reaction Questionnaire (ESRQ). Race outcomes were compared using smartwatch metrics and participant self-assessments (e.g., Likert scales and open-ended responses).ResultsFollowing CB-based VR training, most participants reported using pacing and drafting strategies more frequently during the second race. Self-talk frequency increased, and post-race questionnaires indicated higher motivation ratings. Smartwatch data suggested moderately enhanced pacing consistency compared to baseline for some individuals.DiscussionThese exploratory findings suggest that CB-based VR mental training might help the adoption of certain race strategies and encourage self-talk use withing a coaching context. Results from the study may serve as preliminary reference points for future research aimed at integrating VR tools into complementary coaching approaches.
Quality feedback is essential for supporting student learning in higher education, yet personalized feedback at scale remains costly. Advances in learning analytics and artificial intelligence now enable the automated delivery of personalized feedback to many students simultaneously. At the same time, recent feedback research increasingly emphasizes learner-centered approaches, particularly the role of feedback literacy—students' varying capacities to engage with and benefit from feedback. Despite growing interest, few studies have quantified how feedback literacy affects students' perceptions of feedback, especially in technology-supported contexts. To address this, we examined (1) students' perceptions of personalized, detailed feedback generated via learning analytics and (2) how feedback literacy moderated these perceptions. In a randomized field experiment, teacher education students (N = 196) participated in a week-long computer-supported collaborative learning task on cognitive activation in the classroom. Both groups received automated, personalized feedback: the control group received basic feedback on task completion, while the experimental group received detailed feedback on group processes and the quality of their collaborative statement. The highly informative feedback significantly improved perceptions of feedback helpfulness, enhanced learning insights, and supported self-reflection and self-regulation. Feedback literacy partially moderated these effects, influencing perceptions of feedback helpfulness and motivational regulation.
This paper presents our contribution to the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-Powered Tutors. The objective of this shared task was to assess the quality of conversational feedback provided by LLMbased math tutors to students regarding four facets: whether the tutors 1) identified mistakes, 2) identified the mistake's location, 3) provided guidance, and whether they 4) provided actionable feedback. To leverage information across all four labels, we approached the problem with FLAN-T5 models, which we fit for this task using a multi-step pipeline involving regular fine-tuning as well as model merging using the DARE-TIES algorithm. We can demonstrate that our pipeline is beneficial to overall model performance compared to regular fine-tuning. With results on the test set ranging from 52.1 to 68.6 in F1 scores and 62.2% to 87.4% in accuracy, our best models placed 11th of 44 teams in Track 1, 8th of 31 teams in Track 2, 11th of 35 teams in Track 3, and 9th of 30 teams in Track 4. Notably, the classifiers' recall was relatively poor for underrepresented classes, indicating even greater potential for the employed methodology.
This paper examines the potential of digital games as communication tools to reach global audiences, extending beyond established cultural and geopolitical divides. It shows the empirical data gathered in our EU and UKRI-funded Games Realising Effective and Affective Transformation (GREAT) project, where we collaborated with several organizations to investigate this potential. Namely, a significant case study called Play2Act was undertaken in collaboration with the United Nations Development Programme (UNDP), which forms the focus of this paper. The aims of this study were to find out how much of the world’s population could be reached via digital games and how many citizens would be willing to communicate their climate attitudes in a simple and short survey that was inserted into popular mobile games. Currently, there are 3 billion gamers in the world and the idea of reaching citizens via games to understand their opinions on critical global issues and then passing this information to policy-makers emerged. This is the main objective of our project, as to whether games can act as an effective communication channel between citizens and policy-makers, the context being the climate emergency. Governments do not typically have the opportunity to understand their citizens’ needs fully. The aim of this project is to decrease the barrier and increase representation and democracy. The findings obtained from the Play2Act study suggest that games, moreover their ability to engage, and inherent social dynamics create a unique opportunity to support meaningful dialogue with a large proportion of citizens reached, engaged and completed the surveys. The study engaged with almost 1 million players from every UN recognised country, with only two exceptions, and ca. 181,000 surveys completed, confirming the global reach of games. The next steps are for UNDP to take this information to individual countries with recommendations of appropriate climate policies based on their citizens’ voices, this having huge potential for digital games being policy transformational tools. This research contributes to knowledge on the intersection of technology, culture, and communication and offers valuable insights for policymakers, researchers, and stakeholder groups seeking to leverage digital games for social impact.
[This paper is part of the Focused Collection in Artificial Intelligence Tools in Physics Teaching and Physics Education Research.] Students struggle to acquire the needed energy understanding to meaningfully participate in the energy discourse about socially relevant topics, such as energy transformation or climate change. Identifying students on differing learning trajectories, as well as differences in knowledge used, is essential to help students achieve the needed energy understanding. Collecting and analyzing the longitudinal and fine-grained data necessary for this represents a substantial challenge. However, the use of a digital workbook, which captures all interaction data, has enabled us to collect such data from N=548 students (data from 172 students were analyzed after applying exclusion criteria). Using machine learning and natural language processing, we analyzed the data to identify productive and unproductive learning trajectories and their underlying reasons. The learning trajectories were classified according to the post-test score. To analyze the tasks from the digital workbook, machine learning methods, specifically random forest, and natural language processing, were employed to identify how students on different learning trajectories progress through the unit. The random forest analysis was accurate in distinguishing between productive and unproductive learning trajectories. Furthermore, natural language processing was employed to analyze open-ended responses, which revealed disparities in the knowledge elements that students on productive and unproductive trajectories utilized. The findings of this study indicate that machine learning techniques have the potential to provide valuable insights into student learning trajectories, which can inform the design of instructional units and the feedback provided to teachers and students.