The growing use of artificial intelligence (AI) in education, particularly large language models (LLMs), has increased interest in intelligent tutoring systems. However, LLMs often show limited adaptivity and struggle to model learners' evolving knowledge over time, highlighting the need for dedicated learner modelling approaches. Although deep knowledge tracing methods achieve strong predictive performance, their opacity and susceptibility to bias can limit alignment with pedagogical principles. To address this, we propose Responsible-DKT, a neural-symbolic deep knowledge tracing approach that integrates symbolic educational knowledge (e.g., mastery and non-mastery rules) into sequential neural models for responsible learner modelling. Experiments on a real-world dataset of students' math interactions show that Responsible-DKT outperforms both a neural-symbolic baseline and a fully data-driven PyTorch DKT model across training settings. The model achieves over 0.80 AUC with only 10
Game-based learning environments (GBLEs) incorporate stealth assessments, unobtrusively capturing learners' evolving knowledge and competencies. However, literature have not focused on how these assessments may impact learners' emotional experiences during gameplay. As such, this paper captured undergraduate students' (N = 26) physiological arousal as they played Crystal Island, a microbiology GBLE, to understand how physiology changes over time as a reaction to stealth assessments embedded in the environment. Results revealed that learners experienced greater physiological arousal while completing stealth assessments and that, as time progresses, learners experienced a steep decrease in physiological arousal after they concluded their task. This indicates that embedded stealth assessments may be more physiologically arousing than intended. Overall, while stealth assessments can be a highly effective tool for promoting deeper engagement, they must be designed with emotional regulation in mind.
Background: Immersive virtual reality (IVR) research must (1) consider learning as both an outcome but also a dynamic process and (2) ground IVR design and learning assessment within an empirically tested theoretical model, i.e., the Cognitive Affective Model of Immersive Learning (CAMIL). This model posits that IVR learning may be hindered or supported by several cognitive and affective variables (e.g., agency, cognitive load, situational interest). Aims: We examine the learning processes and outcomes (i.e., declarative knowledge, procedural knowledge, and procedural skills) of Chemistry lab skills through a commercially available IVR game, HoloLab Champions, situated in the CAMIL model. Sample(s): Participants were 46 high school students. Methods: Participants were randomly distributed between two conditions (IVR versus non-IVR video) to complete five mini-labs for approximately 30 min. Results: Participants in the IVR condition had higher self-perceptions of agency and situational interest compared to the non-IVR condition. Several regression models of learning outcomes showed agency was a significant predictor for all outcomes, while situational interest was only a significant predictor for declarative and procedural knowledge. Finally, when examining learners' performance on a practical assessment of Chemistry skills, we found that those in the IVR condition were more accurate than those in the non-IVR condition. Conclusions: Our results empirically support certain assumptions of the CAMIL model, propose theoretical expansions of the model, and highlight some of the limitations of IVR-based learning. We discuss implications for designing intelligent, adaptive scaffolds embedded in IVR, including using multimodal process-based data during learning for real-time feedback.
Self-regulation (SR) has emerged as a pivotal construct in educational research, yet its diverse theoretical accounts often lack comprehensive integration. This paper addresses this by proposing an integrative conceptual framework that consolidates four core research traditions on SR in educational settings, each emphasizing distinct underlying mechanisms: (1) Stable Dispositions (e.g., personality traits; Matthews et al., 2009; Roberts et al., 2007; Song et al., 2020), (2) Limited Resources (e.g., working memory capacity and executive functioning; Friedman Miyake, 2017; Paas et al., 2003; Sweller, 1994), (3) Driving Forces (e.g., motivation, interest, and affect; Eccles Wigfield, 2002; Hidi Renninger, 2006; Trautwein et al., 2019), and (4) Learning Activities (e.g., self-regulated learning, cognitive and metacognitive strategies; Winne Hadwin, 1998; 2008; Zimmerman, 2000). By synthesizing these perspectives, our framework positions SR as a dynamic, interdependent system, emphasizing how interactions and compensatory mechanisms explain individual differences and influence learning outcomes. We further argue for an integrative methodological toolkit, leveraging advanced machine learning and computational models, to capture the multi-level, temporal, and social dynamics inherent in SR. This holistic approach is crucial for advancing a comprehensive and ecologically valid understanding of SR in educational contexts.
Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive judgments? Existing evaluations emphasize problem-solving accuracy, overlooking the fragmented and imperfect reasoning that characterizes human learning. We evaluate LLMs as novices using 630 think-aloud utterances from multi-step chemistry tutoring problems with problem-solving logs of student hint use, attempts, and problem context. We compare LLM-generated reasoning to human learner utterances under minimal and extended contextual prompting, and assess the models' ability to predict step-level learner success. Although GPT-4.1 generates fluent and contextually appropriate continuations, its reasoning is systematically over-coherent, verbose, and less variable than human think-alouds. These effects intensify with a richer problem-solving context during prompting. Learner performance was consistently overestimated. These findings highlight epistemic limitations of simulating learning with LLMs. We attribute these limitations to LLM training data, including expert-like solutions devoid of expressions of affect and working memory constraints during problem solving. Our evaluation framework can guide future design of adaptive systems that more faithfully support novice learning and self-regulation using generative artificial intelligence.
As artificial intelligence becomes increasingly embedded in educational contexts, the ability of AI systems to perceive, interpret, and respond to learners’ affect has shifted from a niche research interest to a central necessity. This workshop introduces Multimodal Affect in AI for Education (MAAI4Ed), bringing advances in multimodal analytics, affective computing, and learning theories. The workshop focuses on (1) AI-driven approaches in detecting and interpreting affective states through multimodal data streams and (2) designing affect-aware AI systems to support emotional regulation and human well-being. By fostering dialogue among researchers, designers, and practitioners, the workshop aims to advance ethically responsible affect-aware AI for Education.
Advances in learning technologies now enable continuous real-time analysis of learner process data. While these innovations have generated interest in adaptive educational technologies, most research mistakenly assumes that the processes and indicators involved are largely domain-general. This overlooks how domain-specific cognitive structures and practices crucially shape learning and the meaning of process data. This paper synthesizes current theory and research regarding real-time process data and adaptive support, emphasizing the urgent need for domain-sensitive approaches. I argue that each domain sets unique learning goals, expertise trajectories, and measurement constraints, fundamentally affecting what adaptivity can achieve. I identify key conceptual, methodological, and analytical challenges. Drawing on Strohmaier et al. (2025), the lead article, and other contributions in this special issue of Learning and Instruction, I highlight opportunities and challenges in current approaches. I also outline directions for developing interpretable, instructionally aligned adaptive educational technologies.
Narrative game-based learning environments (GBLEs) offer learners an exploratory platform for engaging in self-regulated learning (SRL) processes to deepen their understanding of complex STEM topics. However, the affordances of narrative-based GBLEs can allow learners to avoid intended learning materials and processes through (mal)adaptive behaviors, such as gaming the system. The current study examines high school students’ (N = 204) gameplay with Crystal Island, a narrative-based GBLE, and identifies how the proportional time spent gaming the system impacts their learning outcomes. We identified two gaming behaviors that were highly correlated and used principal component analysis (PCA) to combine the behaviors into one principal component (PC) of gaming behaviors. Regression analysis identified a significant negative relationship between PC gaming behaviors and students’ learning gains, indicating that the more time a participant spent gaming the system, the less microbiology content they learned. We discuss future directions for the development of frameworks to distinguish adaptive versus maladaptive behaviors using granular trace data and multimodal data, such as concurrent think-alouds, to gain insight into students’ motivations and intentions behind their behaviors.
In STEM education, game-based learning environments (GBLEs) have become prominent platforms to scaffold students’ self-regulated learning (SRL) strategies. This study challenges the assumption that all engagement in GBLEs is equally beneficial by distinguishing cognitive from behavioral engagement within the Integrative Model of Multidimensional SRL Engagement (IMMSE). We analyzed data from 227 high-school students playing Crystal Island, a narrative-centered microbiology GBLE, to examine how these engagement types during the forethought and performance SRL phases predict learning gains and problem-solving accuracy. Overall, results showed that cognitive engagement significantly predicts learning gains, whereas behavioral engagement does not. This result supports IMMSE’s distinction between surface-level actions and deeper processing, indicating that meaningful learning depends on strategic cognitive engagement. We discuss implications for designing adaptive scaffolds that dynamically balance engagement dimensions across early SRL phases to optimize both efficiency and effectiveness.
The rapid rise of large language model (LLM)-based tutors in K–12 education has led to the misconception that generative models can replace traditional learner modelling and act as general-purpose engines for adaptive instruction. This is especially problematic in K–12 settings, which the EU AI Act classifies as a high-risk domain requiring responsible design. Motivated by concerns surrounding the role of learner modelling in responsible AI-powered tutoring, this study synthesises existing research evidence on key limitations of LLM-based tutoring systems and then presents an empirical case study investigating one critical aspect of these concerns: the accuracy, reliability, and temporal coherence of assessing learners’ evolving knowledge over time. To this end, we compare a deep knowledge tracing (DKT) model with a widely used LLM (with and without fine-tuning) that has demonstrated competitive performance in tutoring-related tasks, using a large-scale open-access dataset. Our findings show that DKT achieves the highest discrimination performance (AUC = 0.83) on next-step correctness prediction and consistently outperforms the LLM across evaluation settings. Although fine-tuning improves the LLM’s AUC by about 8% over the zero-shot baseline, it still remains 6% below DKT and produces higher early-sequence errors, precisely where incorrect predictions would be most harmful for adaptive learner support. Temporal-coherence analyses further reveal that while DKT maintains stable, directionally correct mastery updates, LLM variants display substantial temporal weaknesses, including smooth but wrong-direction updates, and, even after fine-tuning, remain inconsistent and unable to match DKT’s temporal stability. We also illustrate that these issues persist even though fine-tuned LLMs required nearly 198 hours of continuous high-compute training, far exceeding the computational demands of the lightweight DKT model. Our qualitative analysis of multi-skill mastery estimation further shows that, even after fine-tuning, the LLM produced unstable and inconsistent mastery trajectories, whereas DKT maintained smooth and coherent multi-skill updates. Collectively, these findings suggest that LLMs alone are unlikely to achieve the same positive effect sizes observed in decades long intelligent tutoring systems literature. Rather than replacing learner modelling, LLMs may be more appropriately deployed as pedagogical interfaces or content generators paired with dedicated learner modelling components to ensure responsible, accurate, reliable, and pedagogically sound support.
Narrative-driven game-based learning environments (GBLEs) are characterized by their capacity to offer learners autonomy in navigating immersive virtual spaces, facilitating self-directed or independent exploration, and interactive learning experiences with complex content. Such narrative-focused GBLEs could thus provide dynamic interaction with educational content in ways traditional instructional approaches cannot replicate at scale. Prior research shows that restricted navigational movement and in-game interactions in GBLEs have improved learning outcomes. Additionally, prior research done on immersive environments has shown that in-game affordances promote learners’ sense of presence, a fundamental psychological experience of immersion that can lead to significant learning outcomes. However, there is limited empirical understanding of how learners’ initial navigation strategy choices unfold in unrestricted GBLEs. It also remains unclear how these approaches differ based on individual differences such as prior gaming experience and perceived presence. This study identifies two ‘navigational strategies’ used by high school learners (N = 152) during gameplay with a narrative-driven GBLE, Crystal Island, and examines how their prior video game experiences and self-reported sense of presence influenced their learning outcomes. Our results showed that learners with higher video game experience engaged significantly more in independent exploration than those with less video game experience. In contrast, we found that learners who followed pedagogical scaffolding and structured guidance to initiate navigation reported a greater sense of presence during gameplay than those who independently explored the environment. Collectively, these results highlight the importance of diverse individual differences in informing the design of adaptive scaffolding and navigation support within GBLE.
Game-Based Learning (GBL) is a learner-engaging pedagogical methodology, yet adapting games to heterogeneous learners requires transparent, real-time Open Player Models (OPMs). We contribute to the community Open Player Socially Analytical Intelligence (OPSAI), an architecture implementing OPM beyond conceptual frameworks and validated in a GBL application. It decouples gameplay telemetry and analysis from the game engine and automatically derives pedagogically actionable insights, supporting the transparency of computational player models while making them accessible to players. OPSAI comprises three logical layers: a Frontend that both provides the GBL experience and collects information needed for analytics; a stateless Backend that hosts transparent analytics services producing reflective prompts, recommendations, and visualization guides; and a two-tier Log Storage that balances heavy raw gameplay data with lightweight reference indices for low-latency queries. By feeding analytics outputs back into the game interface, OPSAI closes the feedback loop between play and learning, empowering teachers, researchers, and learners alike. We further showcase OPSAI with a full deployment on the Parallel GBL environment, featuring live play traces, peer comparisons, and personalized suggestions, demonstrating a reusable blueprint for future educational games.
The rapid rise of large language model (LLM)-based tutors in K-12 education has led to the misconception that generative models can replace traditional learner modelling and act as general-purpose engines for adaptive instruction. This is especially problematic in K-12 settings, which the EU AI Act classifies as a high-risk domain requiring responsible design. Motivated by concerns surrounding the role of learner modelling in responsible AI-powered tutoring, this study synthesises existing research evidence on key limitations of LLM-based tutoring systems and then presents an empirical case study investigating one critical aspect of these concerns: the accuracy, reliability, and temporal coherence of assessing learners' evolving knowledge over time. To this end, we compare a deep knowledge tracing (DKT) model with a widely used LLM (with and without fine-tuning) that has demonstrated competitive performance in tutoring-related tasks, using a large-scale open-access dataset. Our findings show that DKT achieves the highest discrimination performance (AUC = 0.83) on next-step correctness prediction and consistently outperforms the LLM across evaluation settings. Although fine-tuning improves the LLM's AUC by about 8% over the zero-shot baseline, it still remains 6% below DKT and produces higher early-sequence errors, precisely where incorrect predictions would be most harmful for adaptive learner support. Temporal-coherence analyses further reveal that while DKT maintains stable, directionally correct mastery updates, LLM variants display substantial temporal weaknesses, including smooth but wrong-direction updates, and, even after fine-tuning, remain inconsistent and unable to match DKT's temporal stability. We also illustrate that these issues persist even though fine-tuned LLMs required nearly 198 hours of continuous high-compute training, far exceeding the computational demands of the lightweight DKT model. Our qualitative analysis of multi-skill mastery estimation further shows that, even after fine-tuning, the LLM produced unstable and inconsistent mastery trajectories, whereas DKT maintained smooth and coherent multi-skill updates. Collectively, these findings suggest that LLMs alone are unlikely to achieve the same positive effect sizes observed in decades long intelligent tutoring systems literature. Rather than replacing learner modelling, LLMs may be more appropriately deployed as pedagogical interfaces or content generators paired with dedicated learner modelling components to ensure responsible, accurate, reliable, and pedagogically sound support.
Pediatric emergencies in general and community hospitals, where most children are seen, are infrequent and cognitively demanding, and deviation from resuscitation algorithms is common despite certification. This review treats that as a failure of execution under load rather than a knowledge deficit, and reads together two separately developed literatures: head-mounted augmented reality (AR) guidance, which externalizes the algorithm, and real-time physiological sensing, which estimates the clinician’s cognitive state. Controlled trials of head-mounted guidance report improved guideline adherence and fewer dosing errors in simulation, but effects are inconsistent across outcomes, samples are small, endpoints are process measures, and some interventions slow performance or raise workload. Cognitive-state classification is accurate within individuals but generalizes poorly across them, and the signals index arousal and effort rather than load. The adaptive combination has been proposed but never tested. Adaptive guidance needs controlled comparison against static guidance with cognitive load and performance as co-primary outcomes, models that generalize across clinicians, and patient-level rather than simulation-only endpoints.
High school students consistently struggle with algebra problem-solving, not solely because of computational deficits but primarily because of inadequate metacognitive and self-regulatory strategies. Current intelligent tutoring systems, while demonstrating effectiveness in procedural skill development, frequently fail to adequately support the development of such higher-order thinking processes. This paper examines the cognitive and metacognitive challenges inherent in algebra word-problem solving and presents an initial theoretical and empirically driven design framework for simulated learners as pedagogical agents that model, teach, learn, externalize, and explain metacognitive processes embedded in MetaSim, a new interactive learning environment (ILE) for algebra problem-solving. By making invisible metacognitive processes visible through simulated peer learners who demonstrate both productive and unproductive strategies, students can develop more sophisticated approaches to complex mathematical problem-solving. This paper describes thirteen specific design principles grounded in theories of metacognition and self-regulated learning to guide the design of simulated learners in ILEs.
Game-based learning environments (GBLEs) are educational platforms created to promote and scaffold learners’ self-regulated learning (SRL) strategies while learning complex STEM topics. However, learners can still abuse the intended purposes of GBLEs and engage in ‘gaming the system’, deliberate behaviors employed to forgo the learning process while still completing the task at hand. ‘Gaming the system’ is often investigated within intelligent tutoring systems (ITSs), yet these behaviors have been the subject of limited investigation in narrative-based GBLEs. Thus, this study identified two ‘gaming the system’ behaviors from undergraduate students’ (N = 93) gameplay during the narrative-based GBLE, Crystal Island, and examined how their prior knowledge of microbiology and two agency conditions influenced their frequencies of gaming behaviors and their learning outcomes. Results indicated that students with high prior knowledge engaged in significantly more ‘trial-and-error’ gaming behaviors than those with lower prior knowledge. Moreover, we found that students with partial agency engaged in significantly more “trial-and-error’ gaming behaviors than those will full agency; whereas those will full agency engaged in more ‘guessing’ gaming behaviors than those with full agency. These results indicate the importance for identifying contextually driven gaming behaviors, as well as future GBLE scaffolds to adopt real-time, adaptive scaffolding methods based on learner’s active gameplay and SRL (in)efficiency.
Given the demand for responsible and trustworthy AI for education, this study evaluates symbolic, sub-symbolic, and neural-symbolic AI (NSAI) in terms of generalizability and interpretability. Our extensive experiments on balanced and imbalanced self-regulated learning datasets of Estonian primary school students predicting 7th-grade mathematics national test performance showed that symbolic and sub-symbolic methods performed well on balanced data but struggled to identify low performers in imbalanced datasets. Interestingly, symbolic and sub-symbolic methods emphasized different factors in their decision-making: symbolic approaches primarily relied on cognitive and motivational factors, while sub-symbolic methods focused more on cognitive aspects, learned knowledge, and the demographic variable of gender – yet both largely overlooked metacognitive factors. The NSAI method, on the other hand, showed advantages by: (i) being more generalizable across both classes – even in imbalanced datasets – as its symbolic knowledge component compensated for the underrepresented class; and (ii) relying on a more integrated set of factors in its decision-making, including motivation, (meta)cognition, and learned knowledge, thus offering a comprehensive and theoretically grounded interpretability framework. These contrasting findings highlight the need for a holistic comparison of AI methods before drawing conclusions based solely on predictive performance. They also underscore the potential of hybrid, human-centered NSAI methods to address the limitations of other AI families and move us closer to responsible AI for education. Specifically, by enabling stakeholders to contribute to AI design, NSAI aligns learned patterns with theoretical constructs, incorporates factors like motivation and metacognition, and strengthens the trustworthiness and responsibility of educational data mining.
This commentary highlights three problems that can emerge by integrating Digital Twin Technology (DTT) and Human-Machine Systems (HMS), drawing insights from Human-Technology Interaction, Systems Engineering and Computer Science, and Learning Sciences experts, who participated in the IEEE SMC Society/SMST Workshop on HMS-DTT, hosted at the University of Central Florida. The paper focuses on ethics, human and data interoperability, and trust issues. Rather than providing a traditional literature review, it consolidates contributions from workshop discussions and highlights the need for transparent, reliable systems, standardized data protocols, and ethical frameworks to guide development and implementation. Synthesizing diverse perspectives underscores the importance of interdisciplinary approaches in realizing the benefits of HMS and DTT integration while mitigating potential risks. Overall, this work aims to inform future research agendas and foster responsible innovation by integrating viewpoints across disciplines in this rapidly evolving field.
Artificial intelligence (AI) has demonstrated significant potential in enhancing digital learning by offering personalized and adaptive experiences that meet learners' individual needs. However, while self-regulated learning (SRL) skills are critical for success in digital environments, AI-driven learner models mainly focus on cognitive processes, with limited integration of SRL skills. This systematic review synthesizes research from 1990 to 2024, analyzing digital trace data from various learning platforms to identify which data serve as indicators of the three phases and areas of SRL in K-12 digital learning. Our findings highlight digital trace data that measure (meta)emotion, (meta)motivation, and (meta)cognition across the three SRL phases, while also revealing significant gaps in tracing (meta)motivation and (meta)emotion, particularly in the preparatory and appraisal phases. Although a variety of data traces address the (meta)cognitive area, the results underscore the challenges of meaningfully interpreting the learning process. Despite these challenges, the evolving research on trace data demonstrates substantial potential for integrating adaptive SRL scaffolding into digital learning environments.
Student modeling has been widely investigated to enable adaptive support in game-based learning environments. Two key aspects of student modelling in game-based learning are stealth assessment and goal recognition. Stealth assessment infers students' knowledge and skills without disrupting gameplay, while goal recognition predicts their in-game objectives based on their interactions. Prior work has largely treated these tasks separately, yet students' learning processes, outcomes, and their in-game goals often influence each other. This paper presents a multi-task student modeling framework that enhances stealth assessment and goal recognition by jointly predicting students' learning outcomes and in-game goals. The framework integrates game trace logs and students' written reflections as input features, enabling shared learning across related tasks. We investigate stealth assessment at two levels: post-test score prediction for generalizability and concept-level predictions for more fine-grained insights where domain expertise is available. Empirical evaluations suggest that multi-task learning significantly improves predictive performance for concept-level stealth assessment and goal recognition compared to single-task baselines. These findings highlight the potential of multi-task learning to enhance student modeling that aligns with students' learning trajectories and goals in game-based learning environments.