AI-powered teachable agents represent a promising approach to improving middle school students’ engagement with and understanding of mathematics, given the well-documented challenges students face at this developmental stage. During student-AI interactions, it is essential to determine when the agent should stop to effectively regulate the cognitive and emotional states of the students, which are factors closely linked to productive participation and tutoring strategy use. This raises a key question: who should decide when the teachable agent stops or continues? Empirical evidence highlights the advantages of using LLM-as-a-judge and knowledge-graph–based decision-making, as well as the potential benefits of fixed-turn conversations. Drawing on 64,060 messages from 7,991 conversations across four randomly assigned stopping mechanisms - 8-turns, 16-turns, standalone LLM-as-a-judge, and agent decisions with knowledge graphs (KGs), our experimental results highlight that (a) agent decisions informed by KGs were most effective at detecting when conversations should continue, sustaining learner engagement with the teachable agent; (b) off-topic utterances occurred more frequently under fixed-turn conditions; (c) tutoring strategies such as instructional guidance and elaborating with justification were most prevalent in condition of agents with KGs for decision-making; and (d) 16-turn stopping mechanism led to the smallest amount of learning gain and results from three other stopping mechanisms were comparable. Practical implications for designing effective agent-student interactions are discussed.
Creativity is a critical learning outcome in K-12 computer science, yet assessing it at scale remains challenging. Human-scored approaches, such as the Consensual Assessment Technique (CAT), are resource-intensive and prone to rater variability. Leveraging advances in large language models (LLMs), this study investigates whether GPT-4o can reliably assess creativity in student-generated code. We collected 383 flow-based music programs from 194 upper-elementary students (ages 10-12) between 2022 and 2024. Each artifact was rated by five human experts across four dimensions: originality, complexity, efficiency, and emotional expressiveness. We evaluated three prompting strategies: zero-shot, few-shot with theory-driven exemplars (ECD), and few-shot with human-selected examples. Among them, the ECD-based few-shot prompting yielded the best performance, achieving the lowest mean absolute error (MAE = 0.582) and highest agreement within +/- 1.0 of human scores (82.4%). Zero-shot prompting, while slightly less accurate, achieved the highest correlation with human scores (Spearman's p = 0.53), suggesting its potential for lightweight deployment.
Machine learning (ML) has become integral to online education, enhancing prediction, personalization, and automated assessment. However, algorithmic bias remains a critical barrier to equitable learning, as ML models can systematically under- or over-estimate outcomes for particular demographic groups. Existing fairness approaches—especially those focused on single attributes such as race or gender—fail to capture the complex, intersectional identities that shape students' experiences. Even multi-group fairness methods face key limitations in educational contexts, including computational scalability and difficulty adapting to shifting data distributions and fairness priorities. To address these challenges, this study proposes a reinforcement learning (RL)-based pre-processing framework that dynamically reweights data to optimize both predictive accuracy and multi-group fairness while safeguarding privacy. An AUC-based fairness metric ensures stability as subgroup combinations increase, and explainable AI (XAI) techniques enhance interpretability. Using large-scale data from Algebra Nation and state assessments, results show that the proposed framework achieves improved fairness and stable accuracy, offering a scalable, model-agnostic, and privacy-preserving pathway toward trustworthy AI in education.
Online discussion forums are a key data source in STEM education for analyzing student interaction, knowledge construction, and social learning. From a learning analytics perspective, understanding the temporal dynamics of instructional behaviours, for example, the differences between instructors and peers, can offer valuable insights into how instructional support unfolds in online discussions. However, few studies have examined these interactions with attention to their sequential and temporal organization. This study adopts a learning analytics approach to investigate the behavioural dynamics of instructional support in a large-scale online math discussion forum. We analyzed 83,569 posts from Math Nation, a widely used online learning platform supporting K–12 math education in the United States. Instructional behaviours were automatically coded based on a theoretically grounded coding scheme using transformer-based language models, and their temporal patterns were modelled using multilevel vector autoregression. Network visualizations were employed to illustrate the dynamic relationships among instructional strategies. Our findings show that peers posted more frequently than instructors and were more likely to engage in acknowledgement and feedback behaviours. When examining behaviour proportions and temporal transitions within discussion threads, peer interactions were associated with more diverse behavioural sequences, whereas instructor interactions exhibited more focused and directive patterns. In mixed-participant threads, instructors often responded following peer contributions, which was associated with a reduced likelihood of subsequent direct peer interventions. Rather than making claims about instructional effectiveness or learning outcomes, this work contributes to the field of learning analytics by demonstrating how natural language processing (NLP)-based behaviour modelling can be integrated with temporal interaction analysis to characterize role-based instructional support dynamics in asynchronous, discussion-based STEM learning environments.
As automotive systems increasingly incorporate electronics, software, and sensors, they become complex and widely distributed cyber-physical systems. This complexity makes them more susceptible to cyber-attacks. Nevertheless, despite its great need, the cybersecurity of automotive systems needs to be better understood, even by key stakeholders. This article presents an immersive virtual environment (IVE) platform that enhances the understanding of cybersecurity in automotive systems, focusing on ranging sensor attacks. By utilizing virtual reality (VR), the platform provides a hands-on experience for users to explore and comprehend cyber-attack implications. User studies were conducted to evaluate the effectiveness of the platform, revealing statistically significant improvements in participants’ knowledge, engagement, and self-efficacy related to automotive security. The findings underscore the potential of immersive learning tools in automotive security education, with IVE demonstrating a substantial impact on participants’ comprehension and interest in automotive security.
With the growing integration of artificial intelligence (AI) in education, conversational AI agents are increasingly used to support student learning. This study examines how interactions with AI teachable agents are temporally associated with students' agency and how these associations relate to students' learning outcomes. Analysing 7188 discussion threads containing over 117,000 text utterances, we explore the relationship between AI authority and student agency using classification and regression models. Findings reveal that AI authority is significantly associated with subsequent student agency levels; however, increased student agency does not lead to changes in AI authority. Sequential interaction analysis shows that students initially demonstrate higher agency in response to authoritative AI prompts, though this effect stabilizes over time. In addition, higher student agency is associated with more elaboration and clarification talk but also with increased off-task discussions, which slightly hinder learning gains. These findings underscore the need for balancing structured AI guidance with opportunities for student autonomy. This research contributes critical insights into designing AI-assisted learning environments that foster both engagement and effective learning outcomes.
Generative AI is shifting educational technology from tool support to intelligent collaboration, yet the evolving balance between AI instructional authority and student autonomy in human-AI learning remains under-specified. Static analyses miss the nonlinear trajectories through which interaction unfolds. We analyzed 5,013 valid dialogue sequences to characterize authority-agency dynamics and linked dialogue-based features from 794 students with complete pre- and post-test data to learning gains to examine their association. We obtain continuous, turn-level Authority and Agency scores via large language model regression and model their coupled evolution with vector autoregression, linking temporal dynamics to learning gains. Students exhibited consistently higher agency levels than the authority levels of the teachable agent, aligning with the learning-by-teaching premise. However, authority and agency typically showed task-focused bidirectional convergence rather than a one-way transfer of control. Vector autoregression (VAR) results indicate substantive bidirectional coupling (co-regulation index =0.74 , computed as the mean absolute sum of cross-lag parameters), consistent with students functioning as active partners who shape subsequent AI guidance rather than purely passive recipients of scaffolds. Controlling for prior knowledge and socioeconomic status, greater structural complexity in the AI authority trajectory predicted lower learning gains, whereas concentrated high-intensity instructional peaks at key moments positively predicted outcomes. Overall, effective teachable-agent support appears to be structurally parsimonious while delivering precisely timed reinforcement. These findings frame intelligent instruction as a dynamically negotiated co-regulation system rather than unilateral execution of a preset script, with implications for designing next-generation educational agents that preserve student agency while providing adaptive, well-timed scaffolding.
While prior research has examined student participation in online discussions in various ways, limited studies have investigated how students’ early participation patterns relate to their sustained participation, especially in the online mathematical learning context, where online mathematical discussions are an essential component of effective teaching. Leveraging a dataset comprising more than 80,000 students and over two million online discussion interactions, this study first examined students’ sustained participation in mathematical discussions by analyzing how newcomers to an online discussion board transitioned into long-term participants or disengaged over time. Then, building on the Communicative Ecology Theory (CET), which suggests that individuals’ sustained participation in an online community can be influenced by technical, social, and content-related factors, this study investigated how students’ early participation patterns on these three factors were related to their sustained participation in the discussion board. The findings revealed that students’ earlier participation patterns, especially social participation patterns, predicted their sustained participation in online mathematical discussions. This study contributes to the theoretical understanding of online educational discussions by demonstrating the successful application of CET in online educational communities. It offers practical implications for educators, emphasizing the importance of focusing on the sustainability of student participation in online discussions. Additionally, it provides insights for identifying and supporting students at risk of continued disengagement based on their current participation patterns.
This paper presents a mixed-methods study of a virtual, narrative-driven AI literacy module designed for Algebra 1 students. Grounded in the AI4K12 framework, the module integrated storytelling, real-world datasets, and mathematical modeling to support engagement and conceptual understanding of AI. Eighty middle and high school students from a virtual school participated in the pilot study. Through a sequence of activities, students build, test, and refine models for text classification tasks, applying verbal rules, feature extraction, and algebraic expressions. Quantitative pre- and post- results showed significant gains in AI self-efficacy and math motivation, along with modest improvements in AI conceptual understanding. Subgroup analyses by grade level, gender, race/ethnicity, and schooling context showed broadly consistent patterns of improvement, indicating the module’s relevance across diverse learners. Qualitative findings highlighted the role of interactive model-building, narrative framing, and real-world AI applications in supporting student engagement. Together, these results demonstrate the potential of integrating AI learning into core math instruction through accessible, hands-on, and learner-centered design.
This study examines how two LLM prompting paradigms support qualitative coding of students’ open-ended responses in an AI-integrated learning module on sentiment analysis. The dataset includes approximately 110 students’ responses to three model-revision tasks of increasing abstraction: (Q1) handling negative sentiment words, (Q2) handling multiple sentiment words, and (Q3) weighting mixed sentiment. A human-coded baseline was established through iterative codebook development and three rounds of double coding across two dimensions, Interpretation and Proficiency, achieving high reliability (round-level Cohen’s k ranging from .80 to 1.00). We explored (1) an inductive reverse-engineering pipeline, where the LLM inferred an executable codebook from labeled input-output pairs and was evaluated on held-out test sets, and (2) a deductive few-shot pipeline, where the LLM applied the finalized codebook with 1-, 3-, or 5-shot demonstrations under three context levels (simple, some, full), yielding 27 experimental conditions. Results show that task complexity is the dominant constraint: performance is high for Q1 and Q2 but degrades sharply for the more abstract Q3, especially for Proficiency. Increasing shots generally improves accuracy, yet gains are non-linear and depend on context; “Some” context often matches or exceeds full context. Error analysis reveals systematic failure modes, including proficiency overestimation driven by surface completeness, forced interpretation under ambiguity, keyword-triggered misclassification, weak differentiation among low-information labels, inconsistency blindness, and task-goal misalignment. Findings suggest LLMs can assist with lower-abstraction coding and triage, but complex interpretive judgments require human oversight and error-aware workflow design.
Generative AI-powered teachable agents offer new opportunities in mathematics education, yet systematic investigations of how vocal design affects middle school students' engagement and learning outcomes remain limited. We conducted a three-week 2 × 2 × 3 factorial randomized controlled experiment with 386 middle school students, systematically manipulating three vocal attributes of a generative AI-powered teachable agent: gender congruence (avatar-voice matching vs. mismatching), voice age (youth vs. adult), and emotional tone (sad, calm, cheerful). Engagement was assessed through behavioral coding of ten interaction categories and temporal analysis of dialogue sequences, while learning gains were measured through standardized state mathematics assessments. Analyses employed generalized estimating equations for factorial effects, time-series feature extraction for temporal dynamics, and machine learning models with SHAP-based feature importance for outcome prediction. Results revealed a theoretically significant dissociation between vocal configurations that activate immediate teaching behaviors and those that support long-term achievement. Empathy-evoking, identity-neutral configurations combining gender-mismatched and sad voices significantly enhanced immediate generative processing, including conceptual knowledge demonstration (β = 0.776), answer-giving behaviors, and explanatory responses. Conversely, schema-congruent, approachable configurations combining gender-matched and cheerful voices yielded the highest standardized learning gains (β = 0.997) and also significantly enhanced conceptual knowledge demonstration (β = 0.722), supporting sustained engagement over time. Time-series analysis confirmed that sad emotional tone maintained higher peak and overall levels of instructional behaviors across dialogue sequences. Feature importance analysis identified frustration management as the strongest predictor of standardized learning gains, while explanatory behaviors most strongly predicted conceptual knowledge demonstration. Voice age manipulations showed no significant effects. These findings demonstrate that strategically designed vocal attributes configuring multiple voice dimensions simultaneously activate learning-by-teaching processes more effectively than single-dimensional features. Practically, this means using sad tones with gender-mismatched voices to activate immediate teaching behaviors, and cheerful tones with gender-matched voices to sustain long-term achievement. Future research should explore adaptive designs that transition between these configurations based on learning phases.
Multimodal automated feedback systems are core components of next-generation K--12 mathematics learning interfaces, processing textual responses and handwritten images for large-scale personalized learning. When students upload assignment images in uncontrolled environments, visual privacy information such as faces and hands may be inadvertently captured, potentially correlating with contextual factors including the photo-taking environment, image quality, student effort, assignment type, and device access, thereby raising critical algorithmic fairness concerns in high-stakes educational assessment. This study extends group fairness from traditional demographic attributes to privacy contexts by treating privacy leakage status and type as group-defining characteristics. We propose a component-driven multimodal feedback model that simulates teachers' cognitive processes through four components: Mathematical Element Masking (MEM), Cross-Modal Consistency Verification (CMCV), Question--Answer Interaction (QAI), and Scoring Prototype Contrastive Learning (SPCL). The optimal configuration, MatCha with all components, achieves an MSE of 0.0906 and an \(R^2\) of 0.1686, representing a 257.2\% improvement in explanatory power over the baseline. This performance is highly competitive given the inherent subjectivity of open-ended mathematical scoring. Using a dual-perspective fairness framework that distinguishes model fairness (opportunity allocation) from error fairness (service quality), we identify a privacy leakage paradox: samples containing privacy leakage exhibit superior predictive accuracy yet systematically lower high-score opportunity rates. The overall disparate impact ratio for privacy leakage is 0.7565, falling below the legal threshold of 0.80, while the ratio for facial leakage is 0.8274, approaching the fairness boundary. No substantial gender bias is detected under facial leakage conditions.
Multiple-choice questions (MCQs) are central to instruction and assessment, with distractors revealing student understanding and misconceptions. However, creating high-quality distractors is time-consuming, especially for emerging domains like K–12 AI education. This study explores using generative AI to support distractor creation in a self-paced online module integrating AI and Algebra 1. Five MCQs were selected to compare distractors written by human developers and ChatGPT, using expert reviews and log data from 80 students. Experts rated human distractors higher overall, though AI ones consistently ranked second. Log analysis showed human distractors drew more initial selections, while students who chose AI distractors spent more time engaging without differences in hint use or revisits. Transition patterns across attempts suggest AI-generated distractors can effectively guide students toward correct answers, highlighting their potential for scalable MCQ design.
Generative artificial intelligence enables teachable agents to shift from instructional tools to collaborative learning teammates in K-12 mathematics education. However, prior research has emphasized cognitive outcomes while paying relatively less attention to emotional processes and systematic examination of voice characteristics that carry essential socio-emotional cues for human-AI collaboration. This study integrates control-value theory with social voice theory to construct an analytical framework linking voice characteristics, achievement emotions, engagement patterns, and learning gains. We conducted a three-week randomized experiment on Math Nation with 310 middle school students, collecting demographic data and standardized pre- and post-test scores, yielding 731 dialogue records. Manipulating voice emotion (sad, calm, happy) and using fine-tuned language models to annotate student achievement emotions and engagement patterns, we applied multilevel modeling to analyze group-differentiated associations. Results indicate significant group differences: happy tone was associated with enhanced positive activating emotions in males, while exploratory analyses suggest potential suppressive effects for low-SES students (though this finding requires cautious interpretation given the small subsample, N = 13). Achievement emotions were strongly associated with engagement patterns, with positive activating emotions linked to task-oriented engagement while negative deactivating emotions were associated with disengagement. At the student level, voice dialogue frequency and moderate negative activating emotions significantly predicted posttest scores after controlling for pretest performance. This study provides preliminary empirical evidence for designing inclusive adaptive learning environments where AI systems function as responsive teammates through socio-emotional voice design.
As artificial intelligence (AI) continues to transform education, Multi-Agent Systems (MAS) are emerging as a promising framework for delivering adaptive, personalized, and scalable learning experiences. While traditional MAS have demonstrated capabilities in instructional delivery and basic emotion detection, they remain limited in their ability to dynamically and accurately assess students’ real-time cognitive load and psychological states. Addressing these gaps, this study proposes a next-generation MAS architecture that integrates a Learning-by-Teaching large language model (LLM) agent, a Real-Time Cognitive Load Detection agent based on natural language processing (NLP), and a Questionnaire-Integrated Holistic Analysis agent. To develop and validate this system, we conducted an empirical study within Alter-Math, an AI-augmented mathematics learning platform designed to support dialogic, learning-by-teaching interactions in K-12 mathematics education, collaborating with local middle schools to recruit 155 students. During structured learning sessions, students engaged in teaching mathematical concepts to the LLM agent, generating rich linguistic interaction data, complemented by pre- and post-intervention measures of cognitive load, stress, confidence, and related psychological factors. These multimodal datasets informed the training of NLP-based classifiers for unobtrusive, real-time monitoring of learner states. Results suggest the potential of the proposed MAS framework to enhance the responsiveness and psychological awareness of AI-driven educational systems, paving the way for more personalized and cognitively supportive learning environments.
Despite extensive research on code plagiarism detection in higher education and for programming languages like Java and Python, limited work has focused on K-12 settings, particularly for pseudocode. This study aims to address this gap by building explainable machine learning models for pseudocode plagiarism detection in online programming education. To achieve this, we construct a comprehensive dataset comprising 7,838 pseudocode submissions from 2,578 high school students enrolled in an online programming foundations course, along with 6,300 pseudocode samples generated by three versions of generative pre-trained transformer (GPT) models. Utilizing this dataset, we develop an explainable model to detect AI-generated pseudocode across various assessments. The model not only identifies AI-generated content but also provides insights into its predictions at both the student and problem levels, thus enhancing our understanding of AI-generated pseudocode in K-12 education. Furthermore, we analyzed SHAP values and key features of the model to pinpoint student submissions that closely resemble AI-generated pseudocode. This research offers implications for developing robust educational technologies and methodologies to uphold academic integrity in online programming courses.