The integration of Generative AI (GenAI) into education has raised concerns about over-reliance and superficial learning, particularly in writing tasks in higher education. This study explores whether a theory-driven learning analytics dashboard (LAD) can enhance human-AI collaboration in the academic writing task by improving writing knowledge gains, fostering self-regulated learning (SRL) skills and shaping different human-AI dialogue characteristics. Grounded in Zimmerman's SRL framework, the LAD provided real-time feedback on learners' goal-setting, writing processes and reflection, while monitoring the quality of learner-AI interactions. A quasiexperiment was conducted involving 52 postgraduate students in a human-AI collaborative writing task. The students were divided into an experimental group (EG) that used the LAD and a control group (CG) that did not. Pre- and post- knowledge tests, questionnaires measuring SRL and cognitive load, and students' dialogue data with GenAI were collected and analyzed. Results showed that the EG achieved significantly higher writing knowledge gains and improved SRL skills, particularly in self-efficacy and cognitive strategies. However, the EG also reported increased test anxiety and cognitive load, possibly due to heightened metacognitive awareness. Epistemic Network Analysis revealed that the EG engaged in more reflective, evaluative interactions with GenAI, while the CG focused on more transactional and information-seeking exchanges. These findings contribute to the growing body of literature on the educational use of GenAI and highlight the importance of designing interventions that complement GenAI tools, ensuring that technology enhances rather than undermines the learning process.
As artificial intelligence (AI) transforms education, teachers’ participation in AI-focused professional development (AIPD) is essential but understudied. This mixed-method study explores teachers’ intentions to engage in AIPD. Survey data from 738 in-service teachers in China underwent latent profile analysis based on five indicators: AI attitudes, prior behaviors, performance expectancy, social influence, and personal growth. Three profiles emerged: Enthusiastic Endorsers (59.5%), Balanced Optimists (29.5%), and Cautious Skeptics (11.0%), with significant differences in AIPD intentions (p < .001). Enthusiastic Endorsers showed the highest intent. Thematic analysis of 543 open-ended responses highlighted the importance of practical, hands-on AIPD addressing AI ethics, supported by institutional policies. These findings provide insights into teacher readiness for AI and offer guidance for designing targeted AIPD to support effective and equitable AI integration in education.
Generative AI (GenAI)-supported automated scoring is increasingly applied in language assessment, yet evidence remains limited for classroom tasks targeting nuanced lexical-grammatical structures in vocational high school contexts. This study examines the reliability and validity of a GenAI-supported analytic scoring system for evaluating phrasal verb (PV) use in an English as a second language (L2) task at a vocational high school in Beijing. Student PV use in textbook-aligned sentence tasks was assessed across three analytic dimensions: accuracy, completeness, and naturalness. Three human experts scored 226 sentences with 10 PVs written by 25 students, with criterion scores defined as the mean of the three human scorers’ scores for each analytic dimension and the total score. Three large language models GPT-OSS, GLM-4, and Gemma-3 scored those sentences across two rounds using specified prompts. Aggregated human scoring was highly reliable, whereas single-scorer reliability was only moderate. Among the models, GPT-OSS demonstrated the strongest human-machine alignment and temporal stability, while Gemma-3 exhibited substantial test-retest variability. GenAI-supported scoring can approximate aggregated human judgments, but stability and alignment vary across models. From an argument-based perspective, the findings provide initial validity evidence for the evaluation and generalization inferences underlying GenAI-supported scoring of PV use in classroom-based sentence-construction tasks.
Mathematics education is commonly confronted with the contradiction between large-scale classroom teaching and personalized tutoring, while traditional photo-based question-searching software hinders students' in-depth thinking by directly providing answers. Therefore, this paper firstly proposed a new paradigm for intelligent tutoring termed the u201CDiagnostic Socratic Dialogueu201D model. Then, relying on this model, an intelligent mathematics tutoring system based on Socratic Dialogue was designed. Integrating large language model, OCR technology and Socratic Dialogue, the system replaced directly providing answers with dynamically questioning chains to guide students toward independent thinking, overcoming the limitations of traditional tools. Finally, the system was applied at a primary school. By analyzing system usage log data, student questionnaire data and semi-structured interview transcripts, this paper revealed that the system featured user-friendly operation and prompt feedback, effectively improving students' learning efficiency, mathematical comprehension and learning confidence, with its core educational functions well recognized. The designed intelligent mathematics tutoring system based on Socratic Dialogue in this paper had prominent application value in real teaching scenarios, which could provide feasible solutions to the dilemma of personalized mathematics tutoring and promote the deep integration of artificial intelligence and education, as well as facilitate the paradigm innovation of intelligent tutoring tools.
This study investigates the impact of learner autonomy affordance on students' learning performance and motivation within an ITS context, as well as the moderating effects of learner characteristics. Using traditional AI techniques to ensure experimental control, we developed an ITS teaching the Pythagorean Theorem to junior high school students in mainland China. Three system modes were constructed, each representing a different level of learner autonomy affordance. Through a double-blind randomized controlled trial, participants were randomly assigned to one of these three experimental conditions. We evaluated the impact of learner autonomy on learning performance and attrition and examined the potential moderating effects of learner prior knowledge and self-regulated learning abilities. Results indicated that learner autonomy affordance significantly influenced learner motivation, with this effect moderated by learners' metacognitive self-regulation. Furthermore, learners' intrinsic goal orientation and control of learning beliefs moderated the impact of autonomy affordance on learning performance.
With the rapid development of artificial intelligence, educational assessment is evolving from static testing to more adaptive and process-oriented evaluation. To address the limited interpretability and instructional relevance of existing automated assessment tools, this study designs an intelligent assessment and tutoring system that integrates Item Response Theory (IRT), Large Language Models (LLMs), and pedagogical theories such as the Zone of Proximal Development (ZPD). The system supports an adaptive learning cycle through continuous ability estimation, Chain-of-Thought (CoT)-based feedback generation, and dynamic learning path recommendations. A usability evaluation involving teachers and students yielded a System Usability Scale (SUS) score of 79.36, indicating a high level of usability and positive user experience. The results suggest that the proposed system can effectively support teachers in interpreting learner performance and providing timely, personalized feedback within authentic instructional settings, offering a practical pathway for AI-supported formative assessment in classroom practice.
As Generative AI (GenAI) becomes integral to education, fostering GenAI literacy is critical. However, current assessments largely rely on self-reported scales, lacking insights into how literacy manifests in actual learning processes. This study leverages Learning Analytics (LA) to bridge this gap. We collected interaction logs from 162 university students engaged in a GenAI-assisted abstract writing task. Using Epistemic Network Analysis (ENA), we modeled and compared the questioning strategies of students with varying GenAI literacy levels. Preliminary results reveal distinct interaction signatures: high-literacy students engage in iterative refinement and strategic questioning, while low-literacy students rely on direct generation commands. This work contributes to the workshop by demonstrating how process data can characterize GenAI literacy, paving the way for data-driven literacy assessment and real-time interventions.
Currently, STEM education often emphasizes performance on final exams, which can undermine the development of students' intrinsic motivation and long-term career interests. This focus on assessment results leads to a narrow approach that neglects the broader goals of engaging students in meaningful learning experiences. This study investigates the impact of an intelligent robotics-enabled STEM education project (IRESEP) on the interest levels of elementary school students in Beijing, involving 42 students (aged 10-12, from Grades 4 to 6) and a STEM teacher. The IRESEP was designed using effective teaching strategies such as phased teaching, blended learning, life-oriented teaching, and project-based learning. A mixed-methods approach was employed, combining quantitative and qualitative analyses to examine motivation mechanisms and influencing factors. Quantitative data were gathered through student questionnaires assessing basic information, STEM motivation performance, and influencing factors, analyzed using descriptive statistics, independent sample t-tests, cluster analysis, correlation analysis, and regression modeling. Qualitative insights were derived from interviews with both teachers and students. Results indicated that the IRESEP significantly enhanced students' cognition, interest, and abilities in STEM, while also increasing their willingness to pursue STEM careers. Key factors influencing students' STEM motivation included Teacher Support, Collaboration Skills, and Positive Academic Emotions. These findings suggest that intelligent robotics-based STEM projects should prioritize teacher support, collaborative learning, and positive educational experiences to effectively foster student motivation. The study offers a practical framework for designing engaging STEM programs that not only enhance student interest but also support their career aspirations in primary education.
AI-enabled personalization in STEM education is a transformative paradigm in pedagogical approaches, offering students a learning experience that is dynamically responsive to their nuanced requisites. Owing to the rapid development of AI-enabled personalized STEM education, there is a need for more systematic meta-analyses synthesizing its implementation in K-12 schools and examining its effects on improving educational outcomes. To examine the effect of AI technologies on personalized K-12 STEM education under different conditions, a meta-analysis was conducted by synthesizing 99 effect sizes extracted from 32 randomized controlled trial studies. The screening process followed the PRISMA flow diagram. In the coding process, the PICO tool was adapted as a coding sheet to extract pertinent details of each study. The analysis reveals a medium effect of AI technologies in personalized K-12 STEM education, indicating potential advantages over traditional non-AI methods. Six moderators were examined, with four (i.e., school levels, AI application types, personalized learning model types, and subjects) found to be significant in improving educational outcomes. Specifically, AI-enabled STEM education is more effective in junior and senior high school. Secondly, AR or VR tools show the most significant impact on students’ learning outcomes. Furthermore, learning models combined with classroom use exert more significant effects. Moreover, AI technology demonstrates the most important effects in comprehensive courses. Finally, observed items and GAI usage show no significant moderating effects. This paper advocates for an evidence-based and pedagogically aligned integration of AI-enabled personalized STEM education.
Understanding learners' preferences in educational settings is crucial for optimizing learning outcomes and experience. As artificial intelligence (AI) becomes increasingly integrated into educational contexts, it is crucial to understand learners' preferences between AI and human tutors to support their learning. While AI demonstrates growing potential in education, the phenomenon of algorithm aversion, which is a tendency to favour human decision making over algorithmic solutions, requires further investigation. To explore this issue, an experiment involving 114 university students was conducted to measure learners' preferences for different feedback sources before and after exposure to one of four conditions: no feedback, human tutor feedback, ChatGPT feedback through a free-dialogue user interface, and AI-powered writing analytics tool feedback through a structured interface. Our results revealed a strong initial preference for human tutors. However, the post-task analysis showed an important nuance. While the general preference for human tutors persisted, learners' preference towards the free-dialogue interface (ChatGPT 4.0) of ChatGPT increased, whereas the structured AI interface (AI-powered writing analytics tool) reinforced the preference for human tutors. These findings offer theoretical and practical contributions by extending algorithm aversion theory to educational contexts and demonstrating that appropriate interaction design can mitigate this aversion. The success of free-dialogue interfaces suggests that overcoming algorithm aversion may depend more on creating natural, flexible interaction experiences than purely technical optimization. However, we must also consider that increased preference for AI tools, particularly those with more engaging interfaces, may potentially lead to over-reliance and metacognitive laziness among learners, highlighting the importance of balancing technological support with the development of independent learning skills.Practitioner notes What is already known about this topic? Algorithm aversion exists across various contexts where individuals tend to prefer human over algorithmic decision-making. The introduction of generative AI brings new possibilities for AI-supported learning. What this paper adds? In academic writing tasks, learners show strong initial preference for human tutors over Generative AI feedback. Strong initial preference for human tutors persists even after exposure to generative AI feedback. Different interaction designs lead to divergent preference patterns: Free-dialogue interface increases preference for AI feedback, structured interface reinforces preference for human tutors. Implications for practice and/or policy Algorithm aversion in educational contexts can be mitigated through appropriate interaction design, particularly through natural dialogue interfaces. Design AI educational tools with back-and-forth, conversational interfaces to reduce algorithm aversion.
Help-seeking is an active learning strategy tied to self-regulated learning (SRL), where learners seek assistance when facing challenges. They may seek help from teachers, peers, intelligent tu- tor systems, and more recently, generative artificial intelligence (AI). However, there is limited empirical research on how learners’ help-seeking process differs between generative AI and hu- man experts. To address this, we conducted a lab experiment with 38 university students tasked with essay writing and revising. The students were randomly divided into two groups: one seek- ing help from ChatGPT (AI Group) and the other from an experienced teacher (HE Group). To examine their help-seeking processes, we used a combination of statistical testing and process mining methods, analyzing multimodal data (e.g., trace data, eye-tracking data, and conversa- tional data). Our results indicated that the AI Group exhibited a nonlinear help-seeking process, such as skipping evaluation, differing significantly from the linear model observed in the HE Group which also aligned with classic help-seeking theory. Detailed analysis revealed that the AI Group asked more operational questions, showing pragmatic help-seeking activities, whereas the HE Group was more proactive in evaluating and processing received feedback. We discussed factors such as social pressure, metacognitive off-loading, and over-reliance on AI in these different help-seeking scenarios. More importantly, this study offers innovative insights and evidence, based on multimodal data, to better understand and scaffold learners learning with generative AI.
AbstractBackgroundLanguage assessment plays a pivotal role in language education, serving as a bridge between students' understanding and educators' instructional approaches. Recently, advancements in Artificial Intelligence (AI) technologies have introduced transformative possibilities for automating and personalising language assessments.ObjectivesThis article aims to explore the design and implementation of AI‐enabled assessment tools in language education, filling the research gaps regarding the impact of assessment type, intervention duration, education level, and first language learner/second language learner (L1/L2) on the effectiveness of AI‐enabled assessment tools in enhancing students' language learning outcome.MethodsThis study conducted a systematic review and meta‐analysis to examine 25 empirical studies from January 2012 to March 2024 from six databases (including EBSCO, ProQuest, Scopus, Web of Science, ACM Digital Library and CNKI).ResultsThe predominant design in AI‐driven assessment tools is the structural AI architecture. These tools are most frequently deployed in classroom settings for upper primary students within a short duration. A subsequent meta‐analysis showed a medium overall effect size (Hedges's g = 0.390, p < 0.001) for the application of AI‐enabled assessment tools in enhancing students' language learning, underscoring their significant impact on language learning outcomes. This evidence robustly supports the practical utility of these tools in educational contexts.ConclusionsThe analysis of several moderator variables (i.e., assessment type, intervention duration, educational level and L1/L2 learners) and potential impacts on language learning performance indicates that AI‐enabled assessment could be more useful in language education with a proper implementation design. Future research could investigate diverse instructional designs for integrating AI‐based assessment tools in language education.
Historically, copyright regulations have posed challenges in the dissemination of educational resources, leading to financial barriers, limited access, and outdated materials. Open Educational Resources (OER) offer free access to materials, addressing these issues and promoting equitable learning globally. This study examines the integration of OERs among primary school educators in Zambian community schools. Grounded in Connectivism, Diffusion of Innovations Theory, and the OER Adoption Pyramid, this research explores how personal values, institutional support, technology, OER awareness, and cultural reception influence educators' decisions to adopt OER. Data were collected through in-depth interviews with 34 participants, two focus group discussions, document analysis, and social media data. The study was conducted in Lusaka and Chipata, focusing on participants actively involved in using OER provided by an NGO in collaboration with Zambia's Ministry of Education. The findings highlight the significant role of individual preferences, institutional frameworks, and infrastructure in OER adoption. It emphasises the importance of policy development, professional training, and capacity building to overcome barriers. While this study is limited in scope, it advocates for future research on nationwide OER adoption, policy impact, and sustainability to deepen understanding and foster inclusive OER integration.
In designing an intelligent tutoring system, a core area of the application of AI in education, tips from the system or virtual tutors are crucial in helping students solve difficult questions in disciplines like mathematics. Traditionally, the manual design of general tips by teachers is time-consuming and error-prone. Generative AI, like ChatGPT, presents a new channel for designing general tips. This study utilized prompt engineering and Chain of Thought to summarize general tips for given mathematical problems (one geometry problem and one algebra problem) and their solutions. A Turing test was conducted to compare ChatGPT-generated general tips with human-designed ones. Results from 121 human evaluators, each assessing 6 ChatGPT-generated and 6 human-designed general tips for each of two mathematical problems, showed that the average score for ChatGPT-generated tips is less than that of human-designed tips at a statistically significant level (p < 0.05), and Zero-Shot CoT achieved the best score. However, no evaluator could distinguish the tip types exactly. The average precision, recall and F-value of all ChatGPT-generated tips are less than 40%. AI-generated general tips can serve as a valuable reference for teachers to enhance efficiency and students' mathematical learning.
Role-play activities are considered a useful instructional design in enhancing the speaking performance of foreign language learners. However, in the traditional classroom context, learners may not readily have access to interlocutors for role-play activities. In this study, we proposed a designed method that integrated the GenAI agent into role-play activities and conducted a quasi-experiment at a Chinese university. A total of 53 Chinese students in an English course were assigned to an experimental group (n = 29) that engaged in role-playing with a GenAI agent and a control group (n = 24) that engaged in role-playing with their class peers. The results showed role-play activities with GenAI could facilitate students' speaking performance, and one-way ANCOVA showed there was no significant difference in speaking performance improvement when compared with traditional role-play activities with peers. Moreover, the paired-sample Wilcoxon test showed that the treatment students had significant intrinsic motivation and self-efficacy improvement, which were higher than the control group. The findings provide important implications for the design and implementation of GenAI-integrated role-play activities for foreign language speaking learning.
The computerised adaptive test (CAT) may enable individualised and adaptive homework and assessment which seem difficult to be practised in traditional school settings. The authors developed a web-based mathematics intelligent assessment and tutoring system (MIATS) assessing the learner's math content mastery status by a CAT. This research uses a simulation experiment to simulate the performance of 53 middle school students who used all questions included in the MIATS with their authentic answers in the CAT. The data analysis shows that the trait values measured by a CAT have significantly strong and positive coefficients with the correctness rates measured by a classical test with the same questions from the CAT. The data visualisation of questions used in the CAT and the question sequences demonstrate that all the examinees complete the CAT with differentiated questions and various pathways. This research innovatively provides convincing evidence for the effectiveness and equality of the CAT.
Intelligent tutoring systems (ITSs) aim to deliver personalized learning support to each learner, aligning with the educational aspiration of many countries, including China. ITSs' personalized support is mainly achieved by providing individual prompts to learners when they encounter difficulties in problem-solving. The guiding principles and methods of prompts have been less investigated in previous ITS literatures. Based on relevant learning theories, such as self-regulated learning theory, zone of proximal development, scaffolding and heuristic teaching, we proposed seven guiding principles for designing ITS prompts and designed the guiding and adaptive prompts for the difficult questions in a mathematical ITS, math intelligent assessment and teaching system V2.0. In order to verify the effectiveness of this ITS with the aforementioned prompts, we conducted a 2 x 2 quasi-experiment in a high school, where the experimental group followed a process of "pretest, practice with general prompts and adaptive tutoring, and posttest," while the control group followed a process of "pretest, practice with only general prompts, and posttest." We collected the pre and posttest scores of both the experimental and control groups, and log data from the student model within the ITS for the experimental group students. The data analysis indicated that although the experimental group scored lower than the control group in the pretest, they scored higher in the posttest and spent less completion time. The drilled problems and the prompts provided to the experimental group students were personalized. In conclusion, the design principles for guiding and adaptive prompts in the ITS can provide personalized guidance and support for students, thus effectively improve their performance. Those principles are not only valuable for the subject mathematics but also can contribute significantly to the prompt design of other subjects, thereby bolstering the global pursuit of personalized education.
Currently, data-driven adaptive learning technology has shown tremendous potential in the field of education. However, its opaque u201Cblack boxu201D nature has raised widespread concerns among educational researchers and practitioners. Explainable artificial intelligence (XAI) is believed to have the potential to help learners understand intervention decisions in adaptive learning contexts, thereby enhancing learning outcomes, but there is controversy over its practical effects in educational applications. Therefore, the paper employed meta-analysis to analyze 66 effect sizes from 29 empirical studies. It was found that interpretable AI improved the learning effect of adaptive learning to a moderate degree, with a greater impact on learnersu2019 cognitive and metacognitive dimensions. The facilitation effect of XAI varied due to differences in explanation design, presentation design, and experimental design. Based on research results, the paper proposed that future adaptive learning interventions should remain learner-centered, emphasize the interactivity, readability, and boundaries of learning intervention explanations, so as to further promote the in-depth implementation of XAI in education.
This research integrates online learning supported by students' smart phones and a tutoring system CSIEC into a university course of English as a foreign language, through one semester with 548 students and eight lecturers from 13 classes. The treatment students used the system to learn vocabulary during the specified periods through voluntary participation and natural grouping, while the control students used alternative methods. The online learning activities analysis shows significant difference between the online learning behaviour of the treatment and control group, and demonstrates that the treatment group improved the vocabulary grade with a large effect size and decreased the time spent on completing the quiz. The treatment group improved better than the control group in overall learning, especially in vocabulary mastery and writing examined in regular university tests. The anonymous online survey results reveal that the students were satisfied with the English learning. The reasons leading to the findings are discussed. The implication for MALL research and future work are suggested.