
Providing feedback to move learners forward (i.e., feed-forward feedback) is one of formative assessment's key strategies (Black & Wiliam, 2009; Hattie & Timperley, 2007). Yet, within the context of science experiments, teachers lack adequate information on students' experimental conduct to be able to provide feed-forward feedback that is timely, structured and individualized. In this regard, we can leverage computer vision to detect student actions during science experiments for feedback, thereby creating an additional source of information to support teachers in implementing formative assessment. This study aims to investigate the construction of action recognition models using data collected from 140 ninth grade students as they conduct a science experiment in school laboratories. In two separate model investigations, we 1) compared model performance of various action recognition architectures (C2D, I3D, and SlowFast), and 2) inspected confusion matrix of the best fine-tuned action recognition model. Our results indicate that 1) SlowFast architecture is most suited for recognizing student actions in school laboratories and 2) similar and consecutive student actions, as well as uneven training data distribution, generate confusion for action recognition models. Further review of these findings allowed us to 1) explore suitability of using computer vision as an added information source for formative assessment and 2) consider details of implementing computer vision for formative assessment support. Overall, this study represents a preliminary investigation that uncovered intricacies of building action recognition models in a real-world setting and unveiled the ways in which computer vision can support teachers in implementing formative assessment in science experiments.
In collaborative problem-solving (CPS), group cohesion is widely regarded as an important contributing factor, yet its underlying processes remain insufficiently understood. Traditionally, cohesion has been treated as a static construct, assessed primarily through self-report measures. In contrast, theory conceptualizes cohesion as a dynamic and emergent state, pointing to the need for temporally sensitive approaches that consider self-report evidence alongside process data. This exploratory study examined how post-task survey-based cohesion, selected socially shared regulation (SSR) discourse indicators, and synchronized leaning forward as a behavioral cue can be interpreted in relation to one another across CPS phases. Specifically, we investigated how selected SSR discourse indicators and synchronized leaning forward were configured in relation to survey-based dimensions of cohesion, including task, social, perceived, and emotional cohesion. We also conducted an in-depth qualitative analysis of selected groups differing in knowledge-construction quality to examine how SSR discourse and synchronized leaning forward appeared in interaction. The analysis is intended as an exploratory and interpretive account: synchronized leaning forward is treated as a context-dependent behavioral cue, and the observed patterns are interpreted as descriptive configurations rather than validated, causal, predictive, or generalizable relationships. The findings suggest that post-task survey-based cohesion can be descriptively interpreted alongside phase-specific discourse and behavioral configurations. In particular, early synchronized leaning forward appeared more meaningful when situated within elaborative and regulatory discourse in the selected cases. Overall, the study presents an exploratory and phase-sensitive interpretation of how survey-based cohesion, selected SSR discourse indicators, and context-dependent behavioral cues can be jointly interpreted to understand cohesion-related processes in CPS.
As established technologies (e.g., corpus technology) and emerging ones (e.g., generative AI) reshape language education, teachers need clear, evidence-based approaches to integrate both into practice. This mixed-methods study examined how over 150 language teachers co-developed AI-empowered TPACK for corpus technology (AI-TPACK-CT) during online professional development. The program was designed using design-based learning and guided by the Community of Inquiry (CoI) framework, which views effective online learning through learners’ social and cognitive engagement. We analyzed discussion posts, lesson plans, and interviews to trace teacher engagement, knowledge building, and learning outcomes. To understand participation patterns at scale, we conducted social network analysis (SNA) to map interaction structures, and then selected participants for deeper investigation. We conducted case studies of one central participant and two peripheral participants to explore how these teachers’ knowledge and design work evolved. CoI-based quantitative results indicated high social and cognitive presence, suggesting an environment conducive to shared inquiry and knowledge building. Qualitative analyses of lesson plans and interviews further showed that teachers’ TPACK development followed different trajectories depending on engagement level and professional experience. Overall, the findings link teachers’ CoI interaction patterns, underlying cognitive processes, and observable design outcomes, extending CoI operationalization beyond postings alone. The study contributes (a) an empirically grounded account of co-developing AI-TPACK-CT through design-based learning where corpus knowledge is playing a key role fostered through AI and pedagogy; (b) evidence on how CoI and SNA relate to TPACK growth; and (c) practical implications for scalable, evidence-centered teacher learning that supports collaborative inquiry, reflective practice, and data-informed integration of AI and corpus technology in language teacher professional development.
This study investigates public discourse on AI in education by integrating Latent Dirichlet Allocation (LDA) and fuzzy-set Qualitative Comparative Analysis (fsQCA) to connect semantic patterns with configurational associations. Drawing on 28,137 anonymized comments from Douyin, Weibo, and Bilibili, eight latent topics were extracted and systematically aligned with theoretical constructs—technological capital penetration, micro-pedagogical resilience, and functional transformation at the discourse level—through Jaccard coefficient analysis. Sentiment classification further quantified emotional orientations, producing a dataset combining topic distributions, sentiment scores, and theoretical mappings for configurational modeling. The fsQCA results revealed a Technology–Resilience Co-Evolution Model, in which five sufficient configurations were interpreted as four broader synergistic evolutionary pathways—comprehensive stress, dual-cycle cognition, substitutional restructuring, and multimodal adaptation—that are associated with the discursive reconfiguration of pedagogical perceptions in media contexts. These findings suggest that AI is discursively framed as co-constructive rather than purely facilitative or disruptive, underscoring the asymmetric and multifaceted nature of public interpretations of its educational role. Methodologically, the study establishes a novel bridge between semantic clustering and configurational analysis, advancing innovation in educational text mining. At both the theoretical and practical levels, the study provides a multidimensional framework for understanding AI-related transformations in pedagogical perceptions within social media discourse. It also offers tentative implications for future research and practice on technological investment with human resilience.
As artificial intelligence (AI) increasingly substitutes for human tasks, higher education must address changing skill demands. However, the emotional and cognitive processes activated when students confront AI substitution remain insufficiently examined. This study conceptualizes and implements an AI-Induced Disruptive Learning Environment in which structured human–AI performance contrasts are designed as psychologically salient learning events.Using a three-wave quasi-experimental mixed-methods design, this study examined undergraduate students enrolled in a product design and manufacturing course. The findings indicate that AI substitution exposure was linked to heightened career awareness and stronger acceptance of skill transformation. Although learning anxiety increased temporarily during the intervention, it appeared to function as an informational signal when supported by guided reflection. Structural equation modeling supported a theoretically consistent pattern of associations linking emotional activation, career awareness, skill transformation acceptance, and self-directed learning, while qualitative findings showed a non-uniform progression from disruption to adaptive repositioning.These findings suggest that AI substitution can be understood not only as a technological capability, but also as a pedagogically structured learning event that shapes emotional, cognitive, and behavioral adaptation. The study contributes to AI-in-education research by clarifying how structured exposure to AI-driven disruption, when instructionally scaffolded, may support reflective adaptation and self-directed learning orientation.
As the popularity of online courses in higher education keeps increasing, the importance of student engagement gets highlighted. In an online learning environment, peer interaction gets emphasised when students work in small groups in breakout rooms. In this case study, qualitative content analysis and discourse analysis were used to analyse breakout room recordings (458 min in total) of two small groups in a 14-week synchronous online language course. The study focused on how engagement at small group level seems to emerge from the group's interaction, which is a primordial site for language learning, and what kind of collaborative engagement practices seem to support it. The findings suggest that affective, behavioural, and cognitive engagement often co-occur and that they may either support or hinder each other. Also, the group members may be engaged in diverse ways at the same time. Three collaborative engagement practices were identified: participation in group discussions, making and negotiating proposals, and collaborative writing. Participation in group discussions highlighted the importance of both verbal and nonverbal communication and the importance of equity in participation. When making proposals, the role of tentativeness was emphasised as that seemed to open space for interaction in the small group. Finally, collaborative writing was found to engage students specifically from the language learning perspective. The findings suggest that small group interaction and collaboration play a key role in group level engagement. It is suggested that small group level should be taken into consideration when operationalising the concept of engagement in online learning environments.
Many learners exhibit low engagement and high dropout rates when using traditional video lectures. With the rapid advancement of generative artificial intelligence (GenAI), these systems are now capable of generating coherent, context-aware, and human-like feedback. However, limited empirical evidence exists regarding how the type of feedback and learners' active involvement in processing feedback affect learning outcomes in video lecture contexts. To address this gap, the present study conducted two experiments to examine how the type of feedback (Experiment 1) and agency in error detection (Experiment 2) influence students' motivation, engagement (as indexed by eye movements and learning behaviors), and performance gains. Across two experiments, the findings consistently showed that feedback combining knowledge of results and knowledge of correct response was more effective than feedback providing either component alone in enhancing learners’ motivation, self-regulatory learning processes, and performance gains. Furthermore, when such feedback was coupled with learner-detected errors, learners demonstrated greater attentional engagement with instructional content, more sustained cycles of self-assessment and regulation, and superior performance gains compared with conditions involving GenAI-detected errors or no explicit error detection. Together, these results support feedback theories emphasizing that effective feedback should integrate evaluative and corrective information while engaging learners in generative processes of error detection.
In real classrooms, teachers often display mismatched emotional tones and visual expressions, yet the effects of such incongruency on learning remain underexplored. In this study, we investigated how (in)congruency between a human-like pedagogical agent's (PA's) vocal tone and visual expressions affects students' learning performance and experience in a desktop virtual reality (DVR) environment. Learners were randomly assigned to one of four PA conditions: (1) positive congruency (positive tone and visuals), (2) neutral congruency (neutral tone and visuals), (3) tone-dominant incongruency (positive tone with neutral visuals), and (4) visual-dominant incongruency (positive visuals with neutral tone), where they learned the inner workings of the human eyeball through hands-on DVR activities. Results showed that the positive congruency PA yielded significantly higher retention, positive emotions, and intrinsic motivation than the visual-dominant incongruency condition and outperformed the neutral congruency condition on most measures. Notably, the tone-dominant incongruency produced effects comparable to positive congruency across all outcomes and surpassed visual-dominant incongruency in retention and positive emotion. Mediation analysis further confirmed that the tone-induced positive emotions contributed to improved retention, motivation, and perception of the PA. These findings suggest that in emotionally expressive instructional design, tone of voice may override visual cues in influencing affective and cognitive outcomes, offering implications for the design of emotionally intelligent agents in immersive learning environments.
Large language models (LLMs) are increasingly used in educational technology for automated writing assessment, yet most applications rely on prompt-based scoring and feedback generation, which often lack transparency, reproducibility, and interpretability. This study investigates whether model-internal LLM representations can provide interpretable and reproducible metrics for second language (L2) writing assessment. We derived surprisal and perplexity from next-token prediction to quantify linguistic predictability and embedding-based similarity to measure semantic coherence across sentences. These metrics were computed at the token, sentence, and discourse levels using three pretrained Chinese language models and evaluated on 1196 essays written by Chinese L2 learners across four proficiency levels. Their relationships with 11 established linguistic measures of fluency, lexical sophistication, phraseological complexity, and syntactic complexity were also examined. Results showed that surprisal and perplexity generally decreased with proficiency for the two Traditional Chinese-focused models, indicating greater linguistic predictability in more proficient writing, whereas the multilingual model showed weaker sensitivity. Embedding-based similarity increased with proficiency, reflecting stronger semantic coherence. Combining LLM-derived metrics with classical linguistic features improved proficiency classification and prediction beyond either feature set alone. Correlation and qualitative analyses further demonstrated that the proposed metrics capture complementary aspects of writing while revealing conditions under which their interpretations become less reliable. These findings demonstrate the value of interpretable, model-derived metrics for transparent, reproducible, and scalable AI-supported L2 writing assessment, particularly for underrepresented learner populations and lower-resource languages.
As artificial intelligence becomes increasingly integrated into educational and therapeutic settings, understanding professional user experience with AI-generated interventions is crucial for successful implementation. This study examines how education and therapy professionals experience using social stories created by AI versus human experts for children with special needs. Using a 2×2 experimental design, 227 female professionals in Israel evaluated social stories designed to either promote positive behaviors or reduce challenging behaviors, with authorship randomly attributed to either AI (Claude) or a human speech therapist. All stories were actually AI-generated, allowing isolated examination of authorship perception effects. Results revealed no main effects for creator type or story purpose, but significant Creator × Purpose interactions emerged for understanding, feasibility, and overall user experience. AI-generated stories were rated more favorably for behavior promotion interventions, while human-authored stories were preferred for behavior reduction interventions. User experience ratings remained stable across conditions, suggesting fundamental legitimacy of both approaches. Professionals working with younger children reported significantly better user experience compared to those working with adolescents, while professional experience and field showed no influence on user experience evaluations. These findings challenge assumptions about AI resistance in sensitive professional contexts and reveal the importance of strategic alignment between intervention purpose and perceived creator expertise. The results provide evidence-based guidance for integrating AI tools in special education, suggesting that AI may be particularly valuable for promoting positive behaviors while human expertise remains preferred for addressing complex behavioral challenges.
Generative AI-powered teachable agents offer new opportunities in mathematics education, yet systematic investigations of how vocal design affects middle school students' engagement and learning outcomes remain limited. We conducted a three-week 2 × 2 × 3 factorial randomized controlled experiment with 386 middle school students, systematically manipulating three vocal attributes of a generative AI-powered teachable agent: gender congruence (avatar-voice matching vs. mismatching), voice age (youth vs. adult), and emotional tone (sad, calm, cheerful). Engagement was assessed through behavioral coding of ten interaction categories and temporal analysis of dialogue sequences, while learning gains were measured through standardized state mathematics assessments. Analyses employed generalized estimating equations for factorial effects, time-series feature extraction for temporal dynamics, and machine learning models with SHAP-based feature importance for outcome prediction. Results revealed a theoretically significant dissociation between vocal configurations that activate immediate teaching behaviors and those that support long-term achievement. Empathy-evoking, identity-neutral configurations combining gender-mismatched and sad voices significantly enhanced immediate generative processing, including conceptual knowledge demonstration (β = 0.776), answer-giving behaviors, and explanatory responses. Conversely, schema-congruent, approachable configurations combining gender-matched and cheerful voices yielded the highest standardized learning gains (β = 0.997) and also significantly enhanced conceptual knowledge demonstration (β = 0.722), supporting sustained engagement over time. Time-series analysis confirmed that sad emotional tone maintained higher peak and overall levels of instructional behaviors across dialogue sequences. Feature importance analysis identified frustration management as the strongest predictor of standardized learning gains, while explanatory behaviors most strongly predicted conceptual knowledge demonstration. Voice age manipulations showed no significant effects. These findings demonstrate that strategically designed vocal attributes configuring multiple voice dimensions simultaneously activate learning-by-teaching processes more effectively than single-dimensional features. Practically, this means using sad tones with gender-mismatched voices to activate immediate teaching behaviors, and cheerful tones with gender-matched voices to sustain long-term achievement. Future research should explore adaptive designs that transition between these configurations based on learning phases.
Student self-assessment literacy (SSAL) is crucial for second language (L2) writing proficiency and learner autonomy. However, lower-proficiency English as a foreign language (EFL) learners often struggle to apply assessment criteria independently and remain reliant on teacher feedback. Although generative AI is increasingly used to provide writing feedback, its potential to support SSAL development through explicit evaluative reasoning remains underexplored. In this study, AI chain-of-thought (AI-CoT) refers to learner-visible, criterion-guided evaluative reasoning that externalizes the processes of planning, monitoring, and evaluation during supported self-assessment. In this sense, AI-CoT was used alongside AI-generated feedback to make the reasoning linking assessment criteria, textual evidence, and revision priorities more transparent. This study traces SSAL development across four dimensions among lower-proficiency university EFL learners, analyses the complementary roles of AI-generated feedback and AI-CoT, and proposes a dual-loop interaction framework explaining how teacher-led instructional design and student-led AI-mediated self-assessment jointly shape SSAL. Employing a three-cycle action research design over 12 weeks, this study involved 30 lower-proficiency university EFL learners. Data included think-aloud protocols, semi-structured interviews, drafts with tracked revisions, AI-generated feedback and AI-CoT logs, SSAL checklists, and teacher reflective journals. Findings showed that AI-generated feedback primarily supported criterion application and text revision, whereas AI-CoT supported learners’ comprehension and interpretation by making evaluative reasoning more visible. Across the three cycles, students became increasingly self-directed in their self-assessment, although critical engagement and independent enactment remained uneven. The proposed framework offers pedagogical guidance for integrating generative AI into reflective, criterion-guided self-assessment to support metacognitive growth in L2 writing.
Research on artificial intelligence (AI) in educational leadership is expanding rapidly but remains fragmented. This study synthesises the field through a PRISMA-guided bibliometric analysis of 219 peer-reviewed articles from Web of Science and Scopus and develops an agenda for future research. Using VOSviewer, we mapped the field's evolution, prolific authors and journals, and keyword co-occurrence networks, identifying five thematic domains: foundational technologies, student- and teacher-facing applications, data-driven analytics, systems integration, and algorithmic decision support. Results reveal sharp post-2020 growth, disciplinary dispersion, and concentration in higher education and a small number of regions. Building on these patterns, the study proposes a future research agenda focused on geographic scope, educational level, ethical governance, longitudinal implementation, generative AI governance, and theoretical integration. The study advances the literature by consolidating fragmented insights into a taxonomy that reframes AI as an evolving governance ecosystem and by clarifying the key gaps that should shape future research on responsible AI adoption in educational leadership.
This study examines the intellectual evolution of Artificial Intelligence in Education (AIED) from 2000 to 2024 by integrating scientometric analysis and BERTopic-based semantic modeling. Drawing on 2340 peer-reviewed records from the Web of Science, we mapped publication patterns, journals, institutions, authors, and thematic structures, and interpreted the results through a macro–meso–micro analytical framework. Findings reveal rapid growth in AIED research, with marked acceleration after 2022, alongside a gradual reorientation of research agendas. At the macro level, early work was shaped by online learning and MOOCs, while later research increasingly reflects the importance of data-intensive educational systems, especially learning analytics and educational data mining. These developments raise institutional implications related to data regulation, privacy, transparency, accountability, and ethical oversight. At the meso level, teacher digital competence and organizationally supported professional capacity function as key mediating hubs that connect institutional agendas with curriculum design, assessment, and organizational support. At the micro level, research attention moves beyond performance prediction toward learner engagement, affective experience, and interactive forms of AI-supported learning. By linking these clusters through an institutional–organizational–individual coupling framework, the study shows that AIED knowledge has evolved not as a set of isolated themes, but as a cross-level structure in which institutional priorities, organizational arrangements, and learner experiences mutually shape one another. The synthesis offers implications for data-informed governance, institutional capacity building, teacher professional development, curricular integration, and sustainable support structures for user-centered AI adoption in education.
Rapid technological developments are reshaping expectations regarding teachers' digital competence targets, challenging teacher education programmes to design learning experiences that systematically support its development. This study examines how digital competence targets and teacher education strategies intersect within authentic teacher education practices. Drawing on the European DigCompEdu Framework and the second Synthesis of Qualitative Data (SQD2) model, the study analyses 33 documented teacher education practices implemented across five universities. Using quantitative content analysis combined with co-occurrence and network analysis, the study maps (1) the digital competence targets addressed, (2) the SQD2 strategies enacted, and (3) the patterns of association between digital competence targets and teacher education strategies. The findings indicate that pedagogically oriented competence targets related to teaching practice and professional reflection, particularly Reflective practice (1.3) and Teaching (3.1), are most prominently represented across the documented practices. These digital competence targets consistently co-occur with a coherent set of teacher education strategies, including hands-on experience, pedagogical reasoning, and teacher identity development. In contrast, digital competence targets related to technical problem-solving, digital safety, and technological infrastructure appear only marginally or remain largely implicit within the analysed practices. The findings further underscore the importance of curriculum coherence and pedagogical orchestration in supporting the development of pre-service teachers’ digital competence targets.