
Timing is a key feedback-design feature that may shape how second language (L2 learners notice, interpret, and apply written corrective feedback (WCF). With collaborative digital tools such as Google Docs, teachers can provide synchronous, anchored metalinguistic comments while learners compose, making "immediacy" a practical classroom option. Yet most timing research has been conducted in laboratory contexts, typically focusing on a single structure in one-to-one interaction, limiting generalizability to teacher-to-many classroom ecologies. This classroom quasi-experiment compared immediate versus delayed metalinguistic WCF in three intact L2-French classes (N = 75; CEFR B1). Learners wrote in self-selected dyads in shared Google Docs. The immediate group (n = 26) received metalinguistic comments during drafting; the delayed group (n = 24) received comparable comments on the same drafts one week later and revised in a brief, dedicated revision episode; a comparison group (n = 25) completed the same tasks without WCF during the study. Feedback targeted multiple curricular features within the same instructional cycle: noun-phrase agreement, verb-phrase agreement, and the pass & eacute; compos & eacute;-imparfait contrast. Development was assessed using individual writing tests administered at pretest, immediate posttest, and delayed posttest, with accuracy indexed as errors per 100 words per target category. Overall, delayed feedback showed the most consistent short-term advantage across targets at the immediate posttest, whereas longer-term effects were smaller and target dependent. Results suggest that timing effects in technology-mediated collaborative writing are shaped by classroom workflow constraints and by the processing demands of specific linguistic targets rather than reflecting a uniform "immediate is better" principle.
Research on data-driven learning (DDL) has long emphasized the value of engaging learners with authentic language data through corpora. However, little is known about how learners engage with and regulate their learning in DDL environments supported by generative artificial intelligence (GenAI), particularly in academic writing revision. To address this gap, this mixed-methods study examined how seven Chinese EFL learners engaged cognitively, behaviorally, and affectively with a Corpus-GenAI DDL mode, and how their self-regulatory capacity shaped these processes. Data were collected from students' drafts, GenAI interaction logs, corpus query logs, stimulated recall interviews, reflective journals, and classroom observations. Results showed that skilled self-regulators engaged in deeper processing by evaluating GenAI suggestions, consulting corpora to verify linguistic choices, and using evidence to refine revisions. By contrast, less-skilled self-regulators relied more heavily on surface-level GenAI input, with less verification and weaker strategic planning. Despite these differences, learners generally reported positive affective responses, including curiosity, excitement, and growing confidence in the combined use of corpora and GenAI tools. This study contributes to the literature on DDL and feedback engagement by showing how GenAI-mediated scaffolding may support learners' engagement with corpus consultation during writing revision. It further offers insight into how learners coordinate corpora and GenAI during revision, highlighting the role of self-regulation in shaping learner engagement, evidence-based decision-making, and linguistic awareness. Pedagogically, the findings underscore the value of feedback-rich writing environments that integrate corpora and GenAI to foster learner self-regulation and sustained engagement.
This study investigated whether generative artificial intelligence (GenAI) roleplay practice facilitates the development of professional disagreement strategies among second language (L2) business English learners. Although GenAI tools have attracted growing interest in language education, controlled experimental evidence for their effectiveness in developing specific pragmatic competencies remains scarce, with most existing research reporting learner perceptions rather than measured learning outcomes. Employing a pre-test, post-test, delayed post-test design with three experimental conditions (GenAI roleplay, human peer roleplay, and traditional materials), 72 intermediate-level business English learners completed a three-week intervention targeting the pragmatic competence required to challenge, reject, or oppose ideas while maintaining professional rapport. Written discourse completion tasks assessed participants' use of appropriate disagreement strategies, including hedging devices, epistemic markers, account structures, and alternative proposal formulations. Results from mixed-design ANOVA revealed significant group-by-time interactions. The GenAI condition significantly outperformed the traditional materials condition at post-test with a large effect (d = 1.00, p = 0.001). The Human Peer condition showed a medium-sized advantage over traditional materials (d = 0.55, p = 0.064). The comparison between GenAI and Human Peer conditions did not reach statistical significance (d = 0.47, p = 0.113), indicating that while both interactive conditions outperformed traditional materials, the evidence for GenAI-specific advantages over peer interaction remains inconclusive with the present sample size. These findings provide empirical support for interactive practice as pedagogically valuable for developing L2 pragmatic competence in face-threatening speech acts. For practitioners, the results suggest that GenAI roleplay merits consideration as a complement to traditional pragmatics instruction, particularly where peer practice opportunities are constrained.
Despite growing interest in captions for second language (L2) acquisition, the research on keyword captions (KC) is scarce, and their impact on L2 vocabulary phonological learning remains unexplored. This study examined whether KC provided greater benefits than full captions (FC) and no captions (NC) for intentional L2 vocabulary learning and video comprehension. A total of 131 low-proficiency Japanese learners of Chinese were randomly assigned to one of the three captioning conditions while watching an educational video, followed by tests of phonological form recognition, meaning recognition, and content comprehension. Prior to watching, participants were informed that both vocabulary and comprehension tests would follow, and were allowed to preview the test items. The results showed that learners in the KC group consistently outperformed those in the FC and NC groups across all tests, with statistically significant advantages over FC and NC in both phonological learning and comprehension. These findings offer the first empirical evidence that captions can enhance CL2 (Chinese as a second language) phonological development. The superior performance of KC is likely attributable to both its low textual density (3.3% of the full transcript), which reduced cognitive load and directed learners' attention to essential content, and the intentional vocabulary learning context, which further enhanced learners' engagement with target words. Overall, these results highlight the pedagogical value of well-designed, low-density keyword captions in supporting intentional vocabulary acquisition and audiovisual comprehension in CL2 instruction.
Much research has investigated learners' engagement with automated written corrective feedback (AWCF). However, few have employed eye-tracking technology, which captures real-time cognitive processes during reading, and even fewer have targeted postgraduates. This case study investigated how two English postgraduates used Pigai's AWCF to revise their academic writing during the final draft of their doctoral research proposal. A mixed approach based on the established framework was used to collect and analyze data in behavioral, cognitive, and affective dimensions. Eye-tracking technology was utilized to triangulate cognitive engagement. The findings show that each student's engagement with AWCF differed across three dimensions. Specifically, eye movement data provided objective evidence of cognitive engagement differences: whether their visual attention was guided by color-coded feedback severity. The difference was also reflected in their behavioral and affective engagement, with the student who read top-down but selectively, rather than by color-coded severity, demonstrating higher revision effort and trust. These differences, as well as the observed deviations between the tool's processing mechanism and the conventions of academic writing, show several key implications: develop more genre adaptive automated writing evaluation (AWE) tools, critically evaluate the consistency of the tool and genre before adoption, distinguish writing guidance, cultivate balanced trust in AWCF, and integrate eye tracking technology for personalized learning.
Despite the rising popularity of learning Arabic as a foreign language, there has been limited research addressing potential effects of integrating technology into the learning approach for learning Arabic language, particularly using gamification. Few studies have explored whether gamification can influence factors like language acquisition. To address these gaps, we proposed a gamified analytical thinking skills learning approach (Gamified-ATS) for young learners. We employed a mixed-methods research design to investigate the effects of this approach on young learners' Arabic language skills, motivation, engagement, and learning perceptions. The study involved 52 third-grade students from an elementary school in Indonesia. One class (n = 25) was taught using the Gamified-ATS approach, while another class (n = 27) followed the conventional technology-based learning approach (CTL). The results indicated that integrating gamification with the analytical thinking skills approach significantly improved young learners' Arabic language skills, motivation, and engagement. Additionally, we analyzed the learners' perceptions of the Gamified-ATS approach through epistemic network analysis. Our findings suggest that fostering enthusiasm can greatly enhance engagement and motivation. However, gender-specific preferences emerged: boys benefited more from competitive elements and clear progress metrics, whereas girls were more motivated by interactive and narrative-driven activities.
The integration of multimodal input channels, including text, voice, image interaction, and human-like representations, within Generative AI (GenAI) agents enables immersive speaking experiences by providing linguistic scaffolding and emotional support. Despite their growing use in English as a Foreign Language (EFL) speaking instruction, limited research has examined how vocal and visual design features shape learners' experiences and outcomes. This study adopts a 2 (Time: pre vs post) & times; 3 (Group: control, Agent 1, Agent 2) quasi-experimental design to investigate the effects of alternative voice and visual configurations on speaking performance, affective responses, and interactional experience. Two multimodal GenAI agents were developed: Agent 1 employed an animated teacher image with a standard AI-generated voice, whereas Agent 2 featured a realistic teacher photograph with an AI-cloned teacher voice. Participants were drawn from three university EFL classes: the control group (N = 37) received traditional instruction, while experimental group 1 (EG1, N = 34) interacted with Agent 1 and experimental group 2 (EG2, N = 36) interacted with Agent 2. Over eightweeks, learners completed structured speaking tasks under their respective conditions. Multilevel modelling of performance and questionnaire data showed that both EGs outperformed the CG in speaking gains. EG2 demonstrated greater improvements in self-perceived communicative competence and reductions in speaking anxiety, along with higher perceived interactivity and immersion, whereas EG1 showed a significant decrease in foreign language boredom. Follow-up interviews further illustrated how vocal and visual design features shaped learners' learning experiences. These findings provide insights for the design and pedagogical deployment of multimodal GenAI agents in EFL speaking instruction.
Although considerable attention has been directed toward incorporating technology for writing instruction, existing research has primarily focused on students in mid-elementary grades and beyond. Because studies indicate that gaps in writing achievement may surface during early elementary grades and persist, this study aims to address this pressing issue with an intervention that leverages tablet technology and educational applications for second-graders. Guided by the Simple View of Writing (SVW), this study investigated the effect of the modes of writing (by hand vs. using tablets) and writing instruction on writing achievement for struggling writers from families with low socioeconomic status (SES), a group that often experiences challenges with writing and less access to technology. A quasi-experimental study was conducted with three conditions: a Control group receiving no additional writing intervention beyond their regular English Language Arts instruction at school, a Writing Intervention group experiencing additional traditional writing intervention without the use of technology, and a Technology-Enhanced Writing Intervention group receiving writing intervention utilizing tablets and applications to facilitate ideation and transcription, the main pillars of SVW. Results showed that writing modality is significant; the technology-enhanced group performed better when writing by hand than when using tablets, aligning with prior research. Moreover, the technology-enhanced intervention showed a positive and statistically significant effect for high-attendance intervention participants writing by hand compared to the control group. The study extends the scope of previous research by illustrating the potential of sustained technology-enhanced writing instruction among struggling writers from low SES backgrounds in the early elementary grades.
Despite extensive development of computer-assisted language learning (CALL) tools and Intelligent Tutoring Systems (ITS) addressing grammar, learners often still struggle to apply complex morpho-syntactic rules in spontaneous speech. In particular, Arabic poses challenges in agreement (gender, number, case) that traditional instruction and rule-based feedback have only partially overcome. Moreover, recent reviews reveal that generative-AI research in language learning has focused overwhelmingly on written English tasks, leaving speaking and less-studied languages largely underexplored. To fill this gap, we conducted a longitudinal mixed-methods quasi-experiment comparing a GPT-4-based dialogic tutor to conventional explicit instruction for non-native Arabic learners' oral mastery of agreement rules. Participants (N = 60 intermediate AFL learners) were assigned by intact class to either an AI-mediated speaking practice environment or traditional drill-based instruction. Pre-, post-, and delayed-post oral tests (elicited speech tasks targeting verb-subject, adjective-noun, and subject-predicate agreement) were analysed for accuracy in obligatory contexts, error density per 100 words, and fluency metrics (speech rate, pause ratio, response latency). System log data (feedback events, response times) and learner questionnaires (anxiety, perceived usefulness) provided additional insights. Results from ANCOVAs and mixed-effects models suggested that the AI group outperformed the control on agreement accuracy and showed greater improvements in fluency indices, with these gains largely maintained at a 4-week delay. Error density declined more sharply in the AI condition. Learning analytics indicated that log-derived features (e.g. accuracy, focus, and time-on-task) significantly predicted individual gains. Learners interacting with the AI tutor reported lower speaking anxiety and high technology acceptance, consistent with broader evidence that AI-mediated speaking support can enhance enjoyment and willingness to communicateLearners interacting with the AI tutor reported lower speaking anxiety and high technology acceptance, echoing Zhang et al. (2024) findings of boosted enjoyment and willingness to communicate under an AI speaking assistant. Qualitative comments further suggested that the dialogic agent may have functioned as a safe, personalised practice space. In sum, this study suggests that a generative-AI dialogic tutor may support the transition from explicit rule knowledge to more fluent use of Arabic agreement, thereby potentially scaffolding the proceduralisation of grammar. This work helps address the dearth of speaking-focused GenAI research, is consistent with cognitive load theory, and may extend skill-acquisition accounts by suggesting a possible role for GenAI in lowering the affective filter and enhancing noticing. Pedagogically, we frame our system as an Adaptive Learning Ecosystem that dynamically modulates task difficulty and feedback. We conclude with recommendations for curriculum designers on ethically integrating AI tutors - from scaffolding prompts to protecting student data - to harness GenAI affordances while preserving language-specific complexity.
In English language learning, generative artificial intelligence (GenAI) chatbots can support students' writing by providing context-based examples and guiding idea organization. Yet, improving discourse-level writing remains challenging, as feedback often targets accuracy rather than meaning. This study investigates the impact of GenAI chatbots in writing instruction compared with the traditional method on the communicative competence (CC) of Chinese upper secondary English as a Foreign Language (EFL) students. A quasi-experimental study was conducted over a semester with 154 Grade 11 students. The experimental group (n = 76) worked with DeepSeek, while the control group (n = 78) received traditional teacher instruction. Writing was assessed three times using a four-dimensional analytic rubric with 20 items covering grammatical, sociolinguistic, discourse, and strategic competences. Repeated-measures ANOVA revealed greater improvements in the DeepSeek group across all four competence dimensions, with the most notable gains in grammatical and sociolinguistic competence. Item-level analysis showed the largest increases in sentence structure, character voice, and social awareness. In contrast, idea order improved less in the DeepSeek group than in the control group, suggesting that higher-level organization still requires human mediation for effective development. The findings point to the value of incorporating GenAI chatbots into EFL writing instruction to support CC development. When used collaboratively with teacher guidance, GenAI chatbots can better support discourse and pragmatic skills that require sustained attention.
The use of a shared screen plays a central role in the overall architecture of video-mediated interactions, and it becomes even more prominent when a teacher uses it for pedagogical purposes in online classes. In this study, we investigate an underexplored setting, fully-online, synchronous, video-mediated Turkish as a foreign language classrooms. Using multimodal conversation analysis to examine a dataset of screen-recorded video-mediated L2 classroom interactions consisting of a teacher and a small group of students with turned-off cameras, we focus on a recurrent interactional practice, namely screen-based repair, and outline the methods used by the teacher based on a collection of cases. The findings show that upon the identification of a trouble, whether the trouble source is visible on the screen or not, the teacher initiates, enacts and completes other-initiated other-repair by systematically drawing on diverse screen-based resources, and engages in writing aloud, highlighting aloud, and cursor marking aloud, in doing so, thus deploying the practice of screen-based repair to resolve technical, pedagogical, and interactional troubles. We argue that the findings not only describe the pedagogical context of a less taught language but also bring new insights into video-mediated L2 classroom interaction and online language teaching overall.
This study examines pre-service English teachers' lived experiences of using generative artificial intelligence (GenAI) for multimodal lesson planning. Conducted within a six-week, technology-enhanced pedagogy course in a postgraduate teacher education programme in Hong Kong, the study employed interpretive phenomenological analysis drawing on classroom observations, AI-generated artefacts, and reflective interviews. The findings portray a developmental trajectory in which participants advanced from tentative experimentation toward more confident and critical use of GenAI. Three themes were identified. First, 'GenAI as a collaborative multimodal lesson design partner' showed how participants orchestrated heterogeneous tools to create differentiated and multimodal resources. Second, 'negotiating authorship and professional identity' revealed ambivalence about originality and creative ownership but also a shift from template adoption to adaptive curation, contextualisation, and principled integration. Third, 'navigating challenges and ethical boundaries' highlighted how participants acted as cultural and factual mediators while modelling responsible and transparent AI use to foster critical AI literacy among learners. Framed by the PedAIComp framework, the results show movement from Awareness and Exploration to Integration and, in some practices, Expertise and emerging Leadership. While Innovation was not reached within the intervention, the study points to future opportunities for professional development and collaborative design research to cultivate Innovation competences.
Despite the growing integration of GenAI tools into academic writing, little is known about how learners with different performance levels engage with them. This study examines how EFL learners use GenAI during an academic writing task, conceptualizing AI affordances and constraints through a self-regulated learning (SRL) framework that focuses on learners' questioning and feedback adoption behaviors. Thirty-seven learners initially participated; two who did not use GenAI were excluded, resulting in a final sample of 35. Based on writing scores, the top 40% (n = 14, Group H) and bottom 40% (n = 14, Group L) were selected for comparison. Data sources included assignments, screen recordings, GenAI chat logs, and post-task interviews. Descriptive and inferential statistics were used to compare questioning and adoption patterns, complemented by qualitative content analysis. Both groups used GenAI with similar frequency but differed in how they engaged with it. Group H primarily asked language-focused questions related to translation refinement and academic expression, whereas Group L relied more on GenAI for content generation and comprehension support. Group H also revised AI outputs more extensively. Although both groups viewed language editing as GenAI's primary strength, they differed in how they interpreted its limitations: Group H attributed problems to prompting strategies, whereas Group L emphasized technical constraints. Based on these patterns, the study proposes pedagogical principles for AI-mediated EFL writing.
Generative artificial intelligence (GAI) has gained increasing attention in English as a foreign language (EFL) education, with growing evidence supporting its efficacy in enhancing oral performance. Nonetheless, limited research has examined how GAI's multimodal capabilities shape learner behavior, and how different learning behavior clusters relate to motivation and oral performance dynamics. To address these gaps, this study explored (1) the learning behavior clusters emerging from learners' interactions with the Multimodal GAI (MGAI), and (2) the relationship between these clusters, and associated changes in motivation and oral performance. From a three-week intervention with 60 EFL learners, data were collected on behavior, motivation and performance. K-means clustering identified four distinct learners' behavior clusters: interaction-based taskers, comfort-oriented balancers, output-focused monitors, and resource-driven strategists. We found significant differences across clusters. Resource-driven strategists and output-focused monitors showed greater positive changes in intrinsic motivation, while interaction-based taskers exhibited greater positive changes in extrinsic motivation. In oral performance, resource-driven strategists and output-focused monitors exhibited greater positive changes than both comfort-oriented balancers and interaction-based taskers in terms of fluency and coherence. Qualitative insights (i.e., interview and dialogue data) in each cluster provided illustrative information for these quantitative results. These findings provide pedagogical insights for the integration of MGAI into language learning contexts, highlighting that multimodal support alone does not automatically lead to positive learning changes. Thus, pedagogically guided use of multimodal support is needed to help learners move beyond surface-level MGAI-human interaction and more effectively appropriate MGAI tools to support scaffolded oral learning.
No randomized evidence directly compares immersive simulation, intelligent pairing, and adaptive learning instructional designs within the same postgraduate EFL study using a standardized proficiency endpoint. This multisite randomized controlled trial compared three AI-enhanced instructional designs-structured task bundles delivered via AI mediation and evaluated as instructional designs rather than technology platforms-specifically Immersive Simulation Technology (IST), Intelligent Pairing Systems (IPS), and Adaptive Learning Platform (ALP)-with an active control in required postgraduate TESOL coursework at four East Asian university sites (N = 245) over 12 wk. The preregistered confirmatory outcome was official ETS-administered TOEFL iBT total score; Intercultural Communicative Competence Performance-Oriented (ICC-PO) performance was analyzed as a secondary exploratory outcome. Within the preregistered confirmatory proficiency family, all three Holm-Bonferroni-corrected contrasts were supported: IST outperformed Control and ALP, and ALP outperformed Control on TOEFL iBT total score; IPS proficiency contrasts were exploratory. IST produced the largest adjusted gain (Delta = 36.99 TOEFL points, 95% CI [35.58, 38.44]; Hedges' g = 3.02, 95% CI [2.65, 3.39]; d_resid = 5.67), interpreted as a near-transfer, high-dose, context-specific upper-bound estimate rather than a population-average expectation. Exploratory ICC-PO patterns broadly paralleled the proficiency results but are interpreted cautiously as context-bound performance differences under a partially aligned instrument, with cross-group measurement comparability not established. The study contributes head-to-head randomized evidence for instructional-design selection-not technology novelty-and supports a sequenced pedagogical logic in which ALP, IPS, and IST serve distinct roles in postgraduate EFL instruction.
A substantial body of research has demonstrated the benefits of written corrective feedback (WCF) for L2 written accuracy (see meta-analyses such as Kang & Han); however, comparatively few studies have experimentally examined the effectiveness of feedback in different timing conditions or how learners engage with feedback provided at different times. Yet, understanding when feedback is most beneficial and how feedback timing shapes learners' engagement with corrections is pedagogically essential, as it can help learners maximize their opportunities to notice and integrate corrections into their evolving L2 linguistic system. The present study employs a mixed-methods design to compare learners' engagement with WCF and accuracy improvement under three timing conditions: (i) immediate synchronous WCF (during writing), (ii) delayed synchronous WCF, and (iii) asynchronous WCF. Twenty advanced EFL university students completed two writing tasks, processed unfocused indirect WCF, completed audio-recorded questionnaires on their engagement with WCF, and revised their texts using screencast technology to capture real-time engagement and changes in accuracy. All groups significantly improved accuracy, with no advantage for synchronous or delayed synchronous over asynchronous feedback. All participants were similarly engaged, but different patterns of engagement were found because of different feedback timing conditions. These findings suggest that advanced learners can benefit equally from different feedback timing conditions using unfocused indirect feedback. Methodological and pedagogical implications are drawn.
Research on L2 revision suggests that corpus-based learning can improve students' writing performance, particularly in terms of language accuracy. However, how students use the Internet as a corpus tool during revision remains under-investigated. To address this gap, this study examined how five Chinese first-year EFL undergraduates engaged with the search engine Bing as a revision tool over a semester. Data collected from group discussions, reflection journals, written products, and individual interviews were analyzed. The findings indicate that students employed varied search and incorporation strategies, including refining search queries, directly applying retrieved language, and selectively or diversely engaging sources to support revisions. Importantly, Bing supported not only language-related but also meaning-related revisions. These revisions were shaped by the group's evolving division of labor: students in language-focused roles refined search queries for lexical and grammatical issues, whereas those in meaning-focused roles increasingly drew on diverse sources and incorporated retrieved information to strengthen argumentation. Over time, students also exhibited more positive perceptions of Bing's affordances and demonstrated greater flexibility and sophistication in their search practices. We argue that strategic use of Bing in the context of peer collaboration can facilitate meaningful revisions for broader purposes in L2 writing.