
The present study addresses the question of whether an individual who does not understand a sentence might still be able to repeat it verbatim. To answer this question, we examined paraphrasing and repetition data from two previous studies: Pavlenko et al. (2019), which analyzed L1 and L2 participants' paraphrasing of seven Miranda warning sentences, and Akbary et al. (2023), which compared L1 and L2 participants' paraphrasing and elicited imitation (EI) performance on 30 commonly used EI sentences. We formulated our questions into four predictions, two of which directly addressed whether rote repetition without comprehension occurs at all. Our results confirmed both predictions, identifying 36 (5.6%) of the 646 instances of verbatim sentence repetition in the data as potential cases of repetition without comprehension. However, a broader analysis showed that the evidence for a lack of comprehension was relatively weak and ambiguous. We conclude with recommendations for overcoming the limitations of the present study and resolving the ambiguity of our findings. Esta investigaci & oacute;n aborda la interrogante de si una persona que no entiende una oraci & oacute;n podr & iacute;a aun as & iacute; repetirla literalmente. Para responderla, examinamos datos de tareas de par & aacute;frasis y repetici & oacute;n de dos investigaciones previas: Pavlenko et al. (2019), quienes analizaron c & oacute;mo participantes de L1 y L2 parafraseaban siete oraciones de la advertencia Miranda, y Akbary et al. (2023), quienes compararon el desempe & ntilde;o de participantes de L1 y L2 en par & aacute;frasis e imitaci & oacute;n elicitada (EI) de 30 oraciones de uso com & uacute;n en tareas de EI. Nos planteamos cuatro predicciones, dos de las cuales abordaban directamente si la repetici & oacute;n mec & aacute;nica sin comprensi & oacute;n realmente ocurr & iacute;a. Nuestros resultados confirmaron ambas predicciones e identificaron 36 instancias (5.6%) de las 646 repeticiones literales de oraciones en los datos como posibles casos de repetici & oacute;n sin comprensi & oacute;n. Sin embargo, un an & aacute;lisis m & aacute;s amplio mostr & oacute; que la evidencia de falta de comprensi & oacute;n era d & eacute;bil y ambigua. Concluimos con recomendaciones para superar las limitaciones de la investigaci & oacute;n actual y aclarar nuestros hallazgos.
Generative artificial intelligence (GenAI) is rapidly transforming education, including the ways students learn languages and experience motivation. While recent research has focused mainly on the affordances of GenAI, much less attention has been given to its potential risks for learners' motivation and engagement. This paper addresses this gap by examining how GenAI may both enhance and undermine language learning motivation. We discuss how GenAI can affect the satisfaction or frustration of the three basic psychological needs - autonomy, competence, and relatedness - and how these effects may shift motivational orientations from intrinsic to extrinsic forms. We then explore how GenAI might reduce cognitive engagement, critical thinking, and assessment integrity when used uncritically. We also consider key methodological opportunities and challenges in researching GenAI-mediated language learning motivation. Based on this analysis, we identify key directions for future research agendas, including the balance between autonomy and dependence, sustaining motivation beyond novelty effects, rehumanizing digital learning, and redefining teachers' roles in GenAI-rich contexts. We also call for theoretical renewal, ethical awareness, and methodological innovation in the study of GenAI-mediated motivation. A deeper understanding of these processes will allow future research and practice to harness GenAI's potential responsibly, promoting motivation that is both sustainable and self-determined.
This study investigates the role of working memory (WM) in the development of receptive and productive abilities in Spanish among intermediate second/additional language (L2/A) learners. Participants completed WM assessments and receptive and productive language tasks targeting grammatical gender agreement on articles and adjectives (receptive: acceptability judgment task [AJT]; productive: information gap activity). Results showed group-level improvement for productive performance, but substantial variability in growth for both receptive and productive performance. WM did not significantly predict growth in either ability, suggesting that WM may not strongly influence the development of gender agreement accuracy at this intermediate proficiency level. To extend the analysis beyond behavioral outcomes, an exploratory post hoc analysis examined associations between WM and neural responses during receptive processing. By examining parallel receptive and productive performance and integrating behavioral and neurocognitive approaches, this study highlights the value of interdisciplinary methods moving forward for understanding the role of WM in L2/A development.
This position paper chronicles and critiques my attempts to help establish sport as a field of application for applied linguistics. In this way, it is a nontraditional article: a first-person, autobiographical, position paper. As background to this endeavor, the paper first acknowledges the critical role of collaborative publications and the establishment of a research network. The paper then explores the affordances and complications associated with building a sports-based community of applied linguistics, especially given preexisting scholarship in sports linguistics, as well as considerable scholarship in the related discipline of communication studies. The paper then briefly recounts my own research in sports linguistics with the aim of bringing to consciousness my goals and purpose. Drawing on Bakhtin's imagined audience, I reflect on how these motivations not only shaped my research agenda but, in turn, my conceptions of the relationship between linguistics and sport and applied linguistics. I then turn to Bernstein to further reflect on my classification of sports linguistics - the discipline, its knowledge and knowers - and its influence on my attempts at community building. As a position paper, I conclude with some insights for readers who find themselves engaged in similar attempts at community building.
Understanding how attitudinal prosody is processed in a second language (L2) remains an open question, particularly regarding its neural mechanisms and the role of real-world experiences. Using functional magnetic resonance imaging (fMRI), we examined how native Japanese learners processed attitudinal versus linguistic prosody in their L1 (Japanese) and L2 (English) during a forced-choice judgment task. Across languages, attitudinal and linguistic prosody engaged partially dissociable networks: attitudinal prosody recruited socio-cognitive regions involved in inferring speakers' intentions, whereas linguistic prosody engaged phonological-motor regions. Critically, L2 attitudinal prosody elicited distinct frontal modulation, with the left inferior frontal gyrus and middle frontal gyrus showing stronger prosody-type differentiation in English than in Japanese - indicating greater reliance on controlled interpretive and executive processes during non-native attitudinal prosody comprehension. Individual-difference analyses revealed that informal L2 exposure predicted enhanced activation in the thalamus and left hippocampus, as well as better attitudinal prosody identification. These converging neural and behavioral patterns suggest that socially grounded experience plays an important role in developing sensitivity to attitudinal prosody in an L2. Together, these findings provide novel neural evidence for how L2 learners interpret attitudinal prosody and show that L2 exposure is associated with differences in pragmatic prosody at cognitive and neural levels.
Avoidable research waste - that is, research that is unnecessary, poorly designed, or insufficiently communicated - limits the value of scholarship. Building on our earlier work, where we drew lessons from healthcare research to describe five sources of avoidable waste in applied linguistics research, this article focuses on the fifth source: quality of research reporting. Inadequate reporting limits understanding, constrains evidence-informed practice, undermines efforts to replicate or scrutinize empirical claims, and impedes research synthesis. Drawing on precedents from healthcare, particularly the development and widespread adoption of the Consolidated Standards of Reporting Trials reporting guideline and the work of the Enhancing the Quality and Transparency Of Health Research Network, we illustrate how coordinated, consensus-driven reporting guidelines can improve the transparency, completeness, and usability of published research. A survey of instructions to authors across leading applied linguistics journals reveals fragmented and uneven guidance. While promising examples exist, field-wide reporting standards remain absent. We argue that applied linguistics is well positioned to develop design-specific, applied linguistics-focused reporting guidelines through a collaborative, international process involving methodologists, editors, research synthesists, and practitioners. Such an initiative would represent a critical step toward reducing research waste and enhancing the usability of applied linguistics research.
"Conversational" technologies, products, and services are in the headlines more than ever. But what does it mean to be "conversational?" We address this question through the lens of six decades of empirical research in conversation analysis, which has identified and described the foundational structure and interactional machinery of human sociality. We consider not only how tacit notions of "conversationality" manifest in technologies, products, and services, such as role-play, communication training, and chatbots, but also in research methodologies such as focus groups, semi-structured interviews, and laboratory studies - all of which rarely acknowledge how researcher-participant interaction shapes the data collected. Drawing on a range of examples from different institutional settings, we consider how and whether such technologies can or should leverage "conversation" in ways that reproduce "naturalistic" interactions - and ask what might count as "naturalistic" in this context anyway? We argue that if human conversationality becomes a benchmark, then humans themselves will fail tests derived from normative, not empirical, understandings of how social interaction works.
Generative AI (GenAI) offers potential for English language teaching (ELT), but it has pedagogical limitations in multilingual contexts, often generating standard English forms rather than reflecting the pluralistic usage that represents diverse sociolinguistic realities. In response to mixed results in existing research, this study examines how ChatGPT, a text-based generative AI tool powered by a large language model (LLM), is used in ELT from a Global Englishes (GE) perspective. Using the Design and Development Research approach, we tested three ChatGPT models: Basic (single-step prompts); Refined 1 (multi-step prompting); and Refined 2 (GE-oriented corpora with advanced prompt engineering). Thematic analysis showed that Refined Model 1 provided limited improvements over Basic Model, while Refined Model 2 demonstrated significant gains, offering additional affordances in GE-informed evaluation and ELF communication, despite some limitations (e.g., defaulting to NES norms and lacking tailored GE feedback). The findings highlight the importance of using authentic data to enhance the contextual relevance of GenAI outputs for GE language teaching (GELT). Pedagogical implications include GenAI-teacher collaboration, teacher professional development, and educators' agentive role in orchestrating diverse resources alongside GenAI.
Many language assessments - particularly those considered high-stakes - have the potential to significantly impact a person's educational, employment and social opportunities, and should therefore be subject to ethical and regulatory considerations regarding their use of artificial intelligence (AI) in test design, development, delivery, and scoring. It is timely and crucial that the community of language assessment practitioners develop a comprehensive set of principles that can ensure ethical practices in their domain of practice as part of a commitment to relational accountability. In this chapter, we contextualize the debate on ethical AI in L2 assessment within global policy documents, and identify a comprehensive set of principles and considerations which pave the way for a shared discourse to underpin an ethical approach to the use of AI in language assessment. Critically, we advocate for an "ethical-by-design" approach in language assessment that promotes core ethical values, balances inherent tensions, mitigates associated risks, and promotes ethical practices.
To study the potential of generative AI for generating high-quality input texts for a reading comprehension task on specific CEFR levels in German, we investigated the comparability of reading texts from a high-stakes German exam used as benchmarks for the purpose of this study and those generated by ChatGPT (3.5 and 4). These three types of texts were analyzed according to a variety of linguistic features and evaluated by three assessment experts. Our findings indicate that AI-generated texts provide a valuable starting point for the production of test materials, but they require adjustments to align with benchmark texts. Computational analysis and expert evaluations identified key discrepancies that necessitate careful control of certain textual features. Specifically, modifications are needed to address the frequency of nominalizations, lexical density, the use of technical vocabulary, and non-idiomatic expressions that are direct translations from English. To enhance comparability with benchmark texts, it is essential to incorporate features such as examples illustrating the discussed phenomena and the use of passive constructions in the AI-generated content. We discuss the consequences of the usage of ChatGPT for input text generation and point out important aspects to consider when using generated texts as input materials in assessment tasks. Die Erstellung von Pr & uuml;fungstexten f & uuml;r rezeptive Pr & uuml;fungsteile ist ein zeitaufwendiger und ressourcenintensiver Prozess. Um das Potenzial generativer KI f & uuml;r die Erstellung von Input-Texten f & uuml;r eine high-stakes-Pr & uuml;fung f & uuml;r Deutsch als Fremdsprache zu ermitteln, haben wir Lesepassagen, die von geschulten Autor*innen erstellt wurden, mit ChatGPT-generierten Texten verglichen. Die Texte wurden in Hinblick auf eine Reihe von linguistischen Merkmalen mittels einer computerbasierten Analyse ausgewertet und von drei Testerstellungsexpertinnen beurteilt. Unsere Ergebnisse zeigen, dass KI-generierte Texte einen wertvollen Ausgangspunkt f & uuml;r die Erstellung von Pr & uuml;fungstexten bieten, aber Anpassungen erforderlich sind, um eine Vergleichbarkeit mit den Texten der Autor*innen zu erreichen. Insbesondere sind Modifikationen hinsichtlich der folgenden Aspekte notwendig: Veranschaulichung der dargestellten Inhalte durch Beispiele, lexikalische Dichte, Gebrauch von Fachvokabular, Idiomatik, Nominalisierungen und Passivkonstruktionen. Abschlie ss end diskutieren wir die Konsequenzen der Nutzung von ChatGPT zur Erstellung von Input-Texten.
The emergence of ChatGPT as a leading artificial intelligence language model developed by OpenAI has sparked substantial interest in the field of applied linguistics, due to its extraordinary capabilities in natural language processing. Research on its use in service of language learning and teaching is on the horizon and is anticipated to grow rapidly. In this review article, we purport to capture its nascency, drawing on a literature corpus of 71 papers of a variety of genres - empirical studies, reviews, position papers, and commentaries. Our narrative review takes stock of current research on ChatGPT's application in foreign language learning and teaching, uncovers both conceptual and methodological gaps, and identifies directions for future research.
This empirical study explores three aspects of engagement (affective, behavioral, and cognitive) in language learning within an English as a Foreign Language context in Japan, examining their relationship with AI utilization. Previous research has demonstrated that motivation positively influences AI usage. This study expands on that by connecting motivation with engagement, where AI usage serves as an intermediary construct. A total of 174 students participated in the study. Throughout the semester, they were required to use Generative AI (GenAI) to receive feedback on their writing. To prevent overreliance or plagiarism, carefully crafted prompts were selected. Students were tasked with collaboratively constructing essays during the semester using GenAI. At the end of the semester, students completed a survey measuring their motivation and engagement. Structural Equation Modeling was employed to reaffirm the previous finding that motivation influences AI usage. The results showed that AI usage impacts all three aspects of engagement. Based on these findings, the study suggests the pedagogical feasibility of implementing GenAI in writing classes with proper teacher guidance. Rather than being a threat, the use of this technological tool complements the role of human teachers and supports learning engagement.
Since its inception, the Common European Framework of Reference (CEFR) has become increasingly influential in the field of second language (L2) education. In an effort to define the grammatical structures that English learners acquire at each CEFR level, the English Grammar Profile (EGP) provides a list of over 1,200 structure-level mappings derived from largely manual analysis of learner corpora. Though highly valuable for the design of didactic materials and examinations, the EGP lacks comprehensive quantitative methods to verify the acquisition levels it proposes for the grammatical structures. This paper presents an approach for revisiting the EGP structure-level mappings with empirical statistics. The approach utilizes automatic grammatical construction extraction, a large learner corpus, and statistical testing to empirically determine the level of each structure. The structure-level mappings resulting from our approach show limited agreement with that of the original EGP proposals, suggesting that frequency data alone does not provide enough evidence for the acquisition of the grammatical structures at the levels presented by the EGP.
As social and educational landscapes continue to change, especially around issues of inclusivity, there is an urgent need to reexamine how individuals from diverse linguistic backgrounds are perceived. Speakers are often misjudged due to listeners’ stereotypes about their social identities, resulting in biased language judgments that can limit educational and professional opportunities. Much research has demonstrated listeners’ biases toward L2-accented speech, i.e., perceiving accented utterances as less credible, less grammatical, or less acceptable for certain professional positions, due to their bias and stereotyping issues. Then, artificial intelligence (AI) technology has emerged as a viable alternative to mitigate listeners’ biased judgments. It serves as a tool for assessing L2-accented speech as well as establishing intelligibility thresholds for accented speech. It is also used to assess characteristics such as gender, age, and mood in AI facial-analysis systems. However, these AI systems or current technologies still may hold racial or accent biases. Accordingly, the current paper will discuss both human listeners’ and AI’ bias issues toward L2 speech, illustrating such phenomena in various contexts. It concludes with specific recommendations and future directions for research and pedagogical practices.
This study examined the capacity of ChatGPT-4 to assess L2 writing in an accurate, specific, and relevant way. Based on 35 argumentative essays written by upper-intermediate L2 writers in higher education, we evaluated ChatGPT-4’s assessment capacity across four L2 writing dimensions: (1) Task Response, (2) Coherence and Cohesion, (3) Lexical Resource, and (4) Grammatical Range and Accuracy. The main findings were (a) ChatGPT-4 was exceptionally accurate in identifying the issues across the four dimensions; (b) ChatGPT-4 demonstrated more variability in feedback specificity, with more specific feedback in Grammatical Range and Accuracy and Lexical Resource, but more general feedback in Task Response and Coherence and Cohesion; and (c) ChatGPT-4’s feedback was highly relevant to the criteria in the Task Response and Coherence and Cohesion dimensions, but it occasionally misclassified errors in the Grammatical Range and Accuracy and Lexical Resource dimensions. Our findings contribute to a better understanding of ChatGPT-4 as an assessment tool, informing future research and practical applications in L2 writing assessment.
This paper explores the transformative potential of artificial intelligence (AI), particularly generative AI (GenAI), in supporting the teaching, learning, and assessment of second language (L2) listening and speaking. It examines how AI technologies, such as spoken dialogue systems and intelligent personal assistants, can refine existing practices, offer innovative solutions, and address challenges related to spoken language competencies, as well as drawbacks they present. It highlights the role of GenAI, explores its capabilities and limitations, and offers insights into the evolving role of GenAI in language education. This paper discusses actionable insights for educators and researchers, outlining practical considerations and future research directions for optimizing GenAI integration in the learning and assessment of listening and speaking.
Recent developments in artificial intelligence (AI) in general, and Generative AI (GenAI) in particular, have brought about changes across the academy. In applied linguistics, a growing body of work is emerging dedicated to testing and evaluating the use of AI in a range of subfields, spanning language education, sociolinguistics, translation studies, corpus linguistics, and discourse studies, inter alia . This paper explores the impact of AI on applied linguistics, reflecting on the alignment of contemporary AI research with the epistemological, ontological, and ethical traditions of applied linguistics. Through this critical appraisal, we identify areas of misalignment regarding perspectives on knowing, being, and evaluating research practices. The question of alignment guides our discussion as we address the potential affordances of AI and GenAI for applied linguistics as well as some of the challenges that we face when employing AI and GenAI as part of applied linguistics research processes. The goal of this paper is to attempt to align perspectives in these disparate fields and forge a fruitful way ahead for further critical interrogation and integration of AI and GenAI into applied linguistics.
Against the proliferation of large language model (LLM) based Artificial Intelligence (AI) products such as ChatGPT and Gemini, and their increasing use in professional communication training, researchers, including applied linguists, have cautioned that these products (re)produce cultural stereotypes due to their training data. However, there is a limited understanding of how humans navigate the assumptions and biases present in the responses of these LLM-powered systems and the role humans play in perpetuating stereotypes during interactions with LLMs. In this article, we use Sequential-Categorial Analysis, which combines Conversation Analysis and Membership Categorization Analysis, to analyze simulated interactions between a human physiotherapist and three LLM-powered chatbot patients of Chinese, Australian, and Indian cultural backgrounds. Coupled with analysis of information elicited from LLM chatbots and the human physiotherapist after each interaction, we demonstrate that users of LLM-powered systems are highly susceptible to becoming interactionally entrenched in culturally essentialized narratives. We use the concepts of interactional instinct and interactional entrenchment to argue that whilst human–AI interaction may be instinctively prosocial, LLM users need to develop Critical Interactional Competence for human–AI interaction through appropriate and targeted training and intervention, especially when LLM-powered tools are used in professional communication training programs.