Combining move analysis with multidimensional analysis (MDA), this study investigates salient linguistic co-occurrence patterns ("dimensions") as well as linguistic variations in the rhetorical moves of English research article introductions (RAIs) in applied linguistics written by L1 English and L1 Chinese scholars. Our MDA of a corpus of move segments of RAIs from thirteen applied linguistics journals identifies five functionally interpretable dimensions: grounded evaluation, criticism with caution, stated purposes, author-involved narration of current study versus existing knowledge, and research- versus non-research-related information density. A comparative analysis reveals systematic variation in the linguistic features of specific moves between the RAIs written by L1 English and L1 Chinese scholars. Our findings have useful implications for research and pedagogy of English for academic purposes.
This chapter introduces (web-based) automated text analyzers (ATAs) in applied linguistics research. It begins by briefly surveying key strands of research involving ATAs and outlining three types of text analysis alongside questions commonly addressed by these tools. The core of this chapter presents a conceptual framework for the typology and implementation of ATAs, structured around four continuum dimensions: (1) pre-built corpus platform vs. custom corpus platform , (2) developer-oriented vs. user-oriented , (3) focused vs. versatile , and (4 ) descriptive vs. interpretive . The framework is illustrated through five practical studies showcasing the application of various types of ATAs in applied linguistics research, including L2SCA, Coh-Metrix, Sketch Engine, #LancsBox, Voyant Tools, and Wmatrix3. The chapter also discusses ethical considerations and methodological challenges associated with ATA use. It concludes by outlining future directions for ATA development and research, including improving annotation accuracy, enhancing qualitative interpretability, and expanding analytical capacities across languages and modalities.
Emotional valence and arousal have been shown to influence word recognition and memory, yet evidence from second language (L2) contexts remains mixed and task-dependent. The present study examined whether emotional valence and arousal modulate auditory recognition memory for L2 English words encoded in emotionally valenced audiovisual discourse among Chinese EFL learners at an upper-intermediate proficiency level. Stimulus words were drawn from authentic news videos and pre-rated for emotional and lexical properties. Participants first encountered target words embedded in audiovisual news clips and were later tested on isolated auditory tokens, such that the task indexed recognition memory rather than real-time auditory lexical access. Reaction times and accuracy were analyzed using trial-level generalized additive mixed models, allowing for nonlinear effects while controlling for lexical variables (frequency or familiarity, number of syllables) and random variability across participants and items. Across analyses, emotional effects were not robust. Neither valence nor arousal reliably predicted recognition accuracy. A small valence effect on reaction time emerged only in frequency-controlled models including both correct and incorrect responses, and disappeared when analyses were restricted to correct trials. In contrast, lexical frequency consistently facilitated responses, and substantial variance was attributable to individual- and item-level differences. Overall, emotional effects in L2 auditory word recognition appear weak, conditional, and sensitive to analytical choices.
This study examined the effects of time, genre, and their interaction on large- and fine-grained syntactic complexity (SC) measures in a set of argumentative and narrative essays produced by Chinese English-as-a-foreign-language (EFL) learners at two different proficiency levels over the course of one semester. The goal of the study was to partially replicate and extend Yoon and Polio’s (2017) and Lei et al.’s (2023) studies on these effects, both of which focused on a single group of adult or college-level second language (L2) English learners and employed a set of large-grained indices to gauge SC. This was done to test the generalizability of their findings regarding the effects of time, genre, and their interaction on L2 writing SC to diverse learner groups and obtain a more comprehensive understanding of such effects. Our results revealed significant, differential effects of time, genre, and their interaction on large- and fine-grained SC indices in the two learner groups. These findings contribute to a more nuanced understanding of L2 writing SC development and the effect of genre on this development.
Lexical sophistication has garnered attention across diverse research domains in which language production and text complexity are relevant areas of study. Nevertheless, among the myriad existing lexical sophistication measures, the vast majority do not systematically differentiate different senses of polysemous words but rather treat all senses of a polysemous word as equally sophisticated. To address this limitation, the current study introduces a system that automatically assigns the words in a text to CEFR (i.e., the Common European Framework of Reference for Languages) levels based on their senses used in context, using the English Vocabulary Profile as a reference. We further propose a set of fine-grained sense-aware lexical sophistication indices based on the CEFR levels of word senses and evaluate the extent to which these indices can predict holistic scores of second language (L2) English writing quality using 1,236 exam scripts from the CLC-FCE dataset (Yannakoudakis et al., 2011). The results show that these fine-grained sense-aware indices are more strongly correlated with scores than existing lexical sophistication measures, with three significant predictors explaining 11.8% of the variance in holistic scores. A regression model that combines the new indices with existing ones achieves substantially greater predictive power than models built with either set of indices alone. We discuss the potential implications of our findings for future research in L2 lexical sophistication.
Selecting appropriate texts for second language (L2) learners is essential for effective education. However, current text difficulty models often inadequately classify materials for L2 learners by proficiency levels. This study addresses this deficiency by employing the Common European Framework of Reference for Languages (CEFR) as its foundational framework. A cohort of expert English-L2 educators classified 1,181 texts from the CommonLit Ease of Readability corpus into CEFR levels. A random forest model was then trained using 24 linguistic complexity features to predict the CEFR levels of English texts for L2 learners. The model achieved 62.6% exact-level accuracy across the six granular CEFR levels and 82.6% across the three overarching levels, outperforming a baseline model based on three existing readability formulas. Additionally, it identified shared and unique linguistic features across different CEFR levels, highlighting the necessity to adjust text classification models to accommodate the distinct linguistic profiles of low- and high-proficiency readers.
This study proposes an approach that combines local grammar and construction grammar to accounting for the linguistic realisations of discourse acts, i.e., communicative or rhetorical functions, in academic contexts. It argues that local grammar patterns of discourse acts, representing form-meaning pairings, can be interpreted as constructions, which is illustrated through a case study of exemplification. Using a corpus of linguistics research articles, the study identifies five 'exemplifying' constructions via local grammar analyses, demonstrating that recurrent formmeaning pairings serve as foundational schematic patterns in academic discourse. Along with offering theoretical insights for research on English for academic writing, our findings highlight the need to move beyond discrete lexis and grammar toward a holistic, function-oriented perspective in EAP writing pedagogy. The proposed approach can also be productively extended to other discourse acts beyond exemplification, with potential to enhance both the descriptive adequacy of academic language and the pedagogical effectiveness of EAP writing instruction.
Large language models (LLMs) like OpenAI's GPT models show significant promise in automated writing evaluation (AWE). However, recent research has mainly focused on non-fine-tuned GPT models, with limited attention to fine-tuned models as well as potential factors influencing performance, such as model type, prompting strategy, and dataset characteristics. This study compares six GPT-based approaches for evaluating TOFEL argumentative writing, namely, GPT-3.5 zero-shot, GPT-3.5 few-shot, GPT-4 zero-shot, GPT-4 few-shot, and two fine-tuning methods. We assess the impact of model type (GPT-3.5 vs. GPT-4), prompting strategy (zero-shot vs. few-shot), fine-tuning, class imbalance and dataset shift on performance. Our findings reveal that fine-tuned GPT models consistently outperform non-fine-tuned GPT-4 models, which in turn outperform GPT-3.5 models. Few-shot prompting does not show clear advantages over zero-shot prompting in this study. Additionally, class imbalance and dataset shift negatively affect model accuracy and reliability. These results offer valuable insights into the effectiveness of different GPT-based approaches and the factors that influence their performance in AWE.
The adjustment of syntactic and phraseological complexity is a key consideration in text adaptation. However, research on this topic in the context of Chinese as a second language (CSL) remains limited. Using 700 CSL reading texts graded following the newly issued Chinese Proficiency Grading Standards for International Chinese Language Education , this study examines differences in syntactic and phraseological complexity across texts of varying grade levels using 12 indices and assesses the predictive power of these indices for the grade levels of adapted texts. The results reveal that the 12 indices at the sentence, collocation, and phrase levels significantly differentiate the grade levels of the reading texts, with 11 showing medium to large effect sizes. Whereas all indices exhibit an upward trend overall, their specific patterns of cross‐level changes vary. The strongest predictors of the grade levels of the texts are diversity of total collocations, mean length of sentences, and mean length of noun phrases. We discuss the implications of these findings for establishing syntactic and phraseological complexity benchmarks in CSL teaching materials, adapting CSL learning and assessment texts, and devising effective instructional strategies.
Second language (L2) writing education is undergoing rapid transformation with the increasing integration of artificial intelligence (AI) tools. While research has explored AI applications across diverse educational domains, there remains a lack of systematic reviews addressing the implementation, benefits, challenges, and emerging trends of AI integration in L2 writing education. This review synthesizes 39 empirical studies published between 2019 and 2024, guided by three central questions: (1) What AI tools have been employed, and how have they been integrated into L2 writing education? (2) What benefits and challenges are associated with their use? (3) What emerging trends are shaping the future of AI integration in L2 writing education? Our findings reveal that three major tool types – automated writing evaluation systems, large language models, and specialized writing assistants – are predominantly used to support writing processes, provide feedback, and enhance learner engagement. Most tools were integrated through teacher-guided designs, highlighting the importance of structured mediation. While learners reported gains in grammar accuracy, motivation, and revision ownership, challenges included over-reliance, limited development of higher-order skills, and ethical concerns. Emerging trends point to a shift from surface-level correction toward deeper, metacognitive uses of AI and the embedding of AI in genre-based and reflective pedagogies. However, these developments remain uneven and contingent on instructional design. This review concludes by advocating for theory-informed, equity-oriented integration of AI tools, positioning them not as substitutes for instruction, but as mediational actors within a broader pedagogical ecology.
Research has shown that continuation writing assessment tasks place considerable demands on tester-takers' vocabulary, especially in constructing coherent story plots and employing vivid language. Traditionally, learners have limited access to model continuations, and teacher feedback on vocabulary usage often falls short of guiding learners in selecting and using diverse words appropriate for specific contexts. To address these challenges, this paper introduces a ChatGPTassisted platform designed to facilitate the learning of the meanings, functions, and usage of frequent core words in continuation writing assessment tasks. We further explore the pedagogical possibilities of this platform and discuss its limitations and implications for future research.
Research on task-based second-language (L2) writing has largely focused on investigating how task complexity (TC) affects various aspects of written production. However, this body of research has rarely examined the effect of task repetition, either on its own or in interaction with TC. Additionally, studies have predominantly used traditional frequency-based measures of lexical sophistication, without incorporating more nuanced, sense-aware indices that could offer a deeper understanding of linguistic complexity. This study aims to explore and compare the impact of TC, task repetition, and their interaction on lexical sophistication in L2 writing, using both traditional and sense-aware frequency-based indices. Ninety-six participants completed two argumentative essays on simple and complex tasks in counterbalanced order, twice over four weeks. These essays were analyzed using a set of lexical sophistication indices, both traditional and sense-aware. Repeated-measures MANOVA with TC and time as within-subject variables showed significant main effects of TC and time on overall lexical sophistication, with sense-aware indices being more sensitive and yielding larger effect sizes than traditional word-form-based measures. The theoretical, methodological, and practical implications of our findings are discussed.
Linguistic complexity analysis has played a prominent role in second language (L2) writing research. Such analysis has focused primarily on form-based indices of inherent or relative complexity, often without systematic attention to the effect of the meanings with which linguistic forms are used on their complexity or the rhetorical/pragmatic functions that complex forms are used to convey. This conceptual review article argues for the importance of and delineates the scope of the meaning and function dimensions of linguistic complexity analysis in L2 writing research, reviews the methods and findings of emerging efforts on these dimensions, and discusses how future L2 writing research could attend to these dimensions.
Text simplification is crucial for enhancing student reading achievement. Although Chat Generative Pre-trained Transformer (ChatGPT) and its following generations have exhibited remarkable efficacy in various educational tasks, the similarities and differences between ChatGPT- and expert teacher-simplified texts remain largely unexplored. This study aims to bridge this gap by compiling a comparable corpus consisting of source texts, expert simplified texts, and three sets of ChatGPT-simplified texts generated by typical prompting strategies, namely general, example, and instructive. We then investigated the similarities and differences between sample texts simplified by ChatGPT and those by expert teachers, focusing on 17 linguistic features at the lexical, syntactic, and cohesion levels. The results revealed that significant differences existed between expert- and ChatGPT-simplified texts across multiple linguistic features, while more detailed prompts increased their similarity. These findings have important pedagogical implications, suggesting that with appropriate guidance, teachers can better leverage the potential of ChatGPT for preparing reading materials and make more informed judgments about the value of both ChatGPT- and teacher-simplified texts.
The story continuation writing task (SCWT) has emerged as an effective pedagogical and assessment tool. To date, most studies on linguistic alignment in the SCWT have focused on its learning benefits, with few exploring it from an assessment perspective. The current study addresses this gap by investigating the relationship between linguistic alignment and writing scores in the SCWT to identify the target abilities measured by alignment with the reading passage, which is considered as a unique construct of the SCWT within the theoretical framework of xuargument. Our data consist of 99 writing samples produced for a continuation task by 99 Chinese high-school English as a foreign language learners. Building on previous studies, we measure alignment in three dimensions: lexical, syntactic, and semantic. The study finds that: 1) learners' continuations tend to align closely with the reading passage, primarily through the reuse of content words and syntactic structures and further characterized by a high degree of semantic similarity; 2) higher alignment levels do not necessarily correlate with higher writing quality. Specifically, the overlap of content words, and noun and verb synonyms negatively correlates with writing scores, while semantic similarity has a positive correlation with writing scores. These results suggest that alignment may be perceived differently from pedagogical and assessment perspectives. Furthermore, the construct of linguistic alignment in the SCWT is closely related to the cohesion and creativity of the continued story. The implications of these findings for revising the SCWT rating rubrics are discussed.
Recent studies have shown the potential of GPT models in automated essay scoring (AES) and suggested fine-tuning and integrating linguistic complexity measures as key approaches to enhancing model performance. However, research on the reliability, accuracy and fairness of fine-tuned GPT models in AES across different first language (L1) groups of writers is still limited. Additionally, it is unclear whether combining fine-tuning with linguistic complexity measures can further improve model performance. To address these issues, this study evaluates the performance of a fine-tuned GPT model in rating English essays from four L1 groups of writers. We also examine the impact of integrating linguistic complexity measures by comparing three approaches: GPT ratings only, linguistic complexity indices only, and a combination of both. The fine-tuned GPT model achieved 78.3% exact agreement with human raters overall, though demonstrating lower reliability and accuracy for low-level essays, a minority class in the training data. The model showed a general tendency to overrate essays. Its performance varied across the four L1 groups, with the lowest reliability and accuracy observed for L1 German writers. Among the three approaches, GPT ratings outperformed a set of 22 linguistic complexity measures in evaluating essay quality. Integrating GPT ratings with linguistic complexity measures further resulted in marginal improvement. These findings demonstrate the effectiveness of the fine-tuned GPT approach in AES, highlight the importance of addressing L1-related biases, and suggest directions for future research on GPT-based approaches to AES.