Multivariate approaches to modeling lexical richness and second language (L2) writing proficiency have expanded rapidly, yet this work has focused overwhelmingly on L2 English, and questions about how best to operationalize lexical items in languages with richer morphology remain underexplored. This study builds on existing work by introducing human judgments of productive vocabulary and language use scores for a corpus of L2 Spanish, evaluating different lexical item operationalizations, and examining the relationship between Vocabulary and Language Use scores and indices of lexical richness. The results indicated that optimal operationalizations varied across indices, with lemmas that include full verbal information best capturing lexical diversity, and raw words best capturing most bigram association measures. In linear mixed-effects models, lexical diversity was the strongest predictor of Vocabulary and Language Use scores, while a smaller number of bigram association measures were also significant predictors. These findings suggest that it is beneficial to consider using lexical item operationalizations that reflect the morphology of the target language, and that lexical diversity is a strong predictor Vocabulary and Language Use in L2 Spanish. The study concludes by discussing implications for L2 Spanish instruction and the development of automated productive vocabulary and language use assessment tools.
Argument structure constructions (ASCs) offer a theoretically grounded lens for analyzing second language (L2) proficiency, yet scalable and systematic tools for measuring their usage remain limited. This paper introduces the ASC analyzer, a publicly available Python package designed to address this gap. The analyzer automatically tags ASCs and computes 50 indices that capture diversity, proportion, frequency, and ASC-verb lemma association strength. To demonstrate its utility, we conduct both bivariate and multivariate analyses that examine the relationship between ASC-based indices and L2 writing scores.
Ample research has examined the linguistic characteristics of second language (L2) writing across proficiency scores, with a focus on lexical diversity, sophistication, and density as key dimensions of lexical richness. However, the applicability of these indices to young L2 learners' writings, often characterized by limited vocabulary and constrained output in standardized writing tasks, remained underexplored. To address this gap, this study analyzed the lexical richness of young L2 learners' written productions from two TOEFL Junior Writing tasks (Opinion and Listen-Write tasks), using 37 tailored indices of lexical diversity, sophistication, and density. The results indicated that the lexical characteristics of young L2 learners vary by task score and task type, particularly when assessed through indices such as lexical diversity (e.g., moving-average typetoken ratio) and sophistication (e.g., n-gram strength of association). Nonetheless, incorporating additional measures, such as syntactic complexity or discourse features, may be essential for distinguishing young L2 learners at higher proficiency levels.
Analyzing the relationship between argument structure construction (ASC) use and language learning has been an important area of investigation in second language (L2) studies from a usage-based constructionist approach. Previous studies have shown that advanced L2 learners’ language use demonstrates greater ASC diversity, less frequent ASC-verb combinations, and stronger ASC-verb associations. However, these investigations have been limited by methodological challenges in identifying ASCs and have predominantly focused on the written texts. To address these limitations, we employ a fine-tuned model to automatically extract ASCs from target and reference corpora, considering their semantic aspects. We then calculate ASC-based indices and, both alone and in combination with other lexicogrammatical indices, use them to predict L2 oral proficiency scores assigned by human judges. Our findings show that ASC-based indices alone explain 44% of the variance in scores. When combined with other indices, they provide complementary insights that enhance multivariate modeling of L2 oral proficiency.
Research has demonstrated that features of lexical and lexicogrammatical use are important predictors of productive second language (L2) proficiency (e.g. ). While some features of lexical use have been studied with L2s other than English (e.g. ), multivariate lexical and lexicogrammatical approaches in these L2s are rare. In this study, we extend the use of multivariate approaches to L2 Spanish writing. Our learner data included a subset of the CEDEL2 corpus (), comprised of proficiency scores and 644 descriptive essays written in L2 Spanish by L1 English writers. Correlational analyses were conducted between proficiency scores and indices of lexical diversity (e.g. MTLD), mean word and bigram frequencies, and bigram strength of association (MI, delta). A final regression analysis accounted for 48.3 per cent of the variance in proficiency scores. Following previous L2 English writing research (e.g. ; ), more proficient L2 Spanish writers tended to use a wider variety of lexical items, more strongly associated word combinations, and lexical items that are less frequent in corpora.
Lexical diversity (LD) is an important indicator of second language lexical development. Much research has investigated LD indices, with a focus on learners of English. However, further research is needed in languages that are typologically distinct from English, such as Korean. In this study, we evaluated the reliability and validity of LD indices applied to argumentative writing produced by Korean learners. The results indicated that HD-D, MATTR, and MTLD were reliable across different text lengths and were correlated with holistic proficiency scores. However, the meaningful differences were found across Korean-specific tokenization types related to the way morphemes are processed.
The use of natural language processing tools such as part-of-speech taggers and syntactic parsers are increasingly being used in studies of second language (L2) proficiency and development. However, relatively little work has focused on reporting on the accuracy of these tools or optimizing their performance in L2 contexts. While some studies reference the published overall accuracy of a particular tool or include a small-scale accuracy analysis, very few (if any) studies provide a comprehensive account of the performance of taggers and parsers across a range of written and spoken registers. In this study, we provide a large-scale accuracy analysis of popular taggers and parsers across L1 and L2 written and spoken texts, both when default and L2-optimized models are used. Accuracy is examined both at the feature level (e.g., identifying adjective-noun relationships) and the text level (e.g., mean mutualinformation scores). The results highlight the strength and weaknesses of these tools.
This study evaluates the effectiveness of pre-trained language models in identifying argument structure constructions, important for modeling both first and second language learning. We examine three methodologies: (1) supervised training with RoBERTa using a gold-standard ASC treebank, including by-tag accuracy evaluation for sentences from both native and non-native English speakers, (2) prompt-guided annotation with GPT-4, and (3) generating training data through prompts with GPT-4, followed by RoBERTa training. Our findings indicate that RoBERTa trained on gold-standard data shows the best performance. While data generated through GPT-4 enhances training, it does not exceed the benchmarks set by gold-standard data.
Research has indicated that lexical richness is an important indicator of second language (L2) proficiency. However, most research has examined written, cross-sectional English L2 corpora and does not necessarily indicate how spoken lexical use develops over time or whether observed trends are stable across L2s. This study adds to previous research on the development of spoken vocabulary by investigating lexical features of L2 Spanish learners over a 21-month period, using the LANGSNAP corpus. Multiple lexical richness indices used in previous studies were examined including lexical diversity, word frequency, word concreteness, and bigram strength of association. Linear mixed-effects models were run to examine changes over time. The results suggest that although some features of lexical richness (e.g., word frequency) see meaningful change over time, others (e.g., bigram T score) may not be indicative of L2 oral development.
The current tutorial paper describes a process of developing a custom natural language processing model with a particular focus on a discourse annotation task. After an overview of recent developments in natural language processing (NLP), the paper discusses the development of the Engagement Analyzer (Eguchi & Kyle, 2023), focusing on corpus annotation, the machine learning model, model training, evaluation, and dissemination. A step-by-step tutorial of this process via the spaCy Python package is provided. The paper highlights the feasibility of developing custom NLP tools to enhance the scalability and replicability of the annotation of context-sensitive linguistic features in L2 writing research.
Although lexical diversity is often used as a measure of productive proficiency (e.g., as an aspect of lexical complexity) in SLA studies involving oral tasks, relatively little research has been conducted to support the reliability and/or validity of these indices in spoken contexts. Furthermore, SLA researchers commonly use indices of lexical diversity such as Root TTR (Guiraud's index) and D (vocd-D and HD-D) that have been preliminarily shown to lack reliability in spoken L2 contexts and/or have been consistently shown to lack reliability in written L2 contexts. In this study, we empirically evaluate lexical diversity indices with respect to two aspects of reliability (text-length independence and across-task stability) and one aspect of validity (relationship with proficiency scores). The results indicated that neither Root TTR nor D is reliable across different text lengths. However, support for the reliability and validity of optimized versions of MATTR and MTLD was found.
Responding to the increasing need for automated writing evaluation (AWE) systems to assess language use beyond lexis and grammar (Burstein et al., 2016), we introduce a new approach to identify rhetorical features of stance in academic English writing. Drawing on the discourse-analytic framework of engagement in the Appraisal analysis (Martin & White, 2005), we manually annotated 4,688 sentences (126,411 tokens) for eight rhetorical stance categories (e.g., PROCLAIM, ATTRIBUTION) and additional discourse elements. We then report an experiment to train machine learning models to identify and categorize the spans of these stance expressions. The best-performing model (RoBERTa + LSTM) achieved macro-averaged F1 of .7208 in the span identification of stance-taking expressions, slightly outperforming the intercoder reliability estimates before adjudication (F1 = .6629).
General indices of syntactic complexity (e.g. mean length of T-unit, clauses per T-Unit) have long been used to measure the writing proficiency of adult language learners (Bulté & Housen, 2012; Norris & Ortega, 2006; Ortega, 2003; Wolfe-Quintero et al., 1998). In contrast, a number of recent studies have focused on measuring adult second language writing proficiency using methods rooted in usage-based theories of language learning. The present study extends previous research (e.g., Kyle & Crossley, 2017) by comparing usage-based and general indices of syntactic and lexicogrammatical use in learners of Mandarin. It also extends previous work by investigating whether meaningful, but non-linear trends exist. To do this, it first compares multiple linear regression models to polynomial models for general and usage-based indices respectively. Then it compares the strongest model for each type of index to determine which is better. Consistent with Kyle and Crossley (2017), it finds that in Mandarin usage-based indices are better predictors of proficiency than general indices and that the frequency of the verb decreases over time while the strength of association between verb and VAC increases. Furthermore, non-linear trajectories were found to exist in general indices while usage-based indices were linear.
The measurement of second language (L2) productive lexical proficiency has driven a great deal of research over the past two decades. Research has indicated that more proficient speakers and writers tend to use a wider range of words and that more proficient writers tend to use words that are more sophisticated (less frequent in reference corpora). Research over the past 15 years has also demonstrated that the way words are used in context (i.e., collocation use) is also an important indicator of both written and spoken proficiency. In this study, we extend recent research that has modeled writing proficiency using collocation indices based on grammatical dependencies (e.g., verb-direct object) to spoken contexts. In particular, we model speaking proficiency scores from a large corpus of oral proficiency interview responses using a range of well-known indices of productive proficiency and newly developed grammatical dependency indices. The results indicated that all index types demonstrated small to moderate correlations with speaking proficiency individually but explained a large proportion of the variance when used in a multivariate model that included dependency collocation indices.
While lexical diversity measures related to lexical variety are commonly used evaluate an individual’s vocabulary breadth, many other sub-constructs of lexical diversity, such as evenness remain understudied. In this study, we introduce and evaluate six indices of evenness regarding reliability and validity in L2 writing and speaking contexts. Our findings indicate that a selection of these indices is reliable across different text lengths. However, pinpointing the precise contribution of evenness indices in the prediction of L2 proficiency scores is challenging due to overlap between these indices and established indices of variety.
The current study extends extant research on lexical collocation and L2 English proficiency by analyzing how L2 argumentative writings assessed at different proficiency levels differ in their compositions of weakly and strongly associated collocations. Using a Natural Language Processing pipeline, a total of 640 essays from the ICNALE corpus (Ishikawa, 2018) were analyzed for word pairs that are syntactically related (e.g., Verb-Direct object), and the relationships between the relative proportions of collocations with varying strengths of association (SOA) and vocabulary proficiency scores were examined. A series of regression analyses revealed that collocations from various MI score bins showed distinct patterns of use across proficiency levels, indicating, for example, increases in the use of strongly associated collocations and decreases in the use of repelled collocations. The finding also indicated that band-based MI measures demonstrated better predictive validity than mean MI scores in modeling the vocabulary proficiency score.
This chapter explores the relationship between writing and vocabulary learning. The primary focus of the chapter is on the ways in which writing samples have been used to indicate a writer's productive lexical proficiency. The use of writing as a site for vocabulary learning is also discussed. The chapter includes a critical discussion of the measurement of lexical diversity, lexical sophistication, and the use of collocations in learner writing. Historical perspectives, current research contributions, and future directions in these areas are outlined. For those interested in conducting research in this area, the relative advantages and disadvantages of common research designs and analysis tools are also outlined.
In the realm of language proficiency assessments, the domain description inference and the extrapolation inference are key components of a validity argument. Biber et al.'s description of the lexicogrammatical features of the spoken and written registers in the T2K-SWAL corpus has served as support for the TOEFL iBT test's domain description and extrapolation inferences. In the time since the T2K-SWAL corpus was collected, however, university learning environments have increasingly become technology-mediated. Accordingly, any description of the linguistic features of university language should account for the language produced in technology-mediated learning environments (TMLEs) in addition to non-technology-mediated learning environments (non-TMLEs). Kyle et al. recently began to address this issue by collecting a corpus of TMLE language use, which they then compared to language use in non-TMLEs using multidimensional analysis (MDA). The results indicated both similarities and substantive differences across the learning environments, but the study did not investigate the effects of particular registers on these results. In this study, we build on previous research by investigating lexicogrammatical features of specific spoken and written registers across technology-mediated and non-technology-mediated learning environments.