
While the assumption within theoretical linguistics is that alternation phenomena (e.g., "everyone" vs. "everybody") lead to redundant complexity because they introduce multiple ways of saying the same thing, this assumption has not yet been tested empirically for EFL learners. To fill this gap, we collected thousands of spoken utterances from low-intermediate to advanced learners of English with either Spanish, Chinese or Italian as their L1 background from the Trinity Lancaster Corpus. Using mixed-effects Poisson regression, we analyzed whether the number of alternation phenomena correlates with the relative complexity of utterances, operationalized as being proportional to the number of disfluencies produced. Results indicate that, even for low-intermediate learners of English, alternation contexts do not induce more disfluencies, contrary to commonly held assumptions in theoretical linguistics and in line with similar research on L1 English speakers.
This study explores if and how phraseological use patterns change over a five-year period for 14 learners of second-language (L2) Spanish. This period covers an academic year spent in a target-language environment, followed by a four-year attrition period. In addition to documenting potential change in usage patterns, we examine how peak attainment and continued L2 contact during the attrition period influence phraseological competence. The analysis focuses on one type of word combination, namely noun/adjective pairs, and measures change by looking at the frequency of noun/adjective sequences and the strength of the association between the two words. Results point to stability in phraseological competence, with no significant patterns of attrition being uncovered. These findings are interpreted against the backdrop of the small body of research on L2 lexical and, specifically, phraseological attrition, contributing to what is known about long-term learning trajectories in the lexical domain.
This report presents the Corpus of Secondary School English as a Foreign Language (EFL) Exams (SEEFLEX). In Germany, upper secondary school EFL exams feature recurring tasks targeting diverse text types. The SEEFLEX was developed to investigate how students complete these tasks linguistically and whether they meet the curricular requirements. The corpus contains data from 575 transcribed authentic curriculum-based examinations (1,979 texts, ~625.000 words). The metadata include standardized receptive vocabulary assessments, a cognition scale, the participants’ reading habits, social background, and their language experience and proficiency. Extensive xml mark-up was added to investigate the influence of inter alia source material, structural text features, and selected language mistakes. An online repository provides full-text access as well as ample additional resources, including an interactive Shiny application to investigate register variation in the corpus.
The present study provides a comparative corpus-based analysis of summaries written by three groups: first-language (L1) German writers, second-language (L2) German writers with L1 Dutch, and L2 German writers with other L1s. The aim is to determine whether there are differences in connective use between L1 and L2 writers in summary writing and whether there are L1 Dutch-specific differences. The results show that L2 German writers with non-Dutch L1s use fewer connectives than L1 German writers, whereas L2 German writers with L1 Dutch use more connectives, especially expansion and contingency connectives. In addition, L2 German writers prefer certain connectives (e.g., und (and), weil (because)) and L2 German writers with L1 Dutch aber (but). Overall, this study highlights the importance of (contrastively) analysing summary writing as well as considering under-researched language pairs such as German and Dutch.
Natural language processing (NLP) tools, primarily trained on L1 written English, have achieved remarkable performance, but are rarely used in L2 learner data. This study leverages a rule-based segmenter to automatically segment spoken English discourse by both L1 speakers and learners, presenting novel preparatory data-cleaning steps that combine a state-of-the-art disfluency detector and additional rules to improve segmentation performance. In three successive segmentation tests on data from the Louvain Corpus of Native English Conversation (LOCNEC; De Cock, 2004) and the Louvain International Database of Spoken English Interlanguage (LINDSEI; Gilquin et al. 2010), we achieve an enhanced segmentation performance that is similar for both the L1 and L2 data (.84). Our approach highlights the effectiveness of leveraging existing NLP tools to process disfluent L2 spoken transcripts, facilitating automatic discourse analysis in Learner Corpus Research (LCR). The code for executing our pipeline is publicly available for future research.
This report presents the Corpus of Secondary School English as a Foreign Language (EFL) Exams (SEEFLEX). In Germany, upper secondary school EFL exams feature recurring tasks targeting diverse text types. The SEEFLEX was developed to investigate how students complete these tasks linguistically and whether they meet the curricular requirements. The corpus contains data from 575 transcribed authentic curriculum-based examinations (1,979 texts, similar to 625.000 words). The metadata include standardized receptive vocabulary assessments, a cognition scale, the participants' reading habits, social background, and their language experience and proficiency. Extensive xml mark-up was added to investigate the influence of inter alia source material, structural text features, and selected language mistakes. An online repository provides full-text access as well as ample additional resources, including an interactive Shiny application to investigate register variation in the corpus.
This paper introduces MultiGEC, a dataset for multilingual Grammatical Error Correction (GEC) in twelve European languages: Czech, English, Estonian, German, Greek, Icelandic, Italian, Latvian, Russian, Slovene, Swedish and Ukrainian. MultiGEC distinguishes itself from previous GEC datasets in that it covers several underrepresented languages, which we argue should be included in resources used to train models for Natural Language Processing tasks which, as GEC itself, have implications for Learner Corpus Research and Second Language Acquisition. Aside from multilingualism, the novelty of the MultiGEC dataset is that it consists of full texts — typically learner essays — rather than individual sentences, making it possible to train systems that take a broader context into account. The dataset was built for MultiGEC-2025, the first shared task in multilingual text-level GEC, but it remains accessible after its competitive phase, serving as a resource to train new error correction systems and perform cross-lingual GEC studies.
The present study concerns the effect of lexical complexity on grading of Swedish EFL learners' texts during high-stakes exams. A learner corpus consisting of 142 texts graded by expert raters and 175 texts graded by teachers was analysed to establish if the latter graded in agreement with the former as intended by the Swedish National Agency for Education (SNAE). Four indices of lexical complexity available in TAALED and TAALES were chosen to explore if this is the case. The method includes conducting ordinal regression with interactions to determine the effect of the independent variables on grade and if these variables have the same effect in texts graded by teachers and expert raters. The findings reveal a discrepancy between expert raters and teachers as they appear to consider lexical complexity to a different extent. It was also found that expert raters and teachers graded more in agreement during source-based writing tasks compared to independent writing tasks.
The present study compares the use of adjective intensification in written L2 Italian production in South Tyrolean upper secondary schools with that of young Italian native speakers. By relying on a Diasystematic Construction Grammar approach, it explores the role of learners' L1s, L2 proficiency levels and their linguistic environments as potential variables affecting the use and choice of different intensifying constructions. Results show that a dominant German-speaking linguistic environment is a significant predictor of learners' preferences for a syntactic over a morphological intensification type. Unexpectedly, however, learners of Italian also make heavy use of the intensifying suffix - issimo, an unfamiliar construction in German. Results also show a difference in the diversity of intensification types used by learners compared to native speakers. Learners are limited to the most frequent types and make a very limited use of maximizers, which seem to be a "blind spot".
This study explores the development of L2 phraseological knowledge, focusing on the relationship between L2 proficiency and the use of adjective-noun combinations from the perspective of collocation density and association strength. While a growing body of evidence suggests that more advanced L2 production tends to be characterised by (i) a greater collocation density and (ii) more strongly associated collocations, several studies did not find this trend. The present study draws on the British Council-Lancaster Aptis Corpus and data from four proficiency levels (A2-C of the CEFR) to test a hypothesis about the direction of collocation development in order to contribute to a systematic advancement of knowledge about the development of productive L2 collocation use. The study replicates some of the key findings from previous research, confirming that these trends are generalisable to different samples of L2 speakers and across different modes of communication.
Preview this online first article: Review of Granger (2021): Perspectives on the L2 Phrasicon: The view from learner corpora, Page 1 of 1 < Previous page | Next page > /docserver/preview/fulltext/10.1075/ijlcr.00047.gro/ijlcr.00047.gro-1.gif
Since Paquot (2019), several unresolved issues have persisted regarding the operationalization of phraseological sophistication in L2 complexity research. One of the most crucial concerns relates to the extent to which the commonly used measures of phraseological sophistication (MI scores) fully represent the intended construct. In this study, we draw upon insights from L2 phraseological research to reexamine the conceptualization and operationalization of phraseological sophistication. We conduct new analyses on the learner corpus used in Paquot (2019), using alternative operationalizations of phraseological sophistication that represent different dimensions of sophistication (based on the register specificity of word combinations and their frequency). Results show that measures representing the dimensions of association (MI scores) and register specificity (ratios of academic collocations) correlate with each other. Frequency-based measures, however, pattern very differently, which we attribute to some issues in the way we operationalized frequency of co-occurrence.
The aim of this article is to survey the field of learner corpus research from its origins to the present day and to provide some future perspectives. Key aspects of the field - learner corpus design and collection, learner corpus methodology, statistical analysis, research focus and links with related fields, in particular SLA, FLT and NLP - are compared in first-generation LCR, which extends from the late 1980s to 2000, and second-generation LCR, which covers the period from the early 2000s until today. The survey shows that the field has undergone major theoretical and methodological changes and considerably extended its range of applications. Future developments that are likely to gain ground are grouped into three categories: increased diversity, increased interdisciplinarity and increased automation.
Previous studies of undergraduate writing investigated linguistic variation across (i) assignment types, (ii) disciplines, and (iii) language backgrounds. The combined findings of these studies allowed us to formulate eight hypotheses as to how undergraduate writing is likely to vary across these three variables. Three of the hypotheses are as follows: (a) writing in humanities will have more features of 'academic involvement', while writing in sciences will have more features of 'information density'; (b) assignments such as proposals and procedural recounts will have more features of 'expression of possibility'; and (c) L1 students will use more features of 'information density' than L2 students. In the current study, we test these hypotheses, examining whether the language of undergraduate writing varies in accordance with the expectations from previous research. We use the dimensions identified in Goulart (2024) to examine these hypotheses in a corpus of undergraduate student writing. The results provide support for hypotheses related to disciplines and communicative purposes, but not for those related to language backgrounds.
This paper tests three hypotheses about written vocabulary in child L2 English. Specifically, as children mature, (1) the mean frequency values of the nouns they use increase; (2) the mean frequencies of other parts-of-speech decrease; (3) the use of academic vocabulary increases only in certain types of writing. Using a corpus of writing by children in Norway, hypothesis 1 was confirmed up to the mid-teenage years. The mean frequency values of nouns then decreased. Analysis showed that the early increase is due to decreased repetition of low-frequency topic words. After age 15, frequencies drop as the main source of vocabulary moves from a region around the 150th most frequent lemma to one around the 550th. Hypotheses 2 and 3 were partially confirmed. Mean frequencies of non-nouns decreased in non-stories after Year 9. Non-stories became more academic across school years. Stories had much lower scores overall but also showed an increase at Year 10.
Mean-frequency scores of lexical sophistication are used to evaluate written and spoken language production. They are calculated using word frequencies extracted from a reference corpus. Using mixed-effects regression models, we analyse the strength of the relationship between L2 proficiency and mean-frequency scores in spoken and written texts using reference corpora representing different modes and registers. We control for task and topic effects. We observe that mean-frequency measures of lexical sophistication are considerably more influenced by the mode and register of the reference corpus used to calculate these scores than by language users' proficiency level. Advanced language users produce more frequent vocabulary, typical of the target register, in both spoken monologues and written essays. These results provide evidence in favour of a conceptual and terminological shift from lexical sophistication to register appropriateness (as suggested by Durrant & Brenchley, 2019) to refer to the construct captured by mean-frequency scores of vocabulary use.
The present study aims to verify previous findings on the role of proficiency in the degree of complexity and accuracy of verbal morphemes in written essays from intermediate and advanced L2 Italian learners. In addition, by taking the perspective of usage-based theories on the distributional properties in the input, it investigates the extent to which contingency affects morphological complexity and accuracy by the emergence of cue-outcome associations. Hypotheses on the effects of contingency and proficiency on complexity and accuracy, and on their mutual relationships, are verified by adopting a confirmatory approach and by using Structural equation modeling. Results support the claims for non-universal mechanisms in the acquisition of verbal morphology: proficiency does not affect the morphological complexity of learner texts in the same way in inflectionally rich and poor languages, and the cue-outcome mechanisms of associative learning work to different extents depending on the inflectional systems of the target languages.