Previous studies have revealed conflicting results as regards whether second language (L2) learners decompose morphologically complex words during online reading or process them in a whole-word manner. To resolve such controversies, first language (L1) morphological characteristics have been proposed as one factor leading to variation in L2 performance. Evidence for the argument of L1 morphological effects includes (1) that L2 learners from an L1 with a similar/congruent morphological feature outperform those from an L1 without a similar corresponding morphological feature and (2) that there is an advantage for L2 learners from an L1 with a more complex morphological system. In this paper, we examine the design of these studies and identify two potentially confounding factors that most studies have failed to consider when comparing L2 learners from different L1 morphological backgrounds: cognateness and L1 writing system. We suggest that studies examining L1 morphological effects recruit L2 groups whose L1s differ in their morphological system characteristics but share the same type of writing system, while also stringently excluding cognates in the design of experimental materials.
Lexical proficiency is a multifaceted phenomenon that greatly impacts human judgments of writing quality. However, the importance of collocations' contribution to proficiency assessment has received less attention than that of single words, despite collocations' essential role in language production. This study, therefore, investigated how aspects of collocational proficiency affect the ratings that examiners give to English learner essays. To do so, collocational features related to sophistication and accuracy were manipulated in a set of argumentative essays. Examiners then rated the texts and provided rationales for their choices. The findings revealed that the use of lower-frequency words significantly and positively impacted the experts' ratings. When used as part of collocations, such words then provided a small yet significant additional boost to ratings. Notably, there was no significant effect for increased collocational accuracy. These findings suggest that low-frequency words within collocations are particularly salient to examiners and deserving of pedagogic focus.
The role of memory in language learning has long been of interest to researchers in first and second language acquisition (SLA) (Baddeley, 1999; Ellis, 2001). At an intuitive level, it seems obvious that part of the explanation for individual differences among adults in success at learning a second language (L2) is attributable to differences in memory capacity. In SLA, researchers have focused on short-term rather than longterm memory differences because they think short-term memory is more responsible for differences in language development. The reason for this belief is that short-term memory is an on-line capacity for processing and analyzing new information (words, grammatical structures and so on); the basic idea is that the bigger the on-line capacity an individual has for new information, the more information will pass into off-line, long-term memory. It is an open question whether low-educated second language and literacy acquisition populations (LESLLA) have short-term memory systems that are similar to literate, educated populations, and if so how their working memory capacity can be measured. This paper will survey the literature on this topic, and will make some suggestions about how models of memory (as they have been applied to second language learning) may and may not be applied to LESLLA contexts. The review is organized as follows. First, different models are presented, along with the principal research results and main areas of disagreement among researchers. Section three deals with working memory and second language acquisition research. Finally, section four addresses how these models may or may not be appropriate to LESSLA contexts.
Previous studies on bilingual children have shown a significant correlation between first language (L1) and second language (L2) morphological awareness and a unique contribution of morphological awareness in one language to reading performance in the other language, suggesting cross-linguistic influence. However, few studies have compared advanced adult L2 learners from L1s of different morphological types or compared native speakers with advanced learners from a morphologically more complex L1 in their target-language morphological awareness. The current study filled this gap by comparing native English speakers (analytic) and two L2 groups from typologically different L1s: Turkish (agglutinative) and Chinese (isolating). Participants' morphological awareness was evaluated via a series of tasks, including derivation, affix-choice word and nonword tasks, morphological relatedness, and a suffix-ordering task. Results showed a significant effect of L1 morphological type on L2 morphological awareness. After accounting for L2 proficiency, the Turkish group significantly outperformed the Chinese group in the derivation, morphological relatedness, and suffix-ordering tasks. More importantly, the Turkish group significantly outperformed the native English group in the morphological relatedness task even without accounting for English proficiency. Such results have implications for theories in second language acquisition regarding representation of the bilingual lexicon. In addition, results of the current study underscored the need to guard against the comparative fallacy and highlighted the influential effect of L1 experience on the acquisition of L2 morphological knowledge.
This report introduces the University of Pittsburgh English Language Institute Corpus (PELIC; Juffs et al., 2020 ), a publicly available 4.2-million-word learner corpus of written texts. Collected over seven years in the University of Pittsburgh’s Intensive English Program, these texts were produced by more than 1,100 students with diverse linguistic backgrounds and proficiency levels. Unlike most learner corpora which are cross-sectional, PELIC is longitudinal, offering greater opportunities for tracking development in a natural classroom setting. This potential is illustrated in an overview of the research conducted to date with these data. The report also provides a description of PELIC’s creation and contents, including how the texts have been managed to facilitate natural language processing. Overall, the corpus contributes to the field of learner corpus research by adding to the pool of freely and publicly available learner corpora, supplemented by a useful set of Python tools and tutorials for accessing these data.
Abstract Vocabulary lists of high-frequency lexical items are an important resource in language education and a key product of corpus research. However, no single vocabulary list will be useful for every learning context, with the appropriateness of such lists affected by the corpora on which they are based. This paper investigates the impact of corpus selection on one measure of lexical sophistication, Advanced Guiraud, focusing on two frequency lists originating from an in-house learner corpus (PELIC) and a global learner corpus (Cambridge Learner Corpus). This analysis shows that frequency lists derived from both types of learner corpus can effectively serve as the basis for measuring the development of lexical sophistication, regardless of the specific program of the learners. Therefore, publicly available learner corpus frequency lists can be a reliable resource for stakeholders interested in the lexical gains of language learners.
This article focuses on the role of crosslinguistic patterns with verbs in the mapping of noun phrases/semantic roles to positions in morphosyntax, with a particular focus on second language (L2) development of Spanish se. The data set derives from high school learners of Spanish in the United States under broadly deductive and inductive learning treatments leading to explicit awareness. Using linear mixed effects modeling (LME) and binomial logistic regression, an analysis of high school learners from three schools (total n = 138) showed that learners based their acceptability judgments of aurally presented sentences and written production on verb classes proposed in formal linguistic theory. However, effects of the instructional intervention were limited to production data. No advantage for either deductive or inductive instruction was identified. The data show a clear role for formal linguistic categories in explaining patterns in the data. Implications for fine-tuning instructional intervention and testing of verb classes are discussed.
Vocabulary lists of high-frequency lexical items are an important resource in language education and a key product of corpus research. However, no single vocabulary list will be useful for every learning context, with the appropriateness of such lists affected by the corpora on which they are based. This paper investigates the impact of corpus selection on one measure of lexical sophistication, Advanced Guiraud, focusing on two frequency lists originating from an in-house learner corpus (PELIC) and a global learner corpus (Cambridge Learner Corpus). This analysis shows that frequency lists derived from both types of learner corpus can effectively serve as the basis for measuring the development of lexical sophistication, regardless of the specific program of the learners. Therefore, publicly available learner corpus frequency lists can be a reliable resource for stakeholders interested in the lexical gains of language learners.
This article focuses on the role of crosslinguistic patterns with verbs in the mapping of noun phrases/semantic roles to positions in morphosyntax, with a particular focus on second language (L2) development of Spanish se . The data set derives from high school learners of Spanish in the United States under broadly deductive and inductive learning treatments leading to explicit awareness. Using linear mixed effects modeling (LME) and binomial logistic regression, an analysis of high school learners from three schools (total n = 138) showed that learners based their acceptability judgments of aurally presented sentences and written production on verb classes proposed in formal linguistic theory. However, effects of the instructional intervention were limited to production data. No advantage for either deductive or inductive instruction was identified. The data show a clear role for formal linguistic categories in explaining patterns in the data. Implications for fine-tuning instructional intervention and testing of verb classes are discussed.
Research into vocabulary knowledge often differentiates between breadth (how many words a person knows) and depth (how well the words are known). Both theoretical categories are essential for understanding language learners’ lexical development, but how the different aspects of vocabulary knowledge interconnect has not received the same attention as each individual dimension, especially in terms of productive knowledge. This study analyses lexis from mid-frequency lemmas in the K3–K9 frequency bands from the learner corpus PELIC (The University of Pittsburgh English Language Institute Corpus). Critically for learners, mastery of lexis in this frequency range is essential for achieving the English proficiency required for university study. From these mid-frequency items, a dataset of 7,554 tokens were collected from word families with multiple derivations and manually annotated. The findings showed high rates of collocational and derivational accuracy for the forms learners opted to use. However, compared to expert speaker texts in the Corpus of Contemporary American English (COCA), learners overused the verb forms and underused the noun forms of these lexical items. These patterns provide evidence of the interplay between breadth and depth in learners’ productive vocabulary usage, suggesting that increased lexical depth will naturally lead to greater lexical breadth and vice versa. Pedagogical implications reaffirm the importance of developing learners’ explicit morphological awareness and collocational accuracy. Suggestions for mid-frequency lexical items to prioritize in language learning are also provided, with a view to helping learners achieve academic readiness.
The past 30 years of reading research has confirmed the importance of bottom-up processing. Rather than a psycholinguistic guessing game ( Goodman, 1967 ), reading is dependent on rapid, accurate recognition of written forms. In fluent first language (L1) readers, this is seen in the automatic activation of a word’s phonological form, impacting lexical processing ( Perfetti & Bell, 1991 ; Rayner, Sereno, Lesch & Pollatsek, 1995 ). Although the influence of phonological form is well established, less clear is the extent to which readers are sensitive to the possible pronunciations of a word ( Lesch & Pollatsek, 1998 ), derived from the varying consistency of grapheme-to-phoneme correspondences (GPCs) (e.g., although ‘great’ has only one pronunciation, [ɡɹeɪt], the grapheme within it has multiple possible pronunciations: [i] in [plit] ‘pleat’, [ɛ] in [bɹɛθ] ‘breath’; Parkin, 1982 ). Further, little is known about non-native readers’ sensitivity to such characteristics. Non-native readers process text differently from L1 readers ( Koda & Zehler, 2008 ; McBride-Chang, Bialystok, Chong & Li, 2004 ), with implications for understanding L2 reading comprehension ( Rayner, Chace, Slattery & Ashby, 2006 ). The goal of this study was thus to determine whether native and non-native readers are sensitive to the consistency of a word’s component GPCs during lexical processing and to compare this sensitivity among readers from different L1s.
Languages have formulaic multiword sequences (MWSs) which occur repeatedly in speech and writing (e.g., Nattinger & DeCarrico, 1992; Siyanova-Chanturia & Pellicer-Sanchez, 2018). For learners, then, the production of MWSs is an important element in developing spoken language that is complex, accurate, and fluent. Though the use of MWSs is important for achieving spoken proficiency, it is unclear whether the production of MWSs supports or hinders another aspect of proficiency, lexical variety. This paper is an exploration of the production of MWSs (recurrent trigrams) and the development of lexical variety, found in 2-min speeches (n = 294) from English L2 learners (n = 66) over time in an intensive English program (IEP). Using hierarchical linear modeling and correlation analysis, we found different patterns of development for the two measures. The use of MWSs increased and then decreased while the lexical variety scores slightly decreased and then sharply increased over time in the IEP. Although the impact of MWSs on oral fluency has been studied, this seems to be the first study to consider how MWSs influence lexical variety across development. (c) 2021 Elsevier Ltd. All rights reserved.
Working memory (WM) overall has a positive relationship with second language acquisition (SLA) processes and language testing outcomes. This relationship may be theoretically fraught, however, as researchers investigate how SLA processes and performance tasks can be adjusted to depend less on working memory. Meanwhile, WM assessment procedures have become increasingly sophisticated, but WM is sometimes measured rather simplistically in applied linguistics. Thus, SLA and language testing researchers are asked to consider the latest MW developments from psychology and cognitive science, to report on the reliability and validity of their WM assessments, and to explore newer function-oriented WM tasks.