IntroductionTraditional studies of the population called “heritage speakers” (HS) have treated this group as distinct from other bilingual populations, e.g., simultaneous or late bilinguals (LB), focusing on group differences in the competencies of the first-acquired language or “heritage language”. While several explanations have been proposed for such differences (e.g., incomplete acquisition, attrition, differential processing mechanisms), few have taken into consideration the individual variation that must occur, due to the fluctuation of factors such as exposure and use that characterize all bilinguals. In addition, few studies have used implicit measures, e.g., psychophysiological methods (ERPs; Eye-tracking), that can circumvent confounding variables such as resorting to conscious metalinguistic knowledge.MethodologyThis study uses pupillometry, a method that has only recently been used in psycholinguistic studies of bilingualism, to investigate pupillary responses to three syntactic island constructions in two groups of Spanish/English bilinguals: heritage speakers and late bilinguals. Data were analyzed using generalized additive mixed effects models (GAMMs) and two models were created and compared to one another: one with group (LB/HS) and the other with groups collapsed and current and historical use of Spanish as continuous variables.ResultsResults show that group-based models generally yield conflicting results while models collapsing groups and having usage as a predictor yield consistent ones. In particular, current use predicts sensitivity to L1 ungrammaticality across both HS and LB populations. We conclude that individual variation, as measured by use, is a critical factor tha must be taken into account in the description of the language competencies and processing of heritage and late bilinguals alike.
Polysynthetic languages present a challenge for morphological analysis due to the complexity of their words and the lack of high-quality annotated datasets needed to build and/or evaluate computational models. The contribution of this work is twofold. First, using linguists’ help, we generate and contribute high-quality annotated data for two low-resource polysynthetic languages for two tasks: morphological segmentation and part-of-speech (POS) tagging. Second, we present the results of state-of-the-art unsupervised approaches for these two tasks on Adyghe and Inuktitut. Our findings show that for these polysynthetic languages, using linguistic priors helps the task of morphological segmentation and that using stems rather than words as the core unit of abstraction leads to superior performance on POS tagging.
Unsupervised cross-lingual projection for part-of-speech (POS) tagging relies on the use of parallel data to project POS tags from a source language for which a POS tagger is available onto a target language across word-level alignments. The projected tags then form the basis for learning a POS model for the target language. However, languages with rich morphology often yield sparse word alignments because words corresponding to the same citation form do not align well. We hypothesize that for morphologically complex languages, it is more efficient to use the stem rather than the word as the core unit of abstraction. Our contributions are: 1) we propose an unsupervised stem-based cross-lingual approach for POS tagging for low-resource languages of rich morphology; 2) we further investigate morpheme-level alignment and projection; and 3) we examine whether the use of linguistic priors for morphological segmentation improves POS tagging. We conduct experiments using six source languages and eight morphologically complex target languages of diverse typologies. Our results show that the stem-based approach improves the POS models for all the target languages, with an average relative error reduction of 10.3% in accuracy per target language, and outperforms the word-based approach that operates on three-times more data for about two thirds of the language pairs we consider. Moreover, we show that morpheme-level alignment and projection and the use of linguistic priors for morphological segmentation further improve POS tagging.
With the increasing interest in low-resource languages, unsupervised morphological segmentation has become an active area of research, where approaches based on Adaptor Grammars achieve state-of-the-art results. We demonstrate the power of harnessing linguistic knowledge as priors within Adaptor Grammars in a minimally-supervised learning fashion. We introduce two types of priors: 1) grammar definition, where we design language-specific grammars; and 2) linguist-provided affixes, collected by an expert in the language and seeded into the grammars. We use Japanese and Georgian as respective case studies for the two types of priors and introduce new datasets for these languages, with gold morphological segmentation for evaluation. We show that the use of priors results in error reductions of 8.9% and 34.2 %, respectively, over the equivalent state-of-the-art unsupervised system.
This study examined play interactions of 15 to 48-month-old children ( n = 59) and their caregivers in Lazuri, a UNESCO-rated endangered South-Caucasian ancestral language, and Turkish, a dominant language supplanting Lazuri usage in the community. Child–caregiver dyads played with two toy sets (animal farm and tea set) that provided different contexts for interaction. Participants’ utterances were coded as instances of distinct functional utterance types (e.g. labels, commands, and questions). Our goal was to analyze the influence of play context on child and caregiver functional utterances in order to identify factors that prompted children to use the ancestral language. The animal farm encouraged dyads to engage in labeling objects, while the tea set encouraged use of social language (e.g. comments). The differing contextual affordances led children to produce many more Lazuri labels than expected when playing with the animal farm. Mixed-effects regression analysis indicated that caregivers’ use of Lazuri and their use of labels predicted children’s use of Lazuri, along with child age. Notably, older children were less likely to use the ancestral language than younger children. The children’s strong preference for speaking Turkish highlights the urgency of interventions to ensure language preservation in Laz communities. Interactions that promote labeling may serve as an effective first step in encouraging children’s use of the Lazuri vocabulary.
Interest in the study and preservation of endangered languages has increased in recent decades as indigenous communities face imminent risk of ancestral language (AL) extinction as a consequence of economic and social factors (Grenoble & Whaley, 2006; Nettle & Romaine, 2000). Languages become critically endangered when adults no longer actively communicate with children using the AL and instead rely on a dominant language (DL) (Fishman, 1991). Efforts to study ALs not only help researchers understand linguistic and sociocultural diversity (Evans & Levinson, 2009), but fuel language revitalization projects by engaging community members in language preservation. Laz communities at the eastern end of the Black Sea have experienced intergenerational language shift stemming from industrialization of the regional economy starting from the 1950’s (Hann, 1997). While it is still common to see elders conversing amongst themselves in Lazuri (the AL), concerns about preparing children for school entry, where Turkish (the DL) is the officially sanctioned language, have led parents to forgo usage of Lazuri and converse with children almost exclusively in Turkish. Most schoolteachers working in Laz communities come from other regions of Turkey and do not speak Lazuri. Hence, Laz children must speak Turkish at school; see Figure 1 (left).
This study investigates morphosyntactic restructuring in Heritage Georgian, a highly agglutinative language with polypersonal agreement. Child heritage speakers of Georgian (n = 26, age 3-16) completed a Frog Story narrative task and a lexical proficiency task in Georgian. Heritage speaker narratives were compared to narratives produced by age-matched peers living in Georgia (n = 30, age 5-14) and Georgian children and young adults who moved to the United States during childhood (n = 7, age 9–24). Heritage Georgian speakers produced more instances of non-standard nominal case marking and non-standard verbal subject agreement than their homeland peers. Individual morphosyntactic divergence was predicted by lexical score, but not by oral fluency or age. Patterns of divergence in the nominal domain included overuse of the default case (nominative) as well as over-extension of non-default cases (ergative, dative). In the verbal domain, person agreement was more consistently marked than number. Subject agreement exhibited more divergence from the baseline than object agreement, contrary to previous evidence from similar heritage languages (e.g., Heritage Hindi, Montrul et al., 2012). Results indicate that morphosyntactic production in child Heritage Georgian generally displays the same divergences as adult heritage-language grammars, but language-specific differences also underscore the need for continued documentation of lesser-studied heritage languages.
An eye-tracking experiment in the Visual World Paradigm was conducted to examine the effects of language history on the predictive parsing of sentences containing relative clauses in the first-learned language of fluent bilingual adults. We compared heritage speakers of Spanish (HSs)—who had spent most of their lives immersed in an English-dominant society—to Spanish–English late bilinguals (LBs), who did not begin immersion in an English-dominant society until adulthood. Consistent with studies of monolinguals, the LBs demonstrated a subject/object relative clause processing asymmetry, i.e. a processing advantage during subject relative clauses and a processing disadvantage during object relative clauses. This suggests that the LBs actively predicted the syntactic structure of subject relative clauses, consistent with the active filler hypothesis. The HSs, on the other hand, did not exhibit this processing asymmetry, suggesting less active prediction. We conclude, therefore, that decreased exposure to the first-learned language causes less active prediction in first-language processing, which causes both disadvantages, and interestingly, advantages, in processing speed.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
An abstract is not available for this content so a preview has been provided. Please use the Get access link above for information on how to access this content.
Judith L. Klavans合作论文数Center for Research on Information Access3