Natural Language Processing (NLP) metrics for bilingual code-switching (CS) have, until now, used words as the token level. However, the assumption that any two words constitute an equally likely switch point is erroneous. In spoken language, a major delimiter of CS is a prosodic chunk known as the Intonation Unit (IU). Switch points are far more likely between words at IU boundaries than between words in the same IU. The word as an elementary NLP unit is thus incommensurate with bilingual speech patterns. Here, we put forward an IU-based adaptation of a familiar metric of CS probability. We then compare the token levels on this metric for ten bilingual datasets featuring multi-word CS. Our comparison shows that the currently standard two-significant-figure precision of the word-based metric is insufficient, as the token level compresses the range of values by inflating the universe of CS. More discerning CS probability values can be obtained by normalizing word-based counts using mean IU length.
Objectives: Bringing linguistic experience into code-switching (CS) constraints, a new hypothesis considers cross-language variable equivalence, which arises from within-language variability. Bilingual choices are assessed for Spanish-English CS between clauses, where subordinating conjunctions may not be consistently equivalent.Methodology: Equivalence exists at the main-and-adverbial clause junction, inasmuch as the conjunctions are consistently present and placed the same way in the two languages. Equivalence is variable with main-and-complement clauses, because English complementizer that is mostly absent. Tokens of clause combining were extracted from the prosodically transcribed speech of members of a long-standing community in northern New Mexico who use both languages in their everyday interactions. Bilingual clause combinations were compared with their unilingual counterparts produced by the same speakers, as benchmarks.Data and Analysis: Over 2,000 tokens of clause combining were coded for conjunction, subordinate clause type, prosodic connection, and CS direction for bilingual instances (n = 189).Findings: Bilinguals treat CS with complement and adverbial clauses differently. With complement clauses, the rate of CS is lower, prosodic separation is greater and, most notably, conjunction language choice is more asymmetrical. Spanish complementizer que is overwhelmingly selected over English that. In contrast, choice between causal conjunctions porque and (be)cause is affected by CS direction.Originality: The Variable Equivalence hypothesis states that bilinguals favor CS with the equivalent option from one of the languages that is more frequent and predictable in their combined linguistic experience, considering both languages.Significance: CS constraints are probabilistic (preferred CS sites) rather than categorical (permissible CS sites). The Variable Equivalence hypothesis accommodates variation in actual language use. Methodologically, comparing spontaneous CS with the same speakers' unilingual production allows discovery of CS asymmetries. These asymmetries reveal quantitative bilingual preferences to switch at particular sites.
Code-switching (CS) metrics in NLP that are based on word-level units are misaligned with true bilingual CS behavior. Crucially, CS is not equally likely between any two words, but follows syntactic and prosodic rules. We adapt two metrics, multilinguality and CS probability, and apply them to transcribed bilingual speech, for the first time putting forward Intonation Units (IUs) – prosodic speech segments – as basic tokens for NLP tasks. In addition, we calculate these two metrics separately for distinct mixing types: alternating-language multi-word strings and single-word incorporations from one language into another. Results indicate that individual differences according to the two CS metrics are independent. However, there is a shared tendency among bilinguals for multi-word CS to occur across, rather than within, IU boundaries. That is, bilinguals tend to prosodically separate their two languages. This constraint is blurred when metric calculations do not distinguish multi-word and single-word items. These results call for a reconsideration of units of analysis in future development of CS datasets for NLP tasks.
What is simplification, when may it occur in language contact and does it especially affect discourse-pragmatic aspects? In this chapter we assess parallel but differently variable structures across the languages in contact, in a bilingual speech corpus allowing comparisons of both bilinguals' languages. Spanish and English main-and-complement clauses are analogous but the locus of the variation differs across the languages. There is no corresponding variability in the other language when subjunctive is chosen over indicative in Spanish (variable subjunctive selection) or when presence over absence of the complementizer is chosen in English (variable complementizer presence). Overall rate may be an equivocal measure of contact-induced change, here masking productivity of the subjunctive, as shown by the range of subjunctive-licensing main verbs. Instead, comparisons can rely on the linguistic conditioning of variation, including contextual constraints operationalizing discourse-pragmatic factors, such as grammatical polarity for the Spanish subjunctive and subject form for the English complementizer. Bilinguals' Spanish and English each align with their respective monolingual speech benchmarks. Thus, in the northern New Mexico bilingual community, active bilinguals, who regularly use both languages, display continuity, rather than change, independently in each.
The widespread occurrence of nouns in one language with a determiner in the other, often referred to as mixed NPs, has generated much theorizing and debate. Since both a syntactic account based on abstract features of the determiner and an account highlighting the notion of a Matrix language yield largely the same predictions, we assess how the tenets of each play out in speaker choices. The data derive from a massive corpus of spontaneous nominal mixes, produced by bilinguals in New Mexico, where bidirectional code-switching is the norm. Bilinguals' choices concern (1) NP status (mixed vs. unmixed); (2) mixing type (limited-item vs. multi-word); and (3) language of the noun (here, English vs. Spanish). Results show that the community preference is for mixed NPs, independent of their theoretical felicity as dictated by determiner language properties. As to mixing type, these NPs are mostly constituted of lone nouns, such that the language of the determiner and any associated verb is perforce that of the surrounding discourse. Finally, the overwhelming choice is for English lone nouns incorporated into Spanish, and hence for a Spanish determiner. The language of the determiner thus proceeds, not from abstract linguistic properties, but instead from straightforward adherence to bilingual speech community conventions.
AbstractThis study revisits variable subject pronoun expression in Spanish, bringing to bear insights from cross-linguistic patterns of person-number systems. Based on 2259 tokens from two corpora of Mexican Spanish representing distinct social classes, the study focuses solely on first person plural (1pl) subject pronouns, revealing unique aspects of variablenosotrosexpression. Evidence is offered in favor of a more nuanced measure of switch reference for non-singular grammatical persons through an analysis of the local effects of partial co-referentiality. This measure reconciles the large body of work on switch reference with Cameron’s (1995) measures of reference chains. Additionally, topic persistence — heretofore neglected in prior studies — conditions the variation, with subsequent mentions in the thematic paragraph favoring expressed pronouns. An investigation of clusivity demonstrates that Spanish 1pl subject pronoun expression is sensitive to a distinction grammaticalized in other languages, though subject pronoun rates across clusivities differ from previous results from Peninsular Spanish (Posio 2012). Finally, while subject pronoun expression is generally not sensitive to social factors, distributions of 1pl subjects according to clusivity differ between corpora. Results reveal the style and topic-conditioned differences in contextual distributions that underlie apparent social class differences in subject expression constraints.
Building on studies seeking to position the Romance languages on the cline of grammaticalization, this study targets the evolution of subjunctive into subordination marker in speech corpora of French, Italian, Portuguese and Spanish. By considering the conditioning of variation between subjunctive and indicative in complement clauses, we operationalize parameters of late-stage grammaticalization, and establish measures of productivity. Results show that, with the exception of Spanish, subjunctive selection is constrained neither by contextual elements consistent with its oft-ascribed meanings nor by semantic classes of governors harmonic with such meanings. Instead, in all four languages, lexical bias is the major predictor of subjunctive selection, abetted by structural elements of the linguistic context. The overriding processes are lexical routinization, which is language-particular, with cognate governors displaying idiosyncratic associations with the subjunctive, and structural conventionalization, which is cross-linguistically parallel, with languages differing merely in degree.
The productivity of the Spanish subjunctive in complement clauses is assessed in diachronic perspective, through comparison of variation patterns in three pre-modern texts and a present-day speech corpus. Evidence for the earlier use of the subjunctive as a subordination marker is that it was favored over the indicative in the absence of the complementizer que. Productivity as measured by main-clause verb type frequency remains mostly unchanged, with one expanding niche: a broad class of expressions anchored in ser (e.g., es bueno). A second measure of productivity is the greater favoring of the subjunctive by non-frequent than by frequent main-clause verbs. Nevertheless, a tendency toward a bimodal distribution of main-clause verbs according to their subjunctive rate indicates lexical routinization: some hardly ever take the subjunctive, while others do so categorically, accounting for more than half of all subjunctive occurrences. Structural routinization takes the form of co-occurrence patterns with main-clause negative polarity and interrogatives, both of which increasingly favor the subjunctive. In sum, despite pockets of productivity, use of the subjunctive in complement clauses, as in other Romance languages (Poplack et al. 2018) is largely driven, even in pre-modern texts, by the lexical identity of the main-clause verb, abetted by local structural elements.