This study tests the influence of acoustic cues and non-acoustic contextual factors on listeners' perception of prominence in three languages whose prominence systems differ in the phonological patterning of prominence and in the association of prominence with information structure-English, French and Spanish. Native speakers of each language performed an auditory rating task to mark prominent words in samples of conversational speech under two instructions: with prominence defined in terms of acoustic or meaning-related criteria. Logistic regression models tested the role of task instruction, acoustic cues and non-acoustic contextual factors in predicting binary prominence ratings of individual listeners. In all three languages we find similar effects of prosodic phrase structure and acoustic cues (F0, intensity, phone-rate) on prominence ratings, and differences in the effect of word frequency and instruction. In English, where phrasal prominence is used to convey meaning related to information structure, acoustic and meaning criteria converge on very similar prominence ratings. In French and Spanish, where prominence plays a lesser role in signaling information structure, phrasal prominence is perceived more narrowly on structural and acoustic grounds. Prominence ratings from untrained listeners correspond with ToBI pitch accent labels for each language. Distinctions in ToBI pitch accent status (nuclear, prenuclear, unaccented) are reflected in empirical and model-predicted prominence ratings. In addition, words with a ToBI pitch accent type that is typically associated with contrastive focus are more likely to be rated as prominent in Spanish and English, but no such effect is found for French. These findings are discussed in relation to probabilistic models of prominence production and perception. (C) 2019 The Authors. Published by Elsevier Ltd.
Mixed effects regression models are widely used by language researchers. However, these regressions are implemented with an algorithm which may not converge on a solution. While convergence issues in linear mixed effects models can often be addressed with careful experiment design and model building, logistic mixed effects models introduce the possibility of separation or quasi-separation, which can cause problems for model estimation that result in convergence errors or in unreasonable model estimates. These problems cannot be solved by experiment or model design. In this paper, we discuss (quasi-)separation with the language researcher in mind, explaining what it is, how it causes problems for model estimation, and why it can be expected in linguistic datasets. Using real-linguistic datasets, we then show how Bayesian models can be used to overcome convergence issues introduced by quasi-separation, whereas frequentist approaches fail. On the basis of these demonstrations, we advocate for the adoption of Bayesian models as a practical solution to dealing with convergence issues when modeling binary linguistic data.
Mixed-effects models have emerged as the gold standard of statistical analysis in different sub-fields of linguistics (Baayen, Davidson & Bates, 2008; Johnson, 2009; Barr, et al, 2013; Gries, 2015). One problematic feature of these models is their failure to converge under maximal (or even near-maximal) random effects structures. The lack of convergence is relatively unaddressed in linguistics and when it is addressed has resulted in statistical practices (e.g. Jaeger, 2009; Gries, 2015; Bates, et al, 2015b) that are premised on the idea that non-convergence is an indication that a random effects structure is over-specified (or not parsimonious), the parsimonious convergence hypothesis (PCH). We test the PCH by running simulations in lme4 under two sets of assumptions for both a linear dependent variable and a binary dependent variable in order to assess the rate of non-convergence for both types of mixed effects models when a known maximal effect structure is used to generate the data (i.e. when non-convergence cannot be explained by random effects with zero variance). Under the PCH, lack of convergence is treated as evidence against a more maximal random effects structure, but that result is not upheld with our simulations. We provide an alternative model, fully specified Bayesian models implemented in rstan (Stan Development Team, 2016; Carpenter, et al, in press) that removed the convergence problems almost entirely in simulations of the same conditions. These results indicate that when there is known non-zero variance for all slopes and intercepts, under realistic distributions of data and with moderate to severe imbalance, mixed effects models in lme4 have moderate to high non-convergence rates which can cause linguistic researchers to wrongfully exclude random effect terms.
In some North American English varieties the diphthong /at/ has developed a distinctively higher nucleus before voiceless consonants and also before a flapped /t/. The phenomenon is known as Canadian Raising, as it was first described for Canadian English. We report on variation in the production and perception of this distinction in a group of female and male speakers from the Chicago area. We focus on the context before flapped /t/ and /d/. The production results show that there is a significant difference in the quality of both the nucleus and the off-glide between these two contexts, albeit of a smaller magnitude than the difference observed before word-final voiceless and voiced consonants. In addition, we find a small difference in duration between diphthongs in the two pre-flap contexts. In perception, our subjects were only moderately successful in recognizing words in minimal pairs containing the target diphthong preceding a flap (as in writer vs rider), although with much higher than chance accuracy. A Quadratic Discriminant Analysis model classified the stimuli with substantially greater accuracy than our subjects. We conclude that in this English variety there is a contrast between a higher diphthong [Lambda i] and a lower diphthong [at], but this contrast is only marginal. This study contributes to our understanding of marginal contrasts in production and perception. The understanding of these contrasts has both theoretical and practical relevance. (C) 2017 Elsevier Ltd. All rights reserved.
AbstractIn the Spanish of north-western Spain, word-final /-d/ shows a remarkable variety of phonetic outcomes. Its possible realizations include voiced approximants, voiceless fricatives and voiced and voiceless plosives, in addition to the deletion of the segment. Here we examine this complex pattern of allophony in a corpus of conversational speech, focusing on the effect of the following phonological context. The results show that most commonly /-d/ is either deleted or realized as a voiceless fricative. Voiceless fricatives are found in all phrasal contexts, but with significantly higher frequency before pause than before a vowel, which is consistent with the hypothesis of diachronic extension of the devoicing from the former context to the latter. The devoicing of /-d/ is neutralizing. Voiceless fricative realizations of /-d/ do not differ from those of phonemic /-θ/ either in amount of voicing or in duration. This implies that deletion and devoicing represent two alternative patterns of reduction starting from [ð], since phonemic /-θ/ is not subject to deletion. Whereas the deletion of /-d/ has lexical exceptions, its devoicing does not. Among the majority of /d/-final words, for which deletion is possible, the relative frequency with which they undergo deletion vs. devoicing appears to vary substantially depending on the specific lexical item. That is, both position in phrase and lexical identity probabilistically determine the realization of /-d/. In addition to contributing to our understanding of the synchronic and diachronic phonology of word-final obstruents in Spanish, we consider the extent to which these data, showing variable word-final devoicing, may help us understand the historical evolution of the crosslinguistically common phenomenon of systematic word-final devoicing.
Mixed effects models are widespread in language science because they allow researchers to incorporate participant and item effects into their regression. These models can be robust, useful and statistically valid when used appropriately. However, a mixed effects regression is implemented with an algorithm, which may not converge on a solution. When convergence fails, researchers may be forced to abandon a model that matches their theoretical assumptions in favor of a model that converges. We argue that the current state of the art of simplifying models in response to convergence errors is not based in good statistical practice, and show that this may lead to incorrect conclusions. We propose implementing mixed effects models in a Bayesian framework. We give examples of two studies in which the maximal mixed effects models justified by the design do not converge, but fully specified Bayesian models with weakly informative constraints do converge. We conclude that a Bayesian framework offers a practical--and, critically, a statistically valid--solution to the problem of convergence errors.
Since Bolinger's [1] discovery that pitch cues accentual prominence in English, a tension has arisen between two strategies: equating accent with pitch excursions and relying on perception for identifying accented words. This paper investigates the relation between prominence judgments from untrained listeners and accentual labels produced by trained transcribers. Naive speakers of English, Spanish and French (30 per language) were asked to mark prominent words in excerpts of conversational speech from their native language (between 900-1100 words in each sample). Aggregated prominence scores (P-scores) were compared with experts' ToBI labels for each language. For all three languages, words ToBI-labelled as accented had substantially higher P-scores than unaccented words, and nuclear accents had higher P-scores than prenuclear ones. P-scores also discriminated among several accent types. Predictions from prior research on the relative prominence of accent labels were tested, and findings confirm that English L+H* accents are more likely to be judged as prominent than H* accents, and Spanish L+H* is more likely judged as prominent than L+>H*. However, for French, our prediction that Accentual Phrase-initial Hi is prominence-lending was not confirmed. The results establish the link between tonal accents and perceived prominence in three languages that differ in their use of contrastive prominence at the lexical and phrasal levels.
Nuclear prominence is assigned to a word based on information status in some languages, while its location is fixed at the end of a phrase in others. We test how this difference affects prominence perception, comparing English, Spanish and French, languages that differ in the strength of the link between informational, positional and acoustic prominence. Using the method of Rapid Prosody Transcription, we compare prominence perception in English Spanish and French in relation to phrasal position and word frequency (a correlate of information status), and by directing listeners’ attention to acoustic criteria or to informational (“meaning-based” criteria). Prominence annotations were collected for spontaneous speech excerpts from 30 listeners of each language. Statistical results of mixed-effect regression show that word frequency as an informational factor most strongly influences prominence ratings for English, where prominence is the primary expression of information status. But despite differences in the phrasal location of nuclear prominence among these languages, the structural factor of adjacency to a prosodic boundary uniformly influences prominence perception based on acoustic criteria in all languages. Listeners in all three languages tend to perceive an acoustically-cued structural prominence on the phrase-final word, suggesting the primacy of a structural nuclear prominence in prosodic theory.
Mixed effects models are widespread in language science because they allow researchers to incorporate participant and item effects into their regression. These models can be robust, useful and statistically valid when used appropriately. However, a mixed effects regression is implemented with an algorithm, which may not converge on a solution. When convergence fails, researchers may be forced to abandon a model that matches their theoretical assumptions in favor of a model that converges. We argue that the current state of the art of simplifying models in response to convergence errors is not based in good statistical practice, and show that this may lead to incorrect conclusions. We propose implementing mixed effects models in a Bayesian framework. We give examples of two studies in which the maximal mixed effects models justified by the design do not converge, but fully specified Bayesian models with weakly informative constraints do converge. We conclude that a Bayesian framework offers a practical--and, critically, a statistically valid--solution to the problem of convergence errors.
We investigate the prominence of English words with stress reversal (e.g. èlevátion 2-1 → élevàtion 1-2). We ask what motivates the occurrence of the “early high” (1-2) pattern outside of stress clash contexts, and consider the hypothesis that it marks prominence non-locally. Experiment 1 tests the effect of prominence pattern on memory. Given its markedness and location at phrasal onset, we hypothesize that early high pitch broadly facilitates recall for sentence information. This hypothesis is not confirmed, suggesting that the effect of pitch accent on memory may be restricted to the accented word. In Experiment 2 listeners perform a prominence-rating task on the same patterns. Results show that early high is prominence-lending, but with weaker prominence than the lexical (2-1) stress pattern. The combined findings suggest a hybrid function for early high in marking the beginning of a discourse-level prosodic unit, and in lending prominence to the early high-accented word.
Central Catalan ‘prepalatal’ (postalveolar) consonants show a complex phonological distribution. Whereas in word-internal intervocalic position a four-way opposition obtains, involving a contrast in voice and a fricative/affricate distinction, elsewhere at least one of the two oppositions is neutralized. Position in word determines whether affrication and/or voicing is contrastive. We study the effect of this factor as well as other phonetic factors and style on the allophony of the voiced prepalatals in a large corpus of Central Catalan. The most significant conditioning factor turns out to be the preceding context, whereas position in word per se is not significant either for degree of constriction or for voicing. Thus, we do not find a direct effect of phonological contrastiveness on phonetic variation.
The “fraction of locally unvoiced frames” measure in Praat’s Voice Report (VR) is an automated method of obtaining the percentage of a segment which is voiced, but its accuracy has been called into question due to values that change based on scrolling and zooming in Praat’s viewing window and don’t always match manual voicing segmentation. This study offers statistical support for the accuracy of VR when certain guidelines are followed: (1) use the object window; (2) decrease the time step to increase temporal resolution; and (3) use gender-specific pitch ranges. The closure and frication portions of 277 affricates were analyzed using VR in this way and the results were compared to manual voicing segmentation using paired Wilcoxon tests. The results show that there is no significant difference between VR and manual segmentation, regardless of whether only the closure portion, only the frication portion, or the entire affricate is considered.