Pronunciation variation in many ways is systematic, yielding patterns that a canny listener can exploit in order to aid perception. This work asks whether listeners actually do draw upon these patterns during speech perception. We focus in particular on a phenomenon known as paradigmatic enhancement, in which suffixes are phonetically enhanced in verbs which are frequent in their inflectional paradigms. In a set of four experiments, we found that listeners do not seem to attend to paradigmatic enhancement patterns. They do, however, attend to the distributional properties of a verb's inflectional paradigm when the experimental task encourages attention to sublexical detail, as is the case with phoneme monitoring (Experiment 1a-b). When tasks require more holistic lexical processing, as with lexical decision (Experiment 2), the effect of paradigmatic probability disappears. If stimuli are presented in full sentences, such that the surrounding context provides richer contextual and semantic information (Experiment 3), even otherwise robust influences like lexical frequency disappear. We propose that these findings are consistent with a perceptual system that is flexible, and devotes processing resources to exploiting only those patterns that provide a sufficient cognitive return on investment.
Functional load (FL) is an information-theoretic measure that captures a phoneme’s contribution to successful word identifi- cation. Experimental findings have shown that it can help ex- plain patterns in perceptual accuracy. Here, we ask whether the relationship between FL and perception has larger conse- quences for the structure of a language’s lexicon. Since re- ducing FL minimizes the risk of misidentifying a word in the case where a listener inaccurately perceives the initial phoneme, we predicted that in spoken language, where perceptual accu- racy is important for successful communication, the lexicon will be structured to reduce FL in auditorily confusable initial phonemes more than in written language. To test this predic- tion, we compared FL of all initial phonemes in spoken and academic written genres of the COCA corpus. We found that FL in phoneme pairs in the spoken corpus is overall higher and more variable than in the academic corpus, a natural conse- quence of the smaller lexical inventory characteristic of spoken language. In auditorily confusable pairs, however, this differ- ence is relatively reduced, such that spoken FL decreases rel- ative to academic FL. We argue that this reflects a pressure in spoken language to use words for which inaccurate perception does minimal damage to word identification.
Variation in the amount of information transmitted by speech elements correlates with (and helps explain) a number of speech patterns, such as phonetic reduction during production and sound change. However, the cross-linguistic effect of information content on speech perception is relatively understudied. This study fills this gap and investigates the relationship between perceptual accuracy and two information measures in English, Korean, and Japanese. We computed the informativity (weighted negative contextual predictability) and functional load (the contribution of a speech element to lexical contrast) of /p/, /t/, and /k/ in sub-syllabic speech chunks in English, Korean, and Japanese sub-lexical corpora. Native speakers of the three languages identified consonant clusters in VC.CV stimuli in a perception experiment. A logistic model predicting perceptual accuracy from informativity and functional load found that informativity is generally negatively correlated with perceptual accuracy, while functional load is positively correlated. Furthermore, the strength of this correlation is different for the three languages studied. We interpret the result as reflecting differences in syllable structure across the three languages.
This paper investigates whether compensation for coarticulation in speech perception can be mediated by native language. Substantial work has studied compensation as a consequence of aspects of general auditory processing or as a consequence of a perceptual gestural recovery processes. The role of linguistic experience in compensation for coarticulation potentially cross-cuts this controversy and may shed light on the phonetic basis of compensation. In Experiment 1, French and English native listeners identified an initial sound from a set of fricative-vowel syllables on a continuum from [s] to [integral] with the vowels [a,u,y]. French speakers are familiar with the round vowel [y], while it is unfamiliar to English speakers. Both groups showed compensation (a shifted 's'/'sh' boundary compared with [a]) for the vowel [u], but only the French-speaking listeners reliably compensated for the vowel [y]. In Experiment 2, 39 American English listeners judged videos in which the audio stimuli of Experiment 1 were used as soundtracks of a face saying [s]V, [integral]V, or a visual-blend of the two fricatives. The study found that videos with [integral] visual information induced significantly more "integral" responses than did those made from visual [s] tokens. However, as in Experiment 1, English-speaking listeners reliably compensated for [u], but not for the unfamiliar vowel [y]. The listeners used visual consonant information for categorization, but did not use visual vowel information for compensation for coarticulation. The results indicate that perceptual compensation for coarticulation is a language specific effect tied to the listener's experience with the conditioning phonetic environment. (C) 2015 Elsevier B.V. All rights reserved.
Words which are probable in their morphological paradigms tend to have lengthened affixes (Cohen, 2014; Kuperman et al., 2007). Here, we ask whether listeners use the pattern to aid perception. If so, paradigmatically probable words with lengthened affixes should be perceived more quickly than similarly lengthened improbable words. In two experiments (Experiment 1: phoneme monitoring; Experiment 2: lexical decision), we measured listeners’ reaction time (RT) to 50 English verbs with [-s] suffixes (e.g., looks, breaks). Each suffix was adjusted in duration to either a normalized proportion of stem duration, shortened by 25% of the normalized duration, or lengthened by 25%. In Experiment 2, paradigmatic probability did not affect RT, but short words had slower RTs than normalized or lengthened words, suggesting that generally reduced suffix duration impedes perception. In Experiment 1, RT decreased with increased probability for the long condition, but not for the normalized condition. This suggests that a match between suffix length and paradigmatic probability facilitates perception. However, RT also decreased in the short condition, suggesting that when the stimulus was most difficult to perceive, listeners drew on general probabilistic information to aid perception. Our results further indicate that the effect of paradigmatic probability on perception is task-dependent.
Listeners can shift their attention to different sizes of speech during speech perception. This study extends this claim and investigates if linguistic structure affects this attention. Since English has a larger syllable inventory than Korean and Japanese, each phoneme plays a larger functional role. Also, listeners have different levels of phonological awareness due to the differences in the orthography. We focus on the effect of perceptual attention on the perceptibility of intervocalic consonant clusters (VCCV) and whether it varies cross-linguistically by these structural factors. We first recorded eight talkers saying VC- and CV-syllables and spliced the syllables to create non-overlapping VC.CV-stimuli. Listeners in three language groups (English/Korean/Japanese) participated in a 9-Alternative-Forced-Choice perception task. They identified the CC as one of 9 alternatives (“pt”, “pk”, “pp”, etc.) and in an attention-manipulated condition did the same task while also monitoring for target talkers. The preliminary result shows that Korean listeners showed less perceptual sensitivity to clusters than English listeners. Also, the English listeners showed better perception of syllable coda when prompted to focus on coda only. The result indicates that the linguistic structure of a language can potentially affect the level of perceptual attention that its users give to a linguistic unit.
We compared the integration of three kinds of contextual information in the perception of the fricatives [s] and [∫]. We asked American English listeners to identify sounds on an [s] to [∫] continuum and manipulated (1) the vowel context of the fricative ([Ce], [Co], [Cœ]), (2) the original fricative of the CV ([s] vs [∫]), and (3) the modality of the stimulus (audio-only, AV). There was a large compensation for coarticulation effect on perception—subjects responded with “s” more often when the following vowel was round. Interestingly, and perhaps significantly, perceptual compensation was not as great with the less familiar vowel [œ] even when listeners saw the face. Measurements of lip rounding in these stimuli show that [o] and [œ] have about the same degree and type of rounding over the CV. In a second experiment, we measured reaction time to audio-visual mismatches in these stimuli (again in a fricative identification task). We found that mismatches of audio and video consonant information slowed reaction time, and that vowel mismatches did as well. However, mismatch between [o] and [œ] did not slow reaction time. These data suggest that linguistic experience and stimulus properties affect perception.
This study is on audio-visual perceptual intelligibility of consonants in intervocalic clusters (VC1C2V). Previous studies have yielded inconsistent findings on perceptual salience of different stop consonants and very few have tested salience in clusters. Consequently, it has been unclear as to whether greater or less perceptual salience leads to greater degree of place assimilation. In Korean, labials are often produced with more gestural overlap than velars in C1. I tested whether labials are perceptually more or less salient in both audio and audio-visual conditions. VC and CV syllables spoken by both English and Korean speakers were first embedded in noise and spliced together for non-overlapping VCCV sequences. Korean listeners identified the two consonants in either audio or AV presentations. A confusion matrix analysis for each stop consonant shows that in C1 there is asymmetric improvement with the addition of videos for labial consonants only, while in C2 this asymmetry was not found. The result suggests that listeners make differential use of visual cues depending on place of articulation and syllabic context. Also, the result supports the talker enhancement view of sound change, which assumes that talkers are aware of perceptual salience and enhance (with less gestural overlap) the weak contrast.
UC Berkeley Phonology Lab Annual Report (2012) Neural basis of the Word Frequency Effect and Its Relation to Lexical Processing Shinae Kang January 21, 2013 Introduction A number of experimental results have shown that in terms of cognitive processing com- mon words – words that occur frequently - differ from uncommon words. People perceive common words more accurately and more quickly when listening (Savin, 1963) and also name them more quickly (Oldfield and Wingfield, 1965) while speaking. Also, people produce the common words in more reduced and shortened forms than uncommon words (Whalen, 1991; Gahl, 2008). These results indicate a certain processing difference between the common (or frequent) words and the uncommon (or infrequent) words. This phenomenon has been extensively studied in the psycholinguistic literature, in particular for speech production, as one way of probing the cognitive architecture. Current models in single word production generally agree that the production process consists of multiple cognitive actions (Dell, 1986; Levelt et al. 1999). Broadly speaking, they include conceptulization, retrieval of syntactic and semantic information from the mental lexicon, retrieval of phonological form, assembly of sounds into syllables (syllabification), and finally implementation of speech motor plan in terms of commands to specific muscles to execute the articulation. It is possible that any one, or all, of these activities could be affected by word frequency. Indeed several studies have pointed to an effect of word frequency at different stages of speech production (see e.g. the contrasting accounts of Jescheniak & Levelt, 1994; Gahl, 2008, details on Section 2). However, studies seeking neural evidence of the processing difference by word frequency have so far yielded inconsistent results that cannot be uniformly accounted for. The present study aims to fill this gap by finding a more reliable neural basis of the effect of word frequency using high density intracranial recordings during word reading. In addition, this study analyzes both one- and two-syllable words in order to observe a possible processing difference caused by the number of syllables. Since several studies suggest that word frequency affects processing of words differently depending on the number of syllables (Balota& Chumbley, 1985; Jescheniak& Levelt, 1994), it might be relevant to Data from this study was collected and made available by the lab of Dr. Edward Chang at UCSF. I appreciate suggestions and feedbacks from Edward Chang, Kristofer Bouchard, Nima Mesgarani, Angela Ren, Keith Johnson, Emily Cibelli, and Susanne Gahl, and members of the Chang Lab at UCSF and the Phonology Lab at UC Berkeley.
This study investigates how visual phonetic information affects compensation for coarticulation in speech perception. A series of CV syllables with fricative continuum from [s] to [sh] before [a],[u] and [y] was overlaid with a video of a face saying [s]V, [∫]V, or a visual blend of the two fricatives. We made separate movies for each vowel environment. We collected [s]/[∫] boundary locations from 24 native English speakers. In a test of audio-visual integration, [∫] videos showed significantly lower boundary locations (more [sh] responses) than [s] videos (t[23]=2.9, p<0.01) in the [a] vowel environment. Regardless of visual fricative condition, the participants showed a compensation effect with [u] (t[23] > 3, p<0.01), but not with the unfamiliar vowel [y]. This pattern of results was similar to our findings from an audio-only version of the experiment, implying that the compensation effect was not strengthened by seeing the lip rounding of [y].
This paper reports an experiment testing whether compensation for coarticulation in speech perception is mediated by linguistic experience. The stimuli are a set of fricative-vowel syllables on continua from [s] to [∫] with the vowels [a], [u], and [y]. Responses from native speakers of English and French (20 in each group) were compared. Native speakers of French are familiar with the production of the rounded vowel [y] while this vowel was unfamiliar to the native English speakers. Both groups showed compensation for coarticulation (both t > 5, p < 0.01) with the vowel [u] (more “s” responses indicating that in the context of a round vowel, fricatives with a lower spectral center of gravity were labeled “s”). The French group also showed a compensation effect in the [y] environment (t[20] = 3.48, p<0.01). English listeners also showed a tendency for more subject-to-subject variation on the [y] boundary locations than did the French listeners (Levene's test of equality of variance, p < 0.1). The results thus indicate that compensation for coarticulation is a language specific effect, tied to the listener's experience with the conditioning phonetic environment.