Previous laboratory studies have shown that prosodic structures are encoded in the modulations of phonetic patterns of speech including suprasegmental as well as segmental features. In particular, effects of prosodic context on duration and intensity of syllables and words have been widely reported. Drawing on prosodically annotated large-scale speech data from the Buckeye corpus of conversational speech of American English, the current study attempted to examine whether and how prosodic prominence and phrase boundary of everyday conversational speech, as determined by a large group of ordinary listeners, are related to the phonetic realization of duration and intensity. The results showed that the patterns of word durations and intensities are influenced by prosodic structure. Closer examinations revealed, however, that the effects of prosodic prominence are not the same as those of prosodic phrase boundary. With regard to intensity measures, the results revealed the systematic changes in the patterns of overall RMS intensity near prosodic phrase boundary but the prominence effects are restricted to the nucleus. In terms of duration measures, both prosodic prominence and phrase boundary are the most closely related to the lengthening of the nucleus. Yet, prosodic prominence is more closely related to the lengthening of the onset while phrase boundary lengthens the coda duration more. The findings from the current study suggest that the phonetic realizations of prosodic prominence are different from those of prosodic phrase boundary, and speakers signal different prosodic structures through deliberate modulations of the internal phonetic structure of words and listeners attend to such phonetic variations.
This study investigated the relation between various acoustic features and prominence. Past research has suggested that duration, pitch, and intensity all play a role in the perception of prominence. In our past work, we found a correlation between these acoustic features and speaker agreement over the placement of prominence. The current study was motivated by a need to enrich our understanding of this correlation. Using the Bayesian information criterion, we show that the best model for a feature that cues prosody is not necessarily a single Gaussian. Rather, the best model depends on the feature. This finding has consequences for our understanding of the role of these features in the perception of prosody and for prosody recognition systems.
AbstractThe perception of prosodic prominence in spontaneous speech is investigated through an online task of prosody transcription using untrained listeners. Prominence is indexed through a probabilistic prominence score assigned to each word based on the proportion of transcribers who perceived the word as prominent. Correlation and regression analyses between perceived prominence, acoustic measures and measures of a word's information status are conducted to test three hypotheses: (i) prominence perception is signal-driven, influenced by acoustic factors reflecting speakers' productions; (ii) perception is expectation-driven, influenced by the listener's prior experience of word frequency and repetition; (iii) any observed influence of word frequency on perceived prominence is mediated through the acoustic signal. Results show correlates of perceived prominence in acoustic measures, in word log-frequency and in the repetition index of a word, consistent with both signal-driven and expectation-driven hypotheses of prominence perception. But the acoustic correlates of perceived prominence differ somewhat from the correlates of word frequency, suggesting an independent effect of frequency on prominence perception. A speech processing account is offered as a model of signal-driven and expectation-driven effects on prominence perception, where prominence ratings are a function of the ease of lexical processing, as measured through the activation levels of lexical and sub-lexical units.
This study investigates the relationship between variation due to prosodic prominence and variation due to sound change. We compare two hypotheses: under prominence vowels move in the direction of vowel shift, and under prominence vowels are hyperarticulated, and move to positions more peripheral in the vowel space. These hypotheses make competing predictions for two vowels currently undergoing change in Chicago American English, /ɛ/ and /uw/. Labov [1] reports that in the Chicago variety, which participates in the Northern Cities vowel shift, the most recent changes have affected the vowels /ɛ/ and /ʌ/, both of which have been retracting since about 1980. In addition, as in many other contemporary varieties of English, the vowel /uw/ has been reported to be fronting. This fronting is not necessarily part of the general vowel shift. An acoustic analysis of controlled vowel productions from 20 speakers in their twenties shows that prominence effects are consistent with the hypothesis of prominence as local hyperarticulation, but do not generally support the claim that prominence and vowel shift effects are in the same direction. The findings also reveal /uw/ fronting as a change in progress, with greatly more variation in the front/back dimension than for other vowels, in all prosodic contexts. Other effects of prominence on vowel height are discussed as indicators of future vowel changes in this variety.
We examine the prosodic variation in production with corpus and experimental speech materials. Prosodic transcription of spontaneous speech from the Buckeye corpus of American English (38 speakers, 54 excerpts, 11–55-s duration each) reveals striking inter-speaker variation in the frequency and distribution of prosodic prominences and boundaries, and in acoustic correlates (pitch, intensity, duration, and spectral measures). Inter-transcriber agreement rates also vary systematically across speakers, suggesting inter-speaker differences in the clarity/consistency of prosodic cues. To further explore variability in the phonological and phonetic expression of prosody while holding lexico-syntactic content constant across speakers, we conducted an auditory repetition experiment. Ten American English speakers listened to 32 excerpts (8–15 words each, 4 speakers) from the American English Map Task corpus and reproduced each utterance with the exact words and “in the way the speaker said them”, without text prompts. Preliminary results from prosodic transcription and acoustic analysis show reliable replication of the phonological structures locating prosodic prominences and phrase boundaries, but with variation in the pitch melody and other phonetic details of the prosodic features. These findings shed light on the mapping between the phonological encoding and the acoustic expression of prosodic features, and highlight those acoustic parameters that identify the prosodic signature of individual speakers.
The relationship between syntactic and prosodic phrase structures is investigated in the production and perception of spontaneous speech. Three hypotheses are tested: (1) syntax influences prosody production; (2) listeners' perception of prosodic boundaries is sensitive to acoustic duration; and (3) syntax directly influences boundary perception, (partly) independent of the acoustic evidence for boundaries. Data are from the Buckeye corpus of conversational speech, and the real-time prosodic transcription of those data by 97 untrained listeners. Inter-transcriber agreement codes boundary strength at word junctures, and Boundary scores are shown to be correlated with both the syntactic context and vowel duration of a word. Vowel duration is also correlated with syntactic context, but the effect of syntactic context on boundary perception is not fully explained by vowel duration. Regression analyses show that syntactic clause boundaries and vowel duration are the first and second strongest predictors of boundary perception in spontaneous speech.
Prosody serves an important function in speech communication: prosodic phrasing groups words into pragmatically and semantically coherent smaller chunks and prosodic prominence encodes the discourse-level status and rhythmic structure of a word within a phrase.Acoustic cues to prosody are available from the speech signal and can be used by listeners to recover the pragmatic and discourse meaning intended by speakers.Effects of prosodic context on the duration of consonants and vowels have been widely reported, and this study extends that line of work by examining how prosodic phrase boundary and prominence influence the temporal structure of the monosyllabic CVC word, based on an analysis of speech excerpts from the Buckeye corpus of spontaneous conversational American English.Prosody annotation for these speech materials is obtained from 97 untrained, non-expert listeners.The results confirm findings from prior studies, showing that (1) monosyllabic CVC words are lengthened before a prosodic phrase boundary and under prominence, and (2) all subcomponents of a syllable, that is, the onset, nucleus, and coda of the monosyllabic word, are elongated.The findings further show that (3) the magnitude of lengthening associated with prosody varies as a function of syllable position, and (4) the magnitude of lengthening of subcomponents of monosyllabic CVC words varies as a function of prosodic characteristics.Nucleus duration is most strongly affected by both prosodic prominence and boundary and the onset and the coda of the monosyllabic word is also affected but to a lesser degree.The lengthening effect of prosodic phrase boundary on the coda is larger than the lengthening effect on onset duration while lengthening of the onset under prosodic prominence is larger than lengthening of the coda.The findings indicate that prosodic context shapes the internal temporal structure of the monosyllabic CVC word.
In speech comprehension, listeners attend to variation in multiple acoustic parameters encoding prosodic structure. Given the multiplicity of acoustic cues, we ask whether prosody perception is dependent on any individual cue or whether acoustic redundancy encoding prosody supports robust prosody perception in the absence of an individual cue. The present paper reports on a study of boundary perception in spontaneous speech with and without silent pause as a boundary cue. Prior studies show that in read speech, silent pause is important for boundary perception, while in spontaneous speech, listeners can detect boundaries without pauses. Our study tests the role of pause in boundary perception with two versions of 36 short speech excerpts from the Buckeye Corpus: one with pauses intact and another with all pauses truncated to 20 ms. In real-time transcription tasks based only on auditory impression, boundary locations were marked by 74 subjects for the intact stimuli and by an additional 15 subjects for truncated excerpts. Inter-transcriber agreement was comparable across the intact and truncated conditions. Paired-sample t-tests show significantly higher rates of boundary perception for intact stimuli indicating that silent pause is an important but not necessary cue to boundary perception and cue redundancy allows for robust perception.
Speakers communicate pragmatic and discourse meaning through the prosodic form assigned to an utterance, and listeners must attend to the acoustic cues to prosodic form to fully recover the speaker's intended meaning. While much of the research on prosody examines supra-segmental cues such as F0 and temporal patterns, prosody is also known to affect the phonetic properties of segments as well. This paper reports on the effect of prosodic prominence on the formant patterns of vowels using speech data from the Buckeye corpus of spontaneous American English. A prosody annotation was obtained for a subset of this corpus based on the auditory perception of 97 ordinary, untrained listeners. To understand the relationship between prominence perception and formant structure, as a measure of the 'strength' of the vowel articulation, we measure the steady-state first and second formants of stressed vowels at vowel mid-points for monophthongs and at both 10% (nucleus) and 90% (glide) positions for diphthongs.Two hypotheses about the articulatory mechanism that implements prominence (Hyperarticulation vs. Sonority Expansion Hypothesis) were evaluated using Pearson's bivariate correlation analyses with formant values and prominence scores' a novel perceptual measure of prominence. The findings demonstrate that higher F1 values correlate with higher prominence scores regardless of vowel height, confirming that vowels perceived as prominent tend to have enhanced sonority. In the frontness dimension, on the other hand, the results show that vowels perceived as prominent tend to be hyperarticulated. These results support the model of the supra-laryngeal implementation of prominence proposed in [5, 6] based on controlled "laboratory" speech, and demonstrate that the model can be extended to cover prosody in spontaneous speech using a continuous-valued measure of prosodic prominence. The evidence reported here from spontaneous speech shows that prominent vowels have expanded sonority regardless of vowel height, and are hyperarticulated only when hyperarticulation does not interfere with sonority expansion.
In comprehending speech, listeners are sensitive to the acoustic variation encoding the prosodic structures that mark phrasing and prominence. The present study, based on prosody transcriptions of untrained listeners (74 monolingual American English speakers), tests which acoustic feature or feature combinations cue prosodic prominence, and specifically, whether ordinary listeners perceive prosodic prominence based on changes in acoustic measures in the local context (syntagmatic comparison) or based on the value of an acoustic measure relative to the distribution of that measure across all instances of the specific phoneme in the listener’s experience (paradigmatic comparison). Subjects listened to 36 short excerpts (11–25 s) of spontaneous speech from the Buckeye Corpus, and marked prominent words on a transcript. After evaluating intertranscriber agreement rates, acoustic measures (duration, intensity, subband intensities, F0 max, and formants) were extracted from stressed vowels and normalized within two domains (syntagmatic/paradigmatic). The results show that all the acoustic measures are correlated with perceived prominence and combinations of the acoustic measures account for the variability in listeners’ perception of prominence. Moreover, syntagmatically normalized acoustic measures explain more of the variability, indicating that ordinary listeners perceive prosodic prominence based on local changes in acoustic measures and especially in changes in overall intensity.
I investigate the acoustic correlates of prosodic prominence and boundary, as they are perceived by naïve listeners, in spontaneous speech from American English (Buckeye corpus). Prosodic prominence and phrasing serve different functions in speech communication: prosodic phrase boundaries demarcate speech chunks that typically cohere semantically, while prominences encode focus and possibly also rhythmic structure. The acoustic correlates of prominence and phrase boundary are examined through measures of vowel duration and overall intensity of stressed vowels, to see how those measures correlate, individually or in combination, with naïve listeners' perception of prominence and boundary. The results show that most stressed vowels are lengthened in pre- boundary words (i.e., those final in the prosodic phrase). Prosodic prominence is also cued by increased duration, but in combination with higher overall intensity for some vowels. These acoustic differences associated with perceived prominence and boundary suggest different mechanisms underlying their production. This claim finds support from consideration of the different functions that prominence and boundary play in encoding information structure, and in speech production planning.
This paper examines how ordinary listeners, naive with respect to the phonetics and phonology of prosody, perceive the location of prosodic boundaries that demarcate speech “chunks” and prominences that serve a “highlighting” function, in spontaneous speech (Buckeye corpus). Over 70 naive listeners marked the locations of prominences and boundaries in a real-time transcription task. Fleiss’ multitranscribers’ reliability tests show that naive transcribers are consistent in their perception of prosodic boundaries and prominences. Specifically, we observe higher multi-transcriber agreement scores for boundary marking than for prominence marking. Variation between transcriptions of the same speech excerpt produced by different listeners reveals individual differences in the perception of prominences and boundaries. Variation in Fleiss’ multi-transcribers’ agreement scores for excerpts from different speakers suggests that speakers vary in how they structure an utterance prosodically and/or in how effectively they cue prosodic structure. We also find that nuclear prominences are more consistently perceived by naive listeners than prenuclear prominences. The finding that naive listeners agree well above chance on the location of prosodic events indicates that naive transcription is a valid method for prosody analysis which can augment analysis based solely on expert labeling.
In comprehending speech, listeners are sensitive to the prosodic features that signal the phrasing and the discourse salience of words (prominence). Findings from two experiments on prosody perception show that acoustic and articulatory kinematic properties of speech correlate with native listeners’ perception of phrasing and prominence. Subjects in this study were 114 university-age adults (74 UIUC + 40 Haskins), monolingual speakers of American English who were untrained in prosody transcription. Subjects listened to short recorded excerpts (about 20 s) from two corpora of spontaneous and read speech (Buckeye Corpus and Wisconsin Microbeam Database) and marked prominent words and the location of phrase boundaries on a transcript. Intertranscriber agreement rates across subsets of 17–40 subjects are significantly above chance based on Fleiss’ statistic, indicating that listeners’ perception of prosody is reliable, with higher agreement rates for boundary perception than for prominence. Prosody perception varies across listeners (both corpora) and across speakers (WMD, where perceived prosody varies for the same utterance produced by different speakers). Acoustic measures from stressed vowels (Buckeye: duration, intensity, F1, F2) and articulatory kinematic measures (WMD) are correlated with the perceived prosodic features of the word. [Work supported by NSF.]
0. Introduction This study examines the acoustic correlates of prosodic prominence as perceived by a large number of native listeners of American English who are naïve to the phonetics and phonology of prosody. In English, as in other stress languages, speech utterances are chunked into smaller prosodic phrases, and within a prosodic phrase some words are assigned phrasal stress, which typically marks a word or a phrase as having a focus or as introducing new information into the discourse. We refer to phrasal stress here as prosodic prominence. Speakers convey the information structure of an utterance through prosodic prominence, and listeners must decode the prosodic structure to recover the speaker’s intended meaning in the course of comprehension. Prosodic structures are phonetically implemented in patterns of pitch (a perceptual attribute of fundamental frequency, F0), duration, loudness (a perceptual attribute of the intensity of sound pressure), and spectral modulations including formants. Pitch as a perceptual correlate of F0 is traditionally described as a primary cue for prominence in many languages, including American English (Beckman 1986, Pierrehumbert 1980). Many studies have investigated F0 as a primary cue for prominence in many languages. Terken (1991, 1994) tested the relative importance of the magnitude of F0 changes or F0 maxima in the perception of prominence in Dutch and these properties of F0 worked together in a complex way to cue prominence. Gussenhoven and Rietveld (1988) and Gussenhoven et al. (1997) also examined the relation between F0 maxima and minima and prominence perception in Dutch and showed that the relative distance between pitch peaks as well as the degree of declination of the baseline is important in the perception of prominence. The role of F0 as a primary cue for prominence is, however, still controversial. Other acoustic measures have also been investigated as correlates of prominence, although the definition of prominence varies across studies. For instance, Cooper et al. (1985) showed that prominent words (contrastively accented) have elongat-
Vowel devoicing is a phenomenon that is reported to occur in many languages such as Japanese, Parisian and Montreal French, Turkish and English. This paper investigates vowel devoicing in Korean. A devoiced vowel does not exhibit characteristic vocal tract resonances, and instead is realized as a long interval of aspiration or frication following consonant release, resulting in non-distinct segment boundaries between devoiced vowels and adjacent voiceless consonants. This paper examines temporal and spectral evidence of devoiced vowels and, among other findings, reveals that in Korean devoiced high vowels are not segmentally deleted but phonetically masked, suggesting that vowel devoicing results from the overlap of glottal gestures. This paper also examines the effect of the preceding consonant place and manner, and the height and front/backness of vowels on devoicing.
Repetition disfluencies are among the most frequent type of disfluency in conversational speech, accounting for over 20% of disfluencies, yet they do not generally lead to comprehension errors for human listeners. We propose that parallel prosodic features in the REP and ALT intervals of the repetition disfluency provide strong perceptual cues that signal the repetition to the listener. We report results from a transcription analysis of repetition disfluencies that classifies disfluent regions on the basis of prosodic factors, and preliminary evidence from F0 analysis to support our finding of prosodic parallelism. 1. Acoustic-prosodic correlates of disfluency Disfluency occurs in spontaneous speech at a rate of about one every 10-20 words, or 6% per word count [17], yet this interruption of fluent speech does not generally lead to comprehension errors for human listeners. Recent research has shown that important cues to disfluency can be found in the syntactic and semantic structures conveyed by the word sequence, and in the phonological and phonetic structures signaled by acoustic features local to the disfluency interval. These cues identify the components of the disfluent regionthe reparandum (REP), edit phrase (EDIT), and alteration (ALT)--- and their junctures. Work on automatic disfluency detection has shown that the most successful approach combines both lexical and acoustic features, with explicit models of the lexical-syntactic and prosodic features that pattern systematically with disfluent intervals [1,6].
Repetition disfluencies are among the most frequent type of disfluency in conversational speech, accounting for over 20% of disfluencies, yet they do not generally lead to comprehension errors for human listeners. We propose that parallel prosodic features in the REP and ALT intervals of the repetition disfluency provide strong perceptual cues that signal the repetition to the listener. We report results from a transcription analysis of repetition disfluencies that classifies disfluent regions on the basis of prosodic factors, and preliminary evidence from F0 analysis to support our finding of prosodic parallelism. 1. Acoustic-prosodic correlates of disfluency Disfluency occurs in spontaneous speech at a rate of about one every 10-20 words, or 6% per word count (17), yet this interruption of fluent speech does not generally lead to comprehension errors for human listeners. Recent research has shown that important cues to disfluency can be found in the syntactic and semantic structures conveyed by the word sequence, and in the phonological and phonetic structures signaled by acoustic features local to the disfluency interval. These cues identify the components of the disfluent region--- the reparandum (REP), edit phrase (EDIT), and alteration (ALT)--- and their junctures. Work on automatic disfluency detection has shown that the most successful approach combines both lexical and acoustic features, with explicit models of the lexical-syntactic and prosodic features that pattern systematically with disfluent intervals (1,6).