In a regression study of conversational speech, we show that frequency, contextual predictability, and repetition have separate contributions to word duration, despite their substantial correlations. We also found that content- and function-word durations are affected differently by their frequency and predictability. Content words are shorter when more frequent, and shorter when repeated, while function words are not so affected. Function words have shorter pronunciations, after controlling for frequency and predictability. While both content and function words are strongly affected by predictability from the word following them, sensitivity to predictability from the preceding word is largely limited to very frequent function words. The results support the view that content and function words are accessed differently in production. We suggest a lexical-access-based model of our results, in which frequency or repetition leads to shorter or longer word durations by causing faster or slower lexical access, mediated by a general mechanism that coordinates the pace of higher-level planning and the execution of the articulatory plan.
Function words ( the, that, and, of, . . . ) vary widely in pronunciation. Understanding this variation is essential both for cognitive modeling of lexic al production and for computer speech recognition and synthesis. This study investigates which f a tors affect the forms of function words, especially whether they have a fuller pronunciation (e.g.,Di, Dæt, ænd, 2v) or a more reduced or lenited pronunciation (e.g., D D1t, n, ). It is based on over 8000 occurrences of ten frequent English function words in a four-hour sample from c onversations from the Switchboard corpus. Ordinary linear and logistic regression mode ls w re used to examine variation in the length of the words, in the form of their vowel (basic, f ull, or reduced), and whether final obstruents were present or not. For all these measures, afte r controlling for segmental context, rate of speech, and other important factors, there are stron g independent effects that made function words more likely to be longer or have a fuller form (1) wh en neighboring disfluencies (such as filled pauses uhandum) indicate that the speaker was encountering problems in pla nning the utterance; (2) when the word is unexpected, i.e less predictable in context; (3) when the word is either utterance-initial or utterance-final. Lo oking at the phenomenon in a different way, function words are more likely to be shorter and to have l ess full forms in fluent speech, in predictable positions or multi-word collocations, and u tterance-internally. Also considered are other factors such as sex (women are more likely to use ful ler orms, even after controlling for rate of speech, for example), and some of the differences among the ten function words in their response to the factors.
To date, studies of deceptive speech have largely been confined to descriptive studies and observations from subjects, researchers, or practitioners, with few empirical studies of the specific lexical or acoustic/prosodic features which may characterize deceptive speech.We present results from a study seeking to distinguish deceptive from non-deceptive speech using machine learning techniques on features extracted from a large corpus of deceptive and non-deceptive speech.This corpus employs an interview paradigm that includes subject reports of truth vs. lie at multiple temporal scales.We present current results comparing the performance of acoustic/prosodic, lexical, and speaker-dependent features and discuss future research directions.
This study examines the phonetic and phonological features essential to the investigation of the internal structure of the intonational phrase. This is crucial for identifying and explaining the forms and functions of basic intonational units. There are two primary intonational contour modeling theories [Ladd (1983) and references therein]. One suggests that phrasal contours are the basic units of intonation. Within this theory, contour shapes are associated with particular functions or meanings. In contrast, a more recent theory claims that individual tones (i.e., abstract phonological units) are the basic units of intonation, and intonational contours result from the concatenation of adjacent tones in a phrase. Using 756 utterances from the Switchboard and Buckeye corpora, the present study takes a closer look at the basic units that compose the intonational contour. While the nucleus has long been identified as a functionally important part of an intonational phrase, the head of an intonational phrase has not been considered, in kind. This work examines global and local phonetic and phonological features (e.g., intensity, pitch range, pitch height, downstepped tones) of intonational phrase heads in an attempt to better understand their forms and functions.
Function words, especially frequently occurring ones such as (the, that, and, and of), vary widely in pronunciation. Understanding this variation is essential both for cognitive modeling of lexical production and for computer speech recognition and synthesis. This study investigates which factors affect the forms of function words, especially whether they have a fuller pronunciation (e.g., thi, thaet, aend, inverted-v v) or a more reduced or lenited pronunciation (e.g., thax, thixt, n, ax). It is based on over 8000 occurrences of the ten most frequent English function words in a 4-h sample from conversations from the Switchboard corpus. Ordinary linear and logistic regression models were used to examine variation in the length of the words, in the form of their vowel (basic, full, or reduced), and whether final obstruents were present or not. For all these measures, after controlling for segmental context, rate of speech, and other important factors, there are strong independent effects that made high-frequency monosyllabic function words more likely to be longer or have a fuller form (1) when neighboring disfluencies (such as filled pauses uh and um) indicate that the speaker was encountering problems in planning the utterance; (2) when the word is unexpected, i.e., less predictable in context; (3) when the word is either utterance initial or utterance final. Looking at the phenomenon in a different way, frequent function words are more likely to be shorter and to have less-full forms in fluent speech, in predictable positions or multiword collocations, and utterance internally. Also considered are other factors such as sex (women are more likely to use fuller forms, even after controlling for rate of speech, for example), and some of the differences among the ten function words in their response to the factors.
Multimodal interfaces are designed with a focus on flexibility, although very few multimodal systems currently are capable of adapting to major sources of user or environmental variation. The development of adaptive multimodal processing techniques will require empirical guidance on modeling key aspects of individual differences. In the present study, we collected data from 24 7-to-10-year-old children as they interacted using speech and pen input with an educational software prototype. A comprehensive analysis of children’s multimodal integration patterns revealed that they were classifiable as either simultaneous or sequential integrators, although they more often integrated signals simultaneously than adults. During their sequential constructions, intermodal lags also ranged faster than those of adult users. The high degree of consistency and early predictability of children’s integration patterns were similar to previously reported adult data. These results have implications for the development of temporal thresholds and adaptive multimodal processing strategies for children’s applications. The long-term goal of this research is life-span modeling of users’ integration and synchronization patterns, which will be needed to design future high-performance adaptive multimodal systems.
We investigate the role that lemmas and wordforms (lexemes) play in form variation in lexical production, using a corpus-based methodology which is sensitive to lexical frequency effects. Average durations and reduction frequencies seem to show differences between the surface forms corresponding to different lemmas for the words of, that, and to. But after controlling for such factors as rate of speech, segmental context, neighbouring disfluencies, and, crucially, predictability from neighbouring words, almost all of these differences disappear. The remaining differences do not require phonetic encoding to be directly sensitive to lemma differences even for homophones. Nor do the results suggest a role for lemma frequency in lexical production, despite its key role in lexical comprehension. Instead, a 'multiple lexeme' model of lexical representation, in which different lemmas are differentially linked in the lexicon to different wordforms can account for lemma-based form differences. Our results further suggest that the lexical specifications of these wordforms include more fine-grained phonetic detail than has been previously suggested.
Gaining time to resolve some difficulty in the production of upcoming speech is the primary function of unplanned repetitions. Understanding their duration structure is thus crucial to modeling their production, which surely differs greatly from fluent phrases. When the duration structure of the entire repetition string of unplanned repetitions is examined, strong global dependencies are found. Repetition strings are words/phrases repeated once or more, together with silent and filled pauses optionally occurring next to them. The main effects are that durations of first and second string items, whether repeated words or pauses, are positively correlated with the duration of the rest of the string; items, whether words or pauses, are shorter as they occur later in the string; and strings that begin with a pause average longer than strings that do not. The study is based on an analysis of 503 disfluent repetitions taken from the ICSI phonetically transcribed sample of the Switchboard conversation corpus, extensively checked and recoded. The results help explain local durational dependencies of the repeated words [Bell and Girand, LabPhon 7 (2000)] and imply that the articulation of repetitions is influenced from the onset by the nature and degree of difficulty they address.
This study examines the role of several non–phonetic factors in the reduction of ten frequent English function words ( I, and, the, that, a, you, to, of, it , and in) in the phoneticallytranscribed portion of the Switchboard corpus of spontaneous telephone conversations. Using ordinary linear and logistic regression models, we examined the length of the words and whether their vowels were full or reduced. We show that function words are more likely to be longer or unreduced when they are turn–initial or utterance–final, when the speaker is female (mostly but not completely due to slower rate of speech) and when the word is surprising given the previous or following words. Finally, focusing on finer details of the effect of planning problems on reduction, we show that filled pauses ( uh and um) are the strongest factor in predicting lengthening of a previous function word. The results bear on issues in speech recognition and models of speech production.
An investigation of reduction in the ten most frequent English function words in Switchboard, to be presented at ICSLP’98, found that repetition or following silence or filled pause, which were taken as symptoms of planning problems, were strongly associated with longer durations and lack of reduction, extending earlier results for the definite article [J. E. Foxtree and H. H. Clark, Cog. 62, 151–167 (1997)]. As a followup to this study, the structure of unplanned repetition strings is examined more closely. The study is based on an analysis of the lexical transcriptions of the repetitions of words and short phrases from over 100 h of recorded conversations from the Switchboard corpus. Detailed information about the phonetic form and contexts of repetitions a retaken from a phonetically transcribed sample of 4 h of conversation [S. Greenberg et al., ICSLP 96 Proc. (1996)]. Repetitions are overwhelmingly unplanned; are overwhelmingly function words; and are mostly single repetitions of words, although complex strings with multiple repetitions of words, and short phrases combined with silences and filled pauses are not uncommon. The form of repetition strings, including internal and overall durations, are related to variables of rate and context.
The causes of pronunciation reduction in 8458 occurrences of ten frequent English function words in a four-hour sample from conversations from the Switchboard corpus were examined. Using ordinary linear and logistic regression models, we examined the length of the words, the form of their vowel (basic, full, or reduced), and final obstruent deletion. For all of these we found strong, independent effects of speaking rate, predictability, the form of the following word, and planning problem disfluencies. The results bear on issues in speech recognition, models of speech production, and conversational analysis.
Function words (the, that, and, of, . . . ) vary widely in pronunciation. Understanding this variation is essential both for cognitive modeling of lexical production and for computer speech recognition and synthesis. This study investigates which factors affect the forms of function words, especially whether they have a fuller pronunciation (e.g., , , , ) or a more reduced or lenited pronunciation (e.g., , , , ). It is based on over 8000 occurrences of ten frequent English function words in a four-hour sample from conversations from the Switch- board corpus. Ordinary linear and logistic regression models were used to examine variation in the length of the words, in the form of their vowel (basic, full, or reduced), and whether final obstruents were present or not. For all these measures, after controlling for segmental context, rate of speech, and other important factors, there are strong independent effects that made func- tion words more likely to be longer or have a fuller form (1) when neighboring disfluencies (such as filled pauses uh and um) indicate that the speaker was encountering problems in plan- ning the utterance; (2) when the word is unexpected, i.e less predictable in context; (3) when the word is either utterance-initial or utterance-final. Looking at the phenomenon in a different way, function words are more likely to be shorter and to have less full forms in fluent speech, in predictable positions or multi-word collocations, and utterance-internally. Also considered are other factors such as sex (women are more likely to use fuller forms, even after controlling for rate of speech, for example), and some of the differences among the ten function words in their response to the factors.
Elizabeth Shriberg合作论文数Speech Technology & Research Laboratory (Wednesdays)1
Julia Hirschberg合作论文数Department of Computer Science, Columbia University1