In this experiment, we tested the hypothesis that adult-child differences in cue weighting are influenced by adult-child differences in knowledge of (a) the relative predictability of wordinitial vs. word-final consonants, and (b) of the relationship between predictability and acoustic salience/distinctiveness. We tested our hypothesis using synthetic speech continua with formant transitions varying from /edi/ to /ebi/, which listeners were encouraged to hear as either “Abe E/Ade E” (VC#V context) or as “A bee/A dee” (V#CV context). We tested the extent to which changes in formant transitions influence /d/ vs. /b/ categorisation. Results show that adults were more influenced by transitions cueing word-initial consonants (less predictable in English) than by transitions cueing word-final consonants (more predictable in English), whereas children showed a more balanced pattern, with marginally more influence of transitions cueing word-final consonants. Results are consistent with the view that adults have learned more about the relative predictability of word-initial vs. word-final consonants and have learned that acoustic cues to the less-predictable initial consonants are more distinctive. They therefore weight these cues more heavily than less-distinctive, more contextually predictable, word-final cues.
Acoustic models used for statistical parametric speech synthesis typically incorporate many modelling assumptions.It is an open question to what extent these assumptions limit the naturalness of synthesised speech.To investigate this question, we recorded a speech corpus where each prompt was read aloud multiple times.By combining speech parameter trajectories extracted from different repetitions, we were able to quantify the perceptual effects of certain commonly used modelling assumptions.Subjective listening tests show that taking the source and filter parameters to be conditionally independent, or using diagonal covariance matrices, significantly limits the naturalness that can be achieved.Our experimental results also demonstrate the shortcomings of mean-based parameter generation.
Speech produced in the presence of noise (Lombard speech) is typically more intelligible than speech produced in quiet (plain speech) when presented at the same signal-to-noise ratio, but the factors responsible for the Lombard intelligibility benefit remain poorly understood. Previous studies have demonstrated a clear effect of spectral differences between the two speech styles and a lack of effect of fundamental frequency differences. The current study investigates a possible role for durational differences alongside spectral changes. Listeners identified keywords in sentences manipulated to possess either durational or spectral characteristics of plain or Lombard speech. Durational modifications were produced using linear or nonlinear time warping, while spectral changes were applied at the global utterance level or to individual time frames. Modifications were made to both plain and Lombard speech. No beneficial effects of durational increases were observed in any condition. Lombard sentences spoken at a speech rate substantially slower than their plain counterparts also failed to reveal a durational benefit. Spectral changes to plain speech resulted in large intelligibility gains, although not to the level of Lombard speech. These outcomes suggest that the durational increases seen in Lombard speech have little or no role in the Lombard intelligibility benefit.
Purpose In this study, the authors aimed to investigate how listener training and the presence of intermediate acoustic cues influence transcription variability for conflicting cue speech stimuli. Method Twenty listeners with training in transcribing disordered speech, and 26 untrained listeners, were asked to make forced-choice labeling decisions for synthetic vowel–consonant–vowel (VCV) sequences “a doe” (/ədo/) and “a go” (/əgo/). Both the VC and CV transitions in these stimuli ranged through intermediate positions, from appropriate for /d/ to appropriate for /g/. Results Both trained and untrained listeners gave more weight to the CV transitions than to the VC transitions. However, listener behavior was not uniform: The results showed a high level of inter- and intratranscriber inconsistency, with untrained listeners showing a nonsignificant tendency to be more influenced than trained listeners by CV transitions. Conclusions Listeners do not assign consistent categorical labels to the type of intermediate, conflicting transitional cues that were present in the stimuli used in the current study and that are also present in disordered articulations. Although listener inconsistency in assigning labels to intermediate productions is not increased as a result of phonetic training, neither is it reduced by such training.
The use of live and recorded speech is widespread in applications where correct message reception is important. Furthermore, the deployment of synthetic speech in such applications is growing. Modifications to natural and synthetic speech have therefore been proposed which aim at improving intelligibility in noise. The current study compares the benefits of speech modification algorithms in a large-scale speech intelligibility evaluation and quantifies the equivalent intensity change, defined as the amount in decibels that unmodified speech would need to be adjusted by in order to achieve the same intelligibility as modified speech. Listeners identified keywords in phonetically-balanced sentences representing ten different types of speech: plain and Lombard speech, five types of modified speech, and three forms of synthetic speech. Sentences were masked by either a stationary or a competing speech masker. Modification methods varied in the manner and degree to which they exploited estimates of the masking noise. The best-performing modifications led to equivalent intensity changes of around 5 dB in moderate and high noise levels for the stationary masker, and 3-4 dB in the presence of competing speech. These gains exceed those produced by Lombard speech. Synthetic speech in noise was always less intelligible than plain natural speech, but modified synthetic speech reduced this deficit by a significant amount. (C) 2013 Elsevier B.V. All rights reserved.
The increased vocal effort associated with the Lombard reflex produces speech that is perceived as louder and judged to be more intelligible in noise than normal speech. Previous work illustrates that, on average, Lombard increases in loudness result from boosting spectral energy in a frequency band spanning the range of formants F1-F3, particularly for voiced speech. Observing additionally that increases in loudness across spoken sentences are spectro-temporally localized, the goal of this work is to further isolate these regions of maximal loudness by linking them to specific formant trends, explicitly considering here the vowel formant separation. For both normal and Lombard speech, this work illustrates that, as loudness increases in frequency bands containing formants (e.g. F1-F2 or F2-F3), the observed separation between formant frequencies decreases. From a production standpoint, these results seem to highlight a physiological trait associated with how humans increase the loudness of their speech, namely moving vocal tract resonances closer together. Particularly, for Lombard speech, this phenomena is exaggerated: that is, the Lombard speech is louder and formants in corresponding spectro-temporal regions are even closer together.
Talkers adopt different speech styles in response to factors such as the perceived needs of the interlocutor, environmental noise and explicit instruction. Some styles have been shown to be beneficial for listeners but many aspects of the relationship between speech modifications and intelligibility remain unclear, particularly for prosodic changes. The current study measures the relative intelligibility in noise of speech spoken in 5 speech styles plain, infant-, computer- and foreigner-directed, and shouted and relates listener scores to acoustic/prosodic parameters and quantitative estimates of energetic masking. Intelligibility changes over plain speech correlated well with durational modifications, which included elongations of all segments as well as increases in the number of unfilled pauses. Both mean fundamental frequency and its range displayed great variation across styles but with no clear intelligibility benefits. Energetic masking per unit time was similar in each style but the total amount of speech which escaped masking was a good predictor of word identification rate. These findings suggest that much of the prosody-related intelligibility gain is derived from durational increases.
How do talkers maintain intelligibility when speaking in the presence of a background conversation? The current study identified acoustic and temporal modifications of speech manifested by interlocutors in the face of competing speech, with and without visual contact. Pairs of talkers held free conversations either alone or in the presence of a second pair. Regardless of the availability of visual information, speaking simultaneously with another talker resulted in overall increases in energy, F0, F1 and a decrease in speech rate. Overlapping with the background pair resulted in an increase in energy but no change in the two prosodic parameters F0 and speech rate. By contrast , within-pair overlap led to an increase in F0 and a decrease in rate, and no change in speech level. The absence of visual cues produced a significant reduction in within-pair overlap, which tended to be greater when the background pair was present. These findings emphasize the need to distinguish between Lombard and interactional influences on acoustic parameters, and suggest that adverse conditions such as competing speech or absence of visual cues cause interlocutors to adopt more careful dialogue strategies, perhaps with the aim of reducing energetic and informational masking at the ears of the listener.
Speakers change the way they speak depending on the surrounding environment. When masking noise obstructs the communication channel between interlocutors, they consistently engage in Lombard speech, whose spectral characteristics are well described (e.g. [1, 2, 3]) and are believed to result in better intelligibility through energetic masking reduction (e.g. [4]). However, less is known about how speakers adapt to the temporal characteristics of a fluctuating masker, and whether any such changes aid communication. In this study pairs of speakers were recorded while engaged in a sudoku-solving task in quiet and in several masking conditions. Maskers were either a competing talker or speech modulated noise with identical temporal characteristics, chosen to investigate the informational masking potential of the masker. The silence density of each masker was also manipulated by adjusting the durations of pauses in the masker to 33% or 66% of the overall duration of the masker. In all masking conditions, speakers displayed a reduction of overlap with the masker relative to a baseline computed from the masker and speech produced in quiet (Figure 1). The overlap reduction tended to be larger for less
The move to unit-selection in speech synthesis has resulted in system improvements being made at subtle suband suprasegmental levels. Human perceptual evaluation of such subtle improvements requires a highly sophisticated level of perceptual attention to specific acoustic characteristics or cues. However, it is not well understood what acoustic cues listeners attend to by default when asked to evaluate synthetic speech. It may, therefore, be potentially quite difficult to design an evaluation method that allows listeners to concentrate on only one dimension of the signal, while ignoring others that are perceptually more important to them. This paper describes a pilot study which aims to evaluate multidimensional scaling (MDS) as a possible method of determining what acoustic characteristics of synthetic speech influence listeners’ judgements of the naturalness of the speech. Using distance measures (either real or perceived distances), MDS techniques represent stimuli as points in n-dimensional space. The space is configured so that similar stimuli are close together, while different stimuli are farther apart. Additionally, the dimensions of the space correspond to characteristics of the stimuli which influenced the perceived distances. Our results indicate that MDS techniques should be a useful tool in understanding the complex psychoacoustic processes that listeners undergo when evaluating synthetic speech. This method has allowed us to identify a number of cues that appear to be particularly perceptually salient to listeners evaluating synthetic speech naturalness, namely prosodic cues (in terms of duration and/or intonation) and segmental or unit level cues (in terms of appropriateness of units, or number of units).
The Blizzard Challenge 2008 was the fourth annual Blizzard Challenge. This year, participants were asked to build two voices from a UK English corpus and one voice from a Man- darin Chinese corpus. This is the first time that a language other than English has been included and also the first time that a large UK English corpus has been available. In addi- tion, the English corpus contained somewhat more expressive speech than that found in corpora used in previous Blizzard Challenges. To assist participants with limited resources or limited ex- perience in UK-accented English or Mandarin, unaligned la- bels were provided for both corpora and for the test sentences. Participants could use the provided labels or create their own. An accent-specific pronunciation dictionary was also available for the English speaker. A set of test sentences was released to participants, who were given a limited time in which to synthesise them and submit the synthetic speech. An online listening test was con- ducted, to evaluate naturalness, intelligibility and degree of similarity to the original speaker Index Terms: Blizzard Challenge, speech synthesis, evalua- tion, listening test
Blizzard 2007 is the third Blizzard Challenge, in which participants build voices from a common dataset. A large listening test is conducted which allows comparison of systems in terms of naturalness and intelligibility. New sections were added to the listening test for 2007 to test the perceived similarity of the speaker’s identity between natural and synthetic speech. In this paper, we present the results of the listening test and the subsequent statistical analysis Index Terms: Blizzard Challenge, speech synthesis, evaluation, listening test
Children and adults appear to weight some acoustic cues differently in perceiving certain speech contrasts. There are currently two main theories to explain this difference. One of these is the Developmental Weighting Shift theory, which proposes that children process speech in terms of more global, syllablelike units. Thus, in this view, children should always give more weight than should adults to cues like within-syllable vowel formant transitions, and less weight to across-syllable vowel formant transitions [1]. Other researchers have proposed that children have lower general auditory sensitivity than adults, which impacts on their speech perception. Thus, in this view, children should always give more weight than adults to cues that are longer, louder, or more spectrally informative than the alternative cues [2]. The current study tested these hypotheses in two ways. First, we examined adults’ and threeto seven-year-old children’s weighting of vowel-onset formant transitions in contrasts in which we systematically varied the consonantal context and the spectral distinctiveness of the transition. Second, we examined adults’ and five-year-old children’s weighting of vowel-formant offset transitions in spectrally identical within-monosyllabic-word and across-monosyllabic-word contexts. The results of the study showed that adult-child differences in cue weighting are affected by the segmental context of the cues, the salience of the cues, and the position of the cue in the word. However, neither of the above two theories, either on their own or in combination, can account for all of the observed cue weighting behaviour.
It has been proposed that young children may have a perceptual preference for transitional cues [Nittrouer, S. (2002). J. Acoust. Soc. Am. 112, 711-719]. According to this proposal, this preference can manifest itself either as heavier weighting of transitional cues by children than by adults, or as heavier weighting of transitional cues than of other, more static, cues by children. This study tested this hypothesis by examining adults' and children's cue weighting for the contrasts /saI/-/integral of aI/, /de/-/be/, /ta/-/da/, and /ti/-/di/. Children were found to weight transitions more heavily than did adults for the fricative contrast /saI/-/integral aI/, and were found to weight transitional cues more heavily than nontransitional cues for the voice-onset-time contrast /ta/-/da/. However, these two patterns of cue weighting were not found to hold for the contrasts /de/-/be/ and /ti/-/di/. Consistent with several studies in the literature, results suggest that children do not always show a bias towards vowel-formant transitions, but that cue weighting can differ according to segmental context, and possibly the physical distinctiveness of available acoustic cues.