Preview this article: Review of Isacenko, A. & H.-J. Schädlich (1970) A Model of Standard German Intonation, Page 1 of 1 < Previous page | Next page > /docserver/preview/fulltext/itl.13.06col-1.gif
Users may download and print one copy of any publication from the public portal for the purpose of private study or research You may not further distribute the material or use it for any profit-making activity or commercial gain You may freely distribute the URL identifying the publication in the public portal Take down policy If you believe that this document breaches copyright, please contact us providing details, and we will remove access to the work immediately and investigate your claim.
In spoken dialogue systems, in which humans interact with computers over the telephone, it is essential that the voice output of the system be of high quality. Both the intelligibility and the naturalness of the output should be sufficiently high. There are several techniques for providing a system with speech output, each with its own advantages and disadvantages. This paper discusses a formal evaluation experiment of three speech output techniques. Natural speech was included as a reference condition. The speech was rated on intelligibility and fluency of the output. Additionally, the overall quality of the speech and its suitability for use in a commercial application were assessed. The results reveal significant differences between the techniques. Diphone synthesis still has an inferior quality compared to the other techniques, both in terms of intelligibility and fluency. Conventional phrase concatenation is quite intelligible, but scores less on fluency. IPO's phrase concatenation is by far the best technique.
From previous research we know that prosodic features are perceptually effective in marking boundaries and that a suitable implementation of these features improves the quality of synthetic speech in terms of acceptability. It can further be assumed that listeners use the perceived prosodic information to compute the meaning of the input speech. This paper, therefore, investigates and determines whether a well-phrased utterance, (that is, an utterance with prosodic boundaries in appropriate positions and with appropriate realizations), is easier to comprehend than a poorly-phrased one. To measure this, we designed a method in which a kind of verification task is combined with a question-answering task (“monitoring for the answer”). The stimulus set consisted of structurally ambiguous sentences. The expectation was that when listeners hear a question followed by an appropriately phrased utterance, they will react more rapidly than when the question is followed by an utterance with neutral phrasing. Also, it was expected that in the latter situation reaction times (RTs) will be shorter than if an inappropriately phrased utterance is presented. The results confirmed the expectations: an appropriately phrased utterance always produced the fastest RTs.
Previous research showed that prosodic features are perceptually conspicuous in marking boundaries and that a suitable implementation of these features improves the quality of synthetic speech in terms of acceptability. Assumably, listeners use prosody to compute the informational structure of the input speech. The hypothesis is that differently phrased utterances may lead to differences in cognitive load during comprehension. A method was designed in which a kind of verification task is combined with a question-answering task. The stimulus set consisted of structurally ambiguous sentences. The expectations were as follows: (a) when listeners receive a question followed by an appropriately phrased utterance, they will react more rapidly than when it is followed by an utterance without phrasing; (b) in the latter situation reaction times (RTs) will still be shorter than if an inappropriately phrased utterance is presented. The results confirmed these expectations for the situation where the prosodically intended reading corresponded to the most likely interpretation of the ambiguous sentences. However, when the least likely interpretation was the prosodically suggested one, it appeared that a question followed by an inappropriately phrased utterance gave the same RTs as one without phrasing. But an appropriately phrased utterance still produced the fastest RTs.
From previous research it is known that speakers use the prosodic cues pause and pitch to audibly structure their spoken messages. Listeners, on the other hand, use these phonetic cues to determine the degree of disjuncture in the flow of speech, which supposedly helps them to process the meaning of the utterances. In the research reported here, a professional speaker’s phrasing behavior was modeled in various sets of rules, corresponding to different levels of prosodic boundary strength. These phrasing rules were evaluated as to their acceptability and it appeared that several of them improve the quality of the synthetic speech. The rule set implementing five levels of boundary strength improved this quality more than rule sets with fewer levels. In fact, it appeared that this rule set produces synthetic speech which is prosodically almost as good as a copy-synthesis version with natural prosody.
The purpose of the study presented in this paper and the accompanying paper [Smits et al., J. Acoust. Soc. Am. 100, 3865-3881 (1996)] is to evaluate whether detailed or gross time-frequency structures are more relevant for the perception of place of articulation of prevocalic stop consonants. To this end, first a perception experiment is carried out with "burst-spliced" stop-vowel utterances, containing the Dutch stops /b,d,p,t/ and /k/. From the utterances burst-only, burstless, and cross-spliced stimuli were created and presented to listeners. The results of the experiment show that the relative importance of burst and transitions for the perception of place of articulation to a great extent depends on place and voicing of the stop consonant and on the vowel context. Velar bursts are generally more effective in cueing place of articulation than other bursts. There is no significant difference in the effectiveness of /p/, /t/, and /k/ transitions, while /b/ transitions are more effective than /d/ transitions. The release burst dominates the perception of place of articulation in front-vowel contexts, while the formant transitions are generally dominant in nonfront vowel contexts. The bursts of unvoiced stops are perceptually more important than the bursts of voiced stops.
The purpose of the present study is to find out how the pitch peak heights on two pitch-accented syllables in one utterance relate to different focus conditions. The focus conditions are neutral focus, double contrastive focus, and single contrastive focus on either the first or the second pitch-accented syllable. In Experiment 1, subjects adjusted the height of one of two pitch peaks, so as to make the pitch contour express different focus structures. No systematic relationship was found between different fixed heights of one peak and the adjusted heights of the other, which suggests the existence of target values for focus-related pitch peaks. In Experiment 2, listeners judged which of the four focus structures was most likely represented by a given relation between peak heights. The results show that some pitch contours are ambiguous with respect to focus, but the majority of them is classified unanimously as signalling only one possible focus structure. The present results also shed new light on some unexplained findings of earlier prominence experiments.
The paper describes a system that produces spoken monologues derived from information in a database. The sentences of these monologues are generated from templates of syntactic structures, which may contain open slots in which other elements, usually noun phrases, can be inserted. It is shown how these sentences string together to form a coherent message. This message has to be pronounced correctly, which means, among other things, that it has to be prosodically acceptable. The paper indicates how linguistic information is used to arrive at an acceptable prosodic structure, which, in turn, feeds into a module which takes care of the phonetic realization of the monologue.
One of the possible functions of intonation is its capacity to clarify textual structure. It may indicate, for instance, that a sentence is likely to be the last one in a sequence of statements that build a discourse unit. In order to investigate the perception of melodic cues to ‘‘finality,’’ a series of three listening experiments was performed with short sentences, the intonation of which was manipulated with respect to different melodic variables. The actual testing was done in two ways (1) by pairwise comparison and (2) by absolute rating. A linear least-squares estimation method brought to light that in both tests finality judgments were influenced significantly by differences in pitch register (experiment 1), pitch range (experiment 1), and shape of the pitch contour (experiments 1, 2, and 3). The results of the data analysis suggest strongly that these different variables generally combine additively in producing finality judgments, though the effect of one is sometimes conditional on the value of another.
This paper addresses two main questions: (a) Can listeners assign values of perceived boundary strength to the juncture between any two words? (b) If so, what is the relationship between these values and various (combinations of) suprasegmental features. Three speakers read a set of twenty utterances of varying length and complexity. A panel of nineteen listeners assigned boundary strength values to each of the 175 word boundaries in the material. Then the correlation was established between the variable strength of the perceived boundaries and three prosodic variables: melodic discontinuity, declination reset and pause. The results show that speakers may differ in their strategies of prosodic boundary marking and listeners agree in the perceptual weight they attribute to the prosodic cues.
Three experiments investigated the role of duration and intonation in the expression of emotions in natural and synthetic speech. Two sentences of an actor portraying seven emotions (neutral, joy, boredom, anger, sadness, fear, indignation) were acoustically analyzed. By copying pitch and duration of the original utterances to a monotonous one, it could be shown that both factors were sufficient to express the various emotions. In the second part, rules about intonation and duration were derived and tested. These rules were applied to resynthesized natural speech and synthetic speech generated from LPC-coded diphones. The results showed that emotions can be expressed accurately by manipulating pitch and duration in a rule-based way.
Two male Dutch talkers produced two tokens of each of 20 stop-vowel syllables (/b,d,p,t,k/ followed by /a,i,y,u/). The release bursts were separated from the voiced parts and four types of stimuli were created: Burst-only stimuli (BO), burstless stimuli (BL), stimuli with cross-spliced bursts, where burst and transitions indicate conflicting place-of-articulation information (CS), and original utterances (OR). These stimuli were presented to 20 subjects for identification. Results show the well-known vowel-dependent trade-off between burst cues and transition cues. An attempt was made to predict the confusion matrices from acoustic properties of release bursts and formant transitions using Luce’s similarity-choice model. Most of the variance (89%) in the confusion matrix for the BO condition was explained using a spectral tilt measure and a spectral compactness measure. Much of the variance (80%) for the BL condition was explained using frequencies of F2 and F3 at voicing onset and of F2 in the vowel. Using the same burst and transition measures for the CS condition the explained variance dropped to 40%. Preliminary results show better predictions when acoustic measures are used which integrate over burst and transitions.