In this paper, we show how prosodic information can be used in automatic dialogue systems and give some examples of promising new approaches. Most of these examples are taken from our own work in the VERBMOBIL speech-to-speech translation system and in the EVAR train timetable dialogue system. In a 'prosodic orbit', we first present units, phenomena, annotations and statistical methods from the signal (acoustics) to the dialogue understanding phase. We show then, how prosody can be used together with other knowledge sources for the task of resegmentation if a first segmentation turns out to be wrong, and how an integrated approach leads to better results than a sequential use of the different knowledge sources; then we present a hybrid approach which is used to perform a shallow parsing and which uses prosody to guide the parsing; finally, we show how a critical system evaluation can help to improve the overall performance of automatic dialogue systems. (C) 2002 Elsevier Science B.V. All rights reserved.
During the last years, we have been working on the automatic classi(cid:12)cation of boundaries and accents in the German VERBMOBIL (VM) project (human-human communication, appointment scheduling dialogues). A sub-corpus was annotated manually with prosodic bound- ary and accent labels, and neural networks (NN) trained with a large set of prosodic features were used for auto- matic classi(cid:12)cation. The classi(cid:12)cation of boundaries could be improved markedly with a combination of the NN with a language model (LM) that was trained with manually annotated syntactic-prosodic boundary labels in a much larger sub-corpus. Here we show how a combination of NN with LM along similar lines can be used for an im- provement of accent classi(cid:12)cation as well. For the training of the LM, accents are annotated automatically in the transliteration with the help of a rule{based system that uses part{of{speech (POS) as well as other linguis- tic/phonological information.
In this paper, we describe an approach that allows us to annotate accent position in Germanspontaneous speech with the help of syntactic--prosodic phrase boundary labels (the so--called M labels). The data are taken from the Verbmobil--corpus. Two factors are mainlyrelevant for such a rule--based assignment of accent to a word in a phrase: the part--of--speech of this word, and the position of this word in the phrase. Acoustic--perceptual accentlabels are classified with a...
Florian Gallwitz合作论文数Sympalog Voice Solutions GmbH, Erlangen, Germany1