Prosody is used to improve the performance of the automatic speech translation system VERBMOBIL [8]. In our earlier work we have developed efficient and robust word-based features that describe F0, energy, speaking rate, and pauses. These features were used to classify prosodic events. We achieved the best recognition results with 95-dimensional feature vectors that describe a context of +/2 words [4]. In the experiments presented in this paper we additionally used Part-Of-Speech (POS) flags as features. The POS features are based on a hierarchical POS label system with up to 15 classes. The 95-dimensional acoustic-prosdic feature vectors are augmented with up to 105 POS features that describe a context of up to +/3 words. The new features significantly improved the recognition of phrase boundaries, phrase accents and question mood; the recognition errors could be reduced by up to 16.7%. The POS flags allow a neural network (NN) to learn a simple language model. We show that it is important to include this syntactic knowledge during the classification of the acoustic-prosodic features instead of combining it later. This implies that there is some kind of synergy: The POS information helps to correctly classify the acoustic observations. The results presented in this paper provide an effective way to improve the recognition of prosodic events with almost no computational overhead.
In our previous research, we have shown that prosody can be used to dramatically improve the performance of the automatic speech translation system VERBMOBIL [16]. The methods to classify prosodic events have been developed on the German subcorpus of the VERBMOBIL speech database. In this paper we describe how the methods that we developed on the German subcorpus can be applied to other languages. Experiments show that these methods are suited for English and Japanese, as well. Efficiency problems are addressed and a new set of features is presented. The new set of features facilitates a multilingual module for prosodic processing. We present an architecture for such a multilingual module and discuss the advantages of this approach compared to an approach that uses separate modules for different languages. This multilingual module and the new feature set are evaluated w.r.t. computation time, memory requirement, and classification performance. The results show that the memory requirement can be reduced by 78%, whereas the recognition accuracy does not decrease.
In this paper, we present an integrated approach for recognizing both the word sequence and the syntactic-prosodic structure of a spontaneous utterance. The approach aims at improving the performance of the understanding component of speech understanding systems by exploiting not only acoustic and syntactic information, but also prosodic information directly within the speech recognition process. Whereas spoken utterances are commonly modelled as unstructured word sequences in the speech recognizer, our approach includes phrase (or clause) boundary information in the language model, and provides HMMs to model the acoustic and prosodic characteristics of phrase boundaries and disfluencies. This methodology has two major advantages compared to pure word–based speech recognizers. First, additional syntactic information is determined by the speech recognizer which facilitates parsing and resolves syntactic and semantic ambiguities. Second, the integrated model yields significantly better word accuracies than the traditional word–based approach.
Automatic methods for grapheme-to-phoneme (G2P) and phoneme-to-grapheme (P2G) conversion have become very popular in recent years. Their performance has improved considerably, while at the same time these developments required less input from expert lexicographers. Continuing in this tradition we will present in this paper a data-driven, language-independent approach called MASSIVE(1) with which it is possible to create efficient online modules for automatic symbol mapping. Our framework is solely based on statistical methods for training and run-time and has been optimized for P2G conversion in the context of spoken inquiries to the Semantic Web, an issue researched in the SmartWeb project(2). MASSIVE systems can be trained using a pronunciation lexicon, the output of a phone recognizer or any other suitable set of corresponding symbol strings. Successful tests have been performed on German and English data sets.
Automatic methods for grapheme-to-phoneme (G2P) and phoneme-to-grapheme (P2G) conversion have become very popular in recent years. Their performance has improved considerably, while at the same time these developments required less input from expert lexicographers. Continuing in this tradition we will present in this paper a data-driven, language-independent approach called MASSIVE with which it is possible to create efficient online modules for automatic symbol mapping. Our framework is solely based on statistical methods for training and run-time and has been optimized for P2G conversion in the context of spoken inquiries to the Semantic Web, an issue researched in the SmartWeb project. MASSIVE systems can be trained using a pronunciation lexicon, the output of a phone recognizer or any other suitable set of corresponding symbol strings. Successful tests have been performed on German and English data sets.
In this contribution we look back on the last years in the history of telephone-based speech dialog systems. We will start in 1993 when the world wide first natural language understanding dialog system using a mixed-initiative approach was made accessible for the public, the well-known EVAR system from the Chair for Pattern Recognition of the University of Erlangen-Nuremberg. Then we discuss certain requirements we consider necessary for the successful application of dialog systems. Finally we present trends and developments in the area of telephone-based dialog systems.
Es wird ein Verfahren vorgestellt, den emotionalen Zustand eines Sprechers zu überwachen. Dabei wird für jede Äußerung bewertet, ob sich der Sprecher eher in einem neutralen oder ärgerlichen Zustand befindet. Integriert in ein automatisches Dialogsystem kann dieses Verfahren dazu beitragen, zu verhindern, dass Anrufer bei Verständnisproblemen einfach auflegen, z.B. indem das System vorher den Anruf an einen Call-Center-Agenten weiterleitet.
The paper describes prosodic annotation procedures of the GOPOLIS Slovenian speech data database and methods for automatic classification of different prosodic events. Several statistical parameters concerning duration and loudness of words, syllables and allophones were computed for the Slovenian language, for the first time on such a large amount of speech data. The evaluation of the annotated data showed a close match between automatically determined syntactic-prosodic boundary marker positions and those obtained by a rule-based approach.
The ‘case’ this paper is dealing with is prosody research at the Chair for Pattern Recognition at the University of Erlangen– Nuremberg during the last fifteen years. We want to show how this mirrors the development of prosody research within automatic speech understanding in general. We sketch the realm of prosody in automatic speech understanding and relate the projects conducted to the research topics. This is illustrated in more detail with experimental results obtained within the last two years. Emphasis is put on the interplay between prosodic information and other knowledge sources.
In this paper, we show how prosodic information can be used in automatic dialogue systems and give some examples of promising new approaches. Most of these examples are taken from our own work in the VERBMOBIL speech-to-speech translation system and in the EVAR train timetable dialogue system. In a 'prosodic orbit', we first present units, phenomena, annotations and statistical methods from the signal (acoustics) to the dialogue understanding phase. We show then, how prosody can be used together with other knowledge sources for the task of resegmentation if a first segmentation turns out to be wrong, and how an integrated approach leads to better results than a sequential use of the different knowledge sources; then we present a hybrid approach which is used to perform a shallow parsing and which uses prosody to guide the parsing; finally, we show how a critical system evaluation can help to improve the overall performance of automatic dialogue systems. (C) 2002 Elsevier Science B.V. All rights reserved.
In this paper, we present an integrated approach for recognizing both the word sequence and the syntactic-prosodic structure of a spontaneous utterance. The approach aims at improving the performance of the understanding component of speech understanding systems by exploiting not only acoustic-phonetic and syntactic information, but also prosodic information directly within the speech recognition process. Whereas spoken utterances are typically modelled as unstructured word sequences in the speech recognizer, our approach includes phrase boundary information in the language model and provides HMMs to model the acoustic and prosodic characteristics of phrase boundaries. This methodology has two major advantages compared to purely word-based speech recognizers. First, additional syntactic-prosodic boundaries are determined by the speech recognizer which facilitates parsing and resolve syntactic and semantic ambiguities. Second - after having removed the boundary information from the result of the recognizer - the integrated model yields a 4% relative word error rate (WER) reduction compared to a traditional word recognizer. The boundary classification performance is equal to that of a separate prosodic classifier operating on the word recognizer output, thus making a separate classifier unnecessary for this task and saving the computation time involved. Compared to the baseline word recognizer, the integrated word-and-boundary recognizer does not involve any computational overhead. (C) 2002 Elsevier Science B.V. All rights reserved.
For the classification of boundaries and accents in German and English spontaneous speech in the VERBMOBIL project (speech to speech translation system), we use a large prosodic feature vector; duration features represent the most important feature class. They are computed in three different ways: (1) The word duration is normalized with respect to the ‘expected’ word duration: DURNORM; (2) Duration is normalized as for the number of syllables in the word: DURSYLL; (3) The absolute duration value DURABS of a word is taken. Normally, we use all these feature classes simultaneously. In the present paper, we have a look at the impact of each of these duration classes separately. In addition, we use partof-speech (POS) information as a further knowledge source. It turns out that throughout, the best feature class, if used alone, is DURABS, followed by DURSYLL, and third comes DURNORM. Best results are achieved by using all feature classes together. With POS information, better results can be achieved than without. This effect is larger for accent classification than for boundary classification, and much larger in combination with DURNORM than in combination with DURSYLL or DURABS. These results indicate that especially DURABS does not only encode prosodic but to a large extent syntactic POS information as well: content words are normally more prone to be accentuated than function words, and at the same time, they tend to be longer. This information is of course lost if duration is normalized, as is the case for DURSYLL and DURNORM.
In the focus of this paper is a comparison of the most relevant prosodic features/feature classes for the classification of boundaries and accents in German and in English. Principal components were computed based on a large prosodic feature vector; these principal components were used as predictor variables in a Linear Discriminant analysis as well as in a Classification and Regression Tree. The number of the most relevant principal components was between three and five; for both languages and for boundary and accent classification alike, most important were principal components modelling duration, in combination with energy, followed by pauses and F0.
Nowadays modern automatic dialogue systems are able to understand complex sentences instead of only a few commands like Stop or No.In a call-center, such a system should be able to determine in a critical phase of the dialogue if the call should be passed over to a human operator.Such a critical phase can be indicated by the customer's vocal expression.Other studies prooved that it is possible to distinguish between anger and neutral speech w i t h prosodic features alone.Subjects in these studies were mostly people acting or simulating emotions like anger.In this paper we use data from a so-called Wizard of O z (WoZ) scenario to get more realistic data instead of simulated anger.As shown below, the classi cation rate for the two classes "emotion" (class E) and "neutral" (class :E) is signi cantly worse for these more realistic data.Furthermore the classi cation results are heavily speaker dependent.Prosody alone might t h us not be su cient and has to be supplemented by the use of other knowledge sources such as the detection of repetitions, reformulations, swear words, and dialogue acts.
Florian Gallwitz合作论文数Sympalog Voice Solutions GmbH, Erlangen, Germany10