In this contribution we look back on the last years in the history of telephone-based speech dialog systems. We will start in 1993 when the world wide first natural language understanding dialog system using a mixed-initiative approach was made accessible for the public, the well-known EVAR system from the Chair for Pattern Recognition of the University of Erlangen-Nuremberg. Then we discuss certain requirements we consider necessary for the successful application of dialog systems. Finally we present trends and developments in the area of telephone-based dialog systems.
Im Rahmen dieses Beitrags wird beschrieben, weiche Eigenschaften automatische Sprachdialogsysteme haben sollten, um Anforderungen hinsichtlich einer angenehmen und effizienten Interaktion zu erfüllen. Dabei wird für das Systemdesign der mixed-initiative Ansatz als optimal betrachtet, da die damit erreichbare Benutzerfreundlichkeit für Sprachanwendungen die Akzeptanz von Sprachdialogsystemen bei Firmen und Endbenutzern erheblich verbessern kann. Es werden die Herangehensweise bei der Entwicklung neuer Applikationen sowie das dazu notwendige Knowhow und hilfreiche Werkzeuge beleuchtet, zudem stellen die Autoren die erörterten theoretischen Konzepte anhand von Systemen, die kommerziell eingesetzt werden, beispielhaft dar.
Summary In this publication experiences with commercial spoken dialogue systems are discussed and guidelines for achieving high usability are pointed out. Different from most commercially deployed IVR (Interactive Voice Response) systems, the systems discussed in this paper belong to a new generation of real mixed-initiative spoken dialogue systems, i. e., the user may take the initiative, using full sentences, at virtually any point in time during the dialogue. We use three commercially deployed systems as example applications: the automated switchboard of a large German company, the movie information system operated by Germany´s largest multiplex cinema, and a football Bundesliga information system operated by a German media company.
The monitoring of emotional user states can help to assess the progress of human-machine-communication. If we look at specific databases, however, we are faced with several problems: users behave differently, even within one and the same setting, and some phenomena are sparse; thus it is not possible to model and classify them reliably. We exemplify these difficulties on the basis of SympaFly, a database with dialogues between users and a fully automatic speech dialogue telephone system for flight reservation and booking, and discuss possible remedies.
Es wird eine natürlichsprachliche Lösung vorgestellt, die in der Hauptverwaltung der Sixt AG seit Dezember 2003 im Einsatz ist. Das System zur automatischen Vermittlung von Anrufern wurde komplett von der Sympalog Voice Solutions GmbH auf Basis eigener Technologie entwickelt. Es werden die im Produktivbetrieb erzielten Ergebnisse präsentiert, darüberhinaus werden Probleme und Stolpersteine diskutiert, die bei der Umsetzung der Anforderungen zu lösen bzw. zu umgehen waren.
Two sets of linguistic features are developed: The first one to estimate if a single step in a dialogue between a human being and a machine is successful or not. The second set to classify dialogues as a whole. The features are based on Part-of-Speech-Labels (POS), word statistics and properties of turns and dialogues. Experiments were carried out on the SympaFly corpus, data from a real application in the flight booking domain. A single dialogue step could be classified with an accuracy of 83% (class-wise averaged recognition rate). The recognition rate for whole dialogues was 85%.
In this paper, we show how prosodic information can be used in automatic dialogue systems and give some examples of promising new approaches. Most of these examples are taken from our own work in the VERBMOBIL speech-to-speech translation system and in the EVAR train timetable dialogue system. In a 'prosodic orbit', we first present units, phenomena, annotations and statistical methods from the signal (acoustics) to the dialogue understanding phase. We show then, how prosody can be used together with other knowledge sources for the task of resegmentation if a first segmentation turns out to be wrong, and how an integrated approach leads to better results than a sequential use of the different knowledge sources; then we present a hybrid approach which is used to perform a shallow parsing and which uses prosody to guide the parsing; finally, we show how a critical system evaluation can help to improve the overall performance of automatic dialogue systems. (C) 2002 Elsevier Science B.V. All rights reserved.
In this paper we take a second look at current research issues for conversational dialogue systems addressed in [17]. We look at two systems, a movie information and a stock information system which were built based on the experiences with the train information system EVAR, described in [17].
In diesem Beitrag wollen wir Ansätze zur Interpretation gesprochener Sprache präsentieren. Wir zeigen zunächst, warum die direkte Übertragung der linguistischen Ansätze zur Interpretation von geschriebener Sprache auf gesprochene Sprache in der Regel nicht erfolgreich ist. Dies führt uns zur Entwicklung eines Konzepts für partielles Parsen, welches stochastische Verfahren und prosodische Information verwendet. Wir präsentieren vorläufige Ergebnisse für die einzelnen Wissensquellen, ein Test des Gesamtkonzepts steht noch aus.
In this paper we present an innovative approach to speech understanding which is based on a fine-grained knowledge representation automatically compiled from a semantic network and on iterative optimization.Besides allowing an efficient exploitation of parallelism, any-time capability is provided since after each iteration step a (sub-)optimal solution is always available.We apply this approach to a real-world task, which is a dialog system able to answer queries about the German train timetable.In order to speed up the search for the best interpretation of an utterance we make use of statistical methods, e.g.neural networks, n-grams, and classification trees, which are trained on application relevant utterances collected over the public telephone network.At the moment the real-time factor for interpreting the initial user's utterance is 0.7.
In this paper, we present an overview of the spoken dialogue system EVAR that was developed at the University of Erlangen. In January 1994, it became accessible over telephone line and could answer inquiries in the German language about German InterCity train connections. It has since been continuously improved and extended, including some unique features, such as the processing of out{of{vocabulary words and a exible dialogue strategy that adapts to the quality of the recognition of the user input. In fact, several diierent versions of the system have emerged, i.e. a subway information system, train and ight information systems in diierent languages, and an integrated multilingual and multifunctional system which covers German and 3 additional languages in parallel. Current research focuses on the introduction of stochastic models into the semantic analysis, on the direct integration of prosodic information into the word recognition process, on the detection of user emotion, and on multilinguality and multifunctionality.
This paper presents a new probabilistic approach to semantic analysis of speech. The problem of finding the semantic contents of a word chain is modeled as the problem of assigning semantic attributes to words. The discrete assignment function is characterized by random vectors and its probabilities. By computing the best of all possible statistically modeled assignments, we get the semantic contents of a word chain and along with it a semantic segmentation. The introduced general statistical framework has to deal with incomplete data estimation problems. These are solved applying the Expectation Maximization algorithm. We show that the well-known hidden Markov models result from the suggested theory as a specialization. Experiments prove that this approach works quite well in the domain of train-time-table inquiries for German IC/EC-train connections.
Florian Gallwitz合作论文数Sympalog Voice Solutions GmbH, Erlangen, Germany7
Christian Hacker合作论文数Pattern Recognition Lab of the Friedrich-Alexander University Erlangen-Nuremberg2