This paper introduces a multimodal dialogue system, which facilitates access to the digital home for people who suffer from cognitive disabilities. The user interface is implemented on a smartphone and allows interaction via speech and pointing gestures. A consistent control concept and a meaningful graphical design allow an intuitive handling of the device. The ISO standard URC enables the system to supervise and control heterogeneous home devices. One focal point is collaborative problem solving which supports users with memory deficits in planning and pursuing objectives. Furthermore, the reaction of the system is adapted to user and interaction context.
This paper presents a full-fledged multimodal dialogue system for accessing multimedia content in home environments from both portable media players and online sources. We will mainly focus on two aspects of the system that provide the basis for a natural interaction: (i) the automatic processing of named entities which permits the incorporation of dynamic data into the dialogue (e.g., song or album titles, artist names, etc.) and (ii) general multimodal interaction patterns that are bound to ease the access to large sets of data.
This paper presents a full-fledged multimodal dialogue system for accessing multimedia content in home environments from both portable media players and online sources. We will mainly focus on two aspects of the system that provide the basis for a natural interaction: (i) the automatic processing of named entities which permits the incorporation of dynamic data into the dialogue (e.g., song or album titles, artist names, etc.) and (ii) general multimodal interaction patterns that are bound to ease the access to large sets of data.
This paper presents a full-fledged multimodal dialogue system for accessing multimedia content in home environments from both portable media players and online sources. We will mainly focus on two aspects of the system that provide the basis for a natural interaction: (i) the automatic processing of named entities which permits the incorporation of dynamic data into the dialogue (e.g., song or album titles, artist names, etc.) and (ii) general multimodal interaction patterns that are bound to ease the access to large sets of data.
SmartWeb aims to provide intuitive multimodal access to a rich selection of Web-based information services. We report on the current prototype with a smartphone client interface to the Semantic Web. An advanced ontology-based representation of facts and media structures serves as the central description for rich media content. Underlying content is accessed through conventional web service middleware to connect the ontological knowledge base and an intelligent web service composition module for external web services, which is able to translate between ordinary XML-based data structures and explicit semantic representations for user queries and system responses. The presentation module renders the media content and the results generated from the services and provides a detailed description of the content and its layout to the fusion module. The user is then able to employ multiple modalities, like speech and gestures, to interact with the presented multimedia material in a multimodal way.
The interactive scenarios realized in the two prototypes of Virtual Human require an approach that allows humans and virtual characters to interact naturally and flexibly. In this article we present how the autonomous control of the virtual characters and the interpretation of user interactions is realized in the Conversational Dialogue Engine (CDE) framework. For each virtual and real interlocutor one CDE is responsible for dialogue processing. We will introduce the knowledge needed for the CDE-approach and present the modules of a CDE. The real-time requirement resulted in the integrated processing of deliberative and reactive processing, which is needed, e.g., to generate an appropriate nonverbal behavior of virtual characters.
Natural multimodal interaction with realistic virtual characters provides rich opportunities for entertainment and education. In this paper we present the current VIRTUALHUMAN demonstrator system. It provides a knowledge-based framework to create interactive applications in a multi-user, multi-agent setting. The behavior of the virtual humans and objects in the 3D environment is controlled by interacting affective conversational dialogue engines. An elaborate model of affective behavior adds natural emotional reactions and presence of the virtual humans. Actions are defined in a XML-based markup language that supports the incremental specification of synchronized multimodal output. The system was successfully demonstrated during CeBIT 2006.
In this chapter we give an general overview of the modality fusion component of SmartKom. Based on a selection of prominent multimodal interaction patterns, we present our solution for synchronizing the different modes. Finally, we give, on an abstract level, a summary of our approach to modality fusion.
We present an extension to a comprehensive context model that has been successfully employed in a number of practical conversational dialogue systems. The model supports the task of multimodal fusion as well as that of reference resolution in a uniform manner. Our extension consists of integrating implicitly mentioned concepts into the context model and we show how they serve as candidates for reference resolution.
We present an extension to a comprehensive context model that has been successfully employed in a number of practical conversational dialogue systems. The model supports the task of multimodal fusion as well as that of reference resolution in a uniform manner. Our extension consists of integrating implicitly mentioned concepts into the context model and we show how they serve as candidates for reference resolution.
Current commercial dialog systems show only limited capabilities with regard to the phenomena occurring in spontaneous, natural dialog. Many research prototypes, in contrast, are already able to deal with a great number of phenomena but lack the clarity and maintainability of commercial systems. In this paper we present a framework for developing advanced multimodal dialog systems designed to bridge this gap. Index Terms: multimodal dialogue systems, commercial applications.
We describe extensions to VirtualHuman, a multimodal dialogue system. It uses an interactive game scenario to demonstrate real-time multi-party dialogue between two human users and three virtual characters. The focus is to make the interaction more natural, robust and flexible. Here, we address issues of speech recognition in noisy environments, resolution of spatial references, and enhancements in the character interactions.
SummaryWe provide a formal description of the fundamental nonmonotonic operation used for discourse modeling. Our algorithm—overlay—consists of a default unification algorithm together with an elaborate scoring functionality. In addition to motivation and highlighting examples from the running system, we give some future directions.
Contextual information plays a crucial role in nearly every conversational setting. When people engage in conversations they rely on what has previously been uttered or done in various ways. Some nonverbal actions are ambiguous when viewed on their own. However, when viewed in their context of use their meaning is obvious. Autonomous virtual characters that perceive and react to events in conversations just like humans do also need a comprehensive representation of this contextual information. In this paper we describe the design and implementation of a comprehensive context model for virtual characters.
SummaryWe provide a discription of the robust and generic discourse module that is the central repository of contextual information in SmartKom. We tackle discourse modeling by using a three-tiered discourse structure enriched by partitions together with a local and global focus structure. For the manipulation of the discourse structures we use unification and a default unification operation enriched with a metric mirroring the similarity of competing structures called Overlay. We show how a wide variety of naturally occuring multimodal phenomena, in particular, short utterances including elliptical and referring expressions, can be processed in a generic and robust way. As all other modules of the SmartKom backbone, DiM relies on the a domain ontology for the representation of user intentions. Finally, we show that our approach is robust against phenomena caused by imperfect recognition and analysis of user actions.
This paper presents the results of an elaborate study on pen and speech-based multimodal interaction systems. The performance o f the “COM IC” system is assessed through human factors analyses and evaluation o f the acquired multimodal data. The latter requires tools that are able to m onitor user input, system feedback, and performance of the multimodal system components. Such tools can bridge the gap between observational data and the complex process o f the design and evaluation of multimodal systems. The evaluation tool presented here is validated in a human factors study on the usability o f COMIC for design applications and can be used for semi automatic transcription of multimodal data.
Synchronizing user and system actions in a real-time virtual reality environment is a challenging task. Key components of a dialogue system like speech recognition, discourse processing, speech generation and synthesis all contribute significant delays to the response time. With human interlocutors, however, a continuous flow of conversation is important as any implausible gap may cause confusion with respect to floor management. In this paper, we describe some cases of how to make sensible use of information available at each processing step to be able to give early useful feedback even while processing.
We introduce an approach to multimodal generation of verbal and nonverbal contributions for virtual characters in a multiparty dialogue scenario. This approach addresses issues of turn-taking, is able to synchronize the different modalities in real-time, and supports fixed utterances as well as utterances that are assembled by a full-fledged tree-based text generation algorithm. The system is implemented in a first version as part of the second VirtualHuman demonstrator.
Anselm Blocher合作论文数Deutsches Forschungszentrum fur Kunstliche Intelligenz GmbH2
Alassane Ndiaye合作论文数DFKI - German Research Center for Artificial Intelligence1