This paper presents the results of an elaborate study on pen and speech-based multimodal interaction systems. The performance o f the “COM IC” system is assessed through human factors analyses and evaluation o f the acquired multimodal data. The latter requires tools that are able to m onitor user input, system feedback, and performance of the multimodal system components. Such tools can bridge the gap between observational data and the complex process o f the design and evaluation of multimodal systems. The evaluation tool presented here is validated in a human factors study on the usability o f COMIC for design applications and can be used for semi automatic transcription of multimodal data.
This paper presents the text generation module of SmartWeb, a multimodal dialogue system. The generation module bases on NipsGen which combines SPIN, originally a parser developed for spoken language, and a tree-adjoining grammar framework for German. NipsGen allows to mix full generation with canned text.
SmartWeb aims to provide intuitive multimodal access to a rich selection of Web-based information services. We report on the current prototype with a smartphone client interface to the Semantic Web. An advanced ontology-based representation of facts and media structures serves as the central description for rich media content. Underlying content is accessed through conventional web service middleware to connect the ontological knowledge base and an intelligent web service composition module for external web services, which is able to translate between ordinary XML-based data structures and explicit semantic representations for user queries and system responses. The presentation module renders the media content and the results generated from the services and provides a detailed description of the content and its layout to the fusion module. The user is then able to employ multiple modalities, like speech and gestures, to interact with the presented multimedia material in a multimodal way.
Increased availability of mobile computing, such as personal digital assistants (PDAs), creates the potential for constant and intelligent access to up-to-date, integrated and detailed information from the Web, regardless of one's actual geographical position. Intelligent question-answering requires the representation of knowledge from various domains, such as the navigational and discourse context of the user, potential user questions, the information provided by Web services and so on, for example in the form of ontologies. Within the context of the SmartWeb project, we have developed a number of domain-specific ontologies that are relevant for mobile and intelligent user interfaces to open-domain question-answering and information services on the Web. To integrate the various domain-specific ontologies, we have developed a foundational ontology, the SmartSUMO ontology, on the basis of the DOLCE and SUMO ontologies. This allows us to combine all the developed ontologies into a single SmartWeb Integrated Ontology (SWIntO) having a common modeling basis with conceptual clarity and the provision of ontology design patterns for modeling consistency. In this paper, we present SWIntO, describe the design choices we made in its construction, illustrate the use of the ontology through a number of applications, and discuss some of the lessons learned from our experiences.
In this chapter we give an general overview of the modality fusion component of SmartKom. Based on a selection of prominent multimodal interaction patterns, we present our solution for synchronizing the different modes. Finally, we give, on an abstract level, a summary of our approach to modality fusion.
This paper presents SPIN, a semantic parser developed for sp oken dialog systems. The parser provides a powerful rule lan gu ge for an easy and efficient creation of the rule set. Important featur s of the rule language include order-independent matching , built-in support for referring expressions, rule ordering, constraints and ction functions. On the basis of an example utterance the ad vantages of the introduced features are shown. The increased processing co mplexity caused by the powerful rule language is handled by a new parsing approach that delivers sufficient performance for rule sets that are typical for dialog systems. We also show how the pars er can be used for text generation. The paper closes with an evaluation of t he parser performance showing that the approach is well suit ed for dialog systems. SPIN: Pomenski parser za sisteme govorjenega dialoga V članku je predstavljen SPIN, semantični razčlenjeval nik, ki je bil razvit za sisteme govorjenega dialoga. Razčl enjevalnik ima zmogljiv jezik za tvorjenje pravil, ki enostavno in učinkovito tvor i nabor pravil. Pomembne značilnosti jezika za tvorjenje p ravil so ujemanje ne glede na besedni red, vgrajena podpora referenčnim izra zom, razvrstitev pravil, omejitve in opravilne funkcije. N a podlagi primera izjave so prikazane prednosti vpeljanih lastnosti. Poveč ana kompleksnost procesiranja zaradi zmogljivega jezika z a tvorjenje pravil obvladujemo z novim pristop k skladenjski analizi, ki ima za dosten učinek pri naboru pravil, značilnih za sisteme dia log . Prikažemo tudi, kako je lahko parser uporabljen za tvorjenje besedila . Č nek zaključimo z vrednotenjem delovanja parserja, ki p o aže, da je pristop primeren za sisteme dialoga.
This paper presents a semantic parser for spoken dialogue systems. The parser is designed especially for the analysis of free word order languages by providing a feature called orderindependent matching. We describe how this feature allows writing of rules for free word order languages in an elegant way (using German as example language) and how it increases the robustness against speech recognition errors. As orderindependent matching makes efficient parsing more difficult, we present a new parsing approach which provides efficient processing for rule bases that are, according to our experience, typical for spoken dialogue systems. The key feature of the parsing approach is a fixed application order of the rules to prune irrelevant results. A preliminary evaluation of the parser shows that this approach works very well in real-world dialogue systems.
Experience shows that decisions in the early phases of the development of a multimodal system prevail throughout the life-cycle of a project. The distributed architecture and the requirement for robust multimodal interaction in our project SmartWeb resulted in an approach that uses and extends W3C standards like EMMA and RDFS. These standards for the interface structure and content allowed us to integrate available tools and techniques. However, the requirements in our system called for various extensions, e.g., to introduce result feedback tags for an extended version of EMMA. The interconnection framework depends on a commercial telephone voice dialog system platform for the dialog-centric components while the information access processes are linked using web service technology. Also in the area of this underlying infrastructure, enhancements and extensions were necessary. The first demonstration system is operable now and will be presented at the Football World Cup 2006 in Germany.
This paper presents a semantic parser for spoken dialogue systems. The parser is designed especially for the analysis of free word order languages by providing a feature called order-independent matching. We describe how this feature allows writing of rules for free word order languages in an elegant way (using German as example language) and how it increases the robustness against speech recognition errors. As order-independent matching makes efficient parsing more difficult, we present a new parsing approach which provides efficient processing for rule bases that are, according to our experience, typical for spoken dialogue systems. The key feature of the parsing approach is a fixed application order of the rules to prune irrelevant results. A preliminary evaluation of the parser shows that this approach works very well in real-world dialogue systems.
This document presents the results of an elaborate human factors study on multimodal interaction in bathroom design Author: Louis Vuurpijl, Louis ten Bosch, Stéphane Rossignol, Andre Neumann, Ralf Engel, Norbert Pfleger Reviewers: Els den Os, Lou Boves, Michael White
In this paper we report on ongoing experiments with an advanced multimodal system for applications in architectural design. The system supports uninformed users in entering the relevant data about a bathroom that must be refurnished, and is tested with 28 subjects. First, we describe the IST project COMIC, which is the context of the research. We explain how the work in COMIC goes beyond previous research in multimodal interaction for eWork and eCommerce applications that combine speech and pen input with speech and graphics output: in design applications one cannot assume that uninformed users know what they must do to satisfy the system’s expectations. Consequently, substantial system guidance is necessary, which in its turn creates the need to design a system architecture and an interaction strategy that allow the system to control and guide the interaction. The results of the user tests show that the appreciation of the system is mainly determined by the accuracy of the pen and speech input recognisers. In addition, the turn taking protocol needs to be improved.
The development of an intelligent user interface that supports multimodal access to multiple applications is a challenging task. In this paper we present a generic multimodal interface system where the user interacts with an anthropomorphic personalized interface agent using speech and natural gestures. The knowledge-based and uniform approach of SmartKom enables us to realize a comprehensive system that understands imprecise, ambiguous, or incomplete multimodal input and generates coordinated, cohesive, and coherent multimodal presentations for three scenarios, currently addressing more than 50 different functionalities of 14 applications. We demonstrate the main ideas in a walk through the main processing steps from modality fusion to modality fission.
This paper describes a language understanding module for spoken dialogue systems producing frame based semantic output. The presented approach adapts ideas from production systems to the task of language understanding. It interleaves in a new manner template-driven cascaded word-to-frame transformation with syntactic analysis. The advantages over conventional parsers are the flexible output structure being independent of the syntactic structure in a wide range, the ability to use different levels of syntactic analysis at the same time and better support for relatively free word order languages like German. Other important properties are robustness, the capability to process complex utterances, and the easy creation of knowledge bases. A preliminary evaluation shows promising results.
This paper describes a novel functionality of the VERBMOBIL system, a large scale translation system designed for spontaneously spoken multilingual negotiation dialogues. The task is the on-demand generation of dialogue scripts and result summaries of dialogues. We focus on summary generation and show how the relevant data are selected from the dialogue memory and how they are packed into an appropriate abstract representation. Finally, we demonstrate how the existing generation module of VERBMOBIL was extended to produce multilingual and result summaries from these representations.
This chapter explains the major functionality of the dialog module in Verbmobil. Dialog knowledge is needed for context sensitive speech translation as well as for the automatic generation of dialog result summaries. Our component produces necessary structures for both purposes and stores them in a centrally accessible data repository — the dialog memory. The structures are based on robustly extracted shallow data which are corrected, extended and structured by our dialog processor. We use time and object completion algorithms to collect context data and compute inter-object relations to infer relevance for summarization. The resulting structures are used by the document generator for dialog minutes and summaries, and by the context evaluation module for translation disambiguation.
We present the multilingual summarization functionality for VERB-MOBIL, a speech translation system. We reuse resources of the system to create a summary. After content extraction, we interpret the results in the dialog context. A summary generator provides the input to generation. A first evaluation indicates the feasibility of the approach.
Anupriya Ankolekar合作论文数Human-Computer Interaction Institute at Carnegie Mellon University.1
Anselm Blocher合作论文数Deutsches Forschungszentrum fur Kunstliche Intelligenz GmbH1
Daniel Oberle合作论文数Institute AIFB1