In this paper, we present our framework for Augmented and Mixed Reality touristic applications, which we are currently developing with the city of Saarlouis (Germany). It aims to provide tourists and visitors with a new augmented experience to discover the unique history of the city. Whereas most Augmented and Mixed Reality touristic applications are designed for a single user without any human city guide, our app, called SaAR-Louis, offers a more immersive touristic experience as it displays holographic information via Augmented and Mixed Reality on HoloLens/iPad in addition to the explanations of the tour guide. Moreover, thanks to geolocation and 3D scanning of the surroundings, our application provides tourists with the unique experience of discovering 3D reconstructed ancient/old buildings that no longer exist and offers multimodal interaction with a 3D Avatar that gives insights about the historic buildings. In the following, we will explain the app architecture, the pipeline of creating the 3D virtual revival of the city of Saarlouis, and the challenges behind designing, conceiving, implementing, and deploying the SaAR-Louis app.
The municipal utility supervises all aspects of the city's public utilities, e.g., the power grid. For example, power grid maintenance requires the technician to know where underground cables are installed while having their hands free for maintenance work. Using an immersive headset such as HoloLens provides a technician with these capabilities. This contribution presents DENKI (Data, Energy, Network, Knowledge, Interaction). This HoloLens application supports technicians of the municipal utility "Stadtwerke Saarlouis GmbH" with a detailed overview of the equipment, its utilization, and possibilities to investigate malfunctions. We will exemplify some app functions.
With the emerging Internet of Things (IoT) techniques in smart home applications, artificial intelligence (AI), and highly interoperable IoT s ystems enable the development of context-sensitivemulti-domain services in smart homes [1]. However, while such systems create enormous challenges regarding security and privacy, the IoT practitioners may overlook certain security and privacy concerns such as European Union (EU) General Data Protection Regulation (GDPR). This paper describes the necessities to consider privacy- and security-related challenges for smart living platforms. Core elements of this contribution are a user survey to detect key aspects to fulfill users' expectations and an in-detail description of a Gaia-X-compatible software technology stack for the smart living domain. The concept will be applied to a smart kitchen use case.
Extended Reality (XR) devices have great potential to become the next wave in mobile interaction. They provide powerful, easy-to-use Augmented Reality (AR) and/or Mixed Reality (MR) in conjunction with multimodal interaction facilities using gaze, gesture, and speech. However, current implementations typically lack a coherent semantic representation for the virtual elements, backend-communication, and dialog capabilities. Existing devices are often restricted to mere command and control interactions. To improve these shortcomings and realize enhanced system capabilities and comprehensive interactivity, we have developed a flexible modular approach that integrates powerful back-end platforms using standard API interfaces. As a concrete example, we present our distributed implementation of a multimodal dialog system on the Microsoft Hololens®. It uses the SiAM-dp multimodal dialog platform as a back-end service and an Open Semantic Framework (OSF) back-end server to extract the semantic models for creating the dialog domain model.
We present a multimodal dialogue system that allows doctors to interact with a medical decision support system in virtual reality (VR). We integrate an interactive visualization of patient records and radiology image data, as well as therapy predictions. Therapy predictions are computed in real-time using a deep learning model.
In this applied research paper, we describe an architecture for seamlessly integrating factory workers in industrial cyber-physical production environments. Our human-in-the-loop control process uses novel input techniques and relies on state-of-the-art industry standards. Our architecture allows for real-time processing of semantically annotated data from multiple sources (e.g., machine sensors, user input devices) and real-time analysis of data for anomaly detection and recovery. We use a semantic knowledge base for storing and querying data (http://www.metaphacts.com) and the Business Process Model and Notation (BPMN) for modelling and controlling the process. We exemplify our industrial solution in the use case of the maintenance of a Siemens gas turbine. We report on this case study and show the advantages of our approach for smart factories. An informal evaluation in the gas turbine maintenance use case shows the utility of automated anomaly detection and handling: workers can fill in paper-based incident reports by using a digital pen; the digitised version is stored in metaphacts and linked to semantic knowledge sources such as process models, structure models, business process models, and user models. Subsequently, automatic maintenance and recovery processes that involve human experts are triggered.
Gaze is known to be a dominant modality for conveying spatial information, and it has been used for grounding in human-robot dialogues. In this work, we present the prototype of a gaze-supported multi-modal dialogue system that enhances two core tasks in human-robot collaboration: 1) our robot is able to learn new objects and their location from user instructions involving gaze, and 2) it can instruct the user to move objects and passively track this movement by interpreting the user's gaze. We performed a user study to investigate the impact of different eye trackers on user performance. In particular, we compare a head-worn device and an RGB-based remote eye tracker. Our results show that the head-mounted eye tracker outperforms the remote device in terms of task completion time and the required number of utterances due to its higher precision.
In this paper, we describe a real-time knowledge acquisition and anomaly handling architecture in a maintenance scenario of an industrial cyber-physical production environment. We use the Business Process Model and Notation (BPMN) for modeling and controlling of the maintenance procedure. Automatic handwriting and pen gesture recognition is combined with a networked smart pen that is used on semantically structured paper forms. Here, we discuss our architecture with BPMN-modeled workflows for real-time data processing. All detected anomalies are automatically associated to the corresponding knowledge sources such as the process model, structure model, business process model and user models; linked Semantic Media-Wiki (SMW) pages are built up accordingly. Our architecture provides for a seamless integration of real-time knowledge sources by smart pen technology and real-time CPS processing by a BPMN engine handling the execution and coordination of the modeled interactions.
In this article we propose a generic framework for ontological intention recognition and action recommendation involving distributed and networked active digital object memories (ADOMe) for production factories and their components as well as individually manufactured products. For this purpose, a concrete and complex application scenario in the industrial environment is designed. In addition, the developed approach provides an important contribution to the utilization of ADOMes and their networking and intelligent representation of consumption, savings and savings potential. Furthermore, the processes and tasks of the proactive intention recognition component reflect the behavior and interaction of man and machine or man and object. Such a system can provide context-aware action recommendations without any explicit request by end users. Its prototypical realization is realized under integrated use of specially designed methods and technologies, such as the instrumentation of objects with ADOMes, process and role models, and rule recommendation procedures.
We will show how to build innovative multimodal dialog user interfaces that integrate multiple heterogeneous web services as data sources on the basis of the Ontology-based Dialog Platform (ODP). More specifically, we will describe how to exploit ODP's well-defined extension points and how generic ODP processing modules can be adopted, in order to support a rapid dialog system engineering process. By means of the latest ODP-based educational information system CIRIUS and the ODP workbench, a set of Eclipse-based editors and tools, we demonstrate step-by-step along the generic multimodal dialog processing chain what has to be done for developing a new multimodal dialog user interface for a specific application domain.
This paper describes the formalism of S--TAGs and the parsing algorithm implementedin VERBMOBIL. Furthermore the language covered by the germangrammar is described. Finally we list examples together with the execution timerequired for their processing.31 Motivation
The Automatic Content Linking Device monitors a conversation and uses automatically recognized words to retrieve documents that are of potential use to the participants. The document set includes project related reports or emails, transcribed snippets of past meetings, and websites. Retrieval results are displayed at regular intervals.
The Automatic Content Linking Device (ACLD) is a just-in-time retrieval system that monitors an ongoing conversation or a monologue and enriches it with potentially related documents, including transcripts of past meetings, from local repositories or from the Internet. The linked content is displayed in real-time to the participants in the conversation, or to users watching a recorded conversation or talk. The system can be demonstrated in both settings, using real-time automatic speech recognition (ASR) or replaying offline ASR, via a flexible user interface that displays results and provides access to the content of past meetings and documents.
The Automatic Content Linking Device (ACLD) is a just-in-time multimedia retrieval system that monitors and supports the conversation among a small group of people within a meeting. The ACLD retrieves from a repository, at regular intervals, information that might be relevant to the group's activity, and presents it through a graphical user interface (GUI). The repository contains documents from past meetings such as slides or reports along with processed meeting recordings; in parallel, Web searches are run as well. The acceptance by users of such a system depends considerably on the GUI, along with the performance of retrieval. The trade-off between informativeness and unobtrusiveness is studied here through the design of a series of GUIs. The requirements and feedback collected while demonstrating the successive versions show that users vary considerably in their preferences for a given style of interface. After studying two extreme options, a widget vs. a wide-screen UI, we conclude that a modular UI, which can be flexibly structured and resized by users, is the most sensible design for a just-in-time multimedia retrieval system.
As multimodal data becomes easier to record and store, the question arises as to what practical use can be made of archived corpora, and in particular what tools allowing efficient access to it can be built. We use the AMI Meeting Corpus as a case study to build an automatic content linking device, i.e. a system for real-time data retrieval. The corpus provides not only the data repository, but is used also to simulate ongoing meetings for development and testing of the device. The main features of the corpus are briefly described, followed by an outline of data preparation steps prior to indexing, and of the methods for building queries from ongoing meeting discussions, retrieving elements from the corpus and accessing the results. A series of user studies based on prototypes of the content linking device have confirmed the relevance of the concept, and methods for task-based evaluation are under development.
The AMIDA Automatic Content Linking Device (ACLD) monitors a conversation using automatic speech recognition (ASR), and uses the detected words to retrieve documents that are of potential use to the participants in the conversation. The document set that is available includes project related documents such as reports, memos or emails, as well as snippets of past meetings that were transcribed using offline ASR. In addition, results of Web searches are also displayed. Several visualisation interfaces are available.
The AMIDA Automatic Content Linking Device (ACLD) is a just-in-time document retrieval system for meeting environments. The ACLD listens to a meeting and displays information about the documents from the group's history that are most relevant to what is being said. Participants can view an outline or the entire content of the documents, if they feel that these documents are potentially useful at that moment of the meeting. The ACLD proof-of-concept prototype places meeting-related documents and segments of previously recorded meetings in a repository and indexes them. During a meeting, the ACLD continually retrieves the documents that are most relevant to keywords found automatically using the current meeting speech. The current prototype simulates the real-time speech recognition that will be available in the near future. The software components required to achieve these functions communicate using the Hub, a client/server architecture for annotation exchange and storage in real-time. Results and feedback for the first ACLD prototype are outlined, together with plans for its future development within the AMIDA EU integrated project. Potential users of the ACLD supported the overall concept, and provided feedback to improve the user interface and to access documents beyond the group's own history.
Speech disfluencies are very common in our everyday life and considerably affect NLP systems, which makes systems that can detect or even repair them highly desirable. Previous research achieved good results in the field of disfluency detection but only in subsets of the disfluency types. The aim of this study was to develop a technology that is able to cope with a broad field of dislluency types. A thorough investigation of our corpus led us to a detection design where basic rule-matching techniques are complemented with machine learning and N-gram based approaches. In this paper, we describe the different detection techniques, each specialized on its own disfluency domain and the results we gained.
We present SUVI - Summary Visualizer -, a generic layout-tool which displays a multimodal summary of a meeting in either a story-board style or newspaper-style. The system relies on constraint solving techniques in such a way that the two layouts have been extensively modeled in a series of constraints representing the underlying design knowledge. While the story-board aims to give the reader an overview of the chronological sequence of the meeting, the newspaper-layout focuses on presenting the topics of a meeting depending on their relevance. We also show two methods for connecting the whole AMI meeting corpus as a large input resource for the story-board part of SUVI and present a first end-to-end implementation of our system.
Alassane Ndiaye合作论文数DFKI - German Research Center for Artificial Intelligence5
Jörg Baus合作论文数Intelligent User Interfaces (IUI) group2