EVA1 is describing a new class of emotion-aware autonomous systems delivering intelligent personal assistant functionalities. EVA requires a multi-disciplinary approach, combining a number of critical building blocks into a cybernetics systems/software architecture: emotion aware systems and algorithms, multimodal interaction design, cognitive modelling, decision making and recommender systems, emotion sensing as feedback for learning, and distributed (edge) computing delivering cognitive services.
Technologies that allow autonomous robots and computer systems to quickly recognize and interact with individuals in a group setting has the potential to enable a wide range of personalized experiences. However, existing solutions fail to both identify and locate individuals with enough speed to enable seamless interactions in very dynamic environments that require fast, implicit, non-intrusive, and ubiquitous recognition of users. In this work, we present a hybrid computer vision and RFID system that uses a novel reverse synthetic aperture technique to recover the relative motion paths of an RFID tags worn by people and correlate that to physical motion paths of individuals as measured with a 3D depth camera. Results show that our real-time system is capable of simultaneously recognizing and correctly assigning IDs to individuals within 4 seconds with 96.6% accuracy and groups of five people in 7 seconds with 95% accuracy. In order to test the effectiveness of this approach in realistic scenarios, groups of five participants play an interactive quiz game with an autonomous robot, resulting in an ID assignment accuracy of 93.3%.
Effective teleoperation requires real-time control of a remote robotic system. In this work, we develop a controller for realizing smooth and accurate motion of a robotic head with application to a teleoperation system for the Furhat robot head [1], which we call TeleFurhat. The controller uses the head motion of an operator measured by a Microsoft Kinect 2 sensor as reference and applies a processing framework to condition and render the motion on the robot head. The processing framework includes a pre-filter based on a moving average filter, a neural network-based model for improving the accuracy of the raw pose measurements of Kinect, and a constrained-state Kalman filter that uses a minimum jerk model to smooth motion trajectories and limit the magnitude of changes in position, velocity, and acceleration. Our results demonstrate that the robot can reproduce the human head motion in real time with a latency of approximately 100 to 170 ms while operating within its physical limits. Furthermore, viewers prefer our new method over rendering the raw pose data from Kinect.
We present the main ideas of the recently initiated EU-IST H2020 project “BabyRobot”. The project is a collaborative, multidisciplinary effort for developing and commercially exploiting the next generation of human-robot interaction technologies in order to promote the adoption of robotic systems in educational settings, consumer applications and beyond. Keywords—Child-Robot Communication and Collaboration, Multimodal Interaction, Child-Robot Interaction, Evaluating Child-Robot Interaction, Spoken Dialogue Systems
We describe the conceptual design, architecture, and implementation of a multimodal, robot-child dialogue system in a fast-paced, speech-controlled collaborative game. In Mole Madness, two players (a user and an anthropomorphic robot) work together to move an animated mole character through its environment via speech commands. Using a combination of speech recognition systems and a microphone array, the system can accommodate children's natural behavior in real time. We also briefly present the details of a recent data collection with children, ages 5 to 9, and some of the challenging behaviors the system elicited that we intend to explore.
We examine the effects of coordinated head and eye movement on children's turn-taking behavior in the context of a multiparty game. Twenty-two pairs of children competed in a trivia quiz scenario that is moderated first by a human and later by a robot. We quantify the effects of eyes-only and combined head-eye movements on the turn-taking behavior of the children in both directed and open questions (where either child is free to respond to win the point), and compare the results to performance with the human moderator who uses natural head and eye movements as well as additional cues that can be relevant to turn-taking. We find that coordinated head and eye movement in the robot is a significantly more successful cueing strategy than eye movement alone in directed questions, producing turn regulation that is comparable to the human moderator's more complex behaviors. Further, in open questions, head gaze results in more balanced turn-taking than eye movement alone. Finally, we compare the results for children to comparable studies with adults and discuss the implications for developing computational models of joint-attention in human-agent spoken interactions.
Children's interpersonal synchrony has been related to various benefits in social, mental and emotional development. We explore verbal and acoustic synchrony patterns between pairs of children playing a speech-controlled video game. Verbal features include word timing and duration patterns, while acoustic cues contain prosodic information. Synchrony is captured through a random-effects model taking into account multiple sources of variation and repeated measurements for each pair of children. Our findings indicate the presence of synchrony between participants during game play, which increases as they become more engaged in the game. These results are discussed in relation to personalized human-computer interaction and adaptive game environments.
A system's ability to understand and model a human's engagement during an interactive task is important for both adapting its behavior to the moment and achieving a coherent interaction over time. Standard practice for creating such a capability requires uncovering and modeling the multimodal cues that predict engagement in a given task environment. The first step in this methodology is to have human coders produce "gold standard" judgments of sample behavior. In this paper we report results from applying this first step to the complex and varied behavior of children playing a fast-paced, speech-controlled, side-scrolling game called Mole Madness. We introduce a concrete metric for engagement-willingness to continue the interaction--that leads to better inter-coder judgments for children playing in pairs, explore how coders perceive the relative contribution of audio and visual cues, and describe engagement trends and patterns in our population. We also examine how the measures change when the same children play Mole Madness with a robot instead of a peer. We conclude by discussing the implications of the differences within and across play conditions for the automatic estimation of engagement and the extension of our autonomous robot player into a "buddy" that can individualize interaction for each player and game.
1 We present Mole Madness, a side-scrolling computer game that is built to explore multi-child language use, turn-taking, engagement, and social interaction in a fast-paced speech-operated activity. To play the game, each of the two users controls the movement of the mole on one axis with either go or jump. We describe the game and data collected from 68 children playing in pairs. We then present a preliminary analysis of game-play vs social turn-taking, and engagement through the use of social side-talk. Finally, we discuss a number of interesting problems in multiparty spoken interaction that are encompassed by Mole Madness and present challenges for building an autonomous game player.
In this paper, we present a brief summary of the international workshop on Modeling Multiparty, Multimodal Interactions. The UM3I 2014 workshop is held in conjunction with the ICMI 2014 conference. The workshop will highlight recent developments and adopted methodologies in the analysis and modeling of multiparty and multimodal interactions, the design and implementation principles of related human-machine interfaces, as well as the identification of potential limitations and ways of overcoming them.
In this paper, we describe a project that explores a novel experimental setup towards building a spoken, multi-modally rich, and human-like multiparty tutoring robot. A human-robot interaction setup is designed, and a human-human dialogue corpus is collected. The corpus targets the development of a dialogue system platform to study verbal and nonverbal tutoring strategies in multiparty spoken interactions with robots which are capable of spoken dialogue. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. Along with the participants sits a tutor (robot) that helps the participants perform the task, and organizes and balances their interaction. Different multimodal signals captured and auto synchronized by different audio-visual capture technologies, such as a microphone array, Kinects, and video cameras, were coupled with manual annotations. These are used build a situated model of the interaction based on the participants personalities, their state of attention, their conversational engagement and verbal dominance, and how that is correlated with the verbal and visual feedback, turn-management, and conversation regulatory actions generated by the tutor. Driven by the analysis of the corpus, we will show also the detailed design methodologies for an affective, and multimodally rich dialogue system that allows the robot to measure incrementally the attention states, and the dominance for each participant, allowing the robot head Furhat to maintain a well coordinated, balanced, and engaging conversation, that attempts to maximize the agreement and the contribution to solve the task.This project sets the first steps to explore the potential of using multimodal dialogue systems to build interactive robots that can serve in educational, team building, and collaborative task solving applications.
This project explores a novel experimental setup towards building spoken, multi-modally rich, and human-like multiparty tutoring agent. A setup is developed and a corpus is collected that targets the development of a dialogue system platform to explore verbal and nonverbal tutoring strategies in multiparty spoken interactions with embodied agents. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. With the participants sits a tutor that helps the participants perform the task and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies were coupled with manual annotations to build a situated model of the interaction based on the participants personalities, their temporally-changing state of attention, their conversational engagement and verbal dominance, and the way these are correlated with the verbal and visual feedback, turn-management, and conversation regulatory actions generated by the tutor. At the end of this chapter we discuss the potential areas of research and developments this work opens and some of the challenges that lie in the road ahead.
Furhat [1] is a robot head that deploys a back-projected animated face that is realistic and human-like in anatomy. Furhat relies on a state-of-the-art facial animation architecture allowing accurate synchronized lip movements with speech, and the control and generation of non-verbal gestures, eye movements and facial expressions. Furhat is built to study, implement and validate patterns and models of human-human and human-machine situated and multi-party multimodal communication, a study that demands the co-presence of the talking head in the interaction environment, some-thing that cannot be achieved using virtual avatars displayed on flat screens [2,3]. In Furhat, the animated face is back-projected on a translucent mask that is a printout of the animated model. The mask is then rigged on a 2DOF neck to allow for the control of head movements. Figure 1 shows a snapshot of Furhat in interaction. We will show in this demonstrator an advanced multimodal and multiparty spoken conversational system using Furhat, a robot head based on projected facial animation. Furhat is an anthropomorphic robot head that utilizes facial animation for physical robot heads using back-projection. In the system, multimodality is enabled using speech and rich visual input signals such as multi-person real-time face tracking and microphone tracking. The demonstrator will showcase a system that is able to carry out social dialogue with multiple interlocutors simultaneously with rich output signals such as eye and head coordination, lips synchronized speech synthesis, and non-verbal facial gestures used to regulate fluent and expressive multiparty conversations. The dialogue design is performed using the IrisTK [4] dialogue authoring toolkit developed at KTH. The system will also be able to perform a moderator in a quiz-game showing different strategies for regulating spoken situated interactions.
To provide a spoken interaction between robots and human users, an internal representation of the robots sensory information must be available at a semantic level and accessible to a dialogue system in order to be used in a human-like and intuitive manner. In this paper, we integrate the fields of perceptual anchoring (which creates and maintains the symbol-percept correspondence of objects) in robotics with multimodal dialogues in order to achieve a fluent interaction between humans and robots when talking about objects. These everyday objects are located in a so-called symbiotic system where humans, robots, and sensors are co-operating in a home environment. To orchestrate the dialogue system, the IrisTK dialogue platform is used. The IrisTK system is based on modelling the interaction of events, between different modules, e.g. speech recognizer, face tracker, etc. This system is running on a mobile robot device, which is part of a distributed sensor network. A perceptual anchoring framework, recognizes objects placed in the home and maintains a consistent identity of the objects consisting of their symbolic and perceptual data. Particular effort is placed on creating flexible dialogues where requests to objects can be made in a variety of ways. Experimental validation consists of evaluating the system when many objects are possible candidates for satisfying these requests.
This paper describes a novel experimental setup exploiting state-of-the-art capture equipment to collect a multimodally rich game-solving collaborative multiparty dialogue corpus. The corpus is targeted and designed towards the development of a dialogue system platform to explore verbal and nonverbal tutoring strategies in multiparty spoken interactions. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. The participants were paired into teams based on their degree of extraversion as resulted from a personality test. With the participants sits a tutor that helps them perform the task, organizes and balances their interaction and whose behavior was assessed by the participants after each interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies, together with manual annotations of the tutor's behavior constitute the Tutorbot corpus. This corpus is exploited to build a situated model of the interaction based on the participants' temporally-changing state of attention, their conversational engagement and verbal dominance, and their correlation with the verbal and visual feedback and conversation regulatory actions generated by the tutor.
This project explores a novel experimental setup towards building spoken, multi-modally rich, and human-like multiparty tutoring agent. A setup is developed and a corpus is collected that targets the development of a dialogue system platform to explore verbal and nonverbal tutoring strategies in multiparty spoken interactions with embodied agents. The dialogue task is centered on two participants involved in a dialogue aiming to solve a card-ordering game. With the participants sits a tutor that helps the participants perform the task and organizes and balances their interaction. Different multimodal signals captured and auto-synchronized by different audio-visual capture technologies were coupled with manual annotations to build a situated model of the interaction based on the participants personalities, their temporally-changing state of attention, their conversational engagement and verbal dominance, and the way these are correlated with the verbal and visual feedback, turn-management, and conversation regulatory actions generated by the tutor. At the end of this chapter we discuss the potential areas of research and developments this work opens and some of the challenges that lie in the road ahead.
In this paper, we present Furhat - a back-projected human-like robot head using state-of-the art facial animation. Three experiments are presented where we investigate how the head might facilitate human - robot face-to-face interaction. First, we investigate how the animated lips increase the intelligibility of the spoken output, and compare this to an animated agent presented on a flat screen, as well as to a human face. Second, we investigate the accuracy of the perception of Furhat's gaze in a setting typical for situated interaction, where Furhat and a human are sitting around a table. The accuracy of the perception of Furhat's gaze is measured depending on eye design, head movement and viewing angle. Third, we investigate the turn-taking accuracy of Furhat in a multi-party interactive setting, as compared to an animated agent on a flat screen. We conclude with some observations from a public setting at a museum, where Furhat interacted with thousands of visitors in a multi-party interaction.
In multiparty multimodal dialogue setup, where the robot is set to interact with multiple people, a main requirement for the robot is to recognize the user speaking to it. This would allow the robot to pay attention (visually) to the person the robot is listening to (for example looking by the gaze and head pose to the speaker), and to organize the dialogue structure with multiple people. Knowing the speaker from a set of persons in the field-of-view of the robot is a research problem that is usually addressed by analyzing the facial dynamics of persons (the person that is moving his lips and looking towards the robot is probably the person speaking to the robot).This thesis investigates the use of lip and head movements for the purpose of speaker and speech/silence detection in the context of human-machine multiparty dialogue. The use of speaker and voice activity detection systems in human-machine multiparty dialogue is to help the machine in detecting who and when someone is speaking out of a set of persons in the field-of-view of the camera. To begin with, a video of four speakers (S1, S2, S3 and S4) speaking in a task free dialogue with a fifth speaker (S5) through video conferencing is audio-visually recorded. After that each speaker present in the video is annotated with segments of speech, silence, smile and laughter. Then the real-time FaceAPI face tracking commercial software is applied to each of the four speakers in the video to track the facial markers such as head and lip movements. At the end, three classification techniques namely Mahalanobis distance, naïve Bayes classifier and neural network classifier are applied to facial data (lip and head movements) to detect speech/silence and speaker. In this thesis, three types of training methods are used to estimate the training models of speech/silence for every speaker. The first one is speaker dependent method, in which the training model contains the facial data of testing person. The second one is speaker independent method, where the training model does not contain the facial data of testing person. It means that if the test person is S1 then the training model may contain the facial data of S2, S3 or S4. The third one is hybrid method, where the training model is estimated using the facial data of all the speakers and testing is performed on one of the speaker. The results of speaker dependent and hybrid methods show that the neural network classifier provides the best results. In the speaker dependent method, the accuracies of neural network classifier for speaker and speech/silence detection are 97.43% and 98.73% respectively. However, in the hybrid method, the accuracy of neural network classifier for speech/silence detection is 96.22%. The results of speaker independent method shows that the naïve Bayes classifier provides the best results with an optimal accuracy of 67.57% for speech/silence detection. Sammanfattning Gentemot Talaren Detektering med FaceAPI Facial rörelser i Människa-Maskin Multiparty Dialog I fleraparter med fleramodala dialoginställningar, där roboten är inställd på att interagera med flera personer. Det är en viktig förutsättning för roboten att känna igen att användaren talar till den. Detta skulle göra det möjligt för roboten att uppmärksamma (visuellt) den person roboten lyssnar till (till exempel genom att titta i blicken och på huvudet för att känna igen talaren) och att organisera dialogens struktur med flera personer. Talaren från en upp sättning av personer i roboten synfält är ett forskningsproblem som vanligtvis riktar sig till att analysera dynamiken i ansiktsuttryck för personer (den person som rör på sina läppar och riktar blicken mot roboten är förmodligen den person som talar till roboten). Denna avhandling undersöker användningen av läpp och huvudrörelser i syfte av att upptäcka högtalare och tal/tystnad i samband med människa-maskin flerpartisystem dialog. Användningen av högtalare och röstaktivitetsdetekteringssystem i människa-maskin flerpartisystem dialog är att hjälpa maskinen att upptäcka vem och när någon talar i kamerans synfält. Till att börja med, en video av fyra högtalare (S1, S2, S3 och S4) talar i en uppgift utan dialog med en femte högtalare (S5) genom videokonferenser blir ljud-visuellt inspelat. Sedan tillämpas realtid FaceAPI tracking kommersiell programvara på vardera fyra högtalarna i videon, för att spåra ansiktets markörer som huvud-och läpprörelser. I slutet finns tre klassificeringstekniker nämligen Mahalanobis distans, naiva Bayes klassificeraren och neuralanätverk klassificerare, som tillämpas på ansiktet (läpp och huvudrörelser) för att upptäcka tal/tystnad och talare. I denna avhandling har tre typer av träningsmetoder använts för att uppskatta utbildningsmodellerna för tal/tystnad för varje talare. Den första är en talarberoende metod, där utbildningsmodellen innehåller uppgifter om ansiktsdrag från testpersonen. Den andra är en talaroberoende metod, där träningsmodellen inte innehåller ansiktsdrag från testpersonen. Det innebär att om testpersonen är S1 kan utbildningsmodellen innehålla data om ansiktsdrag från S2, S3 eller S4. Den tredje är en hybrid metod, där utbildningsmodellen beräknas utifrån data från alla talares ansiktsdrag men tester utförs på en av talarna. Resultaten av talarberoende och hybridmetoderna visar att den neurala nätverksklassificeraren ger bästa resultat. Utifrån data från alla talares ansiktsdrag är, noggrannheten på neurala nätverk klassificerare för talare och tal/tystnad upptäckt är 97,43% och 98,73% respektive. I hybridmetoden, är däremot noggrannheten hos neurala nätverksklassificeraren för tal/tystnad detektering 96,22%. Resultaten av talaroberoende metod visar att den naïve Bayes klassificerare ger de bästa resultaten med en optimal noggrannhet på 67,57% för tal/tystnad detektering.