Facial expressions and speech are elements that provide emotional information about the user through multiple communication channels. In this paper, a novel multimodal emotion recognition system based on visual and auditory information processing is proposed. The proposed approach is used in real affective human robot communication in order to estimate five different emotional states (i.e., happiness, anger, fear, sadness and neutral), and it consists of two subsystems with similar structure. The first subsystem achieves a robust facial feature extraction based on consecutively applied filters to the edge image and the use of a Dynamic Bayessian Classifier. A similar classifier is used in the second subsystem, where the input is associated to a set of speech descriptors, such as speech-rate, energy and pitch. Both subsystems are finally combined in real time. The results of this multimodal approach show the robustness and accuracy of the methodology respect to single emotion recognition systems.
In this article the different decision-making that were followed in the design of the expressive robotic head Muecas are explained. Muecas is a system with a humancaricatured shape, equipped with a pair of robotic eyes, eyebrows, neck and mouth. The main goal in the design was to provide the robot with basic skills for an affective humanrobot interaction, where the emphaty and attention plays an important factor. When developing this robotic head, it was necessary to study a number of parameters and characteristics related to human anatomy and psychology, as well as other similar robotic heads or the opinions of experts. All the results of these study are pointed out in this paper. Throughout this work, the step followed in the design in conjuction with the main conclusions drawn from it, are explained. We want that our work helps researchers in this field in their decision making.
The proposed work presents the concept of emotional affordances as an extension of the classical perceptual affordances for human-robot interaction. In this paper, emotional affordances represent the relation between affective elements, such as the objects in the scenario or the emotional states, and the effects and oportunities for the robot’s reactions. Thus, the system proposed in this work describes the use of the affordances in affective scenarios using robotic agents that can recognize and predict the emotional information of the user through elements of the environment and the human natural language (i.e. facial expressions). The major purpose of this paper is to evaluate the proposed conceptual idea through simple experiments to demonstrate the validity and impact of this work.
This paper presents a multi-sensor humanoid robotic head for human robot interaction. The design of the robotic head, Muecas, is based on ongoing research on the mechanisms of perception and imitation of human expressions and emotions. These mechanisms allow direct interaction between the robot and its human companion through the different natural language modalities: speech, body language and facial expressions. The robotic head has 12 degrees of freedom, in a human-like configuration, including eyes, eyebrows, mouth and neck, and has been designed and built entirely by IADeX (Engineering, Automation and Design of Extremadura) and RoboLab. A detailed description of its kinematics is provided along with the design of the most complex controllers. Muecas can be directly controlled by FACS (Facial Action Coding System), the de facto standard for facial expression recognition and synthesis. This feature facilitates its use by third party platforms and encourages the development of imitation and of goal-based systems. Imitation systems learn from the user, while goal-based ones use planning techniques to drive the user towards a final desired state. To show the flexibility and reliability of the robotic head, the paper presents a software architecture that is able to detect, recognize, classify and generate facial expressions in real time using FACS. This system has been implemented using the robotics framework, RoboComp, which provides hardware-independent access to the sensors in the head. Finally, the paper presents experimental results showing the real-time functioning of the whole system, including recognition and imitation of human facial expressions.
In the last decade, affective Human-Robot Interaction has become an interesting topic for researching. Facial expressions are rich sources of information about affective behaviour and have been commonly used for emotion recognition. In this paper, a novel facial expression recognition algorithm using RGB-D information is proposed. The algorithm achieves in real-time the detection and extraction of a set of both, invariant and independent facial features based on the wellknown Candide-3 reconstruction model. Experimental results showing the effectiveness and robustness of the approach are described and compared to previous related works.
Facial expressions are a rich source of communicative information about human behavior and emotion. This paper presents a real-time system for recognition and imitation of facial expressions in the context of affective Human Robot Interaction. The proposed method achieves a fast and robust facial feature extraction based on consecutively applying filters to the gradient image. An efficient Gabor filter is used, along with a set of morphological and convolutional filters to reduce the noise and the light dependence of the image acquired by the robot. Then, a set of invariant edge-based features are extracted and used as input to a Dynamic Bayesian Network classifier in order to estimate a human emotion. The output of this classifier updates a geometric robotic head model, which is used as a bridge between the human expressiveness and the robotic head. Experimental results demonstrate the accuracy and robustness of the proposed approach compared to similar systems.
This paper presents a new system for recognition and imitation of a set of facial expressions using the visual information acquired by the robot. Besides, the proposed system detects and imitates the interlocutor’s head pose and motion. The approach described in this paper is used for human-robot interaction (HRI), and it consists of two consecutive stages: i) a visual analysis of the human facial expression in order to estimate interlocutor’s emotional state (i.e., happiness, sadness, anger, fear, neutral) using a Bayesian approach, which is achieved in real time; and ii) an estimate of the user’s head pose and motion. This information updates the knowledge of the robot about the people in its field of view, and thus, allows the robot to use it for future actions and interactions. In this paper, both human facial expression and head motion are imitated by Muecas, a 12 degree of freedom (DOF) robotic head. This paper also introduces the concept of human and robot facial expression models, which are included inside of a new cognitive module that builds and updates selective representations of the robot and the agents in its environment for enhancing future HRI. Experimental results show the quality of the detection and imitation using different scenarios with Muecas.
Human-Robot Interaction (HRI) is one of the most important subfields of social robotics.In several applications, text-to-speech (TTS) techniques are used by robots to provide feedback to humans.In this respect, a natural synchronization between the synthetic voice and the mouth of the robot could contribute to improve the interaction experience.This paper presents an algorithm for synchronizing Text-To-Speech systems with robotic mouths.The proposed approach estimates the appropriate aperture of the mouth based on the entropy of the synthetic audio stream provided by the TTS system.The paper also describes the cost-efficient robotic head which has been used in the experiments and introduces the use of conversational gestures for engaging Human-Robot Interaction.The system, which has been implemented in C++ and can perform in realtime, is freely available as part of the RoboComp open-source robotics framework.Finally, the paper presents the results of the opinion poll that has been conducted in order to evaluate the interaction experience.
Human-Robot Interaction (HRI) is one of the most important subfields of social robotics. In several applications, text-to-speech techniques are used by robots to provide feedback to humans. In this respect, a natural synchronization between the synthetic voice and the mouth of the robot could contribute to improve the interaction experience. This paper presents an algorithm for synchronizing Text-To-Speech (TTS) systems with robotic mouths. The proposed approach estimates the appropriate aperture of the mouth based on the entropy of the synthetic audio stream provided by the TTS system. The paper also describes the cost-efficient robotic mouth which has been used in the experiments. The system, which has been implemented in C++ and can perform in real-time, is freely available as part of the RoboComp open-source robotics framework. Finally, the paper presents the results of the opinion poll that has been conducted in order to evaluate the overall user experience.