
Hybrid Intelligence (HI) is defined as “the combination of human and machine intelligence” for collaboration, “achieving goals that were unreachable by either [alone]” [2], leveraging the strength of both machine intelligence (such as strong optimization capabilities, effective handling of probabilities and less fallacy for confirmation biases) and human intelligence (such as generalization capabilities, situational understanding and common sense). In their research agenda, Akata and colleagues identify three key challenges for creating HI systems that relate to the interactive process: HI should be adaptive (how can a system learn from and adapt to humans and vice versa), explainable (how to create shared and explained awareness, goals and strategies) and collaborative (how to work in synergy). A popular learning mechanism for robots and AI agents that allows agents to dynamically adapt to their environment and promises interesting opportunities for adaptation in Human-Agent Interaction scenarios is Reinforcement Learning (RL). In RL, an agent learns through exploration and optimization of reward-based feedback for future actions (e.g. through the value-function method, policy search and actor critic approaches, [26]). While RL agents have been able to demonstrate their potential for success in a broad range of narrow tasks, many real-life applications present them with high-dimensional or continuous state-spaces that are not always fully observable which renders exploration and reproduction of actions costly and slow and usually require a considerable amount of domain knowledge or common sense, in order to succeed [19]. A more hybrid approach with Human-in-the-loop Reinforcement Learning, where human users enrich task learning in the form of teaching signals, provides a compelling modification of this setting that allows to shape and improve the learning process through methods such as evaluative teaching, Learning from Demonstration and instruction [9], [21] and can also be streamlined with more specialized approaches such as the TAMER [18] or the COACH architecture [23]. However, past work has shown that simply inserting a human teacher in the RL process, yielded limited success in applications where the reward was provided by teachers [16]. While a teacher can provide insightful expert knowledge or even just common sense to a learning agent, many times they are no experts in interpreting and understanding the behavior of RL-agents. For example, is a suboptimal action of a collaborative RL-agent and exploratory move, (that would help it to explore the state-action space and bring it closer to the optimal policy), or does the agent really “think” that move is indeed optimal (i.e. exploiting its currently believed optimal policy)? This ambiguity in the interpretation of learner behavior in turn results in suboptimal teaching signals (e.g. [16], [27]), effectively creating a misalignment in the teacher-learner loop. In human-human interaction, affective signals play a vital role in synergic interaction and their influence is a core research question in the field of Affective Computing [25], [28]. This project investigates if and how agent affective signals, grounded in the RL process itself, can improve explainability and collaboration in the teacher-learner loop in interactive Reinforcement Learning.
In the social media age, platforms like X (formerly Twitter) are essential for spreading information and shaping public opinion. This initial study proposes a methodology to investigate interactions with Italian opinion leaders, forming clusters of leaders based on user engagement patterns while preserving user anonymity by analyzing the vocabulary eccentricity and sentiment analysis. Our findings reveal that positive sentiment is linked to a more eccentric vocabulary, particularly within specific clusters, indicating that sentiment is deeply linked to the community's shared language. Despite the prevalence of negative sentiment, most clusters feature a minority of leaders with eccentric, positive language. These consistent patterns over time suggest a potential universal behaviour for further study.
Affective computing, a field dedicated to enabling machines to understand human behavior and emotions, is increasingly integrated into various societal domains such as, mental healthcare. The predominant focus of affective computing in Western societies raises potential risks for implementing these technologies in the diverse, global real-world. This paper briefly describes ethical and regulatory concerns of deploying affective computing technologies globally, emphasizing on needs to address privacy, security, and cultural sensitivity concerns.
Emotion recognition is a growing area of research interest in multiple disciplines, including psychology, linguistics, and computer science. The field of affective computing has changed significantly with the advent of deep learning, which has brought in new and innovative techniques such as multimodal learning for emotion recognition. In this position paper, we explore the issues related to ambiguity in emotion annotations arising while designing and utilizing multimodal (specifically speech, vision and/or text) emotion-based datasets for applications modeling perceived human emotions. Our findings are based on observations and lessons learned while performing multimodal emotion recognition experiments focusing on oral history interviews. We categorize the ambiguity issues in emotion annotations into two categories, namely technical (emotion representation, contextual reference, segmentation, skewed annotator sample, and annotator evaluation) and socio-psychological (subject- and annotator-induced) factors. We provide eight recommendations for managing ambiguity in emotion annotation relevant for research on datasets for emotion recognition in the domain of affective computing.
Awe is an emotion characterized by the perception of vastness and the need to accommodate this vastness into one's mental framework. We propose an elicitation scene to induce awe in Virtual Reality (VR), validate it through selfreport, and explore the feasibility of using skin conductance to predict self-reported awe and vastness as labeled from the stimuli in VR. Sixty-two participants took part in the study comparing the awe-eliciting space scene and a neutral scene. The space scene was confirmed as more awe-eliciting. A k-nearest neighbor algorithm confirmed high and low-awe score clusters used to label the data. A Random Forest algorithm achieved 65% accuracy (F1 = 0:56;AUC = 0:73) when predicting the self-reported low and high awe categories from continuous skin conductance data. A similar approach achieved 55% accuracy (F1 = 0:59;AUC = 0:56) when predicting the perception of vastness. These results underscore the potential of skin-conductance-based algorithms to predict awe.
In human-machine explanation interactions, such as tutoring systems or customer support chatbots, it is important for the machine explainer to infer the human user's understanding. Nonverbal signals play an important role for expressing mental states like understanding and confusion in these interactions. However, an individual's expressions may vary depending on other factors. In cases where these factors are unknown, machine learning methods that infer understanding from nonverbal cues become unreliable. Stress for example has been shown to affect human expression, but it is not clear from the current research how stress affects the expression of understanding. To address this gap, we design a paradigm that induces understanding and confusion through game rule explanations. During the explanations, self-perceived understanding and confusion are annotated by the participants. A stress condition is also introduced to enable the investigation of changes in the expression of social signals under stress. We conducted a study to validate the stress induction and participants reported a statistically significant increase in stress during the stress condition compared to the neutral control condition. Additionally, feedback from participants shows that the paradigm is effective in inducing understanding and confusion. This paradigm paves the way for further studies investigating social signals of understanding to improve human-machine explanation interactions for varying contexts.
In this paper, we introduce EMOLight, an innovative AI-driven ambient lighting solution that enhances viewer immersion, especially for the audibly impaired, by dynamically synchronising with the emotional content of audio cues and sounds. Our proposed solution leverages YAMNet deep learning model and Plutchik's emotion-colour theory to provide real-time audio emotion recognition, user-specific customisation, and multi-label classification for a personalised and engaging experience. This synchronisation enriches the viewing experience by making it more engaging and inclusive. This research demonstrates the feasibility and potential of EMOLight, paving the way for a future where technology adapts to diverse sensory needs and preferences, revolutionising the way we experience and interact with media.
In today's era of Human-Robot Interaction (HRI), the ability of robots to understand and respond to human emotions is crucial. Facial Expression Recognition (FER) plays an important role in enabling social robots to engage naturally with humans. However, existing FER systems face challenges such as computational complexity and sensitivity to facial orientation, which limit their practical effectiveness. This paper proposes a novel approach using 3D face orientation to enhance the accuracy and robustness of FER models on social robots. We demonstrated our methodology using the FER2013 dataset, employing graph-based facial landmarks and a lightweight Artificial Neural Network (ANN) architecture. Our results show significant improvements in classifying negative emotions, while also highlighting areas for further refinement across a range of emotional states.
Mental fatigue recognition by heart rate variability analysis in a photoplethysmography (PPG) device in terms of pulse rate variability (PRV) is gathering interest for stress relief and well-being applications. To realize mental fatigue recognition systems, many studies have been conducted, which showed the high accuracy of such systems evaluations. However, there are limitations in generalization due to the lack of consideration of the physiological mechanism. In this study, we developed a framework for investigating mental fatigue recognition using physiologically interpretable deep neural network (DNN) in a wrist-worn PPG device. Furthermore, we developed a novel feature extraction method that captures the temporal accumulation of arousal states to consider the perspective of cognitive resource consumption. The results suggest that the method we developed improved the mental fatigue recognition performance of the wrist-worn PPG device throughout our experiments (an accuracy of 0.90 and an F1 score of 0.91). In addition, it was observed that the method for the wrist-worn PPG device enhanced the contribution of HRV indices, known to be associated with cognitive function. The findings demonstrate that the proposed approach enhances not only accuracy but also the physiological interpretability of the model in PPG devices.
Understanding emotions is key to Affective Computing and in many other research and application fields. Therefore, researchers are working on various strategies to recognize emotions, e.g., emotion recognition from EEG signals - a rapidly growing area of research in recent years. This scientific study of emotions is based on the assumption that emotions can be identified objectively (locationist approach). At the same time, studies have shown that emotions are complex internal experiences. As such, they cannot be objectively identified, nor do they have a unique fingerprint in measurable human behavior, physiology or neurology (constructionist approach). In this paper, we argue for a more sophisticated view on emotions taking other emotional components into account.
Conversational agents are increasingly introduced in mental health and well-being research. One promising application is in assessing students' mental well-being while facilitating self-disclosure activities such as administering questionnaires or assisting in journaling. In these applications, the data collected is mostly multimodal and strategies for fully harnessing the potential of this data is an open challenge. Moreover, most of the machine learning (ML) research within healthcare has been carried out using conventional, baseline models which fail to capitalise on the inherent graphical characteristics of health data. Learning from graph representations encompasses techniques such as feature propagation and aggregation methods to combine feature information and graph structure in order to learn better data representations. Therefore, in this paper we propose a novel graph representation learning strategy to predict student mental health status from multimodal data in a robot assisted journal writing context. Twenty undergraduate / graduate students interacted with the Pepper robot in a dyadic interaction setting over four weeks (one session per week). Graph representation learning was applied on their questionnaire data (mental wellbeing state, mood changes, perceptions toward Pepper, and motivational state) and physiological data (electro-cardiogram). Comparative experimental results demonstrate that graph representation learning based classification enables mental wellbeing assessment with an accuracy of 92.5% which is superior to standard ML methods that do not take into account graphical structures.
Talking to others about one's thoughts and gaining empathy is important to controlling emotions and relieving stress. Unfortunately, suitable conversation partners are not always available. While conversational agents might serve as partners, the technology for accurately interpreting user's statements and unspoken intentions remains imperfect. When agents express explicit empathy based on inaccurate interpretations of the user's message, users are likely to feel discomfort. Existing study suggested that when a single agent provided feedback to a user's utterance with simple motions, users tended to feel empathy from the agent in limited scenes. We explore the hypothesis that using two agents, compared to using one agent, might be more effective on users perception of empathy. This study investigates the relationship between the number of agents and the degree to which users feel empathy from the agent using simple motions in scenarios. The experimental results showed that the number of agents may not affect the degree to which users feel empathy from the agent.
This paper suggests cultivating and modeling individual variation in research on affect and emotion, emphasizing the importance of studying real-world contexts to reveal their individualized nature. We propose a new standard for study design using more flexible and context-aware methods that emphasize modeling of high-dimensional signal arrays within individuals and then inductively learning generalizations rather than presuming their existence. We compare laboratory and real-life settings for the study of affect and emotion and highlight opportunities for studying individual variation structured by context. We discuss current challenges and future directions therein and emphasize the need for interdisciplinary and international collaboration and leveraging new technologies.
This paper discusses the problem of dynamic game narration and proposes a solution conceptualized to enhance player experience by providing real-time, creative and context- aware narration using large language models (LLMs). Based on the literature review, the framework would consist of three interconnected modules to solve the remaining challenges identified: an anomaly detection module, a Retrieval-Augmented Generation (RAG) module, and an emotion detection module. The anomaly detection module would identify relevant events within the game, serving as a trigger for narration. The RAG module would provide necessary context regarding the game's rules and strategies, ensuring accurate and consistent narration. The emotion detection module would gauge the atmosphere of the game and the players' emotions, allowing the LLM to craft a creative and engaging narrative. The framework's effectiveness could be evaluated in the multiplayer card game Chef's Hat, where player interactions and game data can be easily recorded and analyzed. The evaluation metrics would include Coherence, Relevance and Engagement.
This research presents a novel multimodal data fusion methodology for pain behavior recognition, integrating statistical correlation analysis with human-centered insights. Our approach introduces two key innovations: 1) integrating data-driven statistical relevance weights into the fusion strategy to effectively utilize complementary information from heterogeneous modalities, and 2) incorporating human-centric movement characteristics into multimodal representation learning for detailed modeling of pain behaviors. Validated across various deep learning architectures, our method demonstrates superior performance and broad applicability. We propose a customizable framework that aligns each modality with a suitable classifier based on statistical significance, advancing personalized and effective multimodal fusion. Furthermore, our methodology provides explainable analysis of multimodal data, contributing to interpretable and explainable AI in healthcare. By highlighting the importance of data diversity and modality-specific representations, we enhance traditional fusion techniques and set new standards for recognizing complex pain behaviors. Our findings have significant implications for promoting patient-centered healthcare interventions and supporting explainable clinical decision-making.
Reinforcement learning (RL)-based agents have demonstrated remarkable performance in multiplayer card game environments such as Chef's Hat. However, understanding why these agents excel in such dynamic and competitive settings remains a challenging endeavor. In this paper, we propose a novel method, named the “Mullet's Gambit“ to elucidate the strategies employed by RL-based agents within the context of the Chef's Hat card game. This method aims to provide insights into how RL-based agents navigate the complexities of multiplayer dynamics and assess their impact on opponents. By employing Mullet's Gambit, this investigation reveals the unique traits and efficacy of RL-based strategies compared to heuristic methodologies. This leads to the inference that RL-based agents not only acquire the skills to win but also to disrupt their opponents, thereby minimizing their potential actions.
Path planning (PP) is a major topic of concern in the domain of autonomous agents applicable to many different applications such as disasters. Reinforcement learning approaches to path planning have seen significant advancements in recent years. Incorporating the concept of emotions in reinforcement learning agents enhances their ability to take better decisions in a given situation Here we have proposed a methodology that encompasses the concept of regret emotion in the epsilon decay strategy of the Q-learning algorithm that uses a cumulative regret value of experienced regret and anticipated regret which occurs in the agent. Experienced Regret can be modeled as a function that minimizes the error between the optimal action and the actual action taken in a given state. Anticipated regret can be modeled using the concept of temporal difference errors. In the epsilon-greedy strategy, regret emotion can be used to find a balance between exploitation and exploration. If the agent wants to exhibit the best possible behaviour, the agent should explore the environment more during the beginning phase of the learning process and then move on to exploiting the environment during the later phases. However, with the current epsilon greedy algorithm, epsilon decays at a faster rate, which thereby restricts the agent's ability of exploration. In the proposed method, the agent at a particular state would decide whether to exploit or explore based on the intensity of the regret emotion of the agent. Using the proposed methodology the agent will be able to explore the environment more frequently helping the agent to acquire more rewards and learn the environment more rapidly. The proposed method outperforms the traditional epsilon greedy Q-learning algorithm for epsilon decay in terms of the rate at which the decay occurs and the amount of reward that the agent gains.
Facial expressions are essential for non-verbal human communication as they convey behavioral intentions and emotional states. While facial action units (AUs) can occur bilaterally or unilaterally, the existing research in affective computing predominantly concentrates on bilateral expressions, largely due to the lack of datasets with unilateral AU labels. In this study, we present a method for generating unilateral AU labels and assess its efficacy against expert-labeled facial images. Furthermore, we introduce a dedicated model trained on the generated data and evaluate its performance across multiple datasets. Our findings offer insights into feature extraction for unilateral facial expression recognition. This research contributes to advancing the understanding and recognition of nuanced facial expressions, with potential applications in various domains such as healthcare and human-computer interaction.
The Public Speaking project is a Virtual Reality (VR) training system dedicated to the improvement of public speaking skills. It provides users with the opportunity to rehearse their speech in a realistic virtual environment which displays an interactive audience. The experience is customized based on the chosen training scenario, which is determined by the room, participant's role, audience behaviors, and timer. A companion website provides comprehensive feedback reports on speakers' performance, based on their verbal and nonverbal behavior, as well as replay functionalities. In addition to its educational objective, this system is used as a tool to advance research into the development of effective public speaking training environments.
Praising behavior is an important method of communication. An existing study constructed models to predict praising skill, which indicates the degree to which the praise is done well, by using only unimodal behavior such as speech audio or visual behavior of a praiser who gives praise in dyad interactions. To improve prediction performance, a model should be constructed that uses various additional information. In this study, we propose two approaches to predict praising skill highly accurately. The first uses trimodal (multimodal) behaviors extracted from visual, acoustic, and linguistic modalities. The second uses the behaviors of the receiver of praise since the reaction of the receiver should differ depending on how good the praise is. For this study, we collect trimodal features and the degree of praising skill in each praising scene in a dialogue. We construct multiple models to predict the degree of praising skills using various combinations of the trimodal features from the praiser and receiver. The experimental results show that the model that predicts praising skill most accurately uses multiple features related to both verbal and nonverbal behaviors of the praiser and receiver. Therefore, the two approaches of using trimodal behaviors and using features from both the receiver and praiser are effective for predicting praising skills in dyad interactions.