The acoustic and articulatory characteristics of the syllable position lateral allophony in English (clear /l/ in onsets vs. dark /l/ in codas) have been well documented. The present study tests whether speech technology derived methods can be used to evaluate lateral allophony in L2 English production, by combining classic acoustic analyses and automatic speech recognition (ASR). In this study, an ASR system is forced to choose between English and French /l/ acoustic phone models when force-aligning a corpus consisting of read English texts by 43 L2 French learners. The output is correlated with a staple measure for /l/ darkness: the difference between the second and first formants (F2-F1). Results show that segments aligned with the French /l/ acoustic model correspond to "clearer" /l/s (i.e. higher values of F2-F1) suggesting automatic, less time consuming methods of speech processing could be used to identify L1 transfer in L2 production.
This contribution presents a study on the detection of emotions and mixtures of emotions in a corpus, CEMO, collected from a Parisian emergency call center. Our corpus, recorded 'in the wild', is rich in voice diversity (age, accent, number of speakers) and is annotated with an original scheme that represents up to two emotions per segment. Tests on a portion of CEMO's unmixed emotions with systems using audio-specific Transformers adapted to the CEMO corpus obtained a detection score (Accuracy) of 56.7% for 4 classes (fear, neutral, positive, sadness) surpassing those obtained with more classical approaches based on expert prosodic features. Additional tests were carried out on a portion of the CEMO corpus with mixed emotions, highlighting some of the outstanding challenges, in particular how to take into consideration the context of the interaction.
Cette contribution présente une étude sur la détection d’émotions et de mélanges d’émotions dans un corpus collecté dans un centre d’appels d’urgence à Paris (CEMO). Notre corpus, enregistré ‹in the wild›, est riche en diversité vocale (âge, accent, nombre de locuteurs) et est annoté avec un schéma original qui représente jusqu’à deux émotions par segment. Des tests avec des systèmes utilisant des Transformers audio spécifiques adaptés à CEMO sur une partie des émotions non mixtes ont permis d’obtenir un score de détection ( Accuracy ) de 56.7 % pour 4 classes (peur, neutre, positif, tristesse) surpassant ceux obtenus avec des approches plus classiques basées sur des caractéristiques prosodiques expertes. Des tests supplémentaires ont été effectués sur une partie de CEMO avec des émotions mixtes, mettant en évidence certains des défis à relever, en particulier la prise en compte du contexte de l’interaction.
Depuis 2022, ChatGPT d’OpenAI a popularisé l’intelligence artificielle (IA), rendant les technologies numériques essentielles. L’IA, qu’elle soit prédictive ou générative, progresse dans des domaines variés comme la médecine et les chatbots. Désormais, certaines actions humaines, comme le langage, sont séparées de l’intelligence réelle. Les IA génératives imitent le langage sans le comprendre, en utilisant des modèles statistiques complexes. Elles peuvent produire des résultats erronés, soulevant des questions éthiques. La transparence des IA est cruciale pour instaurer la confiance, et leur régulation est nécessaire pour concilier technologie et bien-être humain.
Affective computing develops systems, which recognize or influence aspects of human life related to emotion, including feelings and attitudes. Significant potential for both good and harm makes it ethically sensitive, and trying to strike sound balances is challenging. Common images of the issues invite oversimplification and offer a limited understanding of the moral consequences and ethical tensions. Considering the state-of-the-art shows how pervasive and complex they are. In many areas, the discipline can potentially bring ethically significant benefits and hence has a duty to try. They include making interactions with machines more effective and less stressful, diagnostic and therapeutic roles in emotion-related disorders, intelligent tutoring, and reducing isolation. However, the limits of recognition technology mean that actions are likely to be based on impoverished representations of people's affective state, particularly with certain groups; systems are liable to arouse feelings that are positive, but not well grounded in reality, affectively engaging systems can become addictive and manipulative, and they confer dangerous power on those who control the technology. We offer an overview of those and other particular ethical issues, positive and negative, which arise from the current state of affective computing. It aims to reflect the complexities inherent in both the technology and current ethical discussions. Establishing appropriate responses is a challenge for society as a whole, not only the affective computing community.
Robot-directed speech refers to speech to a robotic device, ranging from small home smart speakers to full-size humanoid robots. Studies have investigated the phonetic and linguistic properties of this type of speech or the effect of anthropomorphism of the devices on the social aspect of interaction. However, none have investigated the effect of the device’s human-likeliness on linguistic realizations. This preliminary study proposes to fill this gap by investigating one phonetic parameter (speech rate) and one linguistic parameter (use of filled pauses) in speech directed at a home speaker vs a humanoid robot vs a human. The data from 71 native speakers of French indicate that human-directed speech shows longer utterances at a faster speech rate and more filled pauses than speech directed at a home speaker and a robot. Speaker- and robot-directed speech is significantly different from human-directed speech, but not from each other, indicating a unique device-directed type of speech.
The ability to automatically infer relevant aspects of human users’ thoughts and feelings is crucial for technologies to intelligently adapt their behaviors in complex interactions. Research on multimodal analysis has demonstrated the potential of technology to provide such estimates for a broad range of internal states and processes. However, constructing robust approaches for deployment in real-world applications remains an open problem. The MSECP-Wild workshop series is a multidisciplinary forum to present and discuss research addressing this challenge. Submissions to this 5th iteration span efforts relevant to multimodal data collection, modeling, and applications. In addition, our workshop program builds on discussions emerging in previous iterations, highlighting ethical considerations when building and deploying technology modeling internal states in the wild. For this purpose, we host a range of relevant keynote speakers and interactive activities.
Robot-directed speech refers to speech to a robotic device (speakers, computers, etc.). Studies have investigated the phonetic and linguistic properties of this type of speech and shown that, humans tend to change their pitch when talking to a robot vs to a human. Parallelly, it has shown that the anthropomorphism of the devices affects the social aspect of interaction. However, none have investigated the effect of the device's human-likeliness on linguistic realizations. This study proposes to fill this gap by comparing the effect of anthropomorphism in speech directed at a speaker vs a humanoid robot vs a human by analyzing the F0 values and range in the three conditions, and how these parameters change throughout the conversation. The data from 52 native speakers of French show that robot-directed speech shares several pitch tendencies with speaker-directed speech, which in its turn is situated between human- and robot-directed speech.
Intelligence Augmentation (IA) has long been understood as a concept that describes how human capabilities are enhanced by technologies to improve their intelligence (often in a collective sense), and therefore improve the outcomes of many tasks. While IA's goal is to keep the human in the loop and design technology around users, the future of IA as being universally beneficial for people has been widely questioned. As IA outputs are increasingly used to train Artificial Intelligence (AI) and AI is used to enable IA, new questions arise around the implications of integrating IA systems in the real world. These concerns span from autonomous agency, to privacy and ethics, workers' rights and the essence of ownership of cognitive output. This workshop seeks to explore the latest research on technologies and policies that allow IA to benefit individuals and groups, both in the private sphere and the workplace.
“Robot-directed speech” désigne la parole adressée à un appareil robotique, des petites enceintes domestiques aux robots humanoïdes grandeur-nature. Les études passées ont analysé les propriétés phonétiques et linguistiques de ce type de parole ou encore l’effet de l’anthropomorphisme des appareils sur la sociabilité des interactions, mais l’effet de l’anthropomorphisme sur les réalisations linguistiques n’a encore jamais été exploré. Notre étude propose de combler ce manque avec l’analyse d’un paramètre phonétique (débit de parole) et d’un paramètre linguistique (fréquence des pauses remplies) sur la parole adressée à l’enceinte vs. au robot humanoïde vs. à l’humain. Les données de 71 francophones natifs indiquent que les énoncés adressés aux humains sont plus longs, plus rapides et plus dysfluents que ceux adressés à l’enceinte et au robot. La parole adressée à l’enceinte et au robot est significativement différente de la parole adressée à l’humain, mais pas l’une de l’autre, indiquant l’existence d’un type particulier de la parole adressée aux machines.
Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcity of emotion datasets collected in the wild and the inability to take into account the dialogue context. The CEMO dataset, composed of conversations between agents and patients during emergency calls to a French call center, fills this gap. The nature of these interactions highlights the role of the conversation’s emotional flow in predicting patient emotions, as context can often make a difference in understanding observed emotional expressions. This paper presents a multi-scale conversational context learning approach for speech emotion recognition, which takes advantage of this hypothesis. We investigated this approach on both speech transcriptions and acoustic segments. Experimentally, our method uses the previous or next information of the targeted segment. In the text domain, we tested the context window using a wide range of tokens (from 10 to 100) and at the speech turns level, considering inputs from both the same and opposing speakers. According to our tests, the context derived from previous tokens has a more significant influence on accurate prediction than the following tokens. Furthermore, taking the last speech turn of the same speaker in the conversation seems useful. In the acoustic domain, we conducted an in-depth analysis of the impact of the surrounding emotions on the prediction. While multi-scale conversational context learning using Transformers can enhance performance in the textual modality for emergency call recordings, incorporating acoustic context is more challenging.
The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in this paper use the CEMO, which was collected in a French emergency call center. Two pre-trained models based on speech and text were fine-tuned for speech emotion recognition. Using pre-trained Transformer encoders mitigates our data’s limited and sparse nature. This paper explores the different fusion strategies of these modality-specific models. In particular, fusions with and without cross-attention mechanisms were tested to gather the most relevant information from both the speech and text encoders. We show that multimodal fusion brings an absolute gain of 4-9% with respect to either single modality and that the Symmetric multi-headed cross-attention mechanism performed better than late classical fusion approaches. Our experiments also suggest that for the real-life CEMO corpus, the audio component encodes more emotive information than the textual one.
Speech Emotion recognition (SER) in call center conversations has emerged as a valuable tool for assessing the quality of interactions between clients and agents. In contrast to controlled laboratory environments, real-life conversations take place under uncontrolled conditions and are subject to contextual factors that influence the expression of emotions. In this paper, we present our approach to constructing a large-scale real-life dataset (CusEmo) for continuous SER in customer service call center conversations. We adopted the dimensional approach for continuous emotion annotation while incorporating contextual information, such as the metadata of the interlocutors and the context of the call. The study also addresses the challenges encountered during the application of the End-to-End (E2E) SER system to the dataset, including determining the appropriate label sampling rate and input segment length, as well as integrating contextual information (interlocutor’s gender and empathy level) with different weights using multi-task learning. The result shows that incorporating the empathy level information can slightly improve the performance of the E2E SER model.
Speech emotion recognition (SER) has received a great deal of attention in recent years in the context of spontaneous conversations. While there have been notable results on datasets like the well-known corpus of naturalistic dyadic conversations, IEMOCAP, for both the case of categorical and dimensional emotions, there are few papers which try to predict both paradigms at the same time. Therefore, in this work, we aim to highlight the performance contribution of multi-task learning by proposing a multi-task, multi-modal system that predicts categorical and dimensional emotions. The results emphasise the importance of cross-regularisation between the two types of emotions. Our approach consists of a multi-task, multi-modal architecture that uses parallel feature refinement through self-attention for the feature of each modality. In order to fuse the features, our model introduces a set of learnable bridge tokens that merge the acoustic and linguistic features with the help of cross-attention. Our experiments for categorical emotions on 10-fold validation yield results comparable to the current state-of-the-art. In our configuration, our multi-task approach provides better results compared to learning each paradigm separately. On top of that, our best performing model achieves a high result for valence compared to the previous multi-task experiments.
Pervasive and ubiquitous sensing technologies have been frequently leveraged in experimental settings within the Ubicomp, ISWC, and HCI community. In the recent past, we have seen advances in the use of ubiquitous sensing technologies to understand and support learning. Recent work investigated a wide range of education-related technologies that detect attention, postures, behaviors, and emotions. However, the promise of delivering these technologies to potential users en masse remains difficult to fulfill. At this time it is important to understand the factors hindering the use of these technologies, to critically evaluate methodologies for technological development, and to discuss in what ways developed technologies can be safely delivered to students. In this workshop, we bring together experts from the fields of ubiquitous and pervasive computing, HCI, and education to discuss and develop an agenda for moving sensing technologies for learning and education from research to practice.
Stéphanie Buisine合作论文数Ecole Nationale Superieure d'Arts et Metiers4