
In this study, we tested two bipolar groups, manic-episode and depressive episode patients and heathy controls over a series of McGurk effect stimuli as well as auditory and visual-only (lip-read) speech stimuli. We hypothesized that the bipolar group’s auditory–visual speech integration should be weaker than the control group. We also predicted the manic-period bipolar individuals should integrate visual speech information more robustly than their depressive-episode counterparts. These hypotheses were not supported, and all groups were found to integrate auditory and visual speech information comparably. However, paradoxically, the depressive-episode individuals were unable to lip-read despite that they were able to integrate the two sources of speech information. Here, we discuss and try to make sense of this behavioural data with solid predictions on how corresponding physiological data can decipher our findings. On top of this discussion, physiological predictions and possibilities are presented.
In this work, we propose a multi-modal emotion recognition model to improve the speech emotion recognition system performance. We use two parallel Bidirectional LSTM networks called acoustic encoder (ENC1) and speech embedding encoder (ENC2). The acoustic encoder is a Bi-LSTM which takes sequence of speech features as inputs and speech embedding encoder is also a Bi-LSTM which takes sequence of speech embeddings as input and the output hidden representation at the last time step of both the Bi-LSTM are concatenated and passed into a classification which predicts emotion label for that particular utterance. The speech embeddings are learned using the encoder-decoder framework as described in [1] using skipgram [2] training. These embeddings are shown to outperform word embeddings(word2vec) in many word similarity benchmarks. Speech embeddings are shown to capture semantic information present in speech and speech embeddings have the capa-bilities to handle speech variabilities which is not possible by plain text. We compare our model with the word embedding based model where we feed Word2Vec to ENC2 and speech features to ENC1 and we observe that the speech embedding based model gives better results compared to word embedding based model. We compare our system to previous multi-modal emotion recognition models which use text and speech features and we get absolute 2.59% improvement over the previous sys-tems[8] on IEMOCAP dataset. We also compare our system performance with different speech embedding dimensions of [50,100,200,300] and we observe the speech embedding of 50 dimensions is achieving 68.59% accuracy.
The 15th International Conference on Auditory-Visual Speech Processing (AVSP2019), 10 augustus 2019
During everyday communication, voice and facial cues are combined. A preference for the auditory or visual channel is chosen automatically whereas, in most of the previous studies, guided attention was used. In our study, we performed a comparison of the visual influence on vocal non-verbal emotions and gender using the same paradigm in the same subjects without instructions on attention direction. The voice for emotions and gender was modeled as a continuum with 11 steps. The validated non-ambiguous images of gender and emotions were presented in the congruent and incongruent with the voice way. Audiovisual performance was assessed with respect to the auditory performance. We observed a small improvement of performance in the congruent audiovisual stimulation both for gender and emotions with a smaller effect for emotions. In the incongruent conditions, face cues strongly dominated the performance with a significantly larger effect for gender. The proportion of the subjects who made their decision on the visual basis was significant only for the gender voice continuum. The strength of facial dominance is significantly different between the identity voice information and emotional prosody. We suggest that face-voice interaction in human may not be the same for linguistic, para-linguistic and identity properties.
There is well-established evidence that visual articulatory information in the face and head aids identification and discrimination of lexical tone. However, the nature and locus of this information is only beginning to be specified. In previous work we identified a predominant role of head motion over face motion in both the perception and production of Cantonese lexical tone, the latter using OPTOTRAK motion tracking. We have now extended the set of OPTOTRAK markers to include the eyebrows and the larynx, and collected data from a corpus of Cantonese, Thai and Mandarin speakers. Here we report on a Thai speaker producing the five Thai tones on four Thai syllables in isolated words and sentences and in normal, whispered, and Lombard speech. Principal components (PCs) for the face (eyebrows, lips, jaw), the larynx and for independent head movement were extracted and linear mixed model analyses of range of PC1 scores revealed good differentiation on the basis of syllable identity and context and speech style. Of particular importance, the five Thai tones were best differentiated by head and larynx motion. So, these results add larynx motion as a possible visible cue for tone perception. Studies across speakers and the three languages will follow.
Prominent theories of consciousness such as Global Workspace Theory propose that consciousness is required for multimodal integration. We tested this proposal with a processing bottleneck known as the unmasked attention blink, which we used with synchronized auditory-visual stimulus streams to delay awareness of the onset of a stimulus in one modality but not the other. Event-related potentials (ERPs) were then used to examine auditory and visual integration processes in the context of the unmasked attentional blink. To index auditory-visual (AV) integration, we recorded ERPs following the presentation of auditory-visual (AV) and unimodal second targets (T2) in AV presentation streams, which were presented during or after the attentional blink period, 200-300 ms or 600-700ms after the onset of first targets (T1) respectively. The results showed that AV and unimodal ERP responses were more similar during the attentional blink than outside of it. This result suggests that AV integration was suppressed and visual and auditory information were processed independently during the attentional blink. AV integration occurred both before and during the time window of the P3 ERP component (300-500 ms), which is well-established as the earliest time window for attentional blink ERP effects. The attentional blink suppressed AV integration only at later stages of processing while that at earlier, pre-P3 latencies was relatively intact. We discuss the implications of this finding for theories linking consciousness and integration.
Automatic emotion recognition is a challenging task since emotion is communicated through different modalities. Deep Convolution Neural Networks (DCNN) and transfer learning have shown success in automatic emotion recognition using different modalities. However significant improvement in accuracy is still required for practical applications. Existing methods are still not effective in modelling the temporal relationships within emotional expressions or in identifying the salient features from different modes and fusing them to improve accuracies. In this paper, we present an automatic emotion recognition system using audio and visual modalities. VGG19 models are used to capture frame level facial features followed by a Long Short Term Memory (LSTM) to capture their temporal distribution at a segment level. A separate VGG19 model captures auditory features from Mel Frequency Cepstral Coefficients (MFCC). The extracted auditory and visual features are fused together and a Deep Neural Network (DNN) with attention is used in classification using majority voting. Voice Activity Detection (VAD) on the audio stream improves performance by reducing the outliers in learning. The system is evaluated using Leave One Subject Out (LOSO) and K-fold cross-validation and our system outperforms state of the art methods on two challenging databases.
In face-to-face communication, emotions, as well as phonemes, are perceived through the integration of visual information on the face and auditory information on the voice. Previous studies suggested that children integrate audiovisual information differently from adults. To uncover the mechanism behind this developmental process, we investigated the gaze patterns of Japanese children aged 5 to 12 and adults when looking at speakers’ faces after being asked to judge their emotions (emotion perception task) or pronounced syllables (phoneme perception task). The results showed that the participants fixated longer on the speaker’s eyes in the emotion perception task than they did in the phoneme perception task. Moreover, the participants’ fixation on the speaker’s eyes increased with age in emotion perception, while it decreased with age in phoneme perception. This finding suggests that children can shift their attention to a face depending on what they are required to judge, and that this attentional shift becomes more sophisticated with age. We discuss the development of audiovisual integration in terms of the relationship between the participants’ gaze and perception.
Children with hearing loss face a range of challenges when listening to and processing speech; in particular, they may process spoken language slowly in comparison to normal-hearing peers [1]. How then can speech processing speed be improved for children with hearing loss? In this study, a phoneme monitoring task was used to assess whether 7-11-year-old children with hearing loss showed faster speech processing when visual speech cues were available compared to auditory-only presentation. Children with hearing loss did receive an audiovisual benefit for processing speed, however this was primarily driven by cases in which the target phoneme in the monitoring task was visually salient. No difference was found between the performance of the children with hearing loss and a control group of children with normal hearing, however the results suggest that children with hearing loss who use hearing aids may receive a greater audiovisual benefit for processing speed than those who use cochlear implants. These findings have implications for practical interventions for children with hearing loss.
This paper presents the acoustic and visual modeling of nine attitudinal expressions that were realized by the German speaking robot SMiRAE which is a speech-enabled version of the non-speaking robotic face MiRAE previously developed at Indiana University. The parameter-oriented acoustic model is based on the German Mary TTS which is part of the speech processing system InproTK. Visual realization of expressions is based on five defined basic emotions of the Facial Action Coding System (FACS) developed by Ekman . B oth models were additionally modified with respect to results of an audio-visual analysis and evaluation of human portrayals of attitudes recorded in our previous work. The plausibility of synthesized attitudinal expressions is shown by an association study in which 18 participants described 54 attitudes in a free association. Basis for a 5-cluster classification was the first four dimensions of a correspondence analysis which accounted 78% of variance in participant perception. Significant correlations were seen between 66 normalized participant descriptions and the robot’s displayed attitudes. For instance, the attitudes admiration and politeness were associated with the terms freundlich and gluecklich , the interrogative attitudes surprise and doubt with the terms fragend, verwundert and skeptisch, the expression uncertatinty was perceived with traurig and besorgt .
Visual speech cues from a speaker's talking face aid speech segmentation in adults, but despite the importance of speech segmentation in language acquisition, little is known about the possible influence of visual speech on infants' speech segmentation. Here, to investigate whether there is facilitation of speech segmentation by visual information, two groups of English-learning 7-month-old infants were presented with continuous speech passages, one group with auditory-only (AO) speech and the other with auditory-visual (AV) speech. Additionally, the possible relation between infants' relative attention to the speaker's mouth versus eye regions and their segmentation performance was examined. Both the AO and the AV groups of infants successfully segmented words from the continuous speech stream, but segmentation performance persisted for longer for infants in the AV group. Interestingly, while AV group infants showed no significant relation between the relative amount of time spent fixating the speaker's mouth versus eyes and word segmentation, their attention to the mouth was greater than that of AO group infants, especially early in test trials. The results are discussed in relation to the possible pathways through which visual speech cues aid speech perception.
Embodied Conversational Agents (ECA) and Interactive Virtual Humans (IVH) based on 3D Virtual and Augmented reality technologies can be immensely beneficial for human to human communications and interpersonal skills training for a range of skills including sales pitching, negotiation, leadership, interviewing, and communicating with empathy and cultural sensitivities. They provide an opportunity to practice the skills needed for communicating with humans in difficult environments, with virtual simulated scenarios. And these complex environments can include health care contexts, sales, retail and customer service contexts, or communication with people with different cultural and ethnic backgrounds. At HCT research Centre at University of Canberra, we have been investigating several new technology frameworks and algorithmic techniques for building ECAs and IVHs based on integrating Artificial Intelligence (AI), deep machine learning, cloud computing, 3D virtual and augmented reality technologies, as training simulators for different application contexts. In this paper, we present the details in terms of the research, technology and implementation for modelling different persona for ECAs and IVHs. A multilevel architecture with adaptable functional modules based on the customized persona for different virtual environments, their implementation using an integrated cloud based software platform and its evaluation is discussed.
The present study investigates the influence of familial sinistrality on audiovisual speech perception for young adults. Incongruent video stimuli dubbed over dichotic and diotic audio stimuli were utilized. Participants’ responses for the right incongruent, left incongruent and video incongruent audiovisual stimuli were analyzed. Results indicated significantly higher proportion of fusion responses demonstrating increased audiovisual interaction in participants with familial sinistrality compared to those without familial sinistrality.
Recent studies have demonstrated that multisensory emotion perception is modulated by culture. Tanaka et al. (2010) showed that Japanese people are more tuned than Dutch people to vocal processing in adults. The present study investigated how such a cultural difference develops in children aged 5-12 years. In the experiment, a face and a voice, expressing either congruent or incongruent emotions, were presented simultaneously on each trial. Participants judged whether the person is happy or angry. The results showed that the rate of vocal responses was higher in Japanese than Dutch in adults, especially when in-group speakers expressed a happy face with an angry voice. The rate of vocal responses was low in both Japanese and Dutch 5-6-year-olds, while it increased over age only in Japanese participants. These results suggest that combinations of facial and vocal emotions have specific meanings and that culture-specific multisensory display rules are acquired with age in childhood.
In the context of developing an expressive audiovisual speech synthesis system, the quality of the audiovisual corpus from which the 3D visual data will be extracted is important. In this paper, we present a perceptive case study on the quality of the expressiveness of a set of emotions acted by a semi-professional actor. We have analyzed the production of this actor pronouncing a set of sentences with acted emotions, during a human emotion-recognition task. We have observed different modalities: audio, real video, 3D-extracted data, as unimodal presentations and bimodal presentations (with audio). The results of this study show the necessity of such perceptive evaluation prior to further exploitation of the data for the synthesis system. The comparison of the modalities shows clearly what the emotions are, that need to be improved during production and how audio and visual components have a strong mutual influence on emotional perception.