Time-Continuous Audiovisual Fusion with Recurrence vs Attention for In-The-Wild Affect Recognition.

Vincent Karas,Mani Kumar Tellamekala,Adria Mallol-Ragolta,Michel F. Valstar,Björn W. Schuller

IEEE Conference on Computer Vision and Pattern Recognition（2022）

引用 9|浏览40

暂无评分

摘要

This paper presents our contribution to the 3rd Affective Behavior Analysis in-the-Wild (ABAW) challenge. Exploiting the complementarity among multimodal data streams is of vital importance to recognise dimensional affect from in-the-wild audiovisual data, as the contribution affect-wise of the involved modalities might change over time. Recurrence and attention are two of the most widely used modelling mechanisms in the literature for capturing the temporal dependencies of audiovisual data sequences. To clearly understand the performance differences between recurrent and attention models in audiovisual affect recognition, we present a comprehensive evaluation of fusion models based on LSTM-RNNs, self-attention, and cross-modal attention, trained for valence and arousal estimation. Particularly, we study the impact of some key design choices: the modelling complexity of CNN backbones that provide features to temporal models, with and without end-to-end learning. We train the audiovisual affect recognition models on the in-the-wild Aff-wild2 corpus by systematically tuning the hyper-parameters involved in the network architecture design and training optimisation. Our extensive evaluation of the audiovisual fusion models indicate that under various experimental settings, compared to RNNs, attention models may not necessarily be the optimal choice for time-continuous multimodal fusion for emotion recognition.

查看译文

关键词

affect,attention,fusion,recognition,time-continuous,in-the-wild

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要