Comprehensive dialogue understanding and effective cross-modal interaction remain challenging in multimodal Emotion Recognition in Conversations (ERC). Existing methods often struggle to efficiently process lengthy conversations, with cross-modal interactions biased towards text modality, resulting in some degree of misrepresentation. Graph Neural Networks (GNNs) offer promise by structuring dialogue sequences into graphs, yet they tend to overlook semantic details in high-frequency signals due to their reliance on low-pass filters. To address these limitations, we propose FrameERC, a novel framework that utilizes graph framelet transforms to analyze dialogues. FrameERC first transforms multimodal features into efficient graph signals using two distinct decoupling strategies. By decomposing the graph signals across a broad frequency spectrum with low- and high-pass filters, FrameERC captures critical emotional subtleties that traditional spatial message-passing GNN models may ignore, thereby enriching detailed emotion recognition. Moreover, we introduce a dual-reminder fusion mechanism to enhance the meaningful contribution of non-textual modalities in ERC, ensuring a comprehensive semantic understanding. Extensive experimental results demonstrate that FrameERC achieves superior performance compared to current state-of-the-art methods on two widely-used multimodal ERC datasets.
更多
查看译文
关键词
Multimodal graph neural networks,Framelet transform,Emotion recognition in conversation