2025 4TH INTERNATIONAL SYMPOSIUM ON ROBOTICS, ARTIFICIAL INTELLIGENCE AND INFORMATION ENGINEERING, RAIIE(2025)
Dalian Minzu Univ
被引用0|浏览7
摘要
Sound source localization in complex acoustic environments is a challenging task. The spectrogram contains rich sound features, including the distribution of frequency components, intensity variations, and harmonic structures, which are crucial for accurate sound source localization. In this paper, we design a multi-scale feature aggregation network (MFAnet) to effectively utilize the sound features in the spectrogram for accurate sound source localization. Specifically, the multi-scale features are extracted by combining the Res2Net architecture with dynamic convolution to capture the rich contextual dependencies across multiple layers. Moreover, this network incorporates a dual attention mechanism, including both spatial and channel attention modules, in the feature extraction stage to help the network focus on the most discriminative parts of the features. Finally, we use the multi-output regression to estimate the 3D coordinates of sound sources. Experimental results indicate that the proposed MFAnet performs well in complex acoustic environments, especially in the case of overlapping sound sources conditions, and is significantly outperform the existing methods. Additionally, we conduct ablation experiments to verify the effectiveness of the proposed method.