ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)
Xidian University
被引用0|浏览2
摘要
Deep joint source-channel coding (Deep JSCC) has emerged as a key technology for semantic communication. However, existing methods are unable to guarantee the fidelity of Regions of Interest (ROI) under severe bandwidth constraints and adverse channel conditions. To address this issue, we propose a text-guided ROI-aware JSCC framework, leveraging textual cues to direct the encoder’s focus on the ROI and incorporate semantic priors for improved robustness. In this framework, a bidirectional multi-scale alignment strategy is introduced to ensure cross-modal consistency and preserve semantic alignment between visual features and text across various feature hierarchies. Experimental results demonstrate that our method achieves 15–30% improvement over baselines under various channel conditions and compression ratios, with the largest gains under low signal-to-noise ratios and bandwidth-constrained conditions.
更多
查看译文
关键词
Semantic Communication,Deep Joint Source–Channel Coding,Image Transmission,Regions of Interest,Cross-modal Alignment