ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)
State Key Laboratory of Multimodal Artificial Intelligence Systems
被引用0|浏览4
摘要
Recently, image restoration (IR) methods have demonstrated promising potential by leveraging pre-trained cross-modal priors to provide explicit text guidance. However, these methods struggle with mixed degradations, as they cannot effectively process multi-instruction prompts, often yielding unsatisfactory outputs. They typically rely on a single text prompt applied iteratively to address the most dominant degradation. The process is computationally intensive and time-consuming. In this paper, we propose a novel IR framework capable of effectively understanding lengthy and complex instruction prompts and restoring multiple degradations simultaneously. Our key innovation is the attribute-aware attention (ATT) module, which disentangles low-level image features into shared basis attribute elements by modeling the intrinsic relationships between degradations and attributes. The ATT module extracts fine-grained degradation cues from multi-level textual descriptions and establishes cross-modal projections in the latent attribute space, generating a set of attention maps. These maps explicitly localize regions associated with each attribute, enhancing the interpretability of the restoration process. Extensive experiments demonstrate that our method outperforms existing methods across various tasks. Furthermore, our model exhibits strong generalization when handling unseen degradations.