Generative diffusion models have recently emerged as a promising paradigm for image super-resolution (SR) due to their superior generative capabilities. However, existing approaches are typically constrained by fixed upscaling factors determined during training, which hinders the direct generation of high-resolution images at arbitrary scales. Adapting these models to different scales usually requires fine-tuning or retraining, resulting in significant computational costs. In this paper, we propose a training-free framework that enables unseen multi-scale super-resolution using a single pre-trained diffusion model. To mitigate structural distortion at large upscaling factors, we replace standard convolution layers in the pre-trained U-Net with a novel Frequency-Separated Convolution. This mechanism incorporates downsampling and dilated convolutions to enhance structural consistency and detail preservation. Furthermore, we introduce Low-Upscale Guided Fusion and Auxiliary-Branch Classifier-Free Guidance to incorporate stable default-scale priors during denoising, improving perceptual fidelity while maintaining structural integrity. Extensive experiments demonstrate the effectiveness and flexibility of our proposed method, yielding relative CLIPIQA gains of 13.2% for LDM and 3.3% for SeeSR under the × 8 setting on DIV2K-Val.
Restoring images degraded by adverse weather remains challenging due to spatially heterogeneous degradations. Many existing weather-specific restoration models rely on weather-agnostic global aggregation, naive cross-scale fusion, and deterministic objectives, which struggle to handle heterogeneous degradations in all-in-one adverse-weather settings. To address these limitations, we propose an Uncertainty-guided Adverse-weather Restoration Network (UAR-Net), a weather-specific AiO framework that integrates a gated transformer with balanced multi-scale skip connections. Specifically, we employ Gated Dual-scale Transformer Blocks (GDTB) to jointly model selective global interactions and multi-scale local structures, a progressive Balanced Multi-scale Skip Connection (BMSC) for balanced multi-scale feature integration, and an Uncertainty-Aware Refinement Head (URH) that performs artifact removal, detail enhancement, and predictive uncertainty estimation. The model is supervised by a Brightness-Aware Energy Loss (BAE-Loss) to encourage accurate reconstruction with well-calibrated uncertainty. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple adverse-weather benchmarks. The codes will open source upon acceptance.
This paper reviews the NTIRE 2026 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a network that reduces one or several aspects, such as runtime, parameters, and FLOPs, while maintaining PSNR of around 26.90 dB on the DIV2K_LSDIR_valid dataset, and 26.99 dB on the DIV2K_LSDIR_test dataset. The challenge had 95 registered participants, and 15 teams made valid submissions. They gauge the state-of-the-art results for efficient single-image super-resolution.
Adverse weather conditions such as rain, haze, and snow significantly degrade image quality, posing challenges for both human perception and physical AI. Existing restoration methods require large computational budgets, struggling to process high-resolution images and handle different degradations. In this paper, we present Frequency Reconstruction via Spectral Harmonization, a novel lightweight all-in-one restoration method that explicitly decomposes feature representations into high- and low-frequency components at each scale of a hierarchical encoder-decoder architecture. By combining spectral decomposition with spatial processing through Fourier-based skip connections, FReSH-IR captures complementary frequency information without sacrificing spatial detail. Our approach achieves similar restoration quality with 80
Nighttime deraining presents unique technical challenges due to complex degradation patterns, including low illumination and intertwined rain streaks that often amplify visual distortions. Existing methods typically formulate the rainy image as the linear combination of the rain streak residues and the clean image, disregarding the nonlinear interactions between rain and low-light conditions. Additionally, the mixed-scale representations of coupled degradations in nighttime rainy scenes have not been sufficiently explored to enhance deraining performance. To address these issues, we propose a mixed-scale Transformer that effectively captures multi-scale degradations and models them with enhanced nonlinearity, thereby formulating nighttime deraining as a nonlinear restoration problem. Specifically, we achieve mixed-scale representations of degradations through both intra-scale and inter-scale feature fusion across stages. We also propose a gated self-attention (GSA) mechanism to enhance the nonlinear modeling of the Transformer blocks. Building on the dense Mixture of Experts (MoE), we develop a dual-MoE feed-forward network (DMFN) that integrates a scale MoE (SMoE) and a nonlinear MoE (NMoE) to facilitate multi-scale and nonlinear feature representations, respectively. Extensive experimental results demonstrate that our method outperforms state-of-the-art methods on both synthetic and real-world datasets. The code is available at https://github.com/tandaily/NDformer .
Specialized image restoration methods have been extensively explored, each targeting a specific type of degradation. However, real-world images often suffer from composite degradations, prompting growing interest in unified restoration approaches. While recent unified models have shown promising results, many are hindered by high computational complexity, limiting their deployment in resource-constrained settings. Motivated by the parameter-efficient design of Low-Rank Adaptation (LoRA), we propose an efficient attention module specifically designed for composite degradation image restoration. The proposed method adopts a dual-branch architecture, where one branch processes features at full resolution, and the other operates with reduced spatial and channel dimensions to improve efficiency. To better adapt to diverse degradation patterns, the latter branch is further divided into two sub-branches, each incorporating dynamic operations guided by local and contextual priors. These context priors are iteratively updated within each module, drawing inspiration from feedback mechanisms in reinforcement learning, thereby enabling the model to effectively perceive and handle multiple degradation types within a unified structure. Additionally, we introduce a multi-scale feed-forward network to further enhance both performance and computational efficiency. Extensive experiments on two composite degradation benchmarks demonstrate that our proposed network, CDIR, achieves state-of-the-art performance with significantly reduced complexity and fast inference speed. In addition, CDIR shows strong adaptability to various task-specific image restoration scenarios, such as dehazing, desnowing, and deraining. It also performs robustly on domain-specific applications such as ultra-high-definition (UHD), remote sensing, and medical image restoration, highlighting its versatility and practical applicability.
Image restoration aims to recover a high-quality image from its degraded counterpart by mitigating distortions introduced during acquisition, transmission, or environmental interaction. Despite the remarkable progress of deep learning–based restoration models, most conventional approaches remain tightly coupled to predefined degradation assumptions and pixel-level supervision, limiting their capability to handle complex and diverse scenarios or user-dependent restoration targets. Recent advances in multimodal large language models (MLLMs) and vision–language models (VLMs) have introduced a new paradigm in which restoration systems can incorporate semantic reasoning, language-driven interaction, and cross-modal knowledge. In these frameworks, language models extend restoration beyond purely visual reconstruction by enabling degradation interpretation, perceptual alignment, and high-level control. In this survey, we present a systematic review of language-integrated restoration frameworks, organizing existing studies through an interaction-centric taxonomy that captures distinct modes of interaction between language models and restoration networks. We investigate how semantic priors, textual guidance, perceptual supervision, and decision-centric mechanisms recast restoration behavior, and analyze the implications of these developments for model design and training strategies. In parallel, we review emerging language-driven image quality assessment approaches that complement traditional evaluation metrics. Finally, we identify unresolved challenges and outline potential research directions toward more robust, efficient, and trustworthy restoration techniques.
All-in-one image restoration has recently attracted considerable attention for its ability to address multiple degradation types within a single, unified framework. However, existing methods often incur substantial computational overhead, especially when incorporating explicit degradation priors via complex auxiliary branches, hindering their practical deployment. In this paper, we propose AdaptIR, an efficient all-in-one image restoration network equipped with adaptive frequency enhancement. Recognizing that different degradations impact distinct frequency subbands and exhibit spatially varying restoration demands, we design an Adaptive Frequency Enhancement Module (AFEM) that couples frequency learning with adaptive convolutions to better capture frequency-aware information. Specifically, AFEM learns pixel-wise adaptive attention weights to modulate the spectra of dynamic convolutions, enabling spatially adaptive and content-aware restoration. Furthermore, we introduce a lightweight backbone featuring a Receptive Field Expansion Module (RFEM), which enlarges the receptive field of a convolutional U-shaped architecture by convolving wavelet-transform coefficients. By integrating the plug-and-play AFEM into the bottleneck of the baseline model, AdaptIR achieves state-of-the-art performance on all-in-one image restoration tasks involving multiple degradations, while maintaining high computational efficiency. Moreover, the proposed model can be readily extended to single-degradation tasks (e.g., dehazing, desnowing, and deraining) and domain-specific applications, including ultra-high-definition (UHD), medical, and remote sensing image restoration.
This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image restoration under diverse real-world degradation conditions, including blur, low-light, haze, rain, and snow. It provides a common benchmark for evaluating the restoration accuracy, robustness, and generalization capability of models across multiple degradation categories within a unified framework. The competition attracted 158 registered participants, and 20 teams were included in the final ranking after their submitted results were successfully reproduced and verified. This report provides a comprehensive analysis of the submitted solutions and corresponding results, highlighting recent advances in real-world all-in-one image restoration. The summarized methods and empirical findings reveal effective design strategies and establish an updated benchmark for future research in real-world low-level vision.
Vision Transformer (ViT) has shown impressive performance in image restoration due to its ability to capture a large receptive field. However, its complexity grows quadratically with input resolution, limiting its applicability for high-resolution images. In contrast, Convolutional Neural Networks (CNNs) are computationally efficient but are constrained by their inherently local receptive fields, which limit their ability to capture long-range pixel relationships. To address these challenges, we propose StarIR, which possesses the efficiency of CNNs while also capturing a large receptive field, similar to Transformers. StarIR incorporates two key innovations: 1) a dual-domain representation learning framework, with one branch processing spatial details and the other focusing on mesoscale interactions in the frequency domain; and 2) a high-dimensional feature fusion mechanism, the Star operation, which fuses information from both domains through element-wise multiplication, thereby enhancing representational capacity without increasing network width and depth. Our Star operation is followed by a channel attention unit to facilitate global feature modeling and enhance channel-wise interactions. Building on our straightforward yet powerful design principles, StarIR achieves state-of-the-art performance across 21 datasets covering six single-degradation image restoration tasks. Furthermore, our model performs favorably against leading algorithms in two all-in-one settings and demonstrates robustness on two composite-degradation datasets. In addition, StarIR extends well to several domain-specific applications, including ultra-high-definition (UHD) imaging, remote sensing, medical imaging, and underwater image enhancement.
As a non-invasive diagnostic tool, the electrocardiogram (ECG) is easily affected by various noises, which poses difficulties in diagnosing heart diseases accurately. However, traditional denoising methods filter out specific frequencies through time-frequency analysis and are limited by threshold setup, while existing deep-learning methods fail to fully exploit the frequency characteristics of ECG signals. To address this problem, we introduce a frequency-guided framework based on convolutional transformer for ECG denoising, called FGCT. We innovatively combine traditional filtering/decomposition techniques and attention-based networks, and propose a frequency-guided multi-head self-attention (FG-MSA) model and a global channel and spatial enhanced convolution (GCSC) network. For the FG-MSA model, we embed frequency domain priors directly to guide the time domain attention process for extracting intra-band dependencies. For the GCSC network, we employ a global channel-spatial attention to capture inter-band dependencies, distinguish signal from noise to reduce the spectrum overlap noise. Besides, to further enhance the correlation between feature maps across different channels, we use the convolutional layer to perform the down-sampling operation instead of a regular pooling layer. We comprehensively compare our FGCT with the state-of-the-art methods, including the traditional rule-based and learning-based methods. Experimental results demonstrate that our method outperforms these baselines on two widely used ECG benchmarks under four representative noise types (baseline wander, electrode motion, muscle artifact, and their mixture).
Image restoration is a critical task in computer vision, as unwanted degradations significantly reduce image quality and adversely affect the robustness of downstream vision systems. Recently, Transformer-based models have demonstrated remarkable performance in image restoration tasks. However, their quadratic computational complexity with respect to input size limits their practical application. In this paper, we develop a computationally efficient convolutional image restoration model inspired by the design philosophy of Transformer blocks. Specifically, we emulate and unify the functionalities of self-attention and feed-forward layers within a single convolution-based module to achieve adaptive and efficient spatial feature learning. Additionally, this module incorporates multiple receptive fields at attention, module, and scale levels, enabling effective management of degradations of varying sizes. Furthermore, to mitigate frequency discrepancies between degraded and sharp image pairs, we widen and deepen the structure of a frequency learning mechanism, which implicitly refines frequency spectra through a simple yet powerful frequency decomposition and calibration approach. Integrating these designs into a U-shaped convolutional backbone results in our efficient and effective network, termed UIRNet, for image restoration. Extensive experiments across single-degradation, all-in-one, and composite degradation restoration tasks validate the effectiveness of our method. Notably, UIRNet achieves state-of-the-art performance on 11 benchmark datasets covering six single-degradation scenarios and performs favorably against leading all-in-one approaches under a five-degradation setting. Importantly, our method also demonstrates robust performance on two challenging composite degradation image restoration benchmarks, as well as in ultra-high-definition (UHD) applications.
Real-world paired image dehazing remains challenging because haze degradation is spatially non-uniform, illumination-dependent, and physically ambiguous even when haze-free references are available. Existing end-to-end restoration networks usually learn a deterministic mapping from a hazy observation to a clean target, while degradation-sensitive feature responses, reverse haze-formation consistency, and cross-domain negative structure remain insufficiently exploited. In this paper, we propose Backbone-Agnostic Stochastic Perturbation Learning (BSPL), a plug-and-play framework for end-to-end real-world image dehazing. BSPL first introduces a Learnable Stochastic Perturbation Modulator (LSPM), which learns input-conditioned channel-wise and spatial-wise perturbation distributions and converts the resulting feature-response discrepancies into adaptive modulation weights. It then develops a Prior-informed Perturbation-guided Reconstruction Module (PPRM), which reuses the learned bottleneck perturbations together with transmission and atmospheric-light priors to reconstruct the hazy observation from the restored result and enforce degradation consistency. Furthermore, we propose a Dual-space Domain-diversified Distribution-aware Contrastive Loss (D^3CL) to regularize both clean restoration and hazy reconstruction spaces with real-world and synthetic negatives. Experiments on five real-world paired benchmarks show that BSPL consistently improves multiple representative backbones with only marginal additional inference overhead.
Time series anomaly detection (TSAD) is crucial for ensuring the reliability of safety-critical systems. While recent multi-view approaches combining time and frequency domains have advanced performance, they still face important representation and fusion bottlenecks. Specifically, conventional linear spectral mappings may introduce representation distortion when decoding complex non-linear frequency shifts. Furthermore, by restricting the feature space to purely numerical modalities, existing deterministic models lack the global operational semantics required to contextualize operational shifts. This semantic void can lead to inter-modal cognitive conflicts and overconfident misclassifications under noisy environments, whereas traditional probabilistic uncertainty estimation methods remain computationally expensive for real-time TSAD. To address these limitations, this paper proposes a Tri-modal Evidential Synergy Network (TES-Net). First, TES-Net leverages a multi-order Kolmogorov-Arnold network alongside large language model (LLM)-extracted global semantics to adaptively decode non-linear spectral fluctuations and bridge the heterogeneous semantic gap. Second, to resolve inter-modal conflicts efficiently, we design a Tri-modal Evidential Fusion module grounded in Dempster-Shafer evidence theory. This mechanism explicitly quantifies modality-level epistemic uncertainty via mass functions in a single deterministic forward pass, dynamically discounting corrupted modalities through expectation aggregation. Finally, a semantic-gated correlation mechanism employs the global textual prior to modulate local inter-variate physical topologies, differentiating genuine anomalies from benign operational transitions. Extensive experiments on benchmark datasets demonstrate that the proposed method achieves competitive performance compared with recent baselines.
Underwater salient object detection (USOD) has attracted increasing attention for underwater scene understanding and vision-guided robotic applications. However, the spatially non-uniform degradation in underwater images causes spatially varying reliability of structural cues: boundary-sensitive responses can enhance object contours but are vulnerable to degradation-induced noise, whereas region-coherent responses improve semantic completeness but may blur object boundaries. Existing methods rarely explicitly consider the spatial variation in structural cue reliability under underwater image degradation. To address this problem, this work proposes SASC-USOD, a novel framework for learning spatially adaptive structural coordination in USOD. The proposed framework constructs two complementary structural representations with different characteristics. A boundary-sensitive representation is obtained by combining fixed Laplacian filtering with a learnable local-detail transformation to enhance discriminative boundary information, while a region-coherent representation is generated through dual-range anisotropic large-kernel contextual aggregation to capture long-range structural consistency. A spatial coordination module is then introduced to estimate the relative reliability of these structural representations and adaptively coordinate their contributions according to image content. Extensive experiments on the USOD10K and USOD benchmarks demonstrate that SASC-USOD consistently outperforms existing methods, reducing MAE by 4.07% and 23.53% compared with the strongest competing method, respectively. Moreover, its lightweight variant runs at 21 FPS on an NVIDIA Jetson TX2 NX, demonstrating its capability for onboard underwater robotic perception.
Multi-modal image fusion (MMIF) aims to integrate complementary information from heterogeneous sensor modalities. However, substantial cross-modality discrepancies hinder joint scene representation and lead to semantic degradation in the fused output. To address this limitation, we propose C2MFuse, a novel framework designed to preserve content while ensuring cross-modality consistency. To the best of our knowledge, this is the first MMIF approach to explicitly disentangle style and content representations across modalities for image fusion. C2MFuse introduces a content-preserving style normalization mechanism that suppresses modality-specific variations while maintaining the underlying scene structure. The normalized features are then progressively aggregated to enhance fine-grained details and improve content completeness. In light of the lack of ground truth and the inherent ambiguity of the fused distribution, we further align the fused representation with a well-defined source modality, thereby enhancing semantic consistency and reducing distributional uncertainty. Additionally, we introduce an adaptive consistency loss with learnable transformation, which provides dynamic, modality-aware supervision by enforcing global consistency across heterogeneous inputs. Extensive experiments on five datasets across three representative MMIF tasks demonstrate that C2MFuse achieves efficient and high-quality fusion, surpasses existing methods, and generalizes effectively to downstream visual applications.
Monocular depth estimation is critical for applications like autonomous driving and robotics. The complementary properties of event and image modality motivate the fusion-based methods for robust depth estimation. However, existing fusion methods rely on convolutional or attention-based architectures, which either struggle with global dependencies or incur high computational cost, limiting their suitability for long-sequence modeling in depth tasks. Besides, effective image-event fusion remains a key challenge, as most existing methods directly fuse features without addressing the domain gap and differences in representational characteristics between raw events and images, leading to semantic bias and degraded performance. In this work, we propose AIMDepth, an Asymmetric Image-Event Mamba framework for monocular depth estimation, built entirely on state space models to ensure linear computational complexity and accurate prediction. To address input-domain misalignment, we introduce a Spectral Cross-modal Prior Guidance (SCPG) module that performs bidirectional prior injection at the input level. To mitigate representational imbalance between sparse events and dense images, we design an Asymmetric Modal-aware Encoder (AME) that allocates separate encoding paths for each modality and facilitates feature-level alignment tailored to their distinct information densities. To further enhance fusion, we develop a Modality-interactive Local Refinement (ModiLocal) module that enables hierarchical interaction and fine-grained alignment through SSM-based modeling. Extensive experiments on public datasets demonstrate that AIMDepth achieves state-of-the-art performance and strong robustness in complex environments.
Image restoration aims to recover high-quality images from their degraded counterparts, which is critical in various domains. The field has increasingly focused on all-in-one solutions that use a single model to address various types of degradation, enhancing generalization capabilities. Many existing approaches leverage different mechanisms to extract high-dimensional features from the spatial domain for encoding degradations. However, high-dimensional features contain noise and redundant information that can depress restoration performance. Moreover, both spatial and frequency domains include distinct information related to different degradation types. Based on these observations, we introduce a novel all-in-one image restoration framework called DRFIR, which utilizes dimensionality reduction technology to effectively harness both spatial and frequency domain information. Specifically, our approach first obtains low-dimensional embeddings from spatial and frequency domains via a dimensionality reduction block based on contrastive learning. Additionally, a novel fusion module is developed to integrate spatial and frequency embeddings with image features to enhance restoration quality. Extensive experimental results demonstrate that our DRFIR model achieves promising restoration performance on a range of benchmark datasets while improving downstream tasks.
This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world all-in-one image restoration under diverse real-world degradation conditions, including blur, low-light, haze, rain, and snow. It provided a unified benchmark to evaluate the robustness and generalization ability of restoration models across multiple degradation categories within a common framework. The competition attracted 124 registered participants and received 9 valid final submissions with corresponding fact sheets, significantly contributing to the progress of real-world all-in-one image restoration. This report provides a detailed analysis of the submitted methods and corresponding results, emphasizing recent progress in unified real-world image restoration. The analysis highlights effective approaches and establishes a benchmark for future research in real-world low-level vision.