Infrared (IR) imaging is indispensable for perception in adverse environments, yet real-world data is often corrupted by dynamically coupled degradations that impair both visual quality and downstream semantic understanding. Although diffusion models offer powerful generative priors, existing approaches remain ill-suited to this setting. Their slow multi-step sampling, reliance on RGB-driven statistics misaligned with IR physics, and the necessity for costly fine-tuning of all model parameters render them impractical for dynamic IR perception. We present a unified diffusion framework that re-formulates IR restoration as a single-step generative process. The core idea is to associate each degraded input with a specific intermediate latent state in the diffusion trajectory, enabling the model to reconstruct the clean image via a single, direct reverse step. Physical realism is further reinforced through an IR-specific spectral regularization that preserves the characteristic energy distribution of thermal emissions. Addressing the diverse and rapidly shifting demands of dynamic IR perception, we further develop a task-aware low-rank adaptation mechanism. This mechanism employs a lightweight prompting hypernetwork to generate compact modulation parameters, facilitating rapid and scalable adaptation ability without retraining the entire network. Comprehensive evaluations demonstrate that our framework attains state-of-the-art restoration performance, preserves reliable semantic structures, and supports rapid adaptation that generalizes effectively across diverse tasks and conditions.
Infrared and visible video fusion is pivotal for robust perceptual systems, aiming to synthesize a comprehensive video stream that leverages both thermal resilience and textured details. However, prevailing methods, by treating video as independent frames, inherently introduce temporal incoherence, such as flickering and ghosting artifacts. While diffusion models possess strong generative priors to remedy this, their iterative nature is prohibitively slow for video. To resolve this fundamental dilemma, we propose a streaming diffusion model for efficient infrared and visible video fusion, termed SDMFusion. Our key insight is to distill the generative prior of a pre-trained diffusion model into a one-step sampling framework, while explicitly modeling temporal dynamics. We design a memory-augmented latent pipeline where a temporal aggregation adapter aligns and propagates cross-frame features to ensure coherence, supported by a dedicated temporal consistency loss. This approach effectively decouples the challenge of achieving high fidelity from maintaining temporal stability. Extensive experiments on four benchmarks demonstrate that our method establishes a new state-of-the-art, generating fused videos with exceptional spatio-temporal consistency at a speed suitable for real-time application.
Low-light image enhancement (LLIE) must simultaneously brighten images and preserve fidelity, a tension rooted in the disparity between low- and high-frequency components. Low frequencies encode illumination and global structure, whereas high frequencies carry textures but are easily corrupted by noise; spatial-domain models that mix them often cause over-smoothing or noise amplification. We propose LWINet, a lightweight wavelet-domain interaction network. A three-level discrete wavelet transform explicitly separates LL from directional bands LH/HL/HH. At each scale, cross-band cross-attention with an asymmetric bidirectional design achieves “structure-guides-detail, detail-sharpens-structure”: LL acts as the query to direct high-frequency restoration, and an Adaptive Alpha Module (AAM) based on amplitude co-occurrence and orientation consistency modulates the HH contribution to constrain LL contours. A dual-domain processing module preserves local sharpness in the spatial domain and reshapes energy in the frequency domain via radial low-pass modeling. Reconstruction adopts residual inverse DWT with progressive reinjection and skip connections. Under a lightweight configuration (776K parameters) and low computational cost, LWINet suppresses noise, preserves details, and improves visual consistency and generalization, achieving favorable results on public benchmarks.
Image exposure correction tasks often involve addressing issues such as overexposure, underexposure, and uneven exposure. These exposure errors can result in loss of image details, color distortion, and significant degradation of image quality. Exposure correction can be divided into two main challenges: recovering structural information like details and colors, and restoring proper illumination levels. To tackle these challenges more effectively, we propose a dual-prompt approach that introduces structural prompts and illumination prompts to guide the network’s learning process. (i) Structural prompts consist of a set of learnable parameters that help the network adopt specific restoration strategies based on different lighting conditions. For varying degrees of overexposed and underexposed inputs, structural prompts guide the network to apply the appropriate structural restoration strategies to better preserve and enhance image details. (ii) Illumination prompts are divided into positive text prompts and negative text prompts. Negative prompts are further categorized into overexposure negative prompts and underexposure negative prompts. We use CLIP (Contrastive Language–Image Pretraining) to align text prompts with images, creating loss functions that accurately assess the exposure state of the image. Compared to directly inputting text descriptions into CLIP, our pre-trained, learnable text prompt parameters more robustly guide the loss function to make precise assessments of the input image’s exposure quality. In summary, our dual-prompt approach effectively addresses the key issues of detail recovery and illumination estimation in exposure correction tasks through the combined action of structural and illumination prompts.
Enhancing low-light images is a longstanding challenge due to varying illumination, amplified noise, and color distortion. Brightening the dark regions without overexposing well-lit areas, while preserving color fidelity and suppressing noise, requires careful, adaptive processing. Existing methods often rely on global adjustments or simplistic grayscale approximations of the signal-to-noise ratio (SNR), resulting in overexposure, residual noise, and distorted colors. These issues stem from the lack of a dynamic, region-aware balance between brightness correction and denoising, as well as the absence of noise-aware guidance in color refinement. To address these limitations, we propose SNRD-Net, a novel SNR-guided dual-enhancement network that follows a divide-and-conquer principle to decouple luminance and chrominance enhancement. We first enhance the luminance channel and compute an SNR map to accurately represent the spatial distribution of noise. This map drives a dual-pathway denoising strategy, offering targeted noise suppression, and is also leveraged through an SNR-guided fusion mechanism to inject noise awareness into the color refinement process. Additionally, to enrich existing low-light image benchmarks, we introduce a high-quality, real-world paired-image dataset of 450 images (available in CR2, PNG, and JPG formats), captured predominantly outdoors in challenging low-light conditions, serving as a dedicated evaluation benchmark. Our proposed method is assessed across nine low-light image benchmarks, encompassing both paired-image and no-reference datasets, including the newly introduced one. Experimental results demonstrate the effectiveness of the proposed method.
Learning-based real image dehazing methods have achieved notable progress, yet they still face adaptation challenges in diverse real haze scenes. These challenges mainly stem from the lack of effective unsupervised mechanisms for unlabeled data and the heavy cost of full model fine-tuning. To address these challenges, we propose the haze-to-clear text-directed loss that leverages CLIP’s cross-modal capabilities to reformulate real image dehazing as a semantic alignment problem in latent space, thereby providing explicit unsupervised cross-modal guidance in the absence of reference images. Furthermore, we introduce the Bilevel Layer-positioning LoRA (BiLaLoRA) strategy, which learns both the LoRA parameters and automatically search the injection layers, enabling targeted adaptation of critical network layers. Extensive experiments demonstrate our superiority against state-of-the-art methods on multiple real-world dehazing benchmarks. The source code will be publicly available upon acceptance.
Effective fusion of acoustic spatial cues from echo signals and informative visual cues from images is essential for accurate multimodal depth estimation, as each modality provides complementary information while also exhibiting inherent limitations. To achieve better fusion of multimodal data, we introduce HM-Net, a hierarchical multimodal depth estimation framework that integrates echo and visual information through both low level feature fusion and high level semantic fusion. During the low level feature fusion stage, we employ two attention-weighted multimodal fusion modules, with one enhancing echo features under visual guidance and the other enhancing visual features under echo guidance. Each module adjusts modality-specific weights based on cross-modal interactions, resulting in adaptive multimodal features. During the high level semantic fusion stage, these multimodal features are processed to generate scene intervals that represent the global structural layout of the scene, scene feature maps that serve as semantic refinements to the global structure delineated by the scene intervals, capturing localized depth cues such as object boundaries, surface continuity, and occlusion relationships. By fusing these two high level semantic representations, we predict the final depth map, which captures the global depth distribution of the scene and enables precise depth estimation. Quantitative and qualitative evaluations on the Replica, Matterport3D, and BatVision datasets demonstrate the exceptional performance of our method in multimodal depth estimation tasks. The source code is available at: https://github.com/henu77/HM_Net .
Underwater images typically suffer from two main types of degradation: reduced visibility caused by scattering and color distortion due to color cast. Most existing deep learning-based enhancement methods adopt end-to-end architectures to address both issues simultaneously. However, this design not only limits the model’s generalization capability but also hinders practical deployment due to excessive computational overhead. To this end, this paper proposes a binarized Decoupled Synergistic Optimization Strategy (DSOS), which explicitly decouples scattering and color cast degradations and performs collaborative optimization through specialized subtask modules. Each subtask learns purer features under the guidance of independent supervised signals, while a cascaded architecture ensures effective global restoration. Furthermore, cross-module collaborative optimization effectively mitigates the performance degradation caused by binarization, achieving a favorable balance between efficiency and high accuracy. Experimental results on multiple publicly available underwater image datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches in both restoration quality and computational efficiency.
Robust 3D perception under adverse weather is critical for autonomous systems. While mmWave Radars are inherently weather-resistant, conventional 2D rotating Radar sensors lack direct elevation resolution, limiting their 3D perception ability. Although 4D imaging radars can provide elevation information, they typically suffer from limited coverage and range. In this work, we exploit a key observation about mechanically rotating 2D mmWave Radars: in each sweep, an overlap exists between adjacent azimuth beam coverage due to the width of the main lobe, which makes the reflected intensity difference imply object materials and geometric shapes, including elevation. With this observation, we propose a method that learns 3D occupancy by disentangling bird’s-eye view (BEV) layout and elevation estimation from one frame Radar scan. Specifically, we partition one sweep into two interleaved subsets, corresponding to overlapping beam directions, and utilize them to infer coarse geometric structure through spatial differences and intensity patterns. Extensive quantitative and qualitative evaluations on two real-world datasets demonstrate that our proposed method outperforms existing baselines. The codes will be publicly available.
Infrared and visible image fusion (IVIF) integrates complementary modalities to enhance scene perception. Current methods predominantly focus on optimizing handcrafted losses and objective metrics, often resulting in fusion outcomes that do not align with human visual preferences. This challenge is further exacerbated by the ill-posed nature of IVIF, which severely limits its effectiveness in human perceptual environments such as security surveillance and driver assistance systems. To address these limitations, we propose a feedback reinforcement framework that bridges human evaluation to infrared and visible image fusion. To address the lack of human-centric evaluation metrics and data, we introduce the first large-scale human feedback dataset for IVIF, containing multidimensional subjective scores and artifact annotations, and enriched by a fine-tuned large language model with expert review. Based on this dataset, we design a domain-specific reward function and train a reward model to quantify perceptual quality. Guided by this reward, we fine-tune fusion networks through Group Relative Policy Optimization, achieving state-of-the-art performance that better aligns fused images with human aesthetics.
Recent advances in learning-based underwater image enhancement have achieved remarkable progress. However, the inherent diversity and complexity of underwater scenes still limit the ability of existing approaches to simultaneously restore fine structural details and global image layouts. To address this challenge, we propose a Resonant Fusion (ReFu) framework that explicitly leverages complementary information in both spatial and frequency domains. Specifically, we design a frequency decomposer and a spatial decomposer to capture high- and low-frequency cues from different perspectives. A resonant fuser is then introduced to adaptively integrate high-frequency resonances for detail refinement and low-frequency resonances for structural consistency. This fine-grained cross-domain fusion significantly improves structural preservation and detail enhancement, thereby generating visually more natural and perceptually friendly underwater images. Extensive quantitative and qualitative evaluations across diverse underwater benchmarks show that ReFu consistently surpasses state-of-the-art methods by a clear margin. Comprehensive ablation studies further validate the effectiveness of each module and prove the necessity of the proposed ReFu mechanism. Our code is available at https://github.com/CircleQa/ReFu-main.
In computer vision, correcting the exposure level is a fundamental task for enhancing the visual quality of observations with inappropriate lightness. However, existing methodologies tend to be impractical because they lack adaptability to unknown scenes due to restricted modeling patterns and struggle to achieve satisfactory efficiency due to complex computational flows. To tackle these challenges, we establish a new practical exposure corrector (PEC) that excels in both quality and efficiency. Specifically, to overcome the limited expressive power of existing modeling patterns, we build a general model with exposure-sensitive compensation to provide an intuitive modeling perspective. We also design a simple but effective exposure adversarial function to catalyze scene-adaptive compensation. Building on the aforementioned key concepts, we develop a stable and robust iterative shrinkage scheme, avoiding the complex inferences encountered in existing studies. Extensive experimental evaluations across eight challenging datasets showcase the strong adaptability of the developed model to unknown environments. The model offers impressive processing speed, requiring only 0.0009 s to handle a 2K image on a device equipped with a GeForce RTX 2080Ti GPU. Experimental analysis of different downstream vision tasks further verifies the flexibility of the model. The code is available at https://github.com/vis-opt-group/PEC .
Underwater instance segmentation is essential for fine-grained scene understanding. However, underwater imagery exhibits a strong domain gap from in-air vision due to severe degradation (e.g., turbidity). Consequently, despite its general segmentation ability, SAM degrades sharply underwater. In this work, we propose BiPA, which effectively adapts SAM to the underwater domain. To be concrete, we construct an underwater SAM with dual prompts and introduce a foreground-attentive injection block to enhance local foreground representation. We formulate dense prompt learning as a bilevel optimization, explicitly capturing the mutual dependency between prompt and model. To make this tractable, we design a two-stage learning strategy. The first stage adapts the dense prompt itself, updating it with Bayesian optimization to learn efficiently. The second stage fine-tunes the model parameters under the frozen optimized prompt, which finally enables effective cross-domain adaptation. Extensive experiments and analyses verify the superiority and efficiency of BiPA. Code will be released if this work can be accepted, fortunately.
Hyperspectral and multispectral image fusion plays a pivotal role in boosting the spatial resolution of hyperspectral data while simultaneously expanding their applicability across downstream tasks, ranging from precision agriculture and environmental monitoring to urban planning and disaster assessment. Over the past few years, deep learning-based methods have achieved remarkable progress, especially in capturing both global contextual information and fine-grained features from hyperspectral imagery. Some of these models attain impressive quantitative accuracy. However, they often neglect the computational footprint, thereby limiting their deployability on resource-constrained platforms. Moreover, few studies have thoroughly investigated how to guarantee the quality and generalizability of refined feature extraction when confronted with ground objects that exhibit variations in scale and shape. To further enrich and complement the existing body of work, this paper introduces a Hyperspectral-Multispectral Fusion Network, termed AMSGL-Net, which simultaneously pursues efficient global feature extraction and robust, scale-aware representation of complex ground objects. AMSGL-Net contains two important components. First, the Edge Feature Enhancement Module (EFEM) leverages spatial attention mechanism to extract sharp, semantically rich edge cues of large-scale objects, which balances fusion quality against computational efficiency. Second, the Adaptive Multi-Scale Feature Extraction Module (AMSFEM) dynamically adjusts its receptive fields, enabling the network to capture subtle textural details for small instances with different scales, while preserving structural integrity for large ones. We conducted experiments on three public datasets and compared with state-of-the-art methods to demonstrate the effectiveness of AMSGL-Net.
Infrared and visible image fusion (IVIF) targets to integrate thermal saliency and rich textures into a single image that is not only visually appealing but also beneficial to downstream vision tasks. However, conventional methods relying on heuristic visual criteria struggle to guarantee task utility. Conversely, task-driven fusion paradigms typically employ fixed-weight scalarization, which suffers from potentially conflicting objectives among heterogeneous tasks, leading to rigid compromises and sub-optimal generalization in multi-task scenarios. To overcome these bottlenecks, this paper proposes a symbiotic evolutionary learning framework for task-adaptive IVIF, termed EvoFuse. Rather than relying on static loss weights, we formulate the fusion-perception correlation from a multi-objective perspective. To structurally instantiate this formulation, we first develop a re-parameterizable fusion architecture that accommodates multi-branch representational capacity during training, yet analytically folds into an ultra-compact single-branch model for efficient inference. To navigate the conflicting multi-task objectives, we introduce an evolutionary search mechanism that dynamically evolves task-aware loss-weight configurations. This enables a mutual adaptation process where the fusion network and task models are jointly updated under a Pareto-inspired non-degradation criterion. Furthermore, a novel saliency discriminative loss is designed to explicitly emphasize semantically crucial regions. Extensive experiments across eleven datasets covering fusion and downstream perception tasks demonstrate that the proposed method achieves competitive or better results in most evaluated metrics, while maintaining efficient inference under the considered task settings.
Reliable perception is a crucial requirement for autonomous driving, yet harsh weather conditions such as rain, fog and snow severely degrade the reliability of the system. While multimodal fusion of Radar, LiDAR, and cameras can enhance robustness, existing approaches typically rely on static fusion strategies that lack adaptability to varying weather environments. To address this limitation, we propose SWG-Fusion, a VLM-assisted multimodal fusion framework for Bird’s-Eye-View (BEV) object detection in harsh weather environments. Our work introduces a soft weather-adaptive guidance mechanism that leverages the Vision-Language Model (VLM) to extract semantic weather descriptions from visual inputs and embed them as adaptive guidance features for modality weighting. A BEV-alignment module is further applied to project perspective image features into the BEV space, enabling unified and spatially consistent fusion with LiDAR and Radar representations. Furthermore, a dual-stream fusion structure is developed to jointly refine modality-specific features and cross-modal interactions, thereby enhancing detection robustness. Extensive experiments on the real-world RADIATE dataset demonstrate that our approach achieves state-of-the-art performance across rain, fog, snow, and nighttime scenes. Ablation studies further show that each module provides a clear and consistent improvement, validating the effectiveness of the overall design for robust perception.
Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics, leading to noisy pseudo-masks and unstable optimization. To address this, we propose a hierarchical VFM-driven knowledge distillation framework that uses a frozen Vision Foundation Model (VFM) during training. We formulate point-supervised learning as a bilevel optimization process: the inner loop adapts a VFM-embedded teacher on reweighted training samples, while the outer loop transfers validation-guided knowledge to a lightweight student to mitigate pseudo-label noise and training-set bias. We further introduce Semantic-Conditioned Affine Modulation (SCAM) to inject VFM semantics into CNN features at multiple layers. In addition, a dynamic collaborative learning strategy with cluster-level sample reweighting enhances robustness to imperfect pseudo-masks. Experiments on diverse challenging cases across multiple ISTD backbones demonstrate consistent improvements in detection accuracy and training stability. Our code is available at https://github.com/yuanhang-yao/semantic-prior.
In endoscopic surgery, a clear and high-quality visual field is critical for surgeons to make accurate intraoperative decisions. However, persistent visual degradation, including smoke generated by energy devices, lens fogging from thermal gradients, and lens contamination due to blood or tissue fluid splashes during surgical procedures, severely impairs visual clarity. These degenerations can seriously hinder surgical workflow and pose risks to patient safety. To systematically investigate and address various forms of surgical scene degradation, we introduce a real-world open-source surgical image restoration dataset covering laparoscopic environments, called SurgClean, which involves multi-type image restoration tasks from two medical sites, i.e., desmoking, defogging, and desplashing. SurgClean comprises 3,113 images with diverse degradation types and corresponding paired reference labels. Based on SurgClean, we establish a standardized evaluation benchmark and provide performance for 22 representative generic task-specific image restoration approaches, including 12 generic and 10 task-specific image restoration approaches. Experimental results reveal substantial performance gaps relative to clinical requirements, highlighting a critical opportunity for algorithm advancements in intelligent surgical restoration. Furthermore, we explore the degradation discrepancies between surgical and natural scenes from structural perception and semantic understanding perspectives, providing fundamental insights for domain-specific image restoration research. Our work aims to empower restoration algorithms and improve the efficiency of clinical procedures. Data and code are available.
Infrared and visible image fusion is pivotal for robust visual perception across all weather conditions and scenes. Although deep learning-based methods have made notable progress, most either assume pre-aligned inputs or rely on implicit feature-space alignment, which fails to fundamentally address the amplification of registration errors and the loss of semantic structure in the fused results. To this end, we propose a universal representation and end-to-end framework for jointly registering and fusing unaligned infrared-visible image pairs, dubbed URMIF. Each image is mapped into modality-invariant (homogeneous) and modality-specific (heterogeneous) features: the invariant "structural skeleton" encodes geometry and semantics to stabilize alignment, while the specific "texture carrier" preserves thermal saliency and visible details to enable complementary fusion. Therefore, we propose a bi-directionally coupled registration-fusion module. This module performs hierarchical deformation estimation from coarse to fine, effectively mitigating visual mismatches caused by complex parallax in real-world scenes. Within this framework, the fusion component acts as the "evaluator" of registration, providing feedback regularization to update the deformation and suppress error accumulation. Furthermore, we introduce a dominant-plane prior as a scene-level constraint, seeding stable global and patch-wise homographies and reconciling cross-modal detail conflicts, to reinforce geometric consistency and semantic reliability. We also release a large-scale dataset comprising 1,500+ unaligned infrared/visible pairs with registration ground truth, spanning diverse illumination conditions and fields of view. Based on this dataset and additional benchmarks, extensive experiments validate that our framework achieves robust alignment and high-quality fusion on misaligned inputs, markedly reducing artifacts and improving the performance of downstream tasks such as detection and segmentation. Code and benchmark are available at https://github.com/ZengxiZhang/URMIF.