Underwater image enhancement (UIE) is a complex non-linear inverse problem. Traditional models rely on linear operators, which fail to fuse multi-domain degradation features effectively. This paper proposes FreqKAN, a novel dual-domain framework centered on Information Fusion principles to bridge frequency and spatial representations. We design the KANs-driven Dual-domain Collaborative Block (KDCB), which decouples global luminance/color (frequency-domain) from local structural details (spatial-domain). By leveraging the non-linear approximation of Kolmogorov-Arnold Networks (KANs), it achieves a more precise reconstruction of the amplitude spectrum than traditional linear mappings. To address the instability of underwater environments, we propose the Adaptive Fusion Block (AFB). Unlike static concatenation, AFB employs a KAN-based gating mechanism to dynamically harmonize complementary information, ensuring a synergistic balance between global color naturalness and local edge fidelity. Extensive experiments on UIEB and UCCS datasets indicate that FreqKAN significantly outperforms state-of-the-art UIE methods, with a 1.15 dB PSNR gain over the second-best method UIE-UnFold on the UIEB-T90 dataset, effectively enhancing the robustness of underwater vision systems.
Maritime optical imaging systems are critical for environment perception and the execution of key maritime operations. However, optical imaging devices always capture degraded images in haze scenes, such as poor visibility, lost detail, and distorted color. To effectively enhance the imaging quality, we propose a feature modulation image dehazing model (FMID), designed to restore the reliability of visual information in haze scenes. The core of our model is the dual attention residual block (DARB) that comprises a feature enhancement channel attention (FECA) module and a fusion complementary spatial attention (FCSA) module. FECA employs multi-scale convolutions to learn dynamic channel weights, effectively aggregating contextual information. FCSA integrates multi-scale features to obtain spatial attention weights and extracts complementary information along horizontal and vertical directions to guide the aggregation, thereby accurately reconstructing structural detail. Extensive experiments on standard and maritime-related datasets demonstrate that FMID can outperform several state-of-the-art methods, enhance imaging quality, and effectively improve the performance of downstream vision-based maritime tasks, including vessel detection and maritime scene segmentation, providing robust technical guarantee for maritime traffic and intelligent port surveillance.
Water-related optical image enhancement (WOIE) poses a significant challenge due to the complex and variable underwater environment. Existing methods often oversimplify the underwater imaging degradation process, neglecting the effects of medium noise and target motion on image feature distribution. Additionally, reliance on reference gradients from original and synthesized ground-truth images leads networks to local optima. To address these limitations, we propose a dynamic gradient-guided network (DGNet) for WOIE. DGNet introduces a feature reconstruction and re-learning (FRR) module based on a channel combination inference (CCI) strategy and a frequency response smoothing (FRS) module. These components reduce the impact of noise and target motion on fine-grained features. During training, DGNet dynamically updates pseudo-labels with predicted images and introduces a dynamic gradient optimization space guided by white balance, enabling the network to escape saddle points and explore a broader solution space. Experiments on multiple public datasets show that DGNet outperforms the latest methods, achieving a PSNR of 25.6 dB and an SSIM of 0.929 on the UIEB dataset. In addition, extensive qualitative experiments have demonstrated that DGNet exhibits precise control over fine-grained features, highlighting its practical utility. The code will be publicly available.
Underwater imaging faces fundamental challenges due to wavelength-dependent light absorption and scattering, which manifest as severe color distortion, contrast attenuation, and the obscuration of high-frequency details. These degradations not only compromise visual fidelity but also impede downstream machine perception tasks. While sequence modeling paradigms demonstrate robust global dependency capturing, their application to image restoration is constrained by the structural mismatch between 1D sequential scanning and 2D spatial locality. To address this, we propose DinoMamba, a novel framework that synergizes the semantic discrimination of DINOv2 with the efficient long-range modeling of Mamba. Specifically, we introduce a semantic-driven token reordering strategy to bridge the spatio-semantic gap. By fusing high-level object priors with visual features, this mechanism actively prioritizes information-rich regions, where foreground semantics outweigh the homogeneous background, thereby optimizing the computational allocation of the Mamba module. This design is further motivated by the physical characteristics of underwater imaging. Due to wavelength-dependent light absorption and scattering, background regions are often severely degraded, exhibiting low contrast and limited structural information, while foreground objects tend to preserve more reliable semantic cues. Leveraging DINOv2 to identify such object-centric regions and guiding the Mamba sequence modeling process to prioritize them allows more effective allocation of modeling capacity toward semantically meaningful areas. Extensive experiments demonstrate that DinoMamba achieves superior performance in both signal-level restoration and downstream perception tasks, validating the efficacy of fusing foundation model priors with state-space sequence modeling.
Underwater imagery is degraded by depth-dependent absorption and scattering, which often introduce color casts and contrast attenuation. Although recent Vision Mamba models provide efficient long-range dependency modeling, their conventional 2D scanning patterns are not explicitly designed to exploit the depth-correlated structure of underwater degradation and may therefore weaken geometry-aware feature dependencies. To address this limitation, we propose Isoline-Guided Evolutionary Mamba (IG-Mamba), a physics-inspired framework that uses a depth-correlated potential prior to organize state-space token propagation. Specifically, we introduce a Topology-Preserving Isoline Scanning mechanism. By leveraging a geometric prior, this mechanism quantizes the scene into discrete iso-potential strata to guide Mamba sequences along geometry-aware orders, thereby preserving local spatial topology while establishing depth-aware long-range dependencies. Furthermore, a Potential Field Evolution method is developed to mitigate the domain discrepancy between terrestrial geometric priors and underwater optical attenuation. Driven by the reconstruction objective, the network adaptively refines the raw geometric anchor into a restoration-oriented optical potential field via a learned residual map. Finally, features are modulated by a transmission-inspired gate motivated by the Jaffe-McGlamery transmission term, enabling spatially adaptive feature reweighting under scattering-dominant conditions. Extensive experiments demonstrate that IG-Mamba achieves strong performance across multiple benchmarks, offering a physics-grounded perspective for dependency modeling in underwater vision.
Novel view synthesis (NVS) under sparse-view settings remains challenging due to incomplete observations, which often lead to overfitting and the loss of fine geometric structures in reconstructed scenes. In this work, we present a Fine-Structure-Aware Gaussian Splatting (FSA-GS) framework that optimizes the distribution of Gaussian primitives for detail-preserving novel view rendering, achieving clearer structural boundaries and more faithful local details. Our framework introduces two key strategies, Fine-Structure-Aware Resampling (FSAR) and Fine-Structure-Aware Splitting (FSAS), which use geometric cues extracted from multi-view images to guide primitive placement in high-frequency regions. To further stabilize training and maintain a compact yet expressive representation, we incorporate a Global and Local Opacity- and Density-driven Dropout (GLOD-D) mechanism that adaptively regulates the stochastic dropout of Gaussian primitives throughout optimization. Extensive experiments on the LLFF, Mip-NeRF360, and Blender datasets show that our method surpasses state-of-the-art approaches in both quantitative accuracy and visual quality.
Underwater images are often degraded in terms of color fidelity, contrast, and structural details as a result of complex light propagation characteristics in aquatic environments. Existing self-attention-based enhancement methods either adopt coarse global attention that fails to capture fine-grained textures/edges, or window-based strategies that sacrifice long-range modeling for local detail enhancement and computational efficiency, leading to a global-local dependency imbalance. To break this bottleneck for underwater 3D reconstruction, we propose a Frequency-Adaptive Transformer (FAFormer) for front-end image enhancement. FAFormer leverages wavelet decoupling to construct frequency-domain feature representations, splitting feature maps into low-frequency global structural and high-frequency local detail components. It then models these components via channel-wise self-attention and local windowed spatial modeling collaboratively, mitigating color distortion while boosting texture and edge expression. Experiments on two visual task datasets show that FAFormer has a certain degree of cross-scenario performance, while its use as a front end for 3D reconstruction can improve multi-view matching stability in severely degraded underwater scenes and help preserve geometric integrity and surface consistency.
Underwater object detection is crucial for marine ecological monitoring and resource exploration. However, low underwater contrast causes severe aliasing between foreground and background in the spatial domain, and conventional methods struggle to effectively decouple their features. Transformer architectures incur high computational overhead. Conversely, over-compressing these models severely degrades their ability to detect heavily camouflaged or small marine organisms. High-quality underwater samples are limited, and existing time-consuming generative strategies are hard to implement efficiently on edge devices. To address these challenges, this paper proposes a resource-efficient framework for underwater object detection. First, we design a Wavelet-Enhanced Feature Pyramid Network that combines a saliency-focus mechanism and a discrete wavelet transform to overcome background noise in both spatial and frequency domains, extracting features of hidden small objects. Second, a data-dependent dynamic token pruning technique removes redundant tokens, effectively mitigating the computational bottleneck without sacrificing essential semantic capacity. Finally, for extreme sample scarcity, we introduce a Feature Correction Module and a two-stage fine-tuning and feature correction strategy, using a high-precision teacher model to guide a compressed student network in adaptively compensating for optical shifts with few samples. Experiments on URPC2020 and DUO demonstrate that our method improves small object detection accuracy while reducing parameter count and computational overhead, striking a good balance between accuracy and inference efficiency.
Underwater Object Detection (UOD) faces significant challenges due to complex degradation factors, such as color shifts caused by light absorption and scattering, spatially varying noise induced by plankton and sea snow, and motion blur resulting from dynamic water currents. Among existing methods, Convolutional Neural Networks (CNNs) are limited by fixed receptive fields, making it difficult to model long-range noise patterns; while Transformers excel at modeling global dependencies, they suffer from high computational complexity and weak capability in restoring fine-grained local features. Neither can effectively address the demands of detecting underwater-specific noise and small objects. To tackle these issues, we propose UOD-Mamba, a state space model (SSM)-based framework for underwater object detection. At its core is the Noise-Aware Dual-path Mamba (NADM) module, which integrates a global-local dual-path fusion strategy to enable both long-range noise modeling and local feature enhancement. The global path balances noise in input features through the Noise-Balanced Preprocessing Module (NBPM) and leverages Mamba's longrange modeling capability to extract global noise patterns; the local path fuses the Underwater Enhanced Multi-scale Attention Module (UEMA) with CSP convolution to model edge and detail features at a finegrained level, thereby compensating for the loss of local information. By explicitly learning the distribution characteristics of underwater noise and capturing the differences between noise and target features, the framework enhances detection robustness in noisy environments. Experimental validation on the DUO and RUOD datasets demonstrates that UOD-Mamba sets a new state-of-the-art in detection performance. It also exhibits advantages in explicit modeling of diverse noises, preservation of local details, and computational efficiency across multi-noise scenarios, enabling effective handling of complex underwater interference environments.
Color distortion and structural degradation in underwater images are classic challenges in underwater image enhancement. The core goal is to restore degraded images to high-quality images with both color and structure that conform to visual perception. However, in the traditional RGB space, these two issues are highly coupled, resulting in existing enhancement methods often neglecting one over the other. To address this challenge, we propose a guided diffusion model based on the principle of decoupling. Our key insight is that in perceptual color spaces such as HSV, color (H, S) and structure (V) are naturally separated. To exploit this property, we first design an adaptive perceptual guidance module, which analyzes the degraded HSV image and generates two orthogonal guidance signals: a color guide and a structure guide, which guide the denoising process of the diffusion model. To ensure that this decoupled guidance is faithfully implemented, we propose a corresponding decoupled loss optimization module, which uses independent loss functions to supervise the final output color and structure. By combining the forward decoupled guidance with the backward decoupled supervision, we construct a closed-loop optimization framework. This framework enables the model to collaboratively optimize color and structure under various degradation scenarios. Extensive experiments demonstrate that our proposed method outperforms existing state-of-the-art approaches in a variety of underwater scenes, particularly those degraded by color casts and haze. Furthermore, it exhibits superior performance on no-reference image quality assessment metrics. The source code is available at https://github.com/zy-world/DCD-UIE.
Underwater noise-coupled degradation blurs details and weakens high-frequency response transfer in the network. We address this by retaining contour-detail priors under noisy degradation to stabilize high-frequency flow. We first revisit identity mapping in residual connections and introduce Adaptive High-Frequency Injection (AHFI) to reconstruct the residual path and validate explicit gradient compensation. On this basis, we propose the CDP framework, incorporating a Contour-Detail Prior Bottleneck (CDPB) to capture contour/detail-biased structural responses before fusion, while residual-path reconstruction stabilizes high-frequency gradient propagation. We further introduce a Contour-Detail Prior Downsampler (CDPD) into the neck for structural sensing and prior calibration before downsampling to stabilize contour-related high-frequency transfer. Experiments on the RUOD dataset show that CDP-YOLO(s) gains 0.5 AP over the baseline while reducing params and GFLOPs by 9.2% and 7.7%. It also reaches an earlier best-checkpoint epoch. These results show a useful cost-performance trade-off for resource-limited underwater robots.
Underwater image enhancement is a highly challenging task, requiring solutions to complex environmental degradation factors such as light attenuation and color cast. Achieving stability in color restoration and precision in texture recovery is key to improving enhancement results. However, existing methods generally lack in-depth modeling of color and texture information and fail to efficiently fuse these two core visual components, significantly limiting the overall performance of the enhancement results. To this end, we propose an innovative Dual-Attention Fusion Net (DuAF) that solves this problem. On a global scale, DuAF introduces explicit semantic consistency constraints to precisely model color features by reconstructing pixel intensity distribution, enhancing sensitivity to color features, and capturing real pixel gradient changes, effectively addressing complex color distortion issues. On a local scale, DuAF dynamically adjusts the perception window, combines optimized attention weights with positional deviations, and deeply models texture information, significantly improving the restoration of texture details. Overall, DuAF significantly improves the stability of color restoration and the clarity of texture details in complex degraded scenes, providing an efficient and comprehensive solution for underwater image enhancement. Our project is publicly available on https://github.com/HuShuteng/DuAF.
Reliable perception is a prerequisite for multi-source information fusion in autonomous marine systems. However, underwater image enhancement (UIE) remains a bottleneck due to complex, coupled degradations that compromise sensor reliability in dynamic ocean conditions. Most existing methods explicitly model individual degradation components, often leading to detail loss and color distortion in autonomous navigation tasks. To address these challenges, we propose DENIR, a degradation-codebook-aware framework providing stable visual inputs for robust perception. Unlike conventional approaches, DENIR treats environmental degradation as a learnable latent factor via a discrete negative codebook. This allows the system to enhance images by repelling degradation-aligned features, effectively suppressing sensor noise. Specifically, we introduce a Degradation-aware Contrastive Repulsion (DCR) loss for latent feature purification and a Negative Codebook Prompt (NegPrompt) module for adaptive enhancement under complex, mixed degradations. Extensive experiments demonstrate that DENIR achieves consistent improvements in pixel fidelity and perceptual quality. Crucially, task-oriented evaluations in segmentation and depth estimation confirm that DENIR significantly enhances data usability for downstream autonomous perception and multi-sensor fusion pipelines.
Light absorption and scattering in water cause underwater images to suffer from severe degradations, including color distortion, reduced contrast, and detail loss. Unlike local texture or structural degradation, these issues are manifested as global color distribution shifts and dynamic range compression. Existing convolution-or attention-based methods primarily focus on spatial dependency modeling, but they fail to capture and constrain such global statistical properties. To address these challenges, we propose HisMamba, a novel underwater im age enhancement network that performs collaborative structural-color modeling. Specifically, we design the SoftHistogram-DHSA module, which employs Gaussian kernel approximation to construct differentiable his tograms, explicitly incorporating global color statistics into end-to-end training. By treating statistical fluctuations as a form of positive noise, the module couples with an attention mechanism to adaptively mitigate color shifts and dynamic range compression. Furthermore, we introduce the Directional Frequency Mamba (DF-Mamba), which models long-range dependencies along horizontal and vertical directions via a row-column separation strategy, while employing frequency-guided modulation to preserve spatial structures and enhance semantic representation. In addition, a lightweight Gated Multi-Scale Fusion (GMSF) module is designed to adaptively in tegrate contextual and fine-grained details with low computational overhead. Extensive experiments on multiple mainstream datasets demonstrate that HisMamba consistently surpasses state-of-the-art methods, achieving best or second-best performance across a wide range of metrics. Overall, the proposed structural-color collaborative modeling paradigm provides a new solution framework for advancing underwater visual perception.
Water-related images are crucial for applications such as marine exploration, underwater robotics, and ecological monitoring. However, due to wavelength-dependent absorption and scattering, these images often suffer from color distortion, low contrast, and structural blurring, impairing downstream visual tasks. While diffusion models offer strong reconstruction capabilities through progressive denoising, their reliance on paired supervision limits their applicability in underwater scenarios. To address this, we propose a pseudo-label guided diffusion framework for water-related image enhancement, enabling high-quality restoration without real ground-truth labels. Specifically, we design a U-Net-based diffusion network with a linear time-step scheduling strategy to progressively recover global structure and local texture. To compensate for the lack of supervision, we introduce a semantic-visual pseudo-labeling mechanism: semantic prompts generated by the large multimodal model LLaVA are matched with multiple enhanced images using Zip-CLIP to produce semantic similarity heatmaps. A lightweight G-CNN fuses these cues into pseudo-labels that guide the diffusion process. Extensive experiments on UIEB, U45, and UCCS datasets demonstrate that our method consistently surpasses existing approaches in terms of sharpness, color fidelity, and semantic structure preservation. These results highlight the effectiveness of integrating multimodal pseudo-supervision with diffusion modeling for robust underwater image enhancement. The code is available at https://github.com/Liujiatong-111/Semantic-Guided-Diffusion-for-Water-Related-Image-Enhancement.
Underwater image quality assessment (UIQA) is hindered by complex degradation and domain shifts across aquatic environments. Existing no-reference IQA methods rely on costly and subjective mean opinion scores (MOS), which limit their generalization to unseen domains. To overcome these challenges, we propose SCUIA, an unsupervised UIQA framework leveraging semantic contrastive learning for quality prediction without human annotations. Specifically, we introduce a vision-language contrastive learning strategy that aligns image features with textual embeddings in a unified semantic space, capturing implicit degradation-quality correlations. We further enhance quality discrimination with a hierarchical contrastive learning mechanism that combines image-specific statistical priors and semantic prompts. A triplet-based inter-group contrastive loss explicitly models relative quality relationships. To tackle cross-domain variations, we develop an unsupervised domain adaptation module that uses local statistical features to guide CLIP fine-tuning to disentangle domain-invariant quality representations from domain-specific noise. This enables zero-shot cross-domain quality prediction without labeled data. Extensive experiments on public UIQA benchmarks demonstrate significant improvements over existing methods, highlighting superior generalization and domain adaptability.
Recent advances in learning-based underwater image enhancement have achieved remarkable progress. However, the inherent diversity and complexity of underwater scenes still limit the ability of existing approaches to simultaneously restore fine structural details and global image layouts. To address this challenge, we propose a Resonant Fusion (ReFu) framework that explicitly leverages complementary information in both spatial and frequency domains. Specifically, we design a frequency decomposer and a spatial decomposer to capture high- and low-frequency cues from different perspectives. A resonant fuser is then introduced to adaptively integrate high-frequency resonances for detail refinement and low-frequency resonances for structural consistency. This fine-grained cross-domain fusion significantly improves structural preservation and detail enhancement, thereby generating visually more natural and perceptually friendly underwater images. Extensive quantitative and qualitative evaluations across diverse underwater benchmarks show that ReFu consistently surpasses state-of-the-art methods by a clear margin. Comprehensive ablation studies further validate the effectiveness of each module and prove the necessity of the proposed ReFu mechanism. Our code is available at https://github.com/CircleQa/ReFu-main.
The 3D reconstruction of vessel hulls is crucial for enhancing safety, efficiency, and knowledge in the maritime industry. Neural Radiance Fields (NeRFs) are an alternative to 3D reconstruction and rendering from multi-view images; particularly, tensor-based methods have proven effective in improving efficiency. However, existing tensor-based methods typically suffer from a lack of spatial coherence, resulting in gaps in the reconstruction of fine-grained geometric structures. This paper proposes a spatial multi-scale weighted NeRF (MDW-NeRF) for accurate and efficient surface reconstruction of vessel hulls. The proposed method develops a novel multi-scale feature decomposition mechanism that models 3D space by leveraging multi-resolution features, facilitating the integration of high-resolution details with low-resolution regional information. We designed separate color and density weighting, using a coarse-to-fine strategy, for density and a weighted matrix for color to decouple feature vectors from appearance attributes. To boost the efficiency of 3D reconstruction and rendering, we implement a hybrid sampling point strategy for volume rendering, selecting sample points based on volumetric density. Extensive experiments on the SVH dataset confirm MDW-NeRF’s superiority: quantitatively, it outperforms TensoRF by 1.5 dB in PSNR and 6.1% in CD, and shrinks the model size by 9%, with comparable training times; qualitatively, it resolves tensor-based methods’ inherent spatial incoherence and fine-grained gaps, enabling accurate restoration of hull cavities and realistic surface texture rendering. These results validate our method’s effectiveness in achieving excellent rendering quality, high reconstruction accuracy, and timeliness.
Underwater environments present significant challenges for the spatial perception of unmanned systems, particularly in tasks such as 3D mapping and navigation, due to the effects of scattering and attenuation. Existing underwater scene restoration methods often lack robust physical priors and comprehensive rendering models, limiting their applicability in large-scale or dynamic environments. To address these challenges, we propose Sea-out NeRF, a novel framework that integrates underwater physical imaging principles into Neural Radiance Fields (NeRF) to enhance the spatial perception capabilities of underwater unmanned systems. Sea-out NeRF simultaneously generates both underwater and clear-air views, thereby extending the perceptual range and facilitating navigation and decision-making in complex aquatic environments. By incorporating physically motivated priors into the volumetric rendering process, the proposed method achieves high-quality scene restoration and supports the flexible synthesis of diverse underwater environments. Extensive experiments on both real-world and synthetic datasets demonstrate the effectiveness of Sea-out NeRF and its significant potential for improving the spatial perception of underwater unmanned systems. The source code is available at https://github.com/Richard-Underwater/Sea-out-NeRF.