Satellite remote sensing plays a fundamental role in observing oceanic processes by providing large-scale, long-term, and continuous measurements. With the increasing availability of multisource satellite data, challenges such as data gaps, complex environmental conditions, and the limitations of conventional retrieval methods have become more evident. In recent years, artificial intelligence (AI) has emerged as a practical and effective approach to address these issues. This article reviews the development of AI techniques in satellite ocean remote sensing, focusing on three main application areas: parameter retrieval, data reconstruction, and image-based ocean phenomenon detection. For geophysical variable retrieval, AI models such as convolutional neural networks (CNNs) and Transformer architectures have improved the accuracy of ocean waves, sea surface, salinity, wind, and ocean color estimates, especially under extreme or noisy conditions. In the field of data reconstruction, AI methods enable the completion of missing data in both surface and subsurface ocean layers, offering finer spatial-temporal resolution and better consistency than traditional interpolation approaches. For image interpretation, deep learning (DL) models have been applied to detect and segment dynamic ocean features such as mesoscale eddies, internal waves, sea ice, and tropical cyclones (TCs), achieving high efficiency and precision. This article also highlights the integration of AI with physical knowledge, the use of multisource fusion, and the trend toward near real-time (NRT) applications. These developments indicate that AI will play an increasingly important role in future satellite-based ocean observation and environmental monitoring.
Accurate 6-DoF relative pose estimation is essential for multi-AUV cooperative tasks. However, pose-annotated underwater data are difficult to obtain, limiting learning-based methods. We present ZUPose, a monocular pose estimator trained entirely on synthetic data and deployed directly in real underwater scenes. A key challenge is prediction noise introduced by the sim-to-real gap and underwater optical degradation. To address it, we adopt an algorithm-data co-design strategy. At the algorithm level, we develop an uncertainty-guided densecorrespondence framework in which the network jointly predicts dense correspondences and per-pixel uncertainty under a Laplace-based probabilistic formulation. The predicted uncertainty acts as a learned scale parameter to model correspondence noise and guide pose optimization. At the data level, we construct a physics-guided simulation pipeline to model underwater optical degradation and generate diverse synthetic images. In real turbid water, ZUPose achieves translation and rotation errors of 6.7 cm and 7.7$^\circ$, with both reduced by about half compared with the best-performing baseline. The method remains stable under overexposure and long-range observation, and dual-AUV navigation experiments further validate its practical viability.
Accurate perception of a diver’s position and orientation by Autonomous Underwater Vehicles (AUVs) is essential for effective human–robot collaboration in underwater environments. However, conventional position and orientation estimation methods that combine deep learning with Perspective-n-Point (PnP) algorithms are primarily designed for rigid objects. In contrast, divers exhibit highly variable postures underwater, with no fixed configuration. To address this limitation, this paper proposes a framework for estimating the six-degree-of-freedom (6-DoF) position and the orientation of a diver. In addition, a novel network architecture, termed “VideoPose5CH,” is proposed. In the proposed framework, temporal sequences of 2D joint coordinates are provided to VideoPose5CH, which then outputs the 3D joint coordinates of the current frame as well as the corresponding refined 2D joint locations. Subsequently, the diver’s 6-DoF position and orientation relative to the camera are further recovered via a PnP algorithm. To mitigate the scarcity of underwater 3D human pose datasets, a land-based 3D human pose dataset augmentation strategy tailored to underwater conditions is further proposed. With this strategy, diver pose estimation accuracy is improved and the robustness of the proposed method across diverse scenarios is enhanced. Experimental results demonstrate that the proposed method can stably estimate the 6-DoF position and orientation of the diver within a distance range of 2.643 m to 11.477 m. The average position errors along the three axes are 7.33 cm, 4.04 cm, and 27.15 cm, respectively, while the average orientation errors are 6.96°, 8.47°, and 2.62°.
Due to the limited and fixed field of view of the onboard camera, the guiding beacons gradually drift out of sight as the AUV approaches the docking station, resulting in unreliable positioning and intermittent data. This paper proposes an underwater autonomous docking visual localization method based on a cage-type dual-layer guiding light array. To address the gradual loss of beacon visibility during AUV approach, a rationally designed localization scheme employing a cage-type, dual-layer guiding light array is presented. A dual-layer light array localization algorithm is introduced to accommodate varying beacon appearances at different docking stages by dynamically distinguishing between front and rear guiding light arrays. Following layer-wise separation of guiding lights, a robust tag-matching framework is constructed for each layer. Particle swarm optimization (PSO) is employed for high-precision initial tag matching, and a filtering strategy based on distance and angular ratio consistency eliminates unreliable matches. Under extreme conditions with three missing lights or two spurious beacons, the method achieves 90.3% and 99.6% matching success rates, respectively. After applying filtering strategy, error correction using backtracking extended Kalman filter (BTEKF) brings matching success rate to 99.9%. Simulations and underwater experiments demonstrate stable and robust tag matching across all docking phases, with average detection time of 0.112 s, even when handling dual-layer arrays. The proposed method achieves continuous visual guidance-based docking for autonomous AUV recovery.
The proliferation of saddle points, rather than poor local minima, is increasingly understood to be a primary obstacle in large-scale non-convex optimization for machine learning. Variable elimination algorithms, like Variable Projection (VarPro), have long been observed to exhibit superior convergence and robustness in practice, yet a principled understanding of why they so effectively navigate these complex energy landscapes has remained elusive. In this work, we provide a rigorous geometric explanation by comparing the optimization landscapes of the original and reduced formulations. Through a rigorous analysis based on Hessian inertia and the Schur complement, we prove that variable elimination fundamentally reshapes the critical point structure of the objective function, revealing that local maxima in the reduced landscape are created from, and correspond directly to, saddle points in the original formulation. Our findings are illustrated on the canonical problem of non-convex matrix factorization, visualized directly on two-parameter neural networks, and finally validated in training deep Residual Networks, where our approach yields dramatic improvements in stability and convergence to superior minima. This work goes beyond explaining an existing method; it establishes landscape simplification via saddle point transformation as a powerful principle that can guide the design of a new generation of more robust and efficient optimization algorithms.
This paper proposes a method for passive detection of autonomous underwater vehicle (AUV) wakes using a cilium-inspired wake sensor (CIWS), which can be used for the detection and tracking of AUVs. First, the characteristics of the CIWS and its working principle for detecting underwater flow fields are introduced. Then, a flow velocity sensor is used to measure the flow velocities of the “TS MINI” AUV’s wake at different positions, and a velocity field model of the “TS MINI” AUV’s wake is established. Finally, the wake field of the “TS MINI” AUV was measured at various positions using the CIWS. By analyzing the data, the characteristic frequency of the AUV’s propeller is identified, which is correlated with the AUV’s rotation speed and the number of blades. Through further analysis, a mapping model is established between the spectral amplitude of the characteristic frequency at different positions and the corresponding wake velocity. By substituting this mapping model into the AUV’s wake velocity field model, the possible position range of the sensor relative to the AUV propeller can be estimated. The research provides a technical foundation for underwater detection and tracking missions based on wake detection.
In this paper, a connecting joint capable of underwater autonomous docking and separation is proposed, which can be used for a reconfigurable articulated underwater robot (RAU robot). The structural design, optimization, and experimental validation of the connecting joint are presented in detail. First, the concept of the RAU robot is introduced, along with its different operational modes and the application scenarios. Second, the specific structural design and basic functions of the connecting joint are described. Third, a dynamic model of the docking process between different vehicles is established and simulated by kinematic simulation software. Through discretely sampling the parameter space, the optimal parameter combination is obtained. Finally, a prototype of the connecting joint is fabricated and functional tests are conducted. The impact forces on the docking rods before and after optimization are compared. The results show that the designed connecting joint can fulfill the functional requirements for autonomous docking of the underwater robot, and the maximum impact force is reduced by 27.08% compared to the one before optimization.
This study introduces a bionic vector wake sensor utilizing micro electro mechanical system (MEMS) technology to facilitate flow field sensing and the coordinated formation of autonomous underwater vehicles (AUVs). Inspired by the lateral line system of fish, the sensor incorporates a piezoresistive transducer featuring a bionic cilia and double-supported beam configuration. Drawing from the hierarchical fluid mechanism of the neural mound in blind cave fish, the sensor's design optimizes a composite package featuring a star-shaped package structure using finite-element method (FEM). Experimental results demonstrate the sensor's high sensitivity and ability to discern vectors. In multi-AUV platform formation tests, the sensor reliably detects wake turbulence from a leading target, showcasing its practicality and potential for engineering applications in complex flow environments.
Quantifying the sustainability of water-energy-food nexus system provides essential support for addressing climate change and resource scarcity. However, prior research has neglected to link the sustainability of waterenergy-food nexus system to the Sustainable Development Goals and the fact that the system itself is in an unsustainable state due to pressure overload. Based on the complex feedback relationships between subsystems, a sustainability assessment framework for water-energy-food nexus system is proposed from a pressure-support perspective by integrating the Sustainable Development Goal 2, 6 and 7. Entropy and catastrophe progression method, kernel density estimation, Markov chain and Tobit model are employed to examine its effectiveness by taking the Yellow River Basin in China as a case. The outcomes indicate that, the pressure index was greater than the support index for water-energy-food nexus system in the provinces along the basin except Shandong province from 2005 to 2021. The sustainable development index within the basin increased from 0.854 in 2005 to 0.917 in 2021, but was less than 1. It is characterized by fluctuating growth, with widening differences between regions. Water-energy-food nexus system was still in an unsustainable state. The ability of the lower reach to achieve sustainability far exceeded that of the upper and middle reaches. From spatial distribution, it showed a spatial pattern of "high in the east and low in the west", with spatial correlation effect showing "club convergence" characteristic. Foreign trade level, technological progress level, environmental regulation capacity, industrialization level, infrastructure level and rural living standard made significant contributions to improving sustainability level with marginal impacts of 0.121, 0.328, 0.604, 0.233, 0.674 and 0.083, respectively, while urbanization level had an inhibitory effect with marginal impact of -0.255. These findings provide policy guidance for the region to enhance sustainable governance of water-energy-food nexus system.
A brand new non-penetrating tunnel thruster (short for NPT thruster) is proposed in this paper. The tunnel structural parameters of the thruster are optimized, and the performance and optimization effect are verified by experiments. First, the design and function of the NPT thruster are introduced. Second, the computational fluid dynamics method is used to calculate the hydrodynamic performance of the NPT thruster and to analyze the static mooring thrust performance. Third, the tunnel structural parameters of the NPT thruster are optimized with the method of the response surface methodology. The pressure distributions and the flow fields on the tunnel surface of the NPT thrusters before and after optimization are compared with simulations. Finally, the mooring static thrust of the NPT thrusters is tested with experiments. The results show that the average increase in the mooring static thrust for the optimized thruster is 12.4%, and the maximum increase can reach 21.79% when the rotational speed is from 3000 rpm to 6500 rpm.
Accurately modeling the system dynamics of autonomous underwater vehicles (AUVs) is imperative to facilitating the implementation of intelligent control. In this research, we introduce a physics-informed neural network (PINN) method to model the dynamics of AUVs by integrating dynamical equations with deep neural networks. This integration leverages the nonlinear expressive power of deep neural networks, alongside the robust foundation of physical prior knowledge, resulting in an AUV model proficient in long-term motion forecasting. The experimental results indicate that this method is capable of effectively extracting AUV system dynamics from datasets, exhibiting strong generalization capabilities and achieving robust long-term motion prediction. Furthermore, a model predictive control method is proposed, using the learned PINN as the predictive model to accurately track the closed-loop trajectory. This research offers novel perspectives on the dynamics modeling of AUVs and has the potential to be applied in other relevant research endeavors.
In recent years, the detection and localization of tiny persons have garnered significant attention due to their critical applications in various surveillance and security scenarios. Traditional multi-modal methods predominantly rely on well-registered image pairs, necessitating the use of sophisticated sensors and extensive manual effort for registration, which restricts their practical utility in dynamic, real-world environments. Addressing this gap, this paper introduces a novel non-registered multi-modal benchmark named NRPerson, specifically designed to advance the field of tiny person detection and localization by accommodating the complexities of real-world scenarios. The NRPerson dataset comprises 8548 RGB-IR image pairs, meticulously collected and filtered from 22 video sequences, enriched with 889,207 high-quality annotations that have been manually verified for accuracy. Utilizing NRPerson, we evaluate several leading detection and localization models across both mono-modal and non-registered multi-modal frameworks. Furthermore, we develop a comprehensive set of natural multi-modal baselines for the innovative non-registered track, aiming to enhance the detection and localization of unregistered multi-modal data using a cohesive and generalized approach. This benchmark is poised to facilitate significant strides in the practical deployment of detection and localization technologies by mitigating the reliance on stringent registration requirements.
The capabilities of AUV mutual perception and localization are crucial for the development of AUV swarm systems. We propose the AUV6D model, a synthetic image-based approach to enhance inter-AUV perception through 6D pose estimation. Due to the challenge of acquiring accurate 6D pose data, a dataset of simulated underwater images with precise pose labels was generated using Unity3D. Mask-CycleGAN technology was introduced to transform these simulated images into realistic synthetic images, addressing the scarcity of available underwater data. Furthermore, the Color Intermediate Domain Mapping strategy is proposed to ensure alignment across different image styles at pixel and feature levels, enhancing the adaptability of the pose estimation model. Additionally, the Salient Keypoint Vector Voting Mechanism was developed to improve the accuracy and robustness of underwater pose estimation, enabling precise localization even in the presence of occlusions. The experimental results demonstrated that our AUV6D model achieved millimeter-level localization precision and pose estimation errors within five degrees, showing exceptional performance in complex underwater environments. Navigation experiments with two AUVs further verified the model’s reliability for mutual 6D pose estimation. This research provides substantial technical support for more complex and precise collaborative operations for AUV swarms in the future.
Based on the application requirements of aquatic-aerial amphibious propellers for the clutch, a gear overrunning clutch was designed. The engagement characteristics of clutch under different structure parameters were optimized for analysis through simulation. The effects of end face chamfer and displacement coefficient of gear tooth on engagement processes, disengagement speed and contact impact force were discussed. The simulation results show that increase of the angle of end face chamfer and displacement coefficient may accelerate engagement and disengagement speed and reduce contact impact force. When the number of chamfers and the coefficient of displacement are as 44° and 0.84 respectively, the engagement characteristics are optimal. If they continue to increase, the engagement and disengagement speed becomes slower and the contact impact force becomes larger. Finally, the gear tooth structure parameters with the optimal engagement characteristic index are obtained, and the correctness and validity of the simulation analysis results were verified by experiments. And the relative errors of the engagement characteristic index of simulation and experiments are less than 10%. The effectiveness of the optimized design process is proven, and the design results meet the requirements for the use of the propellers.
The autonomous underwater vehicle(AUV) pose dataset is difficult to obtain in underwater scenarios.In addition,the existing deep learning-based pose estimation methods cannot be applied in this scenario.Thus,this paper proposes an AUV visual localization method based on synthetic data.In this method,we first build a virtual underwater scene by Unity3D and obtain the rendering data of the known pose through the virtual camera.Then,we realize the style transfer of the rendered image to the real underwater scene through the unpaired image translation work.We also obtain the synthetic underwater pose dataset by combining the pose information of the known rendered image.Finally,we propose a convolutional neural network(CNN) pose estimation method based on local region keypoint projections.The CNN is trained using synthetic data to predict 2D projections of known reference corners.The resulting 2D-3D point pairs obtain the relative positions and pose through the Perspective-n-Point algorithm that is based on random sample consensus.The effectiveness of the proposed method is examined using quantitative experiments on rendered datasets and synthetic datasets,as well as qualitative experiments on real underwater scenes.Our experimental results show that the unpaired image translation can effectively eliminate the gap between the rendered image and the real underwater image.We also find that the proposed local area keypoint projection method can perform more effective 6D pose estimation.
Sound source localization provides an absorbing capability for unmanned aerial vehicles (UAVs) in scenarios, such as search and rescue operations. The shape fusion between the sound array and UAVs forms a special conformal property that is drawing more and more attention. However, the inevitable shadow effect caused by shape fusion seriously degrades the degrees of freedom (DOF) of the array. In this article, a signal reconstruction-based direction of arrival (DOA) estimation method is proposed to address this limitation. First, we establish a restricted signal model for the conformal cylindrical array (CCA), and then based on frequency domain energy detection, the elements are divided into receiving restricted elements and receiving normal elements. Second, according to the position vector of receiving restricted elements, the approximate range of the DOA is roughly estimated to reduce the complexity. Meanwhile, the signals of receiving restricted elements are reconstructed on the basis of receiving normal elements to eliminate the shadow effect and increase the DOF. Finally, in the estimated approximate range of the DOA, the precise DOA is estimated by peak search. We also derive the 2-D Cramer–Rao lower bound (CRLB) for the CCA. Simulations show that the proposed SR-MUSIC-RS method can achieve satisfactory performance with lower complexity, and the root mean squared error is close to that of general signal model with normal elements.
In challenging tasks such as large-scale resource detection, deep-sea exploration, prolonged cruising, extensive topographical mapping, and operations within intricate current regions, AUV swarm technologies play a pivotal role. A core technical challenge within this realm is the precise determination of relative positions among AUVs within the cluster. Given the complexity of underwater environments, this study introduces an integrated and high-precision underwater cluster positioning method, incorporating advanced image restoration algorithms and enhanced underwater visual markers. Utilizing the Hydro-Optical Image Restoration Model (HOIRM) developed in this research, image clarity in underwater settings is significantly improved, thereby expanding the attenuation coefficient range for marker identification and enhancing it by at least 20%. Compared to other markers, the novel underwater visual marker designed in this research elevates positioning accuracy by 1.5 times under optimal water conditions and twice as much under adverse conditions. By synthesizing the aforementioned techniques, this study has successfully developed a comprehensive underwater visual positioning algorithm, amalgamating image restoration, feature detection, geometric code value analysis, and pose resolution. The efficacy of the method has been validated through real-world underwater swarm experiments, providing crucial navigational and operational assurance for AUV clusters.
Hindwing venation is one of the most important morphological features for the functional and evolutionary analysis of beetles, as it is one of the key features used for the analysis of beetle flight performance and the design of beetle-like flapping wing micro aerial vehicles. However, manual landmark annotation for hindwing morphological analysis is a time-consuming process hindering the development of wing morphology research. In this paper, we present a novel approach for the detection of landmarks on the hindwings of leaf beetles (Coleoptera, Chrysomelidae) using a limited number of samples. The proposed method entails the transfer of a pre-existing model, trained on a large natural image dataset, to the specific domain of leaf beetle hindwings. This is achieved by using a deep high-resolution network as the backbone. The low-stage network parameters are frozen, while the high-stage parameters are re-trained to construct a leaf beetle hindwing landmark detection model. A leaf beetle hindwing landmark dataset was constructed, and the network was trained on varying numbers of randomly selected hindwing samples. The results demonstrate that the average detection normalized mean error for specific landmarks of leaf beetle hindwings (100 samples) remains below 0.02 and only reached 0.045 when using a mere three samples for training. Comparative analyses reveal that the proposed approach out-performs a prevalently used method (i.e., a deep residual network). This study showcases the practicability of employing natural images-specifically, those in ImageNet-for the purpose of pre-training leaf beetle hindwing landmark detection models in particular, providing a promising approach for insect wing venation digitization.
Water–air cross-domain vehicles (CDVs) are capable of both flight and underwater navigation, showing broad prospects in marine science, such as underwater observation, disaster response, and rescue operations. It is crucial to investigate the dynamic performance of CDVs hovering above water surfaces to enhance safety and stability. In this study, the performance of a CDV’s ducted propeller hovering at various heights above a water surface was analyzed via computational fluid dynamic (CFD) simulations using the lattice Boltzmann method (LBM) and thrust tests. The results indicate that the air–water mixture formed by the wake of the propeller impacting the water surface is sucked in by the duct, causing the propeller to enter an unstable vortex ring state. At the same rotation speed in the air, the thrust of the propeller system decreases and the required power increases. With an increase in the height of the propeller above the water surface, the thrust and power return to normal. Furthermore, a numerical model was proposed to express the correlation among thrust, propeller rotation speed, and distance from the water surface. This study establishes a foundation for the dynamic modeling of CDVs and can be utilized by other related studies.
水下机器人集群技术是目前水下机器人技术领域的发展热点之一.针对以往基于通信的水下机器人编队存在的编队精度低、队形保持困难等问题,提出了基于水下矢量光图案及视觉定位的水下集群编队方法,并通过水池试验分别验证了视觉定位以及水下密集编队的功能和相关指标.试验结果表明:水下视觉定位能够达到不大于3%的定位精度和不小于2 Hz的定位频率,能够为水下机器人自主航行提供准确连续的控制输入.基于视觉定位的水下集群能够实现相互距离10 m以内的密集编队,编队精度不大于10%,较以往基于通信的编队方法有较大提升.