Non-repetitive scanning LiDARs provide high coverage yet exhibit irregular sampling patterns, which destabilize local features and correspondences. To address this, we propose a novel spatiotemporal unified energy framework that integrates fractional calculus into rigid pose estimation. Spatially, we introduce a Riesz fractional regularization term to impose non-local smoothness constraints on the residual field, mitigating structural inconsistencies. Temporally, we design a Grünwald–Letnikov fractional dynamics solver that leverages long-memory effects of historical gradients to reduce the risk of being trapped in local minima. Comparative experiments on the Stanford 3D, MVTec ITODD, and HomebrewedDB (HB) datasets demonstrate that our method significantly outperforms state-of-the-art geometric and learning-based approaches. Specifically, it maintains a success rate exceeding 90% even under severe sampling perturbations where traditional methods fail. Ablation studies further validate that the introduction of non-local spatial constraints and historical gradient memory significantly reshapes the energy landscape, ensuring robust convergence. This work provides a rigorous theoretical foundation for applying fractional operators to point cloud processing.
ObjectiveConventional multimodal alignment and fusion methods struggle to adapt to complex underwater detection environments. Owing to the divergences in imaging mechanisms between acoustic sonar and optical cameras, cross-modal features are often spatially misaligned, resulting in degraded fusion performance. This issue significantly limits the accuracy of underwater target detection in turbulent and turbid marine environments. To address these challenges, this study proposes an acoustic-optical multimodal fusion detection architecture for underwater perception tasks. MethodAn end-to-end five-layer detection framework is developed, comprising an input layer, a feature extraction layer, a multimodal fusion layer, a target perception layer, and output layers. Two independent ResNet50 branches are employed to extract multi-scale feature representations from sonar and optical images, respectively. A novel spatial alignment module is designed to estimate affine transformation parameters, including scaling and translation factors, enabling pixel-level spatial registration between modalities. Integrated with channel and spatial attention mechanisms, the dynamic weighted fusion module adaptively adjusts the contribution of each modality, thereby suppressing low-quality and noisy features. Furthermore, a hierarchical interactive fusion encoder incorporating DenseNet and a cross-attention mechanism is constructed to achieve deep complementary fusion of multi-scale cross-modal features. The Transformer decoder and Hungarian matching loss inherited from DETR are utilized for end-to-end target classification and bounding-box regression, eliminating the need for additional non-maximum suppression operation. ResultsComparative and ablation experiments are conducted on a self-constructed real-world paired acoustic-optical underwater dataset. The proposed method achieves an mAP50 of 95.6% and an mAP50-95 of 50.7%, consistently outperforming state-of-the-art unimodal detectors (YOLOv12-X, YOLOv13-X, RTDETR) and advanced multimodal fusion methods (DenseFusion, U2Fusion, SwinFusion). Ablation studies confirm the critical contributions of both the spatial alignment module and the dynamic weighted fusion module to overall detection performance. In addition, noise injection experiments demonstrate that the dynamic weighting strategy exhibits strong robustness against speckle noise and random pixel occlusion in challenging underwater environments. ConclusionThe proposed framework effectively mitigates cross-modal spatial misalignment and fusion degradation between sonar and optical imagery, resulting in significant improvements in underwater target detection accuracy. It delivers a feasible fusion paradigm for underwater multimodal perception.
Precise underwater target classification is a prerequisite for autonomous marine operations. Given the limitations of single sensors, multi-sensor fusion is essential. While RGB images provide rich texture information, they suffer from underwater light attenuation. Conversely, sonar provides robust structural features that compensate for RGB degradation. However, effective fusion is hindered by inherent heterogeneity, asymmetric feature degradation, and paired data scarcity. To overcome these limitations, an adapted dual-stream framework named SVFNet is developed for RGB-Sonar (RGB-S) fusion classification. Specifically, cross-modal contrastive learning is incorporated to bridge the inherent heterogeneity, while a Global Query Spatial Attention (GQSA) module is integrated to tackle asymmetric degradation. Finally, to address the extreme scarcity of paired data collected in unconstrained real-world environments, R-S9 is constructed as the first spatiotemporally aligned underwater RGB-S paired benchmark dataset for fusion classification, containing 3732 image pairs. Experiments on R-S9 demonstrate that SVFNet significantly outperforms leading unimodal baselines, SOTA long-tailed classification methods, and advanced RGB-X fusion architectures. These results validate the effectiveness of SVFNet in overcoming the physical limitations of single sensors. The source code is publicly available at https://github.com/VIP-Lab-NEU/SVFNet.
Stabilization control of large spacecraft after capturing massive, non-cooperative targets for On-Orbit Servicing (OOS) presents a critical challenge due to strong attitude-orbit coupling and complex nonlinear dynamics. Traditional Sliding Mode Control (SMC) is often limited by its reliance on fixed parameters and its propensity for actuator-damaging chattering. To address these issues, this paper proposes an innovative intelligent control framework: ”Adaptive Parameter and Residual Reinforcement Learning Sliding Mode Control (APR-RL-SMC)”. Initially, we derive a nine-degree-of-freedom, high-fidelity attitude-orbit coupled dynamic model that covers the complete post-capture dynamics, based on Lagrangian mechanics. Building on this, we design a Deep Reinforcement Learning (DRL) agent to fill the performance gap of traditional methods. Through a ”Dual-Branch Adaptive Regulation Mechanism”, the agent performs real-time parameter self-tuning and proactive residual compensation simultaneously. We train this agent using the Soft Actor-Critic (SAC) algorithm, enabling it to learn optimal control decisions under complex dynamic conditions. Numerical simulations demonstrate the framework’s significant effectiveness. Compared to traditional SMC, APR-RL-SMC reduces the main body’s attitude Integral Square Error (ISE) by 97.04%. Concurrently, it suppresses sporadic chattering, decreasing the tangential thrust control signal’s Total Variation (TV) by 21.91%. This research provides a high-performance intelligent solution for the post-capture stabilization control problem in OOS. Its dual-branch regulation mechanism effectively addresses the shortcomings of traditional controllers in parameter tuning and chattering suppression. As a key on-orbit enabling technology, this method holds significant value for advancing the level of orbital intelligence in complex space missions.
An experimental investigation was conducted to examine the effects of Al-B alloy additives on the ignition, combustion, and agglomeration characteristics of composite propellants. The properties of four Al-B alloys with different B contents (5%, 20%, and 50% B) were comparatively evaluated against pure Al powder using Thermogravimetric-Differential Scanning Calorimetry , ignition research by CO2 laser, combustion experiments under high-pressure , Scanning Electron Microscopy , and X-ray Diffraction. B addition markedly accelerates the oxidation of Al, increases the total heat release, and shortens ignition delay. Regarding combustion performance, Al-B alloy propellants exhibited higher burning rates, with the alloy containing 50% B (Al/B=1:1) showing the most pronounced improvement and a relatively low pressure exponent. In terms of agglomeration suppression, alloys with lower B content (e.g., 5% B) effectively decreased particle size combustion products size. However, increased B content led to greater formation of B2O3 in the combustion products, whose adhesive properties promoted particle agglomeration. Considering a trade-off among ignition delay, burning rate, and agglomeration inhibition, Al-B alloy with 50% B content demonstrates the greatest potential for enhancing the overall performance of composite propellants.
With the continuous development of space technology, spacecraft are progressively becoming larger, often requiring large appendages such as solar panels for long-term missions. The introduction of these large, flexible structures, however, presents new challenges in attitude control and orbit maintenance, which are rooted in the system’s inherent nonlinear dynamics. A dynamic model for a spacecraft with large appendages is established to investigate its operational behavior under the influence of gravitational potential energy. The impact of the appendages on the spacecraft’s attitude response and orbital variations is analyzed, revealing complex attitude-orbit coupling effects. A two-dimensional model is developed using analytical mechanics, and the system’s Lagrangian function is derived. Through the Hamiltonian variational principle, the highly nonlinear and coupled in-orbit operational equations are established. Numerical simulations using a symplectic Runge-Kutta algorithm demonstrate that the oscillation of the appendages’ attitude angles causes periodic fluctuations in the spacecraft’s orbital radius, confirming the strong coupling between the main body, the appendages, and their respective orbits and attitudes. Furthermore, a nonlinear control method combining spring dampers and a PD controller is proposed to provide theoretical support for enhancing the stability and reliability of such complex aerospace systems.
Verification and validation (V&V) of structural digital twins becomes fragile when high-frequency Sim-to-Real mismatch coexists with sensor-fault-induced under-observability, as missing channels undermine reliable diagnosis and model updating. To improve trustworthy V&V under partial observation, we present a physics-informed virtual sensing and closed-loop updating framework that reconstructs masked frequency responses and enables stable constrained FEM-parameter updating. Specifically, we couple a reduced-order FEM (ROM) prior with a Frequency Attention U-Net (FAU) that injects Fourier Neural Operator (FNO) global spectral increments through Transformer-gated fusion, driving constrained parameter updating. Systematic experiments on a scaled reconfigurable spacecraft across multiple geometric configurations and two excitation conditions, 10 Hz and 800 Hz, validate the framework's efficacy. Under 800 Hz excitation, the proposed virtual sensing reduces the log-amplitude main-peak error metric by an average of 76.7 dB under the adopted full-spectrum peak-search protocol. Furthermore, in missing-channel scenarios, the closed-loop updating strategy robustly limits log-amplitude errors on masked channels to 8-10 dB, outperforming spatial interpolation baselines (up to 48 dB error). These improvements reconstruct reliable full-field spectra, supporting physics-consistent interpretation of configuration-dependent joint softening through effective ROM-level parameter trends and providing a foundation for proactive diagnosis and smart maintenance.
This paper addresses safe leader-follower formation control for multiple quadrotor unmanned aerial vehicles (UAVs) under prescribed transient and steady-state error bounds, external disturbances, input-realization errors, and collision-avoidance constraints. A graph-theoretic formation error derived from the pinned communication topology is selected as the constrained variable in prescribed performance control (PPC) and mapped through a symmetric transformation. An extended state observer (ESO) estimates lumped uncertainties in the transformed-error channel. Because pairwise distance constraints have relative degree two with respect to virtual acceleration, a high-order control barrier function-based quadratic program (HOCBF-QP) minimally modifies the nominal outer-loop command. Command filtering, attitude/thrust mapping, and inner-loop attitude tracking are included to account for realization of the virtual safe input by underactuated quadrotors. Under leader reachability, admissible PPC initial conditions, bounded disturbances, and QP feasibility, all closed-loop signals are uniformly ultimately bounded, and the formation error remains within its prescribed envelope. When all QP slack variables are zero and the HOCBF initial conditions hold, the hard safety guarantee is retained. Simulations demonstrate three-dimensional formation tracking, satisfaction of the PPC and minimum-distance constraints, QP computation within the sampling period, and consistent performance in ablation, baseline-comparison, disturbance-sweep, and Monte Carlo test
To meet the design and evaluation requirements of underwater vision-based docking localization, a Webots-based simulation platform for Autonomous Underwater Vehicle (AUV) visual docking localization was designed and implemented to address the high cost of real sea trials, uncontrollable operating conditions, and the difficulty of systematically covering extreme scenarios. An end-to-end simulation of the docking localization pipeline was provided. Visual components—including fiducial markers, underwater illumination and imaging, and occlusion—were modeled in relatively fine detail, while non-vision-dominant factors such as propulsion and hydrodynamics were treated with approximate models to balance visual realism and simulation efficiency. The platform supported multiple types of visual markers, parameterized configuration of underwater lighting and turbidity, and the generation of diverse occlusion scenarios, enabling unified integration and benchmarking of docking localization algorithms. The results showed that the platform offered tunable scene parameters, repeatable conditions, and broad algorithm compatibility, and it effectively revealed performance differences across algorithms for complex combinations of illumination, turbidity, and occlusion. These capabilities reduced the risk and cost of real underwater docking experiments and supported faster iterative improvement of vision-based localization methods.
Fractional-Order Spiking Neural Network (FOSNN) has the characteristic of infinite memory and neural impulses, which can more accurately describe neural network systems and demonstrate higher precision data processing capabilities in artificial intelligence. The neural spiking leads to multiple equilibrium points coexisting in the networks system. Multistability analysis mainly studies the problem of the multiple equilibrium points, which helps to improve the robustness and reliability of the networks. However, fractional calculus and neural spiking increase the theoretical difficulty of multistability and attractivity analysis in neural networks, which is the main motivation to study and discuss. Firstly, for a Hopfield type of FOSNN with pulse activation functions, the solution existence is proved according to Filippov solutions. Secondly, the state space is divided and the sufficient conditions for multistability are proposed and proved by using fixed point theorem, Laplace transform, Mittag-Leffler function monotonicity analysis, etc. Furthermore, the boundedness and global attractivity of FOSNN are discussed based on fractional-order Lyapunov method. Finally, using the fractional-order prediction correction algorithm, some numerical examples are conducted in order to verify the correctness for all proposed results.
ObjectiveWireless sensor networks (WSNs) are essential for ocean monitoring and are widely used in environmental monitoring, target localization, marine resource development, and disaster warning applications. However, WSNs often face challenges such as arbitrary deployment strategies, low effective coverage, and high coverage redundancy, which degrade network performance. To address these issues, this paper proposes a fractional-order chameleon swarm algorithm (FCSA) to optimize the deployment of static WSN nodes. MethodsFirst, an improved Circle chaotic mapping method is employed to enhance population diversity and global distribution, ensuring higher-quality initial conditions for optimization. Next, during the velocity update phase, a fractional-order velocity update strategy is introduced to effectively leverage the historical search experiences of individuals, enhancing the balance between global exploration and local exploitation. Furthermore, the Levy flight mechanism is incorporated into position updates, providing stronger jumping characteristics and adaptability. These improvements enable FCSA to effectively optimize key performance indicators such as coverage rate and coverage redundancy, significantly enhancing deployment efficiency and distribution uniformity for static WSN nodes while ensuring better adaptability to complex environments.ResultsSimulation results demonstrate that FCSA outperforms CSA, CSA-Circle, CSA-Levy, GA, RSO, and eleven other classical optimization algorithms in static node deployment. FCSA achieves a high coverage rate of 0.8018 while significantly reducing coverage redundancy to 0.0078. Additionally, for single optimization tasks, FCSA exhibits the fastest convergence, requiring only 638 iterations to reach a fitness value of 0.191409, significantly outperforming other algorithms. After 30 independent runs, statistical analysis shows that FCSA maintains an extremely fast convergence speed in the early iterations, reaching an optimal fitness value of 0.198222 after 1000 iterations. Among the ten algorithms, FCSA is the only one with a standard deviation of the fitness value below 0.2, indicating superior global search ability, higher convergence accuracy, and better distribution uniformity. It effectively mitigates the issue of uneven node distribution observed in traditional algorithms while maintaining strong stability and robustness. ConclusionIn addressing the WSN static node deployment problem, FCSA effectively optimizes sensor placement, significantly improving monitoring quality through a multi-strategy collaborative optimization approach. The algorithm exhibits strong robustness and adaptability in complex environments. Additionally, FCSA provides an efficient, high-quality deployment solution for ocean monitoring and similar applications, offering strong theoretical and technical support for sensor network optimization and expansion, with significant application potential and practical value.
Aluminum powder is widely used in solid propellants due to its high energy density, but the formation of the aluminum oxide layer on its surface limits its ignition and combustion efficiency. To address this issue, we designed and synthesized fluorinated lipoic acid monomers (FTA-3/7) and mechanically induced their in-situ polymerization to form a uniform coating on the surface of aluminum powder. The uniformly distributed and abundant F atoms within the FTA polymer, on one hand, contribute to the formation of Al-F coordination bond with the aluminum powder, further enhancing the adhesion between the polymer and the aluminum powder. On the other hand, the F atoms improve the ignition, combustion, and water resistance of the aluminum powder. The coated aluminum powder (Al@PFTA-3) exhibited 66.7% increase in emission spectral intensity and 399% increase in heat of combustion compared to pure aluminum. Additionally, the emission intensity of the solid propellant with Al@PFTA-3 increases by 2.4 times compared to the original, with a 20% reduction in self-sustaining combustion time, while its pressure index is lower than that of typical benchmark propellants. This method provides a simple and effective approach to improving the reactivity and stability of aluminum powder, while enhancing its performance in solid propellants.
The Internet of Underwater Things (IoUT) extends distributed sensing to marine environments, where unmanned underwater vehicles (UUVs) serve as mobile nodes for flexible underwater operations. This article investigates the heterogeneous UUV multi-type task planning problem (HUMTTP) for a collaborative system comprising autonomous underwater vehicles (AUVs), underwater gliders (UGs), and bionic manta-ray underwater vehicles (BMUVs). Existing methods generally prioritize execution efficiency but often overlook platform heterogeneity and acoustic communication limitations. To address these issues, this article proposes a hierarchical task planning framework called REASOM-UCLNS. In the task allocation layer, the reward-and energy-aware self-organizing map (REASOM) algorithm incorporates a novel neuron-based reward estimation strategy into winner selection. By integrating the estimated platform-specific task rewards with current-aware energy evaluation, REASOM assesses the execution suitability of heterogeneous UUVs and facilitates the efficient allocation of multi-type tasks. In the route planning layer, the underwater cooperative large neighborhood search (UCLNS) algorithm converts the generated task sets into feasible routes for individual UUVs. UCLNS develops two constraint-specific operators to restore route feasibility when communication connectivity or energy balance violations occur. By integrating targeted constraint handling into the destroy-repair process, UCLNS improves route quality while satisfying operational constraints. Numerical results demonstrate the superior performance of the proposed framework. Finally, a lake experiment involving four heterogeneous AUVs further confirms its practical applicability.
This paper presents a low-power, self-referenced class-C voltage-controlled oscillator (VCO) with dual feedback loops. To reduce the system complexity and power consumption of class-C VCOs, a self-referenced dual-feedback-loop topology is proposed. The first loop is a negative-amplitude-detection loop, which generates a reference voltage Vref without requiring an external circuit, thereby avoiding additional power consumption. The second loop is a common-source-node feedback loop, which adaptively adjusts the bias voltage of the cross-coupled pair to achieve dynamic biasing and ensure oscillation startup. The proposed VCO is implemented a 22-nm CMOS process, consuming only 0.49 mW under a 0.6-V supply voltage, with a frequency tuning range is 36.4 % from 3.6 to 5.2 GHz. Experimental results exhibit phase noise (PN) of-110 dBc/Hz at a 1-MHz offset and a figure of merit (FoM) of-186 dBc/Hz at a 4.4-GHz carrier frequency.
This study investigates electromigration (EM) degradation in copper redistribution layers (RDLs) under electrothermal loading through high-current-density experiments, microstructural characterization, and phase-field (PF) simulations. The resistance increase showed strong temperature dependence consistent with an Arrhenius-type trend, rising by 0.91% after 16h at $55~^{\circ }$ C and by 20.43% after 3 h at $175~^{\circ }$ C. Without temperature control, resistance increased by 384.71% within 207 s at ${2}.{1} \times {10}^{{6}}$ A/cm2. Representative electron backscatter diffraction (EBSD) observations showed grain coarsening, with the mean grain size increasing from 1.31 to $1.71~\mu $ m, together with enrichment near the [101] orientation. Under a representative coupled condition, the thermomigration-to-EM flux ratio was approximately 2.54, indicating an enhanced thermomigration contribution while the electron-wind force remained important. PF simulations further showed that temperature gradients promote void morphological instability and that diffusivity heterogeneity facilitates preferential void nucleation. Overall, RDL degradation evolves from gradual resistance growth to rapid structural failure through the coupled effects of EM, thermomigration, and Joule-heating feedback.
[Objective]To address issues such as unclear application architecture,weak information interac-tion capability,and low efficiency in the collaborative testing of heterogeneous marine unmanned clusters,this paper proposes an agile collaboration technology.This approach enables system-level clusters to quickly re-spond to task demands platform-level unmanned system to integrate swiftly into the cluster,and system-level controllers decouple software and hardware step-by-step for agile task processing and execution.[Methods]First,this study analyzes the task characteristics,network and optimization requirements for the cooperative application scenarios of marine unmanned clusters,and divides them into several node groups according to their functions.Then,it can design the heterogeneous marine unmanned cluster agile cooperative architecture based on the functional node groups,and carry out the load complementation and task coordination through the fusion configuration of the cooperative architecture.Next,based on the application requirements and architec-tural features,the common marine unmanned cluster application tasks are divided into three categories:time priority tasks,sequential execution tasks,and routine operation tasks,and an agile task planning method based on the self-organizing graph algorithm is proposed for the heterogeneous marine unmanned clusters.Finally,a multilevel software-hardware decoupled marine unmanned cluster agile collaborative controller is developed,which interacts with the various systems of the platform in the form of a unified interface.[Results]Accord-ing to the test results of dynamic target detection and tracking on-lake experiments by heterogeneous un-manned clusters,the final average heading deviation angle between each vehicle and the target is 6.7°,which successfully accomplished the cooperative detection and tracking of dynamic targets.[Conclusion]This method can enable marine unmanned systems with significant performance disparities to quickly integrate in-to and implement typical tasks,and has excellent scalability and adaptability,which can help accelerate the de-velopment and practice of marine unmanned clusters.
Accurate indoor positioning remains challenging due to multipath and non-line-of-sight (NLOS) interference in radio signals and drift in inertial sensors. In this paper, we propose a robust Bluetooth low energy (BLE)-inertial measurement unit (IMU) fusion framework based on the cubature Kalman filter (CKF), where the key parameters-process noise Q, measurement noise R, and BLE path-loss exponent n-are jointly optimized in an offline calibration process using a hybrid artificial bee colony-particle swarm optimization (ABC-PSO) algorithm. The optimizer leverages the global exploration ability of ABC and the local refinement capability of PSO, enhanced with adaptive boundary contraction and elite migration to prevent premature convergence. Robustness is further enhanced by RSSI-history calibration, distance penalization for biased measurements, and a Huber-based CKF update with dynamic noise adjustment. In controlled simulations (with idealized Gaussian noise and perfect synchronization), the hybrid optimizer demonstrates stable convergence and consistent parameter identification. In real-world experiments conducted in a typical indoor environment, the same optimized parameters reduce the root mean square error from 3.467 m (baseline CKF) to 2.832 m, corresponding to an 18-36% improvement over the extended Kalman filter (EKF), standard CKF, particle filter (PF), and coordinate transformation with forward-backward smoothing (CFBS). These results indicate that the proposed framework provides measurable robustness enhancement under the tested indoor NLOS conditions. A qualitative stability discussion and convergence analysis are provided. Simulation studies are conducted to analyze optimization convergence under controlled statistical assumptions, while real-world experiments serve as the primary performance evaluation. By unifying offline parameter calibration with robust sensor fusion, this framework provides a practical and scalable solution for enhancing localization reliability in BLE-IMU indoor systems using low-cost sensors.
As an important autonomous sensing node in the underwater internet of things (UIOT), an autonomous under-water vehicle (AUV) may be unable to obtain accurate latitude information during underwater navigation. Environmental disturbances can also cause the strapdown inertial navigation system (SINS) base to undergo angular sloshing and linear disturbance. These effects degrade the performance of conventional polynomial fitting (PF) methods for SINS-based latitude estimation. Moreover, a fixed PF order may lead to overfitting or underfitting of the state vector, resulting in gravity-component distortion and spurious maximization of the angle between gravity vector. This paper proposes a threshold-determined polynomial fitting (TDPF) method for AUV latitude estimation, constructs a prior correlation constraint between gravity component and latitude error, designs a quantitative intensity evaluation model based on velocity and position vector perturbation, and designs a threshold determination criterion for the selection of PF order at different observation times, which improves the accuracy of gravity component amplitude and latitude estimation under the premise of adapting linear disturbance intensity, thereby rapidly and efficiently guaranteeing high-precision latitude estimation for AUVs in UIOT. Simulation and off-line experimental results show that the TDPF method achieves more stable convergence under SINS swing-base conditions with less additional cost, and the latitude estimation accuracy is improved.
This paper focuses on the effective dynamics of a class of stochastic parabolic equations with a fast oscillation driven by fractional Brownian motion and Poisson jumps under non-Lipschitz conditions. This model is designed to describe systems that exhibit long-range dependence and emergency impact. In this work, we demonstrate the tightness of the slow component and the ergodicity of the fast component. Furthermore, it is shown that the slow component converges to the solution of the corresponding effective equation. The averaging principle simplifies the computational complexity, enabling us to bypass the intricacies of the original system and instead focus on analyzing the behavior of the effective equation. In contrast to studies that rely on Lipschitz conditions, the results presented here extend beyond these restrictions by considering non-Lipschitz conditions, which are more flexible and widely applicable.
Underwater docking of autonomous underwater vehicles (AUVs) was typically dependent on the complete visual detection of markers. When markers were only partially visible due to occlusion or departure from the field of view, conventional localization methods based on complete features were rendered ineffective, resulting in the interruption of docking operations. To address this limitation, an enhanced orientation-aware method based on a spatiotemporal attention convolutional neural network (CNN) was proposed in this study. The core of this method was a dual-path feature fusion architecture: discriminative features of visible marker segments were extracted from single frames by the spatial path, while the temporal path was employed to aggregate features across consecutive frames, thereby compensating for the insufficiency of single-frame information. These two pathways were adaptively fused through a spatiotemporal attention module, which was designed to dynamically focus on the most informative cues. Consequently, robust qualitative judgment of the marker’s relative orientation was achieved. Experimental validation conducted in underwater environments demonstrated that stable orientation awareness was maintained by the proposed method even under conditions where the marker was severely off-center or largely obscured. This approach was shown to significantly extend the initial capture range for AUV docking guidance, and the robustness and operational continuity of the system under extreme visual conditions were effectively enhanced.