
Although Kalman filtering and factor graph optimization have been widely used in GNSS/INS integrated navigation, their positioning accuracy can degrade significantly in urban canyon environments because of satellite signal blockage, multipath effects, and non-line-of-sight errors. To improve robustness and accuracy, this study proposes a forward tightly coupled GNSS/INS factor graph optimization method. In the front end, raw GNSS and INS measurements are fused in a tightly coupled framework, and an IGG-III robust weighting model is introduced to suppress abnormal observations. In the back end, GNSS position factors, IMU pre-integration factors, and marginalization factors are constructed within a sliding window to optimize the navigation states while preserving historical information. Experiments on an urban canyon dataset demonstrate that the proposed method improves navigation accuracy compared with conventional EKF and FGO. Further analyses show that the adopted slidingwindow configuration reduces the mean complete-epoch processing time by 26.31% while maintaining comparable positioning accuracy. RFGO also exhibits better performance during a 120 s complete GNSS outage, and the comparison of robust weighting models demonstrates that IGG-III provides the best overall three-dimensional positioning performance among the tested models.
This paper investigates the bipartite output regulation problem (BORP) of nonlinear multi-agent systems (MASs) over static signed networks. Unlike conventional cooperative output regulation problem (CORP), antagonistic couplings among agents result in a bipartite tracking objective, where different groups of followers asymptotically approach references with opposite signs. To address the nonlinear exogenous dynamics under signed topologies, a nonlinear distributed observer (NDO) is proposed by incorporating the sign information of inter-agent interactions into the observer update law. It is shown that the resulting observer remains well-posed and achieves exponential convergence to the desired signed reference trajectory. Moreover, the solvability of the nonlinear BORP is analyzed which is based on the solvability of nonlinear regulator equations. Furthermore, a distributed state-feedback protocol is constructed for nonlinear MASs based on the observer estimates, guaranteeing leader-follower bipartite synchronization of the overall closed-loop system. Finally, numerical simulations involving multiple pendulum systems are presented to demonstrate the feasibility and effectiveness of the proposed approach.
Unmanned Aerial Vehicles (UAVs) and Unmanned Ground Vehicles (UGVs) offer complementary capabilities for long-duration autonomous missions, where UAV recovery by a ground platform can significantly reduce aerial energy consumption. In practical GPS-denied environments, however, such docking-return operations are limited by three coupled difficulties: establishing an initial relative configuration without external localization infrastructure, selecting an energy-efficient docking point under heterogeneous vehicle constraints, and maintaining feasibility when unknown obstacles are encountered. This paper proposes a relative-initialization and online-replanning framework for cooperative UAV–UGV docking-return. A lightweight front-end initialization module is first developed using a minimal three-node UWB configuration, onboard IMU/odometry increments, and a short excitation trajectory to recover a consistent UAV–UGV-base relative geometry. The initialized configuration is then used in an energy-time docking-point planner that incorporates UAV flight energy, UGV Dubins-type kinematic constraints, asynchronous arrival coordination, Field-Of-View (FOV) feasibility, and obstacle avoidance. The proposed method is positioned as a task-level docking-point planning and replanning framework, rather than a complete low-level docking controller or a full localization system. The relative-initialization module is evaluated in simulation, while the docking-point planning and online replanning layers are validated through Monte Carlo simulations, Gazebo tests, and real-world experiments with controlled initial relative configurations. These results demonstrate the effectiveness of the proposed task-level planner in improving energy-time performance, reducing UAV waiting, and maintaining feasible docking-return execution under the considered cluttered GPS-denied settings.
Pose estimation using aerial images captured by Unmanned Aerial Vehicles (UAVs) allows the localisation in GPS-denied scenarios. Several methods based on deep learning approaches with convolutional neural networks (CNN) have become tools for estimating localisation from images. However, building a model that can estimate the pose from a single image needs a large dataset and training time to obtain a result. Besides, the model can be inappropriate in assessing the correct pose in dynamic scenarios with multiple changes. Therefore, we propose a methodology using a binary network with a Continual Learning (CL) strategy to create an estimation model during the same flight mission. Also, we use a submap scheme and multiple models to acquire the UAV’s localisation into different parts of the trajectory. Finally, we use PoseNet, ORB-SLAM2 and single-model for comparison purposes in four scenarios, achieving a percentage error of 14% of the total trajectory and a processing time of 51 ms with our proposed approach.
A pursuit-evasion problem for unmanned underwater vehicles (UUVs) under a nested asymmetric information structure is investigated in this paper within a finite-horizon non-zero-sum linear quadratic Gaussian (LQG) framework. In the considered scenario, the pursuer has access to more complete state information than the evader, which makes the derivation of Nash equilibrium strategies more challenging than in the conventional symmetric-information setting. To formulate the problem, the nonlinear UUV dynamics are linearized and discretized, and the interaction between the pursuer and the evader is modeled as a discrete-time stochastic dynamic game with distinct information filtrations. Based on the stochastic maximum principle, the necessary and sufficient conditions for the existence of a Nash equilibrium are established in the form of forward-backward stochastic difference equations (FBSDEs). To address the structural coupling caused by the nested information pattern, an orthogonal decomposition method is proposed, through which the coupled equations can be decoupled and explicit affine-feedback equilibrium strategies can be derived. Numerical simulations based on the REMUS 100 UUV model are provided to demonstrate the effectiveness of the proposed method.
In modern transportation systems, eco-driving aims to reduce fuel consumption and emissions while maintaining traffic efficiency and safety. Existing eco-driving methods at signalized intersections often rely on accurately prescribed arrival times, which are difficult to obtain in mixed traffic due to the motion uncertainty of human-driven vehicles (HVs) and the variability of traffic-light phases. Moreover, the interaction between connected and automated vehicles (CAVs) and surrounding HVs is often insufficiently modeled in existing approaches. To address these challenges, this paper proposes a potential-game-based eco-driving framework for mixed platoons at signalized intersections. The proposed method employs an artificial potential field (APF) within a recedinghorizon control architecture to optimize the trajectories of CAVs online, while incorporating predicted HV motion as interactive information. The resulting multi-CAV coordination problem is reformulated as an exact potential game, for which the existence of a pure-strategy Nash equilibrium is established, and local minimizers of the induced potential function correspond to local Nash equilibria. Extensive simulation results demonstrate that the proposed framework effectively reduces delay, idling time, and fuel consumption, while enabling smooth and adaptive intersection crossing under dynamic traffic-light conditions.
This study addresses the limitations of discrete multi-agent path planning in Conflict-Based Search (CBS) by introducing a modular enhancement that improves trajectory smoothness while maintaining completeness and conflict-free guarantees. The proposed approach integrates the standard CBS with a post-processing smoothing stage and an adaptive Dynamic Safety Margin (DSM) embedded in the D* Lite low-level planner. DSM dynamically adjusts traversal costs based on the proximity of the obstacle, enabling aggressive smoothing without violating safety or dynamic feasibility. Extensive Monte Carlo simulations demonstrate that the CBS–D* Lite + DSM pipeline reduces the average path length by 8% and enhances smoothness by up to 9% compared with an unsmoothed baseline, with minimal computational overhead. Dynamic executability was validated in a 6-DOF quadcopter simulation using a cascaded PID controller, achieving a tracking RMSE of less than 3.2 cm. Among the five evaluated smoothing methods (moving average, Gaussian filter, cubic spline, gradient descent, and minimum-snap), gradient descent offers an optimal balance of path quality, safety, and runtime. Requiring no changes to the CBS’s core logic, this approach provides a scalable, practical solution for real-time multi-robot systems in UAV swarms and automated warehouses.
High-efficiency area coverage using Unmanned Aerial Vehicle swarms is a fundamental capability for large-scale search-and-rescue missions. In disaster scenarios lacking prior knowledge, where agents rely on local sensing and decentralized decision-making, stochastic exploration strategies provide a practical solution. Asymmetric Lévy flight has recently emerged as a promising theoretical framework for scalable autonomous coverage in complex environments. However, existing studies primarily focus on kinematic optimization in particle-level simulations and rarely consider practicable strategies tailored for real-world swarm deployment. This limitation becomes particularly critical in dense swarms, where execution-level deadlocks frequently arise as agents mutually obstruct access to their intended waypoints, undermining coverage continuity and reliability. When deployed on physical platforms with standard obstacle-avoidance mechanisms, these deadlocks can severely degrade operational performance. To address this challenge, this paper proposes D-ALF, a stochastic swarm coverage planner augmented with a deadlock-aware exploration policy. By integrating an asymmetric collaborative consensus mechanism, the framework enables agents to actively resolve local congestion, maintain continuous and coordinated stochastic exploration, and translate theoretical coverage efficiency into physically realizable swarm operation. Extensive simulations and real-world experiments demonstrate that D-ALF actively accounts for execution-level deadlocks and consistently enhances coverage efficiency, operational reliability, and scalability compared with classical asymmetric Lévy flight strategies.
In response to the quantitative diagnosis challenge of concurrent multi-actuator faults in quadrotor UAVs operating within safety-critical scenarios, such as nuclear emergency response, this paper proposes a diagnostic framework integrating CNN-LSTM and SHAP. Using 34 flight-state and control-related features as inputs, a CNN-LSTM hybrid model (CLL) is developed, where the convolutional module captures fault-induced local transients and the stacked LSTM module models the temporal dynamics of fault propagation, thereby enabling parallel continuous regression of the four actuator efficiency coefficients. Under unified data partitioning and training settings, the proposed CLL achieves MAE 0.0894 and MSE 0.0201 on the test set and remains optimal in both ablation and multi-method comparisons. To evaluate continuous estimation capability beyond discrete training anchors, representative unseen-efficiency cases with eta = 0.35, 0.55, 0.72, and 0.85 are further tested, yielding an average absolute error of 0.031, which confirms stable nonanchor diagnosis behavior. Post-hoc SHAP analysis shows that key features, including yaw angular velocity, Z-axis acceleration, and four-motor control command channels, are ranked consistently with quadrotor dynamic mechanisms, supporting both physical consistency and interpretability. The proposed framework provides a reliable basis for UAV health-state assessment and downstream fault-tolerant control in safety-critical applications.
In robotic mobile fulfillment system (RMFS), the efficient scheduling of multiple automated guided vehicles (AGVs) for rack retrieval and repositioning is a pivotal determinant of overall picking performance. This operation entails a composite decision-making process that involves the joint determination of multi-AGV task set allocation, task sequencing for each AGV, and repositioning strategies for the retrieved racks. As these decision components are tightly coupled, addressing any single subproblem in isolation leads to poor performance. To address this complexity, this paper proposes a hybridization of meta-heuristic algorithm and deep reinforcement learning with temporal-spatial storage constraints (HMDRL-TSC) to jointly optimize these interdependent decisions. Grounded in a bi-level optimization framework, the HMDRL-TSC employs a strategic problem decomposition algorithm. In the upper-level optimization, a general variable neighborhood search (GVNS) algorithm, incorporating generalized travel distance metrics, is utilized to address the task set allocation problem for the multi-AGVs. Subsequently, the lower-level optimization employs deep reinforcement learning (DRL) integrated with a spatiotemporal conflict resolution strategy, which rapidly generates high-quality rack retrieval sequences and eliminates inter-AGV conflicts to ensure solution feasibility. Computational results demonstrate the superior performance of the proposed HMDRL-TSC in solving the rack retrieval and repositioning problem.
Multirotor Unmanned Aerial Vehicles (UAVs) that are fully-actuated offer improved dexterity for aerial manipulation tasks by decoupling attitude and translation. There is also potential for improved agility when the airframe design is optimized, leading to better disturbance rejection. This paper compares three fixed-tilt configurations by evaluating optimal designs for each under different design requirements: payload, flight time and horizontal force during level hover. Three objective functions are used in the numerical optimization: thrust bandwidth, horizontal acceleration and airframe diameter. The results highlight the advantages of a heterogeneous rotor configuration, providing the best trade-off between agility and energy efficiency.
With the rapid advancements in artificial intelligence and control technologies in recent years, uncrewed systems have become increasingly prevalent across various fields. Path planning, a critical technology enabling autonomy in these systems, remains a challenging and active area of research. This review provides a comprehensive overview of the fundamentals of path planning and deep reinforcement learning (DRL), laying the foundation for understanding the potential and limitations of DRL in uncrewed system applications. It systematically reviews DRL methodologies and examines their applications across uncrewed aerial vehicles (UAVs), uncrewed ground vehicles (UGVs), uncrewed surface vehicles (USVs), and heterogeneous platforms, highlighting representative algorithms, real-world deployment scenarios, and diverse operational environments. The review also identifies major challenges in DRL-driven path planning, including real-time adaptability, robustness in complex environments, and scalability across domains. In addressing these challenges, it offers valuable insights and outlines future research directions, emphasizing the need for enhanced efficiency, safety, and generalization to meet the demands of next-generation autonomous systems.
The integration of Unmanned Aerial Vehicles (UAVs) into Industry 4.0 ecosystems is severely hampered by a critical lack of interoperability. This challenge, stemming from diverse platforms, incompatible protocols, and heterogeneous data structures, creates a semantic gap that hinders collaboration between systems. To address this gap, this paper presents a semantic gateway that bridges the Asset Administration Shell (AAS) specification with the Robot Operating System 2 (ROS 2) framework, enabling structured alignment of UAV telemetry data. Unlike existing static adapters, the proposed architecture provides a resilient, asynchronous pattern specifically optimized for the strict constraints of aerial robotics, including high-frequency telemetry, limited onboard computation, and intermittent network connectivity. The gateway’s effectiveness is rigorously evaluated through controlled benchmarks, high-fidelity software-in-the-loop (SITL) simulations, and edge hardware deployments. Results demonstrate that the configuration-driven transformation model maintains internal processing latency below 5 ms, making it transparent to primary flight-control loops, while preserving digital twin consistency during network disruptions without introducing state “ghosting” effects upon reconnection. The evaluation focuses on the semantic consistency of runtime data transformation rather than full cross-system semantic interoperability. Although demonstrated for UAVs, the gateway provides a generalizable architectural pattern for integrating other mobile ROS-based assets into Industry 4.0 ecosystems. This study offers a practical, validated approach to making robotic systems compatible with AAS, indicating potential reduction in industrial integration effort and enabling seamless multi-agent operations.
This paper presents a novel hierarchical homography-based visual servo control strategy for pose regulation of underactuated autonomous underwater vehicles (AUVs). The proposed approach separates kinematic and dynamic control loops to effectively manage visual feedback, nonholonomic constraints, and dynamic uncertainties, enhancing system flexibility. The kinematic loop uses pure image information for motion guidance, employing homography decomposition for yaw angle calculation and refining translational errors. A sigma -scaling technique is applied to deal with the nonholonomic constraints. Then a switching visual stabilization controller generates reference velocities for the AUV. The dynamic loop employs active disturbance rejection control (ADRC) to estimate and compensate for total disturbances (dynamic uncertainties and random water flow disturbances), ensuring robust velocity tracking under stochastic conditions. Physical simulations demonstrate the effectiveness of the proposed method in completing visual stabilization tasks of AUVs under dynamic uncertainties and nonholonomic constraints.
This paper investigates reinforcement learning (RL)-based tracking control for robotic manipulators under event-triggered mechanisms. First, the dynamic model of the manipulator is formulated, accounting for model uncertainties and external disturbances. Then a super-twisting sliding mode surface is designed and embedded into the RL control framework. To minimize tracking errors and control efforts, a performance index function is constructed, and an identifier-critic RL framework is employed to approximate the optimal control policy. An event-triggered mechanism is introduced to update control commands only upon event occurrences defined by a triggering rule, thus reducing communication overhead. Theoretical analysis establishes closed-loop stability and proves the absence of Zeno behavior by ensuring a strictly positive lower bound on the inter-event time. Simulations and experiments are conducted to validate the effectiveness of the proposed strategy.
The SINS/DVS-integrated navigation can effectively correct velocity errors in SINS, while its navigation accuracy is affected by time-varying systematic errors inherent in DVS measurements and the geometric constraint of requiring at least three noncoplanar stars. To address these limitations, a novel SINS/DVS integrated navigation method using both conventional DVS and a newly constructed Cross-Directional Difference DVS (CDD-DVS) is proposed (referred to as SINS/CDD-DVS). In this method, the CDD-DVS measurement is formed by differencing DVS observations from two distinct directions within the same epoch, which simultaneously eliminates common-mode time-varying systematic biases and synthesizes a virtual velocity constraint along a third observation direction using only two stars. Unlike traditional same-direction time differencing, which suppresses useful motion signals along with errors, the proposed cross-directional strategy preserves essential vehicle motion information while effectively canceling correlated systematic biases. Simulation results demonstrate that the proposed SINS/CDD-DVS method reduces total position error by 45.2% compared to conventional SINS/DVS under star-limited conditions, confirming a significant improvement in navigation accuracy of the integrated navigation system.
This paper proposes a centralized control framework for cooperative aerial transportation using multiple unmanned aerial vehicles (UAVs) connected via suspended cables, where the overall UAVs-payload system is treated as a unified dynamic entity. The objective is to transport a rigid object along a predefined trajectory while ensuring stable and coordinated motion. The payload is connected to the UAVs through elastic cables, which introduce coupling effects between the vehicles and the transported object. Due to the highly coupled and uncertain dynamics of the UAVs-payload system and the absence of onboard sensors on the payload, an explicit dynamic model is difficult to obtain and the payload state cannot be directly measured. To address these limitations, a centralized model-free control (CMFC) strategy is adopted, allowing control to be achieved without relying on an accurate model of the system nor on onboard sensing of the payload. Instead, the payload position and orientation are estimated from the measurements provided by the cooperating UAVs. The MFC-based approach is combined with a force allocation scheme to distribute the global control effort among the UAVs in a coordinated manner, enabling cooperative transportation despite modeling uncertainties and limited payload measurements. In addition, force-based allocation metrics are investigated as analytical indicators for comparing cooperative transport configurations with different numbers of UAVs and attachment geometries. Rather than seeking an optimal configuration, this analysis provides insight into the relative efficiency of each setup and supports the selection of suitable attachment layouts or the estimation of the number of UAVs required for a given task. The proposed control architecture is evaluated through numerical simulations, showing accurate trajectory tracking, robustness to disturbances and consistent cooperative behavior of the UAV team. These results highlight the effectiveness of the MFC-based centralized strategy and its relevance for cooperative aerial transportation applications.
The low-altitude economy, defined as airspace at or below 1000m, has become a key strategic growth driver with applications spanning urban air mobility (UAM), autonomous drone delivery, and infrastructure inspection. These operations require robust real-time object detection of ground-level targets from aerial viewpoints, yet existing methods face fundamental limitations: Severe scale variations across operational altitudes (20-500m), complex urban background interference, sparse small object detection of distant vehicles and pedestrians, and stringent real-time requirements for on-board analytics. To address these challenges, this paper presents SFHA-DET, a detection framework designed for resource-constrained UAV platforms in low-altitude scenarios. The framework consists of three innovations: (1) the CMSCNet backbone network based on cross-scale spatial-Frequency Collaborative Bottleneck (CSFCB) modules, which achieves lightweight multi-scale feature extraction through gating mechanisms and spatial-frequency collaborative processing; (2) the hierarchical bidirectional feature fusion network (HBFFN) with multi-resolution adaptive integration modules (MRAIMs) for complementary cross-altitude feature fusion; and (3) the dynamic hierarchical feature interaction network (DHFIN) with ERAM, CFEM, and PSFU sub-modules for sparsified attention computation and multi-scale feature optimization. Extensive evaluation on two representative aerial datasets - VisDrone2019 (urban traffic monitoring) and UAVDT (vehicle detection) - demonstrates that SFHA-DET achieves 29.8% AP at 63 FPS with 44.8GFLOPs and 14.42M parameters. These computational requirements are designed to fall within the published computational envelope of representative edge accelerators such as the NVIDIA Jetson Orin Nano, substantially reducing hardware demands compared to transformer-based solutions that require desktop-grade GPU platforms. However, we emphasize that all reported inference measurements are obtained on an RTX 3090 desktop GPU; direct on-device benchmarking on embedded UAV platforms is left as immediate future work, and the present framework is best characterized as edge-oriented in design rather than verified for edge deployment. The proposed framework provides a computationally efficient method for foundational perception tasks in emerging low-altitude applications.
Drone-perspective small object detection requires processing high-resolution imagery where objects typically occupy fewer than 32 & times;32 pixels, demanding both fine-grained spatial preservation and global context modeling under strict computational constraints. Convolutional neural networks are limited by local receptive fields and cannot effectively model global context, while Transformer-based approaches, despite their global modeling capability, suffer from O(N-2) computational complexity that becomes a severe bottleneck when processing high-resolution inputs. This paper proposes MambaSOD, an end-to-end detection framework with an encoder-decoder architecture based on state space models that achieves linear computational complexity throughout the entire pipeline of feature extraction, multi-scale fusion, and query interaction. On the encoder side, MambaSOD builds a linear-complexity multi-scale representation by coupling a Vision Mamba backbone with a P2 high-resolution enhancement module, which recovers fine-grained texture details through dual-path fusion of shallow image features and up-sampled backbone output, together with a BiFPN that performs weighted bidirectional fusion across five scales. On the decoder side, we replace both self-attention and cross-attention with Mamba-driven query interaction: the Mamba-based Query Self-Interaction (MQSI) module enables implicit inter-query communication through bidirectional state propagation at O(N-q) cost, while the Mamba-based Query-Feature Interaction (MQFI) module reformulates query-feature cross-attention as a sequence modeling problem, reducing its complexity from O(N-q & times; N) to O(N-q + N). Experiments on the VisDrone2019 and UAVDT datasets demonstrate that MambaSOD achieves 23.8% AP, a 52.6% relative improvement over the Vision Mamba baseline, while requiring fewer FLOPs than state-of-the-art detectors such as Cascade R-CNN, ViT, Deformable DETR, and DINO, offering a competitive accuracy-computation trade-off for high-resolution drone-perspective small object detection. Our code is available at https://github.com/LeoHoW6/MambaSOD-Drone-view-small-object-detection-via-a-Mamba-based-query-feature-interaction-framework.git.