The integration of virtual and real environments is driving transformative advancements in unmanned systems, offering new avenues for intelligent collaboration and adaptive deployment. This survey presents a comprehensive framework for virtual-real integration centered on unmanned systems, aiming to unify fragmented research efforts and guide future exploration. By adopting both global and local perspectives, the framework facilitates a cohesive understanding of system components, their interactions, and the structural relationships that underpin seamless coordination. This is further complemented by an application case study and a synthesis of the current limitations. To operationalize this framework, the paper systematically examines enabling technologies, existing constraints, and the collaborative architecture that supports dynamic interaction across physical and virtual domains. It further outlines the evolving trends and construction requirements of practical application scenarios. Finally, key challenges and emerging research opportunities are discussed to inform future work and encourage deeper exploration of this rapidly developing field. Note to Practitioners-The motivation of this work is to explore advancement opportunities for unmanned systems in the trend of fusion between virtual and real worlds, so as to facilitate their better deployment in the real world. The impact of relevant enabling technologies and restrictions, the enhancement of unmanned systems, the involvement of human intelligence, and the deployment of specific application scenarios are all crucial for researchers and practitioners in the field to carry out practical work. This paper can provide researchers and practitioners with a comprehensive reference including the above content and help them utilize the integration of virtual and real environments to enhance the actual deployment of unmanned systems by further focusing on virtual application scenarios. In addition, challenges and future research directions identified in this survey also help identify entry points driving developments in the field.
The robust manipulation of deformable objects (DOs), such as cables, clothes, and food, is essential for developing next-generation robotic systems in industrial, service, and healthcare applications. However, achieving a reliable system for these tasks has been historically challenging. Unlike rigid objects, the infinite-dimensional state space, severe self-occlusions, and complex dynamics of DOs present significant barriers to robotic perception, modeling, and manipulation. The development of data-driven learning and the recent emergence of foundation models have enabled novel techniques for deformable object manipulation (DOM). These techniques, based on data-driven paradigms, can address some of the challenges that analytical approaches in DOM face. However, some existing reviews do not include all aspects of DOM, and some previous reviews do not summarize data-driven approaches adequately. In this article, we survey more than 150 relevant studies and summarize recent advances, open challenges, and new frontiers in perception, modeling, and manipulation of DOs. We regard research in DOM as paving the way toward general-purpose robotic systems, and we outline key future directions for generalizable robotic manipulation. Specifically, we advocate for the algorithmic synergy of vision-language-action models with reinforcement learning and world models, a universal high-DoF dexterous hand with tactile sensors, and comprehensive evaluation at scale.
Eco-driving and energy management strategy (EMS) significantly influence the real-world performance of hybrid electric vehicles (HEVs). Traditional approaches often optimize eco-driving and energy management strategies separately using layered optimization, which frequently results in suboptimal outcomes. Moreover, transferring eco-driving and energy management strategies between heterogeneous HEVs remains challenging. To address these issues, this paper proposes a collaborative optimization strategy for eco-driving and EMS leveraging transfer learning (TL) and deep reinforcement learning (DRL). The proposed framework not only enhances optimization performance but also improves the generalizability of reinforcement learning methods across different HEV types. Firstly, models of heterogeneous hybrid electric vehicles (HEVs) and traffic signalized lights were constructed. Secondly, an eco-driving and energy management agent for fuel cell hybrid electric vehicles (FCHEVs) was constructed through the utilization of the soft actor-critic (SAC) algorithm. Knowledge transfer via TL was then employed to establish the TL-SAC eco-driving and EMS agent for power-split hybrid electric vehicles (PSHEVs). Compared to conventional SAC-based method, the proposed TL-SAC strategy achieved equivalent fuel reductions of 16.10 %, 10.57 %, and 4.76 % under multi-intersection scenarios. Compared with the PMP algorithm, the proposed TL-SAC strategy reduces energy consumption of PSHEV by 11.88 %, 3.73 % and 4.16 % in three generalization testing scenarios respectively, while achieving close travel time and better real-time performance. This study provides a novel and transferable approach to co-optimize eco-driving and EMS, paving the way for broader deployment of diverse HEVs.
Object detection in low-light environments remains a significant challenge, as the intrinsic noise and severe information loss in intensity-based RGB images compromise feature distinctiveness and limit the robustness of existing single-modal detectors. Conventional cross-modal solutions, such as RGB-Event fusion, introduce practical constraints due to their reliance on expensive, specialized hardware. To address these limitations, we propose the Representation-enhanced Pseudo-event Interactive Detector (RPE-IDet), a novel and practical framework that achieves robust complementary fusion using only single RGB inputs. Our primary innovation is a dedicated Gaussian optical flow-based algorithm to computationally synthesize a “pseudo-event” representation, thereby unlocking the benefits of motion-related information without additional sensors. Furthermore, we design the complementary representation interactive fusion module (CRIFM), which pioneers a deep, two-stage interaction fusion strategy. CRIFM first enhances intra-representation quality via dual-domain feature recalibration (DDFR) and then achieves maximum synergy through the dual-branch interactive fusion (DBIF) module, which implements sophisticated channel and spatial cross-injection mechanisms, i.e., DERA and DEMA. Extensive experiments across three challenging low-light benchmarks demonstrate the superior performance of RPE-IDet. Under the RGB and pseudo-event complementary fusion setting, our method achieves 73.5% mAP on the TunnelDefect dataset, surpassing state-of-the-art cross-modal fusion detectors by a margin of 2.6%, while maintaining high computational efficiency (7.5M parameters, 17.2 GFLOPs). The results robustly validate the efficacy and practicality of our complementary representation fusion paradigm for real-world low-light applications.
Accurate prediction of pedestrian trajectories is crucial for safe motion planning of autonomous vehicles in urban environments. Many existing temporal and generative models are constrained by the long-tailed distribution of the data, which limits their ability to handle random or irregular pedestrian movements. Moreover, few studies have addressed the problem of scoring and probabilistic evaluation of predicted trajectories, despite their importance for downstream decision-making tasks. To address these issues, we propose a regularization-randomization network (R-RNet). The core regularization-randomization (R-R) module enables flexible trajectory prediction across diverse scenarios, while the probability predictor provides trajectory scoring and probability estimation to enhance reliability and utility in subsequent tasks. Besides, a self-attention mechanism is utilized to enhance prediction performance by capturing features from the distribution of the goals. The experimental results on the ETH and the UCY datasets show that R-RNet is capable of making reliable evaluations on output trajectories and achieves competitive results in terms of average displacement error and miss rate, while maintaining a lightweight architecture. Extensive experiments and analyses underscore the critical importance of both regularization and randomization operations. The source code is released at https://github.com/RoboSafe-Lab/R-RNet
This paper investigates the potential of Vision-Language Models (VLMs) to enhance Human-Vehicle Interaction (HVI) in Autonomous Driving (AD) scenarios, particularly in interactions between vehicles and other traffic participants, with a focus on rationality and safety in external HVI. Leveraging recent advancements in large language models, VLMs demonstrate remarkable capabilities in understanding real-world contexts and generating significant interest in HVI applications. This paper provides an overview of AD, HVI, and VLMs, along with the historical context of large language model applications in HVI. The HVI discussed herein involves dynamic game processes encompassing perception and decision-making between vehicles and traffic participants, such as pedestrians. Furthermore, we examine the perceptual challenges associated with applying VLMs to HVI and compile relevant datasets. This research fills a gap in the existing literature by systematically analyzing the current status, challenges, and future opportunities of VLM applications in HVI. To advance VLM integration in AD, various implementation strategies are discussed. The findings highlight the potential of VLMs to transform HVI in AD, improving both passenger experience and driving safety. Overall, this study contributes to a comprehensive understanding of VLM applications in HVI and provides insights to guide future research and development.
The integration of large language models (LLMs) with embodied agents has improved high-level reasoning capabilities; however, a critical gap remains between semantic understanding and physical execution. While vision-language-action (VLA) and vision-language-navigation (VLN) systems enable robots to perform manipulation and navigation tasks from natural language instructions, they still struggle with long-horizon sequential and temporally structured tasks. Existing frameworks typically adopt modular pipelines for data collection, skill training, and policy deployment, resulting in high costs in experimental validation and policy optimization. To address these limitations, we propose ROSClaw, an agent framework for heterogeneous robots that integrates policy learning and task execution within a unified vision-language model (VLM) controller. The framework leverages e-URDF representations of heterogeneous robots as physical constraints to construct a sim-to-real topological mapping, enabling real-time access to the physical states of both simulated and real-world agents. We further incorporate a data collection and state accumulation mechanism that stores robot states, multimodal observations, and execution trajectories during real-world execution, enabling subsequent iterative policy optimization. During deployment, a unified agent maintains semantic continuity between reasoning and execution, and dynamically assigns task-specific control to different agents, thereby improving robustness in multi-policy execution. By establishing an autonomous closed-loop framework, ROSClaw minimizes the reliance on robot-specific development workflows. The framework supports hardware-level validation, automated generation of SDK-level control programs, and tool-based execution, enabling rapid cross-platform transfer and continual improvement of robotic skills. Ours project page: https://www.rosclaw.io/.
Multistage cable routing requires a robot to successfully navigate a cable through a series of clips, which is a challenging task. Due to the unpredictability of cable deformation and the complexity of aligning cables with clip openings, traditional model-based or imitation learning methods are difficult to achieve satisfactory results. To address the difficulty, we propose a novel reinforcement learning-based routing framework guided by a vision-language model (VLM). The core of the proposed framework is to establish the description of the spatial position constraints between cables and clips in VLMs to produce a high-level plan that incorporates grasping and routing points, which then guides the local routing strategy. The local routing strategy is trained by human-in-the-loop reinforcement learning, enabling the model to learn a strategy that reduces deformation when interacting with clips. Our method is validated on a Franka robot across multistage cable routing tasks, and it outperforms some baselines in terms of reliability and generalization.
Establishing trustworthy safety assurance for autonomous driving systems (ADSs) requires evidence that failures arise from avoidable system deficiencies rather than unavoidable traffic conflicts. Current adversarial simulation methods can efficiently expose collisions, but generally lack mechanisms to distinguish these fundamentally different failure modes. Here we present CARS (Context-Aware, Responsibility-attributed Scenario generation), a framework that integrates responsibility attribution directly into adversarial scenario generation. CARS combines context-aware adversary selection with a generative adversarial policy optimized in closed-loop simulation to construct collision scenarios that are both physically feasible and diagnostically attributable. Across benchmark datasets spanning heterogeneous national traffic environments, CARS consistently discovers feasible collision scenarios with high attribution rates under multiple regulation-prescribed careful and competent driver models. By coupling adversarial generation with normative responsibility assessment, CARS moves simulation testing beyond collision discovery toward the construction of interpretable, regulation-aligned safety evidence for scalable ADS validation.
Driver intention prediction has the potential to greatly improve the ability of autonomous vehicles (AVs) to effectively handle risky driving behaviors, thereby ensuring driving safety. Conventional data-driven approaches for driver intention prediction models typically involve gathering extensive driver-related data, which raises significant privacy concerns. With the development of the Internet of Vehicles (IoV), federated learning (FL) has emerged as a prominent privacy-preserving learning paradigm, garnering considerable attention. However, FL encounters challenges in driver intention prediction due to the heterogeneity of driver client data and the limited computational resources of vehicles. To address these challenges, this paper proposes the FedPMR framework, comprising a computationally efficient model for predicting driver intentions. Moreover, to tackle the problem of data heterogeneity, it leverages personalized mixture representation to provide a personalized model adapted to the local data distribution of each driver client. We conducted extensive experiments on the Brain4Cars dataset, achieving an F1-score of 95.24% and a comprehensive evaluation metric of 0.9663, exceeding state-of-the-art. The experimental results demonstrate that the proposed FedPMR effectively addresses the challenges encountered when applying FL to driver intention prediction.
This paper considers the distributed cooperative control of physically interconnected multi-agent networks in the presence of nonlinear and unknown uncertainties that is one of the key problems in robot swarms and autonomous vehicle fleet. A distributed adaptive control framework is established, under which edge-based adaptive algorithms and node-based adaptive algorithms are designed, respectively. In the designed protocols, a linear term is used to force the states of all agents to be the same and a nonlinear term inspired by the sliding-mode control is introduced to suppress the effect of the undesirable uncertainties. Compared with the exiting related works, the algorithms presented here are applicable to physically interconnected systems, and they are scalable without requiring any global information and robust to various uncertainties. The results are also extended to the case of non-matching uncertainties. The effectiveness of the protocols is verified by an example of multiple coupled vehicles. demonstrating their potential to enhance the robustness and adaptability of robotic systems in dynamic and complex operational environments. Note to Practitioners-This research addresses the challenge of achieving consensus in physically interconnected multi-agent networks, such as autonomous vehicle fleets or robotic swarms, particularly in environments with uncertainties. Traditional methods often require global information, which can be impractical. Our study introduces a distributed adaptive control framework that allows these systems to operate effectively without global data. This framework uses edge-based and node-based adaptive algorithms to enhance system stability and adaptability. In this paper, we improve reliability and efficiency in coordinating multi-agent systems. For example, autonomous vehicles can maintain formation and synchronize movements despite environmental unpredictability. Similarly, robotic swarms can adapt to changing conditions during tasks like search and rescue. The fully distributed nature of our algorithms reduces computational and communication overhead, making them scalable for large-scale applications. Future work will explore extending these protocols to directed graphs and integrating event-triggered communication to further optimize performance.
This article introduces an innovative motion planning algorithm for autonomous mobile robots, specifically focusing on quadrotor unmanned aerial vehicles (UAVs), utilizing a gradient descent-enhanced frontend and backend architecture. A trajectory planning algorithm is proposed for the front-end part. It relies on backend optimization feedback and memorized jump points. The algorithm builds on the jump point search (JPS) algorithm and introduces an obstacle table and jump point table. A new heuristic function is proposed, which emphasizes the weight of obstacle proportion in order to avoid getting stuck in local optimal paths. In the backend trajectory optimization part, a backend space-time trajectory optimization method based on gradient descent is proposed, and an optimization objective function is designed to ensure the smoothness and safety of the UAV trajectory. The simulation results show that the algorithm proposed in this article has significant advantages for improving real-time performance and environmental adaptability compared with the method based on ESDF and the EGO-planner. The actual flight experiments show that the proposed algorithm can avoid UAVs getting stuck in local optima during path planning. Notably, the proposed methodology also holds promise for application in path planning for other autonomous robots.
Object detection is widely applied in various fields, including intelligent surveillance, intelligent urban systems, and environmental perception for Advanced Driver Assistance Systems (ADAS). However, in road scenes, the presence of numerous objects and mutual occlusions leads to challenges in feature extraction and reduces detection accuracy for occluded objects. This paper proposes an occlusion-aware object detection algorithm based on YOLO11 to address these issues. Specifically, the input image is segmented into superpixels, and corresponding reactors are generated for each superpixel. A KNN-based adjacency graph is then constructed to extract graph features. Furthermore, we introduce an occlusion attention mechanism to guide the superpixel graph in focusing on occluded regions and computing occlusion loss. Our approach is evaluated on the KITTI dataset. Experimental results demonstrate that, compared to the state-of-the-art detection algorithm YOLO11, our method achieves a 10.28% improvement in mAP50, effectively enhancing detection performance under occlusion scenarios.
Intelligent driving aims to handle dynamic driving tasks in complex environments, while driver behavior onboard is less focused. In contrast, an intelligent cockpit mainly focuses on interacting with a driver, with limited connection to the driving scenarios. Since the driver onboard could affect the driving strategy significantly and thus have nonnegligible safety implications on an autonomous vehicle, a cockpit-driving integration (CDI) is generally essential to take the driver’s behavior and intention into account when shaping the driving strategy. However, no comprehensive review of current existing CDI technologies is conducted despite the significant role of CDI in safe driving. Therefore, we are motivated to summarize the state-of-the-art of CDI methods and investigate the development trends of CDI. To this end, we identify thoroughly current applications of CDI for the perception and decision-making of autonomous vehicles and highlight critical issues that urgently need to be addressed. Additionally, we propose a lifelong learning framework based on evolvable neural networks as solutions for future CDI. Finally, challenges and future work are discussed. The work provides useful insights for developers regarding designing safe and human-centric autonomous vehicles.
With the increasing popularity of unmanned systems in various fields, especially the proposal and application of "embodied intelligence", the related intelligent algorithms and methods have been constantly updated and iterated. However, the application scenarios of these methods becoming more and more complex, which makes the real experimental verification process corresponding to these algorithms more and more difficult. It will possibly result in higher time and monetary costs. With the continuous development of information physics systems and the emergence of digital twins technology, simulation verification methods based on virtual reality fusion are constantly being proposed. Most of them focus on model building and symmetric mapping. In this paper, we propose a dual machine mapping method to achieve the mapping of physical scenes to virtual scenes of different spatial scales. Through experiments, it has been verified that the platform has stable mapping performance under indoor motion capture system positioning and can successfully utilize information from virtual environments to train path planning algorithms.
Coverage control is a foundational domain within multi-agent systems, which has recently undergone substantial advancements. Model-based and optimization-based control methods have achieved new breakthroughs and applications, while data-driven and learning-based algorithms in coverage control have also yielded significant results. This review examines the evolution and application of coverage control algorithms in multi-agent systems. Moreover, this review then focuses on how contemporary algorithms build on traditional control strategies and integrate cutting-edge data-driven and adaptive learning techniques. Finally, it explores potential future directions, emphasizing the importance of interdisciplinary approaches in overcoming existing challenges and seizing new opportunities in coverage control deployment. This review aims to be a valuable resource for researchers and practitioners, guiding continued exploration in this field.
With the increasing demand for energy conservation and emission reduction, the fuel cell hybrid electric vehicle (FCHEV), as a new energy vehicle that includes two power sources-the fuel cell pack and the battery pack-will play an important role in future urban transportation. This article focuses on the car-following problem and studies the car-following optimization control strategy of FCHEV. A speed planning module based on model predictive control (MPC) and an energy management module based on adaptive soft actor critic (SAC) algorithm are designed to improve driving comfort and optimize vehicle's energy consumption while ensuring safe car-following. Finally, simulation verification was conducted under the car-following conditions of Urban Dynamometer Driving Schedule (UDDS) and China light-duty vehicle test cycle (CLTC). Simulation results showed that under the car-following optimization control strategy proposed in this paper, the FCHEV can simultaneously balance safety car-following and driving comfort, achieving almost 95.18% and 91.02% of hydrogen economy of dynamic programming (DP) in UDDS and CLTC car-following conditions, respectively, while reducing hydrogen consumption by 38.05%, and 30.87% compared with the alternating direction method of multipliers (ADMMs). The real-time performance of the proposed car-following optimization strategy is compared and validated under typical car-following driving conditions as well.
Adversarial example generation is crucial for accelerating the testing of autonomous vehicles (AVs). With the introduction of end-to-end AVs, adversarial images are becoming more important than adversarial behavior attacks, as perception and decision-making no longer have explicit information transmission. Generative AI-based methods show promising results when generating new images. However, they are usually hardly controllable and struggle with semantic consistency in generated images. In this paper, we focus on lane detection as an example and integrate a control network along with a bootstrapped language-image pretraining model into the stable diffusion framework to enhance the controllability and realism of generated adversarial examples. Leveraging the PBLMD dataset, we successfully generate diverse adversarial examples such as occluded and worn lane lines. Experimental results demonstrate that these examples reduce the lane detection performance of YOLOP by an average of 11% while maintaining high fidelity, demonstrating the approach's efficacy.
In the domain of autonomous vehicles, the human-vehicle co-pilot system has garnered significant research attention. To address the subjective uncertainties in driver state and interaction behaviors, which are pivotal to the safety of Human-in-the-loop co-driving systems, we introduce a novel visual-tactile detection method. Utilizing a driving simulation platform, a comprehensive dataset has been developed that encompasses multi-modal data under fatigue and distraction conditions. The experimental setup integrates driving simulation with signal acquisition, yielding 600 minutes of driver state and behavior data from 15 subjects and 102 takeover experiments with 17 drivers. The dataset, synchronized across modalities, serves as a robust resource for advancing cross-modal driver behavior detection algorithms.
Driver models are crucial for the safety assessment of autonomous vehicles (AVs) because of their role as reference models. Specifically, an AV is expected to achieve at least the same level of safety performance as a careful and competent driver model. To make this comparison possible, quantitative modeling of careful and competent driver models is essential. Thus, the UNECE Regulation No. 157 proposes two driver models as benchmarks for AVs, enabling safety assessment of AV longitudinal behaviors. However, these two driver models are unable to be applied in non-car-following scenarios, limiting their applications in scenarios such as highway merging. To this end, we propose a careful and competent driver model for highway merging (CCDM2) scenarios using interpretable reinforcement learning-based decision-making and safety constraint control. We compare our model’s safe driving capabilities with human drivers in challenging merging scenarios and demonstrate the ”careful” and ”competent” characteristics of our model while ensuring its interpretability. The results indicate the model’s capability to handle merging scenarios with even better safety performance than human drivers. This model is of great value for AV safety assessment in merging scenarios and contributes to future reference driver models to be included in AV safety regulations.