Self-supervised multi-frame methods have currently achieved promising results in depth estimation. However, these methods often suffer from mismatch problems due to the moving objects, which break the static assumption. Additionally, unfairness can occur when calculating photometric errors in high-freq or low-texture regions of the images. To address these issues, existing approaches use additional semantic priori black-box networks to separate moving objects and improve the model only at the loss level. Therefore, we propose FlowDepth, where a Dynamic Motion Flow Module (DMFM) decouples the optical flow by a mechanism-based approach and warps the dynamic regions thus solving the mismatch problem. For the unfairness of photometric errors caused by high-freq and low-texture regions, we use Depth-Cue-Aware Blur (DCABlur) and Cost-Volume sparsity loss respectively at the input and the loss level to solve the problem. Experimental results on the KITTI and Cityscapes datasets show that our method outperforms the state-of-the-art methods.
With the rapid development of automated terminals, it has greatly contributed to the rapid economic growth. However, the rapid development of terminals has led to a rapid increase in the data that needs to be stored. Storing all the massive terminal data without compression requires a huge amount of storage space, while online transmission consumes a large amount of network bandwidth and puts a lot of pressure on the data transmission lines. In practice, the collected sequence data are not always variable and valuable, and may contain a lot of redundant data, which is detrimental to the analysis efficiency and storage space of the data. Meanwhile, for data with different attributes, the compression methods adopted should be different. To address such issues, this paper proposes a fusion lossy and lossless data compression method based on signal attributes, where the lossy compression algorithm adopts compression perception and the lossless compression algorithm adopts LZW coding. For important data, only the lossless compression method is adopted to prevent the loss of important information; for general data, the fusion lossy and lossless compression method is adopted to improve the compression rate.
Driving scene understanding is to obtain comprehensive scene information through the sensor data and provide a basis for downstream tasks, which is indispensable for the safety of self-driving vehicles. Specific perception tasks, such as object detection and scene graph generation, are commonly used. However, the results of these tasks are only equivalent to the characterization of sampling from high-dimensional scene features, which are not sufficient to represent the scenario. In addition, the goal of perception tasks is inconsistent with human driving that just focuses on what may affect the ego-trajectory. Therefore, we propose an end-to-end Interpretable Implicit Driving Scene Understanding (II-DSU) model to extract implicit high-dimensional scene features as scene understanding results guided by a planning module and to validate the plausibility of scene understanding using auxiliary perception tasks for visualization. Experimental results on CARLA benchmarks show that our approach achieves the new state-of-the-art and is able to obtain scene features that embody richer scene information relevant to driving, enabling superior performance of the downstream planning.
Semi-supervised semantic segmentation aims to maximize the training performance for a limited annotation cost. Existing methods such as cross pseudo supervision have shown excellent performance, yet ignore potential information interactions between labeled and unlabeled data, and suffer from misleading incorrect pseudo labels. This paper takes two ways to improve each of these shortcomings. Firstly, we perform feature-level mixing and cross-decoupling using labeled and unlabeled data to establish potential interactions between the two types of data. Secondly, an uncertainty-aware loss re-weighting method based on information entropy is used to mitigate the negative effects of incorrect pseudo labels. Experimentally, our method further improves the previous cross pseudo supervision method with competitive performance on PASCAL VOC 2012 dataset under various data partition protocols.
LiDARs and RGB cameras are commonly used sensors in autonomous driving vehicles. However, the high-resolution LiDAR is too expensive, limiting its large-scale application in commercial autonomous vehicles. The low-resolution LiDAR is more affordable and it can approximate the perception level of high-resolution LiDAR combined with corresponding images. In this paper, we propose a hierarchical cross-attention Transformer in dual-branch to predict a dense depth map. The hierarchical architecture builds a feature pattern at all scales and the cross-attention modules fuse the features from different modalities at multiple feature levels. Furthermore, we develop a depth refinement stage to amend the dense depth map predicted by the fusion stage. The proposed method is evaluated on the indoor NYUDepthV2 dataset and outdoor KITTI Odometry dataset. The experiments demonstrate its effectiveness and accuracy compared with the present state-of-the-art methods.
Precise localization is crucial for autonomous vehicles. State-of-the-art methods using 3D light detection and ranging (LIDAR) scanner to achieve high-precision. But the features extracted by these methods are not effective in all scenarios. This paper proposes a localization method based on multiple LiDAR features to avoid the possible failure of a single feature in certain environments. In particular, we employ curb, reflection intensity and height LiDAR features. The curb feature is used to constrain the localization results within the road environment, while the reflection intensity feature represents the road marker information, and the height feature records the 3D environmental information surrounding the vehicle. We match real-time sensor measurements with map information by fusing these features, and subsequently combine GPS/IMU data to provide real-time localization information. Experiments show that the multiple feature matching of laser sensors proposed here can effectively improve the accuracy and robustness of the localization of autonomous vehicles under various environments.
Traffic light perception is crucial for autonomous driving. In this paper, a three-step method for traffic-light detection, recognition and understanding is developed to guide self-driving vehicles to pass crossroads. The first step adopts a state-of-the-art deep neural object detection architecture to detect traffic lights. The second step designs a novel four-channel convolutional neural network to classify traffic lights. The last step develops a spatiotemporal trajectory analysis method to filter out false positives and to guide self-driving vehicles. The proposed method is evaluated on two datasets, and experiment results show that it can efficiently perceive traffic-light states and perform better than baseline considered. Furthermore, real vehicle testings are conducted, which demonstrate the effectiveness of the proposed method.
Some machine learning algorithms have shown a better overall recognition rate for facial recognition than humans, provided that the models are trained with massive image databases of human faces. However, it is still a challenge to use existing algorithms to perform localized people search tasks where the recognition must be done in real time, and where only a small face database is accessible. A localized people search is essential to enable robot–human interactions. In this article, we propose a novel adaptive ensemble approach to improve facial recognition rates while maintaining low computational costs, by combining lightweight local binary classifiers with global pre-trained binary classifiers. In this approach, the robot is placed in an ambient intelligence environment that makes it aware of local context changes. Our method addresses the extreme unbalance of false positive results when it is used in local dataset classifications. Furthermore, it reduces the errors caused by affine deformation in face frontalization, and by poor camera focus. Our approach shows a higher recognition rate compared to a pre-trained global classifier using a benchmark database under various resolution images, and demonstrates good efficacy in real-time tasks.
This paper presents a trajectory planning method for autonomous vehicles considering the uncertainty caused by other traffic participants. The presented method employs a two-step architecture. In the first step, Bézier curves based approach generates a set of trajectory candidates. In the second step, an online selector is used to select the most appropriate one. In this step, an uncertainty model considering the intentions of other traffic participants is used to evaluate the collision probability. The uncertainty model generates probabilistic trajectories to express the potential motions of guest vehicles under different intentions. Compared with the previous work, the uncertainty model in our trajectory planning method takes the uncertainty of guest vehicle's intentions into account rather than focusing on the current heading and velocity. Simulation results demonstrate the feasibility and safety of the proposed method in different scenarios.
The effective detection of curbs is fundamental and crucial for the navigation of a self-driving car. This paper presents a real-time curb detection method that automatically segments the road and detects its curbs using a 3D-LiDAR sensor. The point cloud data of the sensor are first processed to distinguish on-road and off-road areas. A sliding-beam method is then proposed to segment the road by using the off-road data. A curb-detection method is finally applied to obtain the position of curbs for each road segments. The proposed method is tested on the data sets acquired from the self-driving car of laboratory of VeCaN at Tongji University. Off-line experiments demonstrate the accuracy and robustness of the proposed method, i.e., the average recall, precision and their harmonic mean are all over 80%. Online experiments demonstrate the real-time capability for autonomous driving as the average processing time for each frame is only around 12 ms.
Stixel based obstacle detection plays an important role in obstacle detection yet with the limitation of assuming a flat road. This assumption would mistake a sloping road in front of the flat one as obstacle. In this paper, we extend the dynamic stixel world with the cubic B-spline based non-flat road estimation in the v-disparity map. The parameters of B-spline are tracked over time using a Kalman filter. We also add color histogram to the stixel for a better representing. Experimental results in KITTI benchmark show the stixel building efficiency of our method in distinguishing between an uprising road and obstacles. Meanwhile, our method achieves an average 3.2% accuracy gain for obstacle detection, and can realize real-time performance about 72 ms per frame on our local computer containing an Intel is-7500T CPU with 2.7 GHz frequency and GTX 1080Ti GPU used for stereo matching.
The rotor flux will be changed to obtain the maximum ratio between torque and stator current,under the maximum toque per ampere (MTPA) criteria.Since flux changes with reference torque,not a constant as in conventional vector control algorithm,nonlinear magnetic saturation effect cannot be ignored.In this paper,a control system considering magnetic saturation effect is proposed.Optimal reference rotor flux is found using linear search algorithm,while the rotor flux is estimated using reduced order observer.The observer-error is taken into consideration,when designing flux controller using sliding mode algorithm and torque controller using backstepping algorithm,to ensure the global stability.Simulation and experimental results show the effectiveness of the proposed algorithm.
It is crucial for autonomous vehicles to navigate at intersections. The accurate location of intersections and the orientation of each branches are necessary for decision making and path planning. In this paper, a unified method is proposed to estimate orientations of each branch at an intersection and to locate the position of the intersection. First, based on vehicle dynamics, a densifying method is used to obtain dense point-cloud data using 3D-LIDAR sensor. Then, according to the data of Open Street Map, the regions-of-interest are extracted and the points are interpolated to transform into an elevation image. Finally, a support vector regression model is employed to estimate the position and orientation of each branch and a fusion method is used to locate the intersection. The experimental results demonstrate the accuracy and robustness of the proposed algorithm.
Rotor flux is an important parameter in vector control of induction motor.The flux is usually estimated from measured three-phase current,which is often being influenced by noise.The accuracy of observer is related with motor parameters.If the flux level varies with motor working points,the nonlinear magnetic saturation effect will cause variation of mutual inductance.Then,the amplitude and position of rotor flux will deviate from real value,deteriorating the control performance.In this paper,observer algorithm for state estimation and parameter identification,based on extended Kalman Filter (EKF) and forth-order polynomial magnetic saturation model,is proposed.The stator current works as feedback,to correct the predicted value from induction motor model.The simulation and experimental results show the effectiveness of the proposed algorithm in reducing the influence of magnetic nonlinear effect on control system.
Hybrid electric vehicles offer potential for reducing the oil consumption and the generation of greenhouse gases, but they have a high battery price and short battery longevity. This paper presents an energy management strategy for power-split hybrid electric vehicles, which seeks not only to reduce the gasoline consumption but also to prolong the battery life. The instantaneous battery usage is penalized by its influence on the battery lifespan. This influence is defined by a new concept of battery ageing, i.e. the battery-fading index, which represents the ageing rate of the battery in various conditions. The multi-objective optimization problem is achieved by a model predictive control framework. The results indicate that the proposed energy management strategy effectively reduces battery fading with only a relatively small increase in the fuel consumption.
Three hierarchical algorithms are proposed to get the passable driving road detection for unmanned ground vehicle (UGV). The innovation of this paper is the self-made directional texture which will be produced if the inverse homography transform is exerted on the image. This self-made texture is firstly used to remove the image corresponding to the objects above road. In order to avoid the influence of other color object, the k-means clustering is used to get the seed for region growing. Finally, the directional texture is secondly used to judge whether the image holes after region growing are the road or not. The final experimental results show the proposed algorithm can remove the non-road objects, can and fill the stained color region and can get the rational driving road.
In a complex environment, simultaneous object recognition and tracking has been one of the challenging topics in computer vision and robotics. Current approaches are usually fragile due to spurious feature matching and local convergence for pose determination. Once a failure happens, these approaches lack a mechanism to recover automatically. In this paper, data-driven unfalsified control is proposed for solving this problem in visual servoing. It recognizes a target through matching image features with a 3-D model and then tracks them through dynamic visual servoing. The features can be falsified or unfalsified by a supervisory mechanism according to their tracking performance. Supervisory visual servoing is repeated until a consensus between the model and the selected features is reached, so that model recognition and object tracking are accomplished. Experiments show the effectiveness and robustness of the proposed algorithm to deal with matching and tracking failures caused by various disturbances, such as fast motion, occlusions, and illumination variation.
环境感知以及导航定位是无人驾驶汽车(以下简称无人车)技术的关键组成部分.针对驾驶环境进行定义和分类,提出与环境相互匹配的传感器组合方法.在此基础上,着重介绍传感器技术以及环境感知技术,比较各技术优缺点,并结合导航与定位对无人车组成架构进行概括介绍,并对未来无人车环境感知技术进行展望.
Environment perception is essential for autonomous driving technology. The curb is a prominent feature of urban roads and therefore is a significant part of environment perception. In this paper, a real-time curb detection and tracking method is proposed for Unmanned Ground Vehicles (UGVs). The proposed curb detection algorithm uses the the surrounding environment data provided by a 3D-LIDAR sensor to extract the curb position based on its spatial features. The curb tracking algorithm is proposed to predict and update the curb position with respect to the current vehicle states in real time. The performance of the proposed method is verified through extensive experiments with a UGV driving on campus roads. The experimental results demonstrate the accuracy and robustness of the proposed method.