Detecting Electric Vehicle Charging Stations (EVCS) is attracting increasing attention for autonomous and assisted driving of electric vehicles. One of the main challenges is the scarcity of EVCS detection data. Thus, we propose a camera-based EVCS detection dataset. The dataset is composed of two parts: The first part contains images labeled at a fine-grained level of categories with eight classes. The second part contains images of 13 additional EVCS types annotated at a supercategory level as “electric vehicle charging station”. The images are annotated with bounding boxes and object masks, together with a visibility level for each EVCS instance. For evaluation, protocols considering both fine-grained and supercategory EVCS detection, including fine-tuning, prompt tuning, and zero-shot detection are proposed. Four baseline methods, including both closed-set and open-set detectors, are evaluated. Our evaluation reveals that with our dataset, closed-set detectors can be trained with a reasonable performance. Also, it shows that tuning the open-set detector to work with EVCS at a fine-grained level while preserving its ability to detect common objects forms an interesting research direction. This dataset is the first dataset for camera-based electric vehicle charging station detection. The dataset is accessible here: https://evcs.viscoda.com .
The current road condition is a crucial factor regarding road safety of the ego-vehicle and other road users. Road condition estimation provides essential input data for friction estimation which is used for autonomous and automated driving systems. Camera-based approaches are still far from being practical and other sensors dominate the field of friction estimation. This is due to the limited performance of current approaches and the lack of datasets for the incorporation of learning-based methods.We propose a novel dataset for a special scenario of road condition, the coverage with snow. It is the first large-scale dataset for camera-based road classification of snow-covered roads with different types of snow coverage. The dataset consists of road patches in bird's eye view perspective and ground truth annotation for the current snow coverage type. It is combinable with RoadSaW [4], a dataset for road surface and wetness estimation, leading to a holistic road condition dataset with 15 categories. The baseline evaluation employs state-of-the-art, real-time capable approaches for classification and uncertainty estimation with RBF (Radial Basis Function) networks. Our experiments demonstrate that the proposed data opens new challenges in the field of camera-based road condition estimation.
Structure-from-motion (SfM) is still the preferred solution for 3D perception of rigid geometry in vehicles using monocular cameras. SfM provides the motion estimation of the ego-vehicle and the localization of obstacles. For the computation of 3D structure from monocular video, camera motion is required. The motion is provided by the movement of the vehicle. To avoid accidents at the beginning of the drive, the obstacle detection is needed before the vehicle starts to move, especially for autonomous vehicles. The idea followed in this paper is the execution of a small vehicle maneuver (micro maneuver) to compute the surrounding scene using SfM. The maneuver leads to the desired camera motion and is executed before starting to drive without any driver interaction. Then, obstacles are detected and considered for any driving action. Two micro maneuver types executed on a standing vehicle are under consideration: (a) steering the front wheels and (b) using the vehicle's handbrake and engine torque. We analyze the generated camera motion and the reconstructed scene. Since the resulting camera motion is small, state of the art keyframe selection techniques are compared. The application of obstacle detection using the 3D point cloud showcases the accuracy of the reconstructed scene. Based on the experiments, the most promising methodology is identified.
Automated driving is one of the most promising technologies for improving road safety. In real driving scenarios, knowledge about the road friction is crucial. For the estimation of the road friction, two properties are of main interest: the road surface type and the road condition. We propose a novel large-scale dataset to enable camera-based road surface and wetness estimation. It consists of video data captured by in-vehicle cameras and ground truth for the current surface type and wetness which is determined by the MARWIS (Mobile Advanced Road Weather Information Sensor). The wetness measurements are associated to high-resolution bird’s eye view road image patches, derived from a calibrated sensor setup. Additionally, data for different distances to the vehicle is provided. The dataset is evaluated with state-of-the-art real-time capable approaches for road condition classification and uncertainty estimation. The results provide a valid baseline, but also demonstrate limitations of the generalization performance. The dataset enables new possibilities for future research on camera-based road friction estimation. It is the first dataset including accurate measurements for the wetness in real driving scenarios.
This chapter provides an overview of state-of-art radio- and image-based positioning tailored to the localization of connected autonomous vehicles. It evaluates the achievable accuracy of video-based vehicle positioning. Therefore, the accuracy achieved in the 5GCAR Lane Merge Coordination use case, using a set of monocular cameras for vehicle positioning, is compared to the accuracy achieved in literature. The chapter addresses the technology and performance requirements of legacy solutions along with the details of time-based, angular-based, and video-based positioning. Some of the proposed methods closely follow the long term evolution and the new radio standards; others show more innovative solutions that can achieve a high level of accuracy. Video-based positioning requires additional infrastructure (camera installation), but the camera locations are flexible. In the 5GCAR project, a mast with 10 m height is chosen for installation. Finally, the chapter concludes with a comprehensive set of simulations highlighting the pros and cons of each solution.
Automated driving is regarded as the most promising technology for improving road safety in the future. In this context, connected vehicles have an important role regarding their ability to perform cooperative maneuvers for challenging traffic situations. We propose a benchmark for automated cooperative maneuvers. The targeted cooperative maneuver is the vehicle lane merge where a vehicle on the acceleration lane merges into the traffic of a motorway. The benchmark enables the evaluation of vehicle localization approaches as well as the study of cooperative maneuvers. It consists of temporally synchronized multi-view video streams, highly accurate camera calibration, and ground truth vehicle descriptions, including position, heading, speed, and shape. For benchmark generation, the lane merge maneuver is performed by human drivers on a test track, resulting in 85 lane merge data sets with various traffic situations and video recording conditions.
Future 5G systems have set a goal to support mission-critical Vehicle-to-Everything (V2X) communications and they contribute to an important step towards connected and automated driving. To achieve this goal, the communication technologies should be designed based on a solid understanding of the new V2X applications and the related requirements and challenges. In this regard, we provide a description of the main V2X application categories and their representative use cases selected based on an analysis of the future needs of cooperative and automated driving. We also present a methodology on how to derive the network related requirements from the automotive specific requirements. The methodology can be used to analyze the key requirements of both existing and future V2X use cases.
Cooperative maneuvers are of high interest within many V2X applications. The implementation of cooperative maneuvers require the accurate localization of the vehicles. Accurate localizations of the ego-vehicle will be provided by the next generation of connected cars using 5G. Until all cars participate in the network, unconnected cars have to be considered as well. These cars are localized via static cameras positioned next to the road. The scope of this paper is the implementation and evaluation of a system which provides the detection, tracking, and localization of vehicles for a cooperative maneuvers scenario. The application is the lane merge of vehicles where the vehicle localizations are used for the planning of trajectories. The observed vehicles are equipped with GNSS RTK units for their self-localization which is the basis for the accuracy evaluation of the localization provided by the camera system.
Cooperative and connected V2X applications drive further improvements of advanced driver assistance systems (ADAS) and automated driving. Three different use cases showing new applications are demonstrated within the 5GCAR project: lane merge coordination, cooperative perception for maneuvers of connected vehicles, and vulnerable road user protection. These use case representatives on the one hand will be demonstrated mid-2019 on a test track, and on the other hand serve as platforms for evaluating key performance indicators (KPIs) of the developed implementations. Keywords—Connected car; Automated driving; 5G; V2X communication; Cooperative vehicular applications;
For the trajectory planning in autonomous driving, the accurate localization of the vehicles is required. Accurate localizations of the ego-vehicle will be provided by the next generation of connected cars using 5G. Until all cars participate in the network, un-connected cars have to be considered as well. These cars are localized via static cameras positioned next to the road. To achieve high accuracy in the vehicle localization, the highly accurate calibration of the cameras is required. Accurately measured landmarks as well as a priori knowledge about the camera configuration are used to develop the proposed constrained multi camera calibration technique. The reprojection error for all cameras is minimized using a differential evolution (DE) optimization strategy. Evaluations on data recorded on a test track show that the proposed calibration technique provides adequate calibration accuracy while the accuracies of reference implementations are insufficient.
Motion segmentation is the task of classifying the feature trajectories in an image sequence to different motions. Hypergraph based approaches use a specific graph to incorporate higher order similarities for the estimation of motion clusters. They follow the concept of hypothesis generation and validation. For the sampling of hypotheses, a high probability of selecting clean samples, i.e. samples consisting of points from the same cluster, is desired. Many approaches use spatial proximity to build an auxiliary graph for the sampling. But, spatial proximity is often not sufficient to capture the main affinities for motion segmentation. Thus, we introduce a simple but effective model for incorporating motion-coherent affinities into the auxiliary graph. The evaluation on two state of the art benchmarks shows that the hypotheses generated from the resulting hypergraphs lead to a significant decrease of the segmentation error. Additionally, less computation time is required due to a reduced hypergraph complexity.
The extraction of scale invariant image features is a fundamental task for many computer vision applications. Features are localized in the scale space of the image. A descriptor is build for each feature which is used to determine the correspondence to a second feature, usually extracted from a second image. For the evaluation of detectors and descriptors, benchmark image sets are used. The benchmarks consist of image sequences and homographies which determine the ground truth for the mapping between the images. The repeatability criterion evaluates the detection accuracy of the detectors while precision and recall measure the quality of the descriptors.Current data sets provide images with resolutions of less than one megapixel. A recent data set provides challenging images and highly accurate homographies. It allows for the evaluation at different image resolutions with the same scene content. Thus, the scale invariant properties of the extracted features can be examined. This paper presents a comprehensive evaluation of state of the art detectors and descriptors on this data set. The results show significant differences compared to the standard benchmark. Furthermore, it is shown that some detectors perform differently on different resolutions. It follows that high resolution images should be considered for future feature evaluations.
Sparse bundle adjustment (SBA) is the state of the art method for simultaneously optimizing a set of camera poses and 3D points. The multibody bundle adjustment optimizes the static scene and the moving rigid object(s). The result is one camera path representing the main camera motion and virtual camera path(s) for each of the independently moving objects in the scene. The bundle adjustment for the multibody problem is performed in a joint optimization. Main motion (static scene) and object motion(s) are included in the optimization such that the sparse algorithm of SBA can still be applied, even when enforcing the constraint that main camera and object camera share the same intrinsic parameters. The joint optimization approach enables weighting the resulting error for each of the motion models and therefore influences the optimization process. Our experiments with synthetic and natural image data show that an appropriate weighting leads to more accurate camera parameters.
The scale invariant feature operator (SFOP) detects circular features from an image using a spiral shape model. Special cases of the spiral model are junctions and circular symmetric shapes. The spatial localization is determined with subpixel accuracy which is obtained by an interpolation of the structure tensor in the scale space. For the interpolation, SFOP uses a 3D quadratic function. This leads to suboptimal solutions since the structure tensor surrounding a feature does not show the shape of a 3D quadratic. The aim of this paper is to improve the localization of the features detected by SFOP. A Difference of Gaussians function is proposed for the signal approximation which leads to improved precision values and to more accurate features. The proposed method improves the localization such that 72.5% of the features increase their precision. Hence, more features are extracted while increasing their repeatability by up to 9% on standard benchmarks.
The detection of scale invariant image features is a fundamental task for computer vision applications like object recognition or re-identification. Features are localized by computing extrema of the gradients in the Laplacian of Gaussian (LoG) scale space. The most popular detector for scale invariant features is the SIFT detector which uses the Difference of Gaussians (DoG) pyramid as an approximation of the LoG. Recently, the alternative interest point (ALP) detector demonstrated its strength in fast computation on highly parallel architectures like the GPU. It uses the LoG scale space representation for the localization of interest points. This paper evaluates the localization accuracy of ALP in comparison to SIFT. By using synthetic images, it is demonstrated that both localization approaches show a systematic error which is dependent on the subpixel position of the feature. The error increases with the scale of the detected feature. However, using the LoG instead of the DoG representation reduces the maximum systematic error by 77 %. For the evaluation with natural images, benchmark data sets are used. The repeatability criterion evaluates the accuracy of the detectors. The LoG based detector results in up to 16 % higher repeatability. The comparisons are completed with a reference feature localization which uses a signal based approach for the gradient approximation. Based on this approach, a new feature selection criterion is proposed.
In recent decades, cold atom experiments have become increasingly complex. While computers control most parameters, optimization is mostly done manually. This is a time-consuming task for a high-dimensional parameter space with unknown correlations. Here we automate this process using a genetic algorithm based on differential evolution. We demonstrate that this algorithm optimizes 21 correlated parameters and that it is robust against local maxima and experimental noise. The algorithm is flexible and easy to implement. Thus, the presented scheme can be applied to a wide range of experimental optimization tasks.
Visual effect creation as used in movie production often require structure and motion recovery and video segmentation. Both techniques are essential to integrate virtual objects between scene elements. In this paper, a new method for video segmentation is presented. It incorporates 3D scene information from the structure and motion recovery. By connecting and evaluating discontinued feature tracks, occlusion and reappearance information is obtained during sequential camera and scene estimation. The foreground is characterized as image regions which temporarily occlude the rigid scene structure. The scene structure is represented by reconstructed object points. Their projections onto the camera images provide the cues for regions classified as foreground or background. The knowledge of occluded parts of a connected feature track is used to feed the object segmentation which crops the foreground image regions automatically. Two applications are presented: the occlusion of integrated virtual objects and the blurred background effect. Several demonstrations on official and self-made data show very realistic results in augmented reality.
Benchmark data sets consisting of image pairs and ground truth homographies are used for evaluating fundamental computer vision challenges, such as the detection of image features. The mostly used benchmark provides data with only low resolution images. This paper presents an evaluation benchmark consisting of high resolution images of up to 8 megapixels and highly accurate homographies. State of the art feature detection approaches are evaluated using the new benchmark data. It is shown that existing approaches perform differently on the high resolution data compared to the same images with lower resolution.
The segmentation of foreground objects in camera images is a fundamental step in many computer vision applications. For visual effect creation, the foreground segmentation is required for the integration of virtual objects between scene elements. On the other hand, camera and scene estimation is needed to integrate the objects perspectively correct into the video.
This paper demonstrates how to effectively exploit occlusion and reappearance information of feature points in structure and motion recovery from video. Due to temporary occlusion with foreground objects, feature tracks discontinue. If these features reappear after their occlusion, they are connected to the correct previously discontinued trajectory during sequential camera and scene estimation. The combination of optical flow for features in consecutive frames and SIFT matching for the wide baseline feature connection provides accurate and stable feature tracking. The knowledge of occluded parts of a connected feature track is used to feed a segmentation algorithm which crops the foreground image regions automatically. The resulting segmentation provides an important step in scene understanding which eases integration of virtual objects into video significantly. The presented approach enables the automatic occlusion of integrated virtual objects with foreground regions of the video. Demonstrations show very realistic results in augmented reality.