Pedestrians are the most vulnerable road users, with around 23% of world road traffic fatalities. To prevent such traffic collisions, the Pedestrian Collision Warning System (PCWS) alerts the driver before an imminent collision. In order to protect worldwide pedestrians, the PCWS should take into account different pedestrian crossing behaviors and different road structures, especially pedestrians with risky behavior on unstructured environments. Since oblique crossing is the usual crossing way of pedestrians with risky behaviors, recognizing the pedestrian walking direction is one of the key factors to consider. However, existing systems focus more on detecting pedestrians rather than recognizing their walking direction. This paper highlights the safety of different kinds of pedestrians by presenting a novel approach used to estimate pedestrian's orientation from a single-frame. Next, the estimated orientation is integrated into the proposed PCWS system. Our method involves Capsule Network based technique trained on pedestrian images. For this purpose, a new pedestrian orientation dataset taken from real city-scenes named SafeRoad was created, using a single camera mounted on a moving vehicle. TUD Multiview Pedestrian and Daimler datasets are thereafter used as a benchmark to evaluate the proposed approach. Experimental results show that Capsule Network exceeds significantly the accuracy performance on the three datasets compared to Convolutional Neural Network algorithms.
One of the main reasons of intersection accidents is related to drivers’ behavior. Differences on age, gender or even personality affected drivers speed and level of aggressiveness which may cause huge problems and fatal crashes in intersections. Those differences assembled with others, such us vehicles type and the approach speed, may lead to different reactions while facing the yellow signal, which cause dilemma zone in signalized intersections. In this paper, we will report different authors views concerning drivers’ behavior in both signalized and unsignalized intersections with a focus on the major predictors of aggressiveness.
Vision systems that provide a 360-degree view are becoming increasingly common in today’s vehicles. These systems are generally composed of several cameras pointing in different directions and rigidly connected to each other. The purpose of these systems is to provide driver assistance in the form of a display, for example by building a Bird’s eye view around the vehicle for parking assistance. In this context, and for reasons of cost and ease of integration, such cameras are generally not synchronized. If non-synchronization is not a problem when it comes to display only, it poses significant issues for more complex computer vision applications (3D reconstruction, motion estimation, etc.). In this article, we propose to use a network of asynchronous cameras to estimate the motion of the vehicle and to find the 3D structure of the scene around it (for example for obstacle detection). Our method relies on the use of at least three images from two adjacent cameras. The poses of the cameras are independently estimated by conventional visual odometry algorithms. Then we show that it is possible to find the absolute scale factor by hypothesizing that the motion of the vehicle is smooth. The results are then refined through a local bundle adjustment on the scale factor and 3D points only. We evaluated our method under real conditions on the KITTI database, and we showed that our method can be generalized to a larger network of cameras thanks to a system developed in our lab.
Thousands of people are dying every year due to road accidents; in fact 23% of world fatal accidents are pedestrians related, where 40% of them occur in Africa as reported by the World Health Organisation (WHO). Predicting the walking direction of a pedestrian could help to avoid an eventual accident. Existing studies can not handle pose and orientation transformations of the input object contrary to our proposed method. This paper describes a novel approach to determine the pedestrian orientation using Capsule Networks (CapsNet) based scheme. CapsNet are a new deep learning architecture that overcome some limitations of the existing studies, they are group of neurons invariant to rotation and affine transformations, which represent a specific interest to this work. Capsule Networks predicts the walking directions of pedestrians to prevent such mortal accidents, using four main walking directions (front, back, left and right).For this purpose, a new pedestrians dataset gathered from the most popular cities in Morocco is collected to be studied and used as a proof of the proposed approach. To enhance this proposed approach, we evaluated it using Daimler dataset and compared it to Convolutional Neural Networks (CNN) architectures. Experimental results reveal that the performance of the proposed approach reaches an accuracy of 97.60% on daimler dataset and 73.64% on our Moroccan collected dataset.
Nous proposons un systeme de vision, base sur un reseau de cameras non-synchronisees permettant d'estimer le mouvement d'un vehicule et la structure de l'environnement 3D a l'echelle absolue. L'algorithme propose necessite au moins trois images prises par au moins deux cameras. Les poses relatives des cameras sont estimees par les methodes classiques d'estimation du mouvement. Les deplacements absolus sont calcules en supposant que les trajectoires entre deux vues consecutives d'une meme camera sont rectilignes et que les cameras sont calibrees hors ligne. Une etape d'optimisation par ajustement de faisceaux est realisee pour affiner l'estimation des facteurs d'echelle et les positions des points 3D. Notre methode est evaluee a partir de donnees reelles issues de la base de donnees KITTI.
In this paper, we present a simple algorithm for obstacle detection, road surface extraction and tracking using Kalman filter and u-v-disparity images. The proposed approach is based on the use of an unsynchronized camera system and the use of sparse maps instead of dense ones due to the unsynchronization constraint. First, a sparse disparity map is computed from two images then the u-v-disparity images are built from it. Road and obstacles are extracted using a modified Hough transform. Our experimental results on real images show the efficiency of our algorithm.
In this paper we present an unsynchronized camera network able to estimate the motion and the structure with accurate absolute scale. The proposed algorithm requires at least three frames: two frames from one camera and a frame from a neighbouring camera. The relative camera poses are estimated with classical Structure-from-Motion and the absolute scales between views are computed by assuming straight trajectories between consecutive views of one camera. We propose a final optimisation step to refine only the scale and the 3D points. Our method is evaluated in real conditions on the KITTI dataset. We show quantitative evaluation through comparisons against GPS/INS ground truth.
This paper presents a visual odometry with metric scale estimation of a multi-camera system in challenging un-synchronized setup. The intended application is in the field of intelligent vehicles. We propose a new algorithm named “triangle-based” method. The proposed algorithm employs the information from both extrinsic and intrinsic parameters of calibrated cameras. We assume that the trajectory between two consecutive frames of a camera is a linear segment (straight trajectory). The relative camera poses are estimated via classical Structure-from-Motion. Then, the scale factors are computed by imposing the known extrinsic parameters and the linearity assumption. We verify the validity of our method both in simulated and real conditions. For the real world, the motion trajectory estimated for image sequence of two cameras from KITTI dataset is compared against the GPS/INS ground truth.
In this paper we propose a new methodology of development dedicated to cooperative ADAS. This methodology led us to implement a new framework for prototyping a communicating ADAS system. Within this framework, we combine the data from multiple modules: a vision module, a V2V communication module and Geo-localization GPS module, in order to accomplish a cooperative warning system. To achieve this goal, we have developed a prototyping system based on the principle of augmented reality, in which we replay real data and change the characteristics of the communication system. The GPS data and routing protocols were crucial elements for V2V communication simulation made with ns-2 simulator. We conducted different scenarios on real experimental platform consists of LaRA vehicles. Multiple results are presented to show up the compatibility and the performance efficiency of real-time multi sensors in an integrated framework for collision avoidance applications. The implementation of the warning system was used to estimate the number of pre- collisions detected in both real and simulated situations. The difference between these two situations was analyzed for several scenarios corresponding to different road situations. The results showed that the simulation of V2V communica- tion provide additional data that improve the implementation of these new ADAS and to assess their performance.
Cet article propose une nouvelle méthodologie de développement dédiée aux systèmes ADAS coopératifs. Cette méthodologie nous a conduit à mettre en œuvre un nouveau cadre de prototypage des systèmes ADAS communicants. Dans ce cadre, nous combinons les données de plusieurs modules, un module de vision, un module de communication V2V et un module de géolocalisation GPS, pour réaliser un système d’alerte coopératif. Afin d’atteindre cet objectif, nous avons développé un système de prototypage basé sur le principe de la réalité augmentée, dans lequel nous pouvons rejouer des données réelles et modifier les caractéristiques du système de communication. Les données du système de géolocalisation GPS et les protocoles de routage ont été des éléments primordiaux pour la simulation de la communication V2V réalisée avec le simulateur ns‐2. Nous avons effectué différents scénarios réels sur la plate‐forme du prototype LaRA composée de véhicules instrumentés. Plusieurs résultats sont présentés pour illustrer la compatibilité et l’efficacité de l’intégration des données réelles issues de plusieurs capteurs dans ce nouveau système de prototypage pour les applications d’alerte. La mise en œuvre du système d’alerte a permis d’estimer le nombre de pré‐collisions détectées dans deux situations, une réelle et une simulée. L’écart entre ces deux configurations a été étudié et analysé pour plusieurs scénarios qui correspondent aux différentes situations routières. Les résultats ont montré que les simulations de communications V2V fournissent des données complémentaires qui améliorent la mise en œuvre de ces nouveaux ADAS et permettent d’évaluer leurs performances.
The aim of this reported work is to estimate travel time on urban road sections. The concerned roads are the ones with no dedicated sensing infrastructure and for which the only potentially available source of data is sparse GPS probe data. Presented is a new approach to solve this problem by using an applied particle filter method and historical probabilities distributions of travel time per road section.
Edge detection is an indispensible initial step in many contour-based computer vision applications like edge-based obstacle detection, edge-based target recognition, etc. The performance of these applications is highly dependent on the quality of edges detected in the initial step. Most of the edge detectors used in these applications only detect boundaries separating two regions with high intensity gradient. However, certain computer vision applications require detection of low contrast boundaries. This paper presents a statistical operator for detecting low contrast boundaries. The proposed operator is highly suited for obstacle detection systems for poor visibility conditions. To evaluate its edge detection capability under normal and low contrast conditions, it is tested on a dataset of 40 object images and on MARS/PRESCAN dataset containing foggy virtual images. The quantitative evaluations using Matthew's correlation coefficient and Pratt's figure of merit indicate that the proposed method outperforms other edge detectors.
In this paper, we propose two different applications in the area of V2V communications. First, we present a method for better car tracking using GPS information shared through the V2V communication and a vision system in order to support accurate positioning. To accomplish this, we propose to use particle filtering techniques, and when GPS data is unavailable, or of poor quality, we couple GPS data with vision data collected from the vehicles. Second, we present a new simulated framework for prototyping the whole process by combining embedded data, vision data and V2V simulations to progress toward an anti-collision application. This framework could provide a better understanding of road security by studying the impact of V2V communications, thereby improving the quality of perception systems and adding new features for "ADAS". Our experimentations have been conducted in different scenarios on a fleet of vehicles moving and communicating in realtime conditions. The obtained results demonstrate the consistency of our method whenever the GPS is unavailable. Moreover, they prove the feasibility and the performance efficiency of such real-time multisensory fusion to provide an integrated framework for collision avoidance. (C) 2012 Elsevier Ltd. All rights reserved.
This paper presents a practical testing of two different methods to estimate the travel time in urban areas. The purpose behind this testing is to validate the behavior of each method regarding the road aspect in urban areas. The first method is based on Monte Carlo Method and the second one is based on adaptive estimation from probes. Both methods were modified to be adapted to our case and also to the nature of our data. The paper also describes an experiment with real-world data that was used in the testing of the two methods. Moreover it contains the architecture of the system used in order to make the tests. This work yeilded interesting results based on real-world experiments which give clear feedback about the application of the two methods to compute the travel time estimation per road section that can be used for processing the historical database as well as real time data. In general this work is a suitable validation of the two methods and encouraging for our future perspectives.
This paper presents an application of the Sequential Monte Carlo that will help to increase the accuracy of travel time estimations in our historical data. Our estimation filter is based on the Monte Carlo Method and was modeled in such a way as to be applicable to our new kind of data in order to estimate travel time per section of road. We took into consideration the delay time while changing the sections to symbolize the delay due to traffic lights or crossroads. We worked on an urban zone of Rouen, a French city, to evaluate our application. In this application, information is collected from a specific GPS system that warns drivers of the location of both fixed and mobile speed radars. Unlike the classical GPS system, this system is characterized by the data flow frequency where the GPS data is received from the probe vehicles at one minute intervals. After receiving the data we apply the map matching method in order to correct the GPS errors. Also, our geo-referencing system has special features; each road or section of road is formed by nodes and segments, and the intersection between each section is called a PUMAS points. The PUMAS Points are GPS coordinate points on a digital map which can be propagated or moved without cost, providing total flexibility to mesh a city or rural area. Over all the performance of the filter estimator is around 85% if we set our threshold at 50%.
Here we present an approach of meaningful curve identification with its depth estimation by chaining of the edge points, to locate and track the obstacles with stereo matching for automatic vehicle navigation. We use a self adoptive and nonlinear principle of extended declivity to obtain the edge points (horizontal declivities) in the images. These edge points include lots of noise and hence matching is not effective directly. The large size of the matching problem does not allow us to use effective matching algorithm properly. We use basic assumptions of continuity in the shape of expected obstacles to reduce the problem size and match less number of features effectively. Vertical chaining is used to obtain features which can be used for the tracking or stereo and obtain obstacles in the region of interest. These newly proposed curves are defined with their features and a matching algorithm is used to obtain results.
In this paper, we present a new fast method for matching stereo images acquired by a stereo sensor embedded in a moving vehicle. The method consists in exploiting the matching results obtained in one stereo pair (frame) for computing the disparity map of the following stereo pair. This can be achieved by finding a temporal relationship, which we named association, between consecutive frames. The disparity range of the current frame is deduced from the disparity map of the preceding frame and the association between the two frames. Dynamic programming technique is considered for matching the image features. The proposed approach is tested on virtual and real stereo image sequences and the results are satisfactory. The method is fast and able to provide about 20 millions disparity maps per second on a HP Pavilion dv6700 2.1GHZ.
This paper presents a real-time stereo image sequences matching approach dedicated to intelligent vehicles applications. The main idea of the paper consists in integrating temporal information into the matching scheme. The estimation of the disparity map of an actual frame exploits the disparity map estimated for its preceding frame. An association between the two frames is searched, i.e. temporal integration. The disparity range is inferred for the actual frame based on both the association and the disparity map of the preceding frame. Dynamic programming technique is considered for matching the image features. As a similarity measure, a variance-based cost function is used. The proposed approach is tested on virtual and real stereo image sequences and the results are satisfactory. The method is fast and able to provide about 20 millions disparity maps per second on a HP Pavilion dv6700 2.1 GHZ.