Human trajectory prediction has garnered significant attention due to its critical role in applications such as autonomous vehicles, service robots, and advanced surveillance systems. As this research area evolves rapidly, numerous research centers are initiating projects dedicated to this domain. In this survey, we provide a comprehensive review of existing methods for human trajectory prediction, emphasizing the key challenges and outlining future research directions. We review a broad range of existing studies and propose a taxonomy that categorizes these methods based on their motion modeling approaches, output types, and situational awareness (including interactions and contextual information). We also discuss the datasets commonly used and the performance metrics applied in this field. Additionally, we address the limitations of current state-of-the-art techniques and offer suggestions for future research.
Skeletal motion prediction aims to forecast future movement based on 3D skeleton sequences, crucial for applications such as autonomous driving and virtual reality. However, anticipating the motion of 3D articulated objects is challenging due to their inherent non linearity and stochastic nature. Existing approaches often represent the skeleton as a set of 3D joints, which unfortunately ignores joint relationships and anatomical constraints. Moreover, conventional recurrent neural networks struggle with capturing long-term dependencies in motion contexts. To address these limitations, we propose encoding anatomical constraints through Lie algebra representation, integrating self-attention in transformer networks. Our Motion-Lie Transformer architecture, leveraging Transformers with self-attention, preserves human motion kinematics. Empirical evaluations on datasets like Human3.6M, GTA-IM, and PROX promise competitive performance and accurate 3D human pose estimation.
Enhancing the convective heat transfer is essential in solar applications for efficiency improvement. This can be achieved using Vortex Generators (VGs). Previous studies have focused on distinguishing the best VGs design while considering single VGs row configuration. In this paper, a study is conducted on investigating the effect of using multiple VGs rows configuration with varying longitudinal pitch (LP) separating the rows at Reynolds number Re = 2000. Then, the effect of increasing the air flow rate is studied by varying Re in the range of 2000-10,000. The SST k-omega model is chosen to model the turbulence at high Re. Validation is performed by comparing the numerical results to numerical and experimental data from the open literature. The best configuration is obtained when using five VGs rows with LP = 3H which increases the thermal enhancement factor by 69 % and 90 % with respect to empty channel configuration at Re = 2000 and 10,000 respectively. Then, local analyses are performed to better understand the physics behind the heat transfer enhancement in the best configuration. Finally, Nusselt number and friction factor correlations are developed to be representative of a dynamic model of photovoltaic/thermal (PVT) systems with vortex generators.
In complex real-world decision problems, ensuring safety and addressing uncertainties are crucial aspects. In this work, we present an uncertainty-aware Reinforcement Learning agent designed for risk-sensitive applications in continuous action spaces. Our method quantifies and leverages both epistemic and aleatoric uncertainties to enhance agent's learning and to incorporate risk assessment into decision-making processes. We conduct numerical experiments to evaluate our work on a modified version of Lunar Lander with variable and risky landing conditions. We show that our method outperforms both Deep Deterministic Policy Gradient (DDPG) and TD3 algorithms by reducing collisions and having significant faster training. In addition, it enables the trained agent to learn a risk-sensitive policy that balances performance and risk based on a specific level of sensitivity to risk required for the task.
Photo-Voltaic/Thermal (PVT) system performance is defined by two main factors, the electric power generated from the Photo-Voltaic (PV) module, and the thermal power that is extracted from the PVT module. To increase the energy output, several control techniques can be applied. In the present work, a Economic model predictive Control (EMPC) strategy is used to enhance the performance of the PVT system. A dynamic model for the PVT system is developed using the Modelica language in the Dynamic Modeling Laboratory software (DYMOLA). Then, EMPC controller is defined in Matlab/Simulink. Two geometrical cases for the duct side of the PVT system are studied as different heat intensification techniques. First, an empty channel is considered and then vortex generators (VGs) are inserted into the channel. Simulations are carried out with summer and winter days in the north of France with two energy use scenarios referred to as no heat recovery (NHR) and heat recovery (HR) scenarios. The results showed that when using an EMPC controller with a heat recovery scenario the energy gain increases by 174% and 234% for empty channel and for vortex generator geometrical cases respectively. In order to better analyze the obtained results, cell temperature and mass flow rate are plotted for all the studied scenarios as a function of time. Finally, power generation as a function of irradiance is plotted in order to distinguish when the benefits of cooling out-weight its cost.
: Real-world decision problems, such as Domestic Hot Water (DHW) production, require the consideration of multiple, possibly conflicting objectives. This work suggests an adaptation of Deep Q-Networks (DQN) to solve multi-objective sequential decision problems using scalarization functions. The adaptation was applied to train multiple agents to control DHW systems in order to find possible trade-offs between comfort and energy cost reduction. Results have shown the possibility of finding multiple policies to meet preferences of different users. Trained agents were tested to ensure hot water production with variable energy prices (peak and off-peak tariffs) for several consumption patterns and they can reduce energy cost from 10.24 % without real impact on users’ comfort and up to 18 % with slight impact on comfort.
Anticipating human motion based on given sequences is a challenging and crucial task in computer vision and machine learning, enabling machines to understand human behaviors effectively. Precise prediction of human pose and motion trajectory holds great significance for various applications, including autonomous driving, robotics, and virtual reality. This paper presents a novel approach to address the interconnected tasks of estimating human motion, represented as 3D poses or 2D trajectories, and predicting future motions using 2D images and human pose/position sequences jointly. We propose an encoder-decoder architecture that leverages Transformer networks with a self-attention mechanism, utilizing visual context features, combined with an LSTM to model human motion kinematics. Our approach demonstrates consistent and remarkable improvements over existing methods, both quantitatively and qualitatively. Extensive experiments conducted on diverse public datasets, such as GTA-IM and PROX for 3D human pose estimation, and ETH and UCY combined datasets for 2D trajectory prediction, showcase that our method substantially reduces prediction errors compared to the current state-of-the-art methods.
One of the key factors in achieving an autonomous vehicle is understanding and modeling the driving environment. This step requires a considerable amount of data acquired from a wide range of sensors. To bridge the gap between the Roadway and Railway fields in terms of datasets and experimentation, we provide a new dataset called RailSet as the second large dataset after Railsem19, specialized in Rail segmentation. In this paper we present a multiple semantic segmentation using two deep networks UNET and FRNN trained on different data configuration involving RailSet and Railsem19 datasets. We show comparable results and promising performance to be applicable in monitoring autonomous train’s ego perspective view.
Understanding the driving environment is one of the key factors in achieving an autonomous vehicle. In particular, the detection of anomalies in the traffic lane is a high priority scenario, as it directly involves vehicle's safety. Recent state of the art image processing techniques for anomaly detection are all based on deep learning of neural networks. These algorithms require a considerable amount of annotated data for training and test purposes. While many datasets exist in the field of autonomous road vehicles, such datasets are extremely rare in the railway domain. In this work, we present a new innovative dataset relevant for railway anomaly detection called RailSet. It consists of 6600 high-quality manually annotated images containing normal situations and 1100 images of railway defects such as hole anomaly and rails discontinuity. Due to the lack of anomaly samples in public images and difficulties to create anomalies in the railway environment, we generate artificially images of abnormal scenes, using a deep learning algorithm named StyleMapGAN. This dataset is created as a contribution to the development of autonomous trains able to perceive tracks damage in front of the train. The dataset is available at this link.
Recently, most state-of-the-art anomaly detection methods are based on apparent motion and appearance reconstruction networks and use error estimation between generated and real information as detection features. These approaches achieve promising results by only using normal samples for training steps. In this paper, our contributions are two-fold. On the one hand, we propose a flexible multi-channel framework to generate multi-type frame-level features. On the other hand, we study how it is possible to improve the detection performance by supervised learning. The multi-channel framework is based on four Conditional GANs (CGANs) taking various type of appearance and motion information as input and producing prediction information as output. These CGANs provide a better feature space to represent the distinction between normal and abnormal events. Then, the difference between those generative and ground-truth information is encoded by Peak Signal-to-Noise Ratio (PSNR). We propose to classify those features in a classical supervised scenario by building a small training set with some abnormal samples of the original test set of the dataset. The binary Support Vector Machine (SVM) is applied for frame-level anomaly detection. Finally, we use Mask R-CNN as detector to perform object-centric anomaly localization. Our solution is largely evaluated on Avenue, Ped1, Ped2, and ShanghaiTech datasets. Our experiment results demonstrate that PSNR features combined with supervised SVM are better than error maps computed by previous methods. We achieve state-of-the-art performance for frame-level AUC on Ped1 and ShanghaiTech. Especially, for the most challenging Shanghaitech dataset, a supervised training model outperforms up to 9% the state-of-the-art an unsupervised strategy.
Anomaly detection in surveillance videos is the identification of rare events which produce different features from normal events. In this paper, we present a survey about the progress of anomaly detection techniques and introduce our proposed framework to tackle this very challenging objective. Our approach is based on the more recent state-of-the-art techniques and casts anomalous events as unexpected events in future frames. Our framework is so flexible that you can replace almost important modules by existing state-of-the-art methods. The most popular solutions only use future predicted informations as constraints for training a convolutional encode-decode network to reconstruct frames and take the score of the difference between both original and reconstructed information. We propose a fully future prediction based framework that directly defines the feature as the difference between both future predictions and ground truth informations. This feature can be fed into various types of learning model to assign anomaly label. We present our experimental plan and argue that our framework’s performance will be competitive with state-of-the art scores by presenting early promising results in feature extraction.
Object tracking is an important proxy task towards action recognition. The recent successful CNN models for detection and segmentation, such as Faster R-CNN and Mask R-CNN lead to an effective approach for tracking problem: tracking-by-detection. This very fast type of tracker takes into account only the Intersection-Over-Union (IOU) between bounding boxes to match objects without any other visual information. In contrast, the lack of visual information of IOU tracker combined with the failure detections of CNNs detectors create fragmented trajectories. Inspired by the work of Luc et al. that predicts future segmentations by using Optical flow, we propose an enhanced tracker based on tracking-by-detection and optical flow estimation in vehicle tracking scenario. Our solution generates new detections or segmentations based on translating backward and forward results of CNNs detectors by optical flow vectors. This task can fill in the gaps of trajectories. The qualitative results show that our solution achieved stable performance with different types of flow estimation methods. Then we match generated results with fragmented trajectories by SURF features. DAVIS dataset is used for evaluating the best way to generate new detections. Finally, the entire process is test on DETRAC dataset. The qualitative results show that our methods significantly improve the fragmented trajectories.
Route planning system in bus network is very important on providing passengers a better experience for public transportation services. This work presents a new on-line dynamic path planning algorithm in a stochastic and time-dependent bus network, with the objective of least expected travel time. Firstly, an initial optimal path is found when a passenger leaves from the starting location at time t for the destination location. Secondly, when the passenger reaches a bus stop, the real time traffic condition is checked to decide whether the optimal path should be modified and re-calculated. If so, the shortest path algorithm is re-applied to find the new path. The simulation scenario is executed to show the performance and efficiency of the proposed dynamic path planning system.
The traffic flow measurement is one of the most important components in the traffic management systems. The existing traditional measurement methods are highly time-consuming and costly to continuously gather the required data, such as loop detectors and video cameras. However the travel duration provided by the emerging Floating Car Data (FCD) on Google Maps offers a novel way to estimate traffic flows. Therefore, this work presents a novel multi-model for urban traffic flows by applying a Gaussian Process Regressor (GPR) tuned using machine learning method based on FCD. The FCD on roads, requested through the Google Maps API, only provides information as congestion and travel duration. Traffic flows is estimated with GPR, including different models built by aggregating together data from days sharing similar configuration. The aggregation is performed manually or using unsupervised classification. At last, a series of experiments are conducted to compare the estimated traffic flow and the real one from actual sensors data. The obtained results show that, the proposed modeling can always reproduce and capture the tendency of real traffic flow. The aggregation permits effectively to increase the performance and to conclude on the capability of the approach to replace traditional loop detectors for the measurement of traffic flows.
Route planning systems in bus networks play a key role in providing a better experience to passengers for public transportation services. The existing systems satisfy the multicriteria in two ways. One way is firstly to find the k-shortest paths, then these paths are compared according to the given multi-criteria to select the ultimate path. Another way is to search for a set of Pareto paths with multi-criteria. However, both of the above systems only consider users' preferences qualitatively but not quantitatively. As a result, they can not integrate multicriteria as a single criterion based on users' quantitative preferences. Therefore, in this paper, a new integrated multi-criteria route planning method is proposed for searching the optimal path in a bus network according to passengers' quantitative references. Multi objectives being considered in this work are arrival time, walking time and number of bus transfers. A set of simulation scenarios is executed to show the performance of proposed new system.
In this article, we introduce a fast, accurate and invariant method for RGB-D based human action recognition using a Hierarchical Kinematic Covariance (HKC) descriptor. Recently, non singular covariance matrices of pattern features which are elements of the space of Symmetric Definite Positive (SPD) matrices, have been proven to be very efficient descriptors in the field of pattern recognition. However, in the case of action recognition, singular covariance matrices cannot be avoided because the dimension of features could be higher than the number of samples. Such covariance matrices (non singular and singular) belong to the space of Symmetric Positive semi-Definite (SPsD) matrices. Thus, in order to classify actions, we propose to adapt kernel methods such as Support Vector Machines (SVM) and Multiple Kernel Learning (MKL) to the space of SPsD matrices by using a perturbed Log-Euclidean distance (Arsigny et al., 2006). The mathematical validity of this perturbed distance (called Modified Log-Euclidean distance) for SPsD is therefore studied. The offline experiments are conducted on three challenging benchmarks, namely MSRAction3D, UTKinect and Multiview3D datasets. A fair comparison demonstrates that our approach competes with state-of-the-art methods in terms of accuracy and computational latency. Finally, our method is extended to an online scenario and experiments on MSRC12 prove the efficiency of this extension.
Over the last few decades, action recognition applications have attracted the growing interest of researchers, especially with the advent of RGB-D cameras. These applications increasingly require fast processing. Therefore, it becomes important to include the computational latency in the evaluation criteria. In this paper, we propose a novel human action descriptor based on skeleton data provided by RGB-D cameras for fast action recognition. The descriptor is built by interpolating the kinematics of skeleton joints (position, velocity and acceleration) using a cubic spline algorithm. A skeleton normalization is done to alleviate anthropometric variability. To ensure rate invariance which is one of the most challenging issues in action recognition, a novel temporal normalization algorithm called Time Variable Replacement (TVR) is proposed. It is a change of variable of time by a variable that we call Normalized Action Time (NAT) varying in a fixed range and making the descriptors less sensitive to execution rate variability. To map time with NAT, increasing functions (called Time Variable Replacement Function (TVRF)) are used. Two different Time Variable Replacement Functions (TVRF) are proposed in this paper: the Normalized Accumulated kinetic Energy (NAE) of the skeleton and the Normalized Pose Motion Signal Energy (NPMSE) of the skeleton. The action recognition is carried out using a linear Support Vector Machine (SVM). Experimental results on five challenging benchmarks show the effectiveness of our approach in terms of recognition accuracy and computational latency.
In recent years, due to road congestion issues and environmental concerns, multimodal transport has received an increased attention from policymakers. Fortunately, they can now rely on advanced computer tools. In particular, there are many tools for optimizing travel time in a multimodal context. In this article, we defend the thesis that the simulation tool remains essential to define an appropriate transport policy. Indeed, there are many profiles of public transport users and not all of them rely only on time to build their itinerary. Other qualitative criteria such as the cleanliness of the vehicle, the number of visitors to the stations, safety, etc. are also taken into account. The integration of user profiles in a multimodal simulation is not trivial. This article compares different traffic simulators and simulation platforms for multimodal scenarios in urban environments. In particular, the focus of the study is on two candidates: SUMO and GAMA.