
Real-time prediction of road adhesion conditions ahead of the vehicle is critical for emergency braking, stability control, and trajectory planning in advanced driver-assistance systems (ADAS). To overcome the limitations of conventional vehicle-dynamics-based methods, which mainly rely on tire–road contact responses and lack preview perception capability, this article proposes a real-time road adhesion prediction method based on onboard lidar point clouds. First, point cloud data were collected from four typical road surface conditions, including dry asphalt, wet asphalt, dry concrete, and wet concrete. Random sample consensus (RANSAC)-based ground segmentation and reflectivity filtering were then applied to remove nonroad points and highly reflective outliers. Subsequently, a speed-adaptive dynamic region of interest (ROI) was constructed by incorporating a vehicle safety distance model, enabling preview perception of the target road surface ahead. On this basis, the processed point clouds were converted into multichannel pseudoimages, and a convolutional neural network (CNN)-vision transformer classification model was developed to enhance local reflectance texture feature extraction and global context modeling. Finally, based on the empirical mapping relationship between road surface type and friction coefficient, the road surface recognition results were converted into road adhesion prediction results. Experimental results show that the proposed method achieves a better tradeoff between recognition accuracy and real-time performance. Compared with conventional vehicle-dynamics-based estimation methods, the proposed method significantly reduces the mean absolute error and mean absolute percentage error in both homogeneous road and transition road scenarios, verifying its effectiveness for real-time forward road adhesion prediction in ADAS.
Autonomous vehicles (AVs) rely on accurate perception systems to interpret their surroundings and make real-time driving decisions. While cameras and lidars have been primarily used for perception, they are vulnerable to adverse weather conditions. In contrast, radar that uses microwave signals shows very robust performance under challenging weather conditions. Recently commercialized, 4D radar simultaneously delivers range, azimuth, elevation, and Doppler measurements of surrounding objects, proving that it can be a useful alternative to lidar and a strong complement to camera sensors. This tutorial article presents a comprehensive overview of 4D radar for AVs, from data processing to artificial intelligence (AI)-driven techniques, such as object detection, drivable area detection, odometry, and sensor fusion. In addition, we survey publicly available 4D radar datasets and analyze current research trends, offering insights into how AI techniques can enhance the use of emerging 4D radar technology. To facilitate practical learning, we provide hands-on materials, including code implementations and example datasets, at https://github.com/kaist-avelab/radar-tutorial. As a result, this article provides researchers and practitioners with a clear understanding of the unique advantages of 4D radar and the state-of-the-art methods being developed to harness its capabilities for safer and more reliable autonomous driving.
Provides society information that may include news, reviews or technical notes that should be of interest to practitioners and researchers.
Provides society information that may include news, reviews or technical notes that should be of interest to practitioners and researchers.
Special events in venue areas and their influences on venue crowds have been a major focus for researchers. However, existing methods often ignore the distinctions between abnormal event-driven patterns and the temporal influence of multimodal event features. Without distinguishing abnormal flows, training learning models on raw data leads to overfitting and reduced prediction accuracy, especially under special events. To address the aforementioned issues, we propose a novel framework named Retrieval-Augmented Forecasting with Multimodal Agent (RAMA), which contains: 1) a multimodal agent that separates abnormal crowd flows and extracts event-contextual features, 2) a contrastive representation learning method to model intrinsic interactions between multimodal event features and abnormal flows, and 3) a two-stage diffusion-based forecasting model that synergizes normal pattern prediction with retrieval of the most relevant abnormal flows as contextual references. Our method automatically distinguishes normal and anomaly dynamics while enabling deep analysis of event-driven influences, offering differentiated guidance for forecasting. Comprehensive experiments demonstrate significant 38.2% relative accuracy improvements, especially for large-scale special events.
Transportation infrastructure is increasingly instrumented, yet data and models remain fragmented across design, construction, monitoring, and operations, which hinders closed-loop decision making. We reframe roads, bridges, tunnels, and hubs as embodied infrastructure agents embedded in physical and social environments. We propose a building information modeling (BIM)-enabled lifecycle framework that maps multimodal sensing and social signals to BIM semantics for traceable state estimation and lifecycle memory, executes latency-critical monitoring and control at the edge while performing long-horizon simulation and policy optimization in the cloud, and enables auditable collaboration via multiagent orchestration. The framework provides a practical pathway from perception to accountable action for infrastructure intelligent transportation systems.
Multi-modal Passenger Flow Forecasting at Hub Airports aims to predict the short-term origin-destination (OD) flow of outbound passengers from airports to various urban regions and its distribution across multiple transport modes. Accurate prediction is essential for efficient airport operations and for maintaining the stability of surrounding transportation systems. While recent deep learning approaches have shown potential in OD forecasting tasks, most focus on single-mode flow and neglect heterogeneous data sources, limiting their ability to model passenger flow and modal allocation in hub airports. To address these limitations, we propose M2F-Net, a Multi-Source and Multi-Modal Flow Forecasting Network that integrates a cross-modal encoder with a Universal Opportunity Model (UOM)-based decoder. Within the M2F-Net, a time-aware allocation matrix jointly models the spatiotemporal and modal flow distribution, while the encoder learns short-term temporal patterns and regional traffic states from historical OD data and road speeds. The decoder generates a dynamic travel-probability matrix to mask the encoder output, mitigating dynamic sparsity and guiding flow prediction. We construct a comprehensive benchmark based on real-world datasets collected from Beijing Capital International Airport and the Beijing Municipal Commission of Transport (June-August 2023), including total outbound flow, multi-modal OD flow, and roadnetwork speed. Experiments on this real-world dataset show that M2F-Net consistently outperforms strong baselines, and ablation results confirm the benefits of multi-source integration and our architectural design.
Urban metro passenger flow exhibits complex spatiotemporal dependencies and pronounced nonstationarity, driven by fine-grained interstation interactions, long-range cross-line couplings, multiscale periodicity, and holiday-induced perturbations. Existing methods often rely on predefined topological structures or static propagation mechanisms, making it difficult to capture spatial correlations that vary across time and operating conditions. To address this limitation, this article proposes a dynamic spatiotemporal hybrid attention (DST-HA) model for metro passenger flow prediction. Along the spatial dimension, DST-HA develops a topology-agnostic interaction-context dual-branch hybrid attention module in which a node-level gating mechanism adaptively reconstructs internode attention weights from the current passenger flow state and fuses node-level interactions with network-level context. This design enables the learning of time-varying spatial dependencies without requiring any prior adjacency matrix. Along the temporal dimension, DST-HA introduces a context-aware spatiotemporal recurrent unit (CA-STRUs) that injects time-of-day and weekday/holiday information into gating and attention computations through linear transformations, thereby enabling adaptive memory updates under varying temporal patterns. Comprehensive experiments on real-world metro passenger flow datasets from Hangzhou and Shanghai show that DST-HA consistently outperforms strong representative baselines across multiple operating scenarios, especially under high-pressure conditions, demonstrating its practical value for congestion warning, capacity adjustment, and real-time metro operations.
Technological advancements have significantly enriched in-vehicle interactive experiences, compelling drivers to process an increasing volume of information. In navigation scenarios, drivers must receive and interpret navigational cues to execute appropriate driving behaviors, potentially elevating their cognitive workload. This article seeks to clarify how drivers’ multisensory cognition affects their route choices and driving behavior. To address this, we first develop a cognitive model that hypothesizes the interactions between multisensory information processing and driver behavior. This model guides the design of a driving simulator experiment, where participants engage in multiple driving sessions simulating diverse information presentation scenarios. After each session, the participants are required to choose between a standard and an alternative route. Information is presented from either a single source or multiple sources, with multisource information varying in consistency, conflict, or delay. After each trial, the participants report their confidence in the traffic information and main information sources. We use a Bayesian model to measure the tendency to choose the alternative route. The results demonstrate the dominance of visual elements in information processing, with auditory information exhibiting lower interference. The route selection in conflict and delay scenarios is influenced by memory effects and descriptive elements, such as signage and road markings. Moreover, the average vehicle speeds are significantly lower in conflict and delay conditions compared to consistent information scenarios, indicating increased mental workload. This effect is likely attributable to attentional dysregulation, where competing stimuli induce cognitive conflict and necessitate the suppression of irrelevant information. The results offer valuable perspectives on the intricate interplay among multisensory information processing, driver decision making, and performance metrics in navigation contexts.
Intersections account for a significant proportion of bicycle-to-vehicle accidents, where early recognition of a cyclist’s turning intention is essential for collision avoidance and trustworthy warnings. We present a model-based framework that combines map-derived routes with sensor data to estimate maneuver probabilities from position, heading, speed, and yaw rate. Focusing on left/right turns versus going straight, we evaluate how detection accuracy evolves under different temporal horizons and reference points. Results show that incorporating trajectory prediction improves early detection, enabling detection of right-turn intentions up to 3 s before the conflict point. Position provides the most reliable late-stage confirmation within the final 2 s (more than 80% accuracy), while speed provides earlier but less stable cues. These findings reveal the variable-wise, stage-dependent reliability of kinematic variables, providing interpretable guidance for intelligent transportation systems safety systems, including advanced driver assistance systems and vehicle-to-everything applications to protect cyclists at intersections.
Traffic flow distribution constitutes a critical component of stage planning in railway marshalling yards, directly influencing operational efficiency and service reliability. To more accurately capture real-world yard operations, this article develops a two-stage optimization framework that explicitly incorporates the stochastic nature of train uncoupling times. In the first stage, a dynamic traffic flow distribution model is formulated to optimize train uncoupling and marshalling sequences, with three objectives: maximizing the number of successfully formed departure trains, maximizing the aggregate departure train priority, and minimizing deviations from the original uncoupling sequence of arriving trains. In the second stage, a static traffic flow distribution model is constructed to minimize the number of source trains required for departure formation and determine the specific wagon resources for each departure train. A modified differential evolution algorithm is developed to solve the dynamic model, while a hybrid sine-cosine whale optimization algorithm is designed for the static model. Using operational data from the Shijiazhuang South Marshalling Station, the proposed framework achieves over a 5% improvement in key performance metrics and demonstrates superior computational efficiency. Comparative evaluations confirm that this integration of hybrid intelligent optimization strategies significantly enhances yard shunting efficiency and increases the fulfillment rate of traffic flow distribution plans.
Provides society information that may include news, reviews or technical notes that should be of interest to practitioners and researchers.
Provides society information that may include news, reviews or technical notes that should be of interest to practitioners and researchers.