Driver interaction intent prediction is one of the core technologies enabling intelligent cockpits to transition from passive response to proactive service. Existing intent prediction methods mainly model for single service functions and are deficient in generalization of cross-functional prediction and generalized temporal prediction models predominantly neglect the spatio-temporal (ST) causal relationships between dynamic evolution of scenes and interaction intents, resulting in compromised prediction efficacy within complex environments. To address the above limitations, we propose a driver interaction intent prediction framework with scene semantic fusion in intelligent cockpits. The methodology involves three key stages: first, extracting temporal patterns from cockpit interaction sequences via LSTM networks, second, converting driving dynamics data into structured scene semantic text through parameter efficient fine-tuning of large language models (LLM), and third, establishing collaborative representations of interaction behavioral features and scene information based on scene-guided attention fusion mechanism to achieve accurate driver interaction intent prediction. Experimental results demonstrate that our method has superior performance compared with other methods, with Precision, Recall, F1-score, and Accuracy values of 0.8416, 0.7842, 0.8045, and 0.7948, respectively. In particular, the introduction of scene semantic information can effectively trace the motivating source of the driver interaction intent generation and effectively enhance the interpretability of the model. This research establishes a technical foundation for implementing personalized proactive services in intelligent cockpits.
Learning and simulating the decision processes of real-world human drivers is a key research direction in autonomous driving (AD). As the core of AD, existing decision systems typically face challenges in cross-scene generalization and decision interpretability: it requires understanding diverse dynamic driving scenarios and formulate transparent strategies that earn broad user trust. We propose VLM-Driver, a Vision-Language Model (VLM) framework with human-like chained driving decision thought, designed to progressively achieve global scene observation, high-level behavior planning, and low-level motion planning based on full-view driving videos, multi-round question queries and optional environmental perception information. Specifically, the full-view videos are resized to unified base resolution using the AnyRes strategy, and video features are extracted by a vision encoder. These video features are then mapped to the text embedding space via a vision-language projector and jointly fed into large language model (LLM) backbone with tokenized text embeddings. During this process, we introduce a Bilinear Interpolation method to efficiently compress the number of video features, ensuring an optimal balance between model performance and computational cost. Meanwhile, a dedicated motion head is designed to output refined future waypoints, improving the model's motion planning efficiency. Additionally, we construct a novel multimodal driving instruction dataset to support VLM-Driver training and introduce a decision-oriented training strategy to further enhance its chained reasoning capability. Extensive experiments show that VLM-Driver excels in scene observation and behavior planning, demonstrating high human-like consistency, and significantly outperforms existing baseline methods in motion planning, achieving SOTA performance. VLM-Driver also enable to handle unseen complex driving scenarios, exhibiting robust cross-scene generalization.
Partially automated driving systems still rely on human drivers to monitor the environment and resume control when needed. Due to reduced vigilance and situation awareness during automation, drivers may misjudge the level of risk and the system’s capability to handle sudden conflict scenarios, which can undermine overall safety. Effectively communicating risk-predictive information generated by the automation through the Human–Machine Interface (HMI) can help drivers better understand system behavior and adjust their trust accordingly. In this study, we conducted a 3 × 2 × 2 within-subjects experiment involving intersection conflict scenarios (HMI: baseline vs. risk-alert vs. multimodal alert; risk level: low vs. high; turning direction: left vs. right). Thirty participants took part in the simulation-based experiment, with psychological measures, peripheral biosignals, neural activity, and driver-initiated takeover behavior recorded throughout the experiment. Results showed that risk-predictive information amplified the difference in drivers’ subjective risk evaluations between high-risk and low-risk conditions, and promoted more effective cognitive control in high-risk scenarios. The multimodal alert HMI (integrating visual and auditory risk prediction information) increased situational trust and reduced takeover frequency. High-risk and right-turn conditions led to higher perceived risk, lower trust, and more frequent takeovers. These findings provide new insights into how risk-predictive HMI designs influence drivers’ performance in partially automated driving, contributing to safer and more effective human–automation interaction.
Object detection is a cornerstone of environmental perception in advanced driver assistance systems(ADAS). However, most existing methods rely on RGB cameras, which suffer from significant performance degradation under low-light conditions due to poor image quality. To address this challenge, we proposes WTEFNet, a real-time object detection framework specifically designed for low-light scenarios, with strong adaptability to mainstream detectors. WTEFNet comprises three core modules: a Low-Light Enhancement (LLE) module, a Wavelet-based Feature Extraction (WFE) module, and an Adaptive Fusion Detection (AFFD) module. The LLE enhances dark regions while suppressing overexposed areas; the WFE applies multi-level discrete wavelet transforms to isolate high- and low-frequency components, enabling effective denoising and structural feature retention; the AFFD fuses semantic and illumination features for robust detection. To support training and evaluation, we introduce GSN, a manually annotated dataset covering both clear and rainy night-time scenes. Extensive experiments on BDD100K, SHIFT, nuScenes, and GSN demonstrate that WTEFNet achieves state-of-the-art accuracy under low-light conditions. Furthermore, deployment on a embedded platform (NVIDIA Jetson AGX Orin) confirms the framework's suitability for real-time ADAS applications.
The driver’s situation awareness (SA) at the moment of a takeover request (TOR) is essential for safe control transition in conditionally automated driving. To address limitations in existing methods for SA quantification and prediction, this study proposes an interpretable framework based on the analysis of SA’s formation process and critical role during takeovers. First, an eXtreme Gradient Boosting (XGBoost) model was constructed to predict takeover performance (TOP). SHapley Additive exPlanations (SHAP) was employed to provide local interpretability and identify the contribution of readily measurable SA Level 1 to TOP, based on which SA Level 2 and 3 were operationalized, thereby enabling the continuous quantification of overall SA. Subsequently, another XGBoost model was developed to predict the quantified SA, and SHAP’s global interpretability was used to optimize model performance and analyze the key features’ main and interaction effects on SA. A takeover simulation experiment was conducted to construct the necessary dataset. Results indicated that the quantified SA was able to match the reasonable distribution patterns of SA across different takeover results. The XGBoost model outperformed four comparison models in SA prediction. In addition, the key features’ main and interaction effects on SA were broadly supported by existing studies, confirming the rationality and effectiveness of the SA prediction model. This research provides support for driver state monitoring and the formulation of takeover strategies in conditionally automated driving.
In a mixed traffic environment consisting of connected and automated vehicles (CAVs) and human-driven vehicles (HDVs), effective platooning cooperation behavior is crucial for enhancing traffic safety and efficiency. The platooning process consists of two stages: platooning formation and platoon driving. In the formation of decision-making, the selection of interaction targets is indispensable. This study proposes a novel integrated co-opetition index (ICI) to assess the cooperation and competition behavior during the platooning process of CAVs and HDVs, which is applicable to optimizing CAV platooning strategies for HDVs in mixed traffic. First, using the Waymo dataset, cooperative movement features are extracted from same-lane and cross-lane platooning events between CAVs and HDVs. Four evaluation metrics are proposed based on the correlation of features: following intention index, lane change competition index, stable following index, and interaction friendliness index. An integrated co-opetition index (ICI) is then constructed to characterize the risks and benefits of the dynamic platooning process between CAVs and HDVs. Validation in various traffic scenarios shows that compared to the single-threshold indicator: time to collision, the composite-threshold indicator: surrogate safety measures, and game theory equilibrium point. ICI demonstrates significant advantages in platooning success rate, platooning efficiency, and traffic efficiency improvement. Under ICI guidance, the average platooning success rate increases by 5.84
Every vehicle needs to frequently handle interactions with other road users in order to travel successfully. Enabling autonomous vehicles to interact naturally, in a manner similar to human drivers, is an important challenge in the field. This paper proposes a decision-making approach for autonomous vehicles that integrates Social Value Orientation (SVO) into a reinforcement learning framework to enable interactive behaviors in complex merging scenarios. We defined a quantitative calculation method for the SVO value, used it for reward shaping, and aimed to investigate the impact of SVO on agent behavior patterns. We employed the Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) algorithms to train our agents, and validated our findings in the highway-env simulator. The simulation results indicate that, with the proposed decision-making approach, agents complete tasks in merging scenarios while following the intended interaction behavior patterns, displaying modes ranging from self-centered to altruistic. This confirms that the SVO-based reward function is both concise and capable of effectively guiding agents to achieve a diverse range of anticipated interaction behaviors.
Background: Functional near-infrared spectroscopy (fNIRS) is portable and robust against electromagnetic interference, making it promising for brain-computer interfaces in intelligent cockpits. However, fNIRS-based motor imagery (MI) decoding for multi-command in-vehicle control remains underexplored. Objective: This study decodes five direction-related MI commands for vehicle-window control and establishes a reproducible framework linking task design, neurophysiological validation, and model development. Methods: Electroencephalography (EEG) source localization first identified task-related cortical regions. Guided by these results, frontal and prefrontal fNIRS signals from 20 participants were acquired and analyzed using group-level general linear models. We propose Spatio-Temporal-Frequency Graph Convolutional Transformer (STF-GCNFormer), an end-to-end architecture integrating spatiotemporal extraction, frequency-domain modeling, attention-based fusion, graph convolution-Transformer modeling, and attention pooling. Decoding performance was evaluated using within-subject stratified five-fold cross-validation, baseline comparisons, and ablations. Results: EEG and fNIRS analyses consistently implicated the frontal-prefrontal network. Across the five conditions, 7 channels survived Benjamini-Hochberg false discovery rate correction, mainly involving Brodmann areas 9, 10, and 46. Across the Up-Down, Left-Right, and Action-Hold contrasts, 9 significant channel-level effects were identified. STF-GCNFormer achieved 91.11 f 0.75% accuracy, 91.56 f 0.80% macro-averaged F1 score, 93.39 f 0.68% precision, and 88.89 f 0.88% Cohen's kappa, with all classification metrics reported as mean f standard deviation across five folds. Ablation results confirmed the complementary contributions of spatiotemporal, frequency-domain, and graph-structural modeling. Conclusions: This study demonstrates accurate and stable offline fNIRS-MI decoding for five-command in-vehicle function control under controlled conditions and provides quantitative neurophysiological and model-level evidence for brain-vehicle interaction.
This paper addresses the challenge of collaborative platooning between connected automated vehicles (CAVs) and human-driven vehicles (HDVs) in mixed traffic environments. It proposes a platoon formation strategy based on quantifying co-opetition relationships and using a selfattention mechanism. By constructing an ICI-enhanced crossvehicle matrix (ICI-CVM), the approach comprehensively evaluates the degree of cooperation and competition between CAVs and HDVs, quantifying indicators such as following intention, lane-changing competitiveness, stable following behavior, and interaction friendliness to form an integrated coopetition index (ICI). Building on this foundation, the paper designs a CAV leadership qualification assessment mechanism that considers factors such as traffic density, reverse influence area length, platoon capacity, and positional advantages to screen potential lead vehicles. It also introduces a self-attention mechanism to optimize the platoon selection process, achieving global cooperative decision-making. Simulation experiments demonstrate that the proposed method of self-attention mechanism based on ICI (ICI-ATT) significantly outperforms traditional methods in terms of traffic capacity, fuel economy, and safety. With a CAV penetration rate of 70%, traffic capacity increases by 15.5% and fuel consumption decreases by 18.4%, while also achieving optimal platoon size and stability. This research provides theoretical support and technical approaches for efficient coordination in mixed traffic flows.
To address the critical challenge of risk perception and assessment for autonomous vehicles in dynamic interactive environments, this study proposes a semi-supervised spatiotemporal interaction risk cognition network with attention mechanism (SS-SIRCN), inspired by the behavioral adaptation patterns of biological groups under external threats. First, by thoroughly analyzing the dynamic interaction characteristics exhibited by typical biological collectives when exposed to risk, the study reveals the underlying patterns of trajectory changes influenced by external danger. Then, an attention-based spatiotemporal risk cognition network is designed to establish a mapping between driving behavior features and potential driving risks.Finally, a semi-supervised learning framework is employed to enable risk assessment for autonomous vehicles using only a small amount of labeled data.Experimental results on real-world vehicle trajectory datasets demonstrate that the proposed method achieves a risk prediction accuracy of 90.76%, outperforming other baseline models in performance.
Accurate pedestrian trajectory prediction plays a critical role in enhancing the safety and reliability of autonomous driving. Traditional methods often overlook the subjective intentions of pedestrians, limiting their ability to address the inherent uncertainty in trajectory prediction. Moreover, they typically use centralized data aggregation from heterogeneous environments, raising concerns over privacy breaches. To tackle these challenges, we propose a context-guided trajectory prediction framework tailored for autonomous driving. It leverages scene semantics to generate a dual-channel destination representation, comprehensively describing pedestrians' potential destinations. A conditional diffusion model is introduced to generate diverse yet accurate future trajectories by conditioning on both pedestrian destinations and social interactions. To protect data privacy, we leverage federated learning with a novel multi-factor aggregation scheme that dynamically adjusts each client's contribution based on data quality, quantity, and training dynamics. Experiments on benchmarking datasets validate that our method achieves superior and privacy-aware trajectory prediction performance.
Quantifying uncertainty significantly enhances the reliability of perception in autonomous vehicles and provides more comprehensive environmental information for downstream modules. However, most existing perception methods lack the capacity to effectively estimate the associated uncertainty. To address this gap, we propose a BatchEnsemble-based network for uncertainty quantification in 3D object detection using point cloud data. Specifically, a BatchEnsemble-based convolutional layer is designed to reduce the memory overhead associated with ensemble-based paradigms. Building upon this, a series of probabilistic object detection networks are constructed by directly modeling object attributes using multivariate Gaussian distributions, thereby enabling the parallel extraction of both object features and their associated variances. Subsequently, an uncertainty-aware fusion strategy is introduced to integrate and filter multiple detection results based on an uncertainty quantification metric—namely, the Uncertainty Index—thereby yielding more reliable and comprehensive outputs. The proposed method is validated on the KITTI dataset. Experimental results demonstrate its competitive accuracy performance and effectiveness across various scenarios, including objects of differing detection difficulty, identification of false-positive results, and under adverse conditions such as snowy weather and sensor degradation.
To standardize the reliability test evaluation schedule of passenger cars powertrain, a multi-condition varying-speed accelerated test criteria correlated with the statistical data of target vehicles is developed. With this test criteria, a real-vehicle test is conducted to evaluate a trial-manufacture car equipped with a newly developed transmission. Regarding the fracture failure of the driving and driven gears of the final drive, a heat treatment quality inspection is conducted on the fractured gears. From the meshing state of the gears, it is observed that machining precision errors caused an excessively large bearing clearance of the helical gear in the differential housing, with the meshing area of the driving gear tilting to one side and generating eccentric loads under uneven stress. The fatigue damage of the gears is calculated by the rainflow cycle matrix of the gears obtained with the rotating rainflow counting method. The calculation results show that largest loads transmitted by the final drive gears in varying-speed test condition 2, which was the main condition for fractures of the driven gear. Under this condition, the driving gear also produced bending fatigue failure due to the significant-amplitude cyclic loading caused by the eccentric loads.
The road surface friction coefficient is a key factor in the decision-making and control strategies of autonomous driving systems. This study presents a groundbreaking method for estimating the road surface friction coefficient using light detection and ranging (LiDAR) point cloud data, enhancing autonomous vehicles' prospective and high-precision perception. Data from eight road types formed a robust dataset. Cloth simulation filtering (CSF) and the random sample consensus (RANSAC) algorithm extracted road point clouds accurately. Gaussian filtering then removed reflectivity outliers. Given the correlation among reflectivity, distance, and incident angle, the road surface was segmented for comprehensive feature extraction. A designed deep neural network (DNN) model, trained rigorously with the dataset, achieved road recognition. Using statistical knowledge of road materials and peak friction coefficients determined the road's friction coefficient. Validation showed the algorithm identifies road types with over 99.62% accuracy, at 55 ms per cycle. This ensures real-time, high-precision estimation of the peak friction coefficient, a major boost for autonomous driving systems.
This paper addresses the challenge of collaborative platooning between connected automated vehicles (CAVs) and human-driven vehicles (HDVs) in mixed traffic environments. It proposes a platoon formation strategy based on quantifying co-opetition relationships and using a self-attention mechanism. By constructing an ICI-enhanced cross-vehicle matrix (ICI-CVM), the approach comprehensively evaluates the degree of cooperation and competition between CAVs and HDVs, quantifying indicators such as following intention, lane-changing competitiveness, stable following behavior, and interaction friendliness to form an integrated co-opetition index (ICI). Building on this foundation, the paper designs a CAV leadership qualification assessment mechanism that considers factors such as traffic density, reverse influence area length, platoon capacity, and positional advantages to screen potential lead vehicles. It also introduces a self-attention mechanism to optimize the platoon selection process, achieving global cooperative decision-making. Simulation experiments demonstrate that the proposed method of self-attention mechanism based on ICI (ICI-ATT) significantly outperforms traditional methods in terms of traffic capacity, fuel economy, and safety. With a CAV penetration rate of 70%, traffic capacity increases by 15.5% and fuel consumption decreases by 18.4%, while also achieving optimal platoon size and stability. This research provides theoretical support and technical approaches for efficient coordination in mixed traffic flows.
As autonomous driving technology advances, researchers are focusing on utilizing expert priors to improve the agents for learning-based decision-making in autonomous vehicles. Expert priors have various carriers, and the existing technology primarily utilizes expert priors derived from demonstration data and interaction data. This paper proposed a deep imitative reinforcement learning method for decision-making in autonomous vehicles, synergizing the expert priors in both demonstration data and interaction data. The gradient projection technique was adopted to mitigate gradient conflicts between the demonstration and interaction data during the training phase, thus preventing learning stagnation and enhancing agent performance. Furthermore, we deployed the proposed decision-making method on real autonomous vehicles. An augmented reality experiment was conducted with random virtual traffic flows from the simulator. The simulation and experiment results demonstrated that the proposed method enhanced training efficiency and safety performance, and preliminarily overcame sim-to-real challenges.
This study investigated the effects of sitting postures on the enhancement of occupants’ situation awareness through seat vibrotactile interaction system in high-level automated driving scenarios during a non-driving related task. Furthermore, addressing the limitation of fixed, non-adaptive vibrotactile signal generation in existing studies, an occupant body pressure monitoring-based seat vibrotactile interaction system (OBPM-SVIS) was developed, and it was compared with the original system lacking the pressure monitoring function. The Wizard of Oz experimental approach was employed to simulate automated driving scenarios in real-road. Eighteen participants received signals indicating vehicle upcoming behaviors, including left turn, right turn, acceleration, and deceleration. The participants adopted five sitting postures, which were upright, left-leaning, right-leaning, backward-leaning, and forward-leaning. These signals were presented in both static and dynamic patterns. Effects were comprehensively evaluated based on correct response rate, reaction time, situation awareness rating technique (SART) score, rating scale mental effort (RSME) score, and user experience questionnaire (UEQ) score. Results indicated that non-upright sitting postures adversely affected the transmission of vibrotactile signals and increased the difficulty of signal recognition. Static pattern demonstrated less robustness to changes in sitting posture compared to dynamic pattern. The OBPM-SVIS effectively mitigated the adverse effects of sitting posture changes. This study provides a reference for optimizing vibrotactile interaction systems that convey information about the vehicle’s upcoming behaviors or takeover requests to occupants in high-level automated driving.
Owing to the shortage of computing resources for autonomous vehicles and redundant modeling among similar tasks, multi-task models have become a feasible solution. The multi-task prediction model of autonomous vehicles refers to the realization of trajectory, behavior, and risk predictions through a multi-task deep neural network. However, whether the multi-task prediction networks can effectively share information between multiple inputs and whether the shared representations are interpretable remains a concern. To address the aforementioned concerns, this study proposes a multi-source multi-dimensional model interpretation (M3-interpretation) method for multi-task prediction neural network (MPNN). The MPNN proposed in this paper is designed with a structure that emphasizes a “task-specific pipeline as the main, high-level semantic information sharing as the supplement”. Then, based on the information entropy theory, this study creatively extends the information bottleneck attribution method to M3 and uses feature masks to display fine-grained interpretation results. Comparison and ablation experiments using naturalistic trajectory datasets indicated that the proposed model has better prediction performance than single-task models. In addition, fine-grained attribution analysis was conducted on specific behaviors in temporal, spatial, and feature dimensions to explore the laws that affect behavioral inference in MPNN.