In driver attention zone prediction tasks, accurately identifying and locating the driver’s attention zone is crucial. Traditional models have significant limitations in complex driving scenarios due to their failure to fully utilize multidimensional driving environment information. To address these issues, this paper proposes a multi-attention feature fusion network (MAFF-HRNet) for driver attention region prediction. The proposed network combines high-resolution feature extraction with bimodal RGB–semantic inputs, multi-scale feature fusion, attention-based feature refinement, and temporal modeling. The experimental results on the DR(eye)VE dataset show that MAFF-HRNet improves driver attention region prediction under the current evaluation protocol. These results indicate that semantic scene information, multi-scale spatial representation, and temporal context are beneficial for generating more accurate driver attention heatmaps in complex driving scenes.
End-to-end autonomous driving is an emerging technology and a prominent research focus in both industry and academia. By integrating perception, localization, decision-making, and control into a single model, end-to-end systems aim to streamline the traditional modular pipeline while introducing new challenges in safety validation and interpretability. Unlike existing surveys that predominantly catalog algorithms, this review proposes a novel function-oriented taxonomy by categorizing architectures into perception-integrated and planning-integrated paradigms. Beyond the architectural dimension, the analysis delves into critical safety and interpretability, emphasizing the fundamental gap between theoretical design and the reliability required for real-world deployment. Industrial applicability is examined through real-world examples of data closed-loop workflows and simulation testing, addressing practical constraints in latency and computing resources. Finally, the review addresses critical challenges, particularly long-tail data scarcity and the deficiency in human-like decision-making and outlines future directions toward achieving robust autonomy.
The deployment of conditionally automated vehicles raises safety concerns, as drivers often engage in non-driving-related tasks (NDRTs), delaying takeover responses. This study investigates driver state monitoring (DSM) using multimodal physiological and ocular signals from the TD2D (Takeover during Distracted L2 Automated Driving) dataset, which includes synchronized electrocardiogram (ECG), photoplethysmography (PPG), electrodermal activity (EDA), and eye-tracking data from 50 participants across ten task conditions. Tasks were reassigned into three workload-based categories informed by NASA-TLX ratings. A unified preprocessing and feature extraction pipeline was applied, and 25 informative features were selected. Random Forest outperformed Support Vector Machine and Multilayer Perceptron models, achieving 0.96 accuracy in within-subject evaluation and 0.69 in cross-subject evaluation with subject-disjoint splits. Sensitivity analysis showed that temporal overlap had a stronger effect than window length, with moderately long windows (5–8 s) and partial overlap providing the most robust generalization. SHAP (Shapley Additive Explanations) analysis confirmed ocular features as the dominant discriminators, while EDA contributed complementary robustness. Additional validation across age strata confirmed stable performance beyond the training cohort. Overall, the results highlight the effectiveness of physiological and ocular measures for distraction detection in automated driving and the need for strategies to further improve cross-driver robustness.
To address the issue of poor yaw stability in distributed-drive electric pickup trucks at medium-to-high speeds, particularly under the influence of continuously varying tire forces and road adhesion coefficients, a novel Kalman filter-based method for estimating the road adhesion coefficient, combined with a Tube-based Model Predictive Control (Tube-MPC) algorithm, is proposed. This integrated approach enables real-time estimation of the dynamically changing road adhesion coefficient while simultaneously ensuring vehicle yaw stability is maintained under rapid response requirements. The developed hierarchical yaw stability control architecture for distributed-drive electric pickup trucks employs a square root cubature Kalman filter (SRCKF) in its upper layer for accurate road adhesion coefficient estimation; this estimated coefficient is subsequently fed into the intermediate layer’s corrective yaw moment solver where Tube-based Model Predictive Control (Tube-MPC) tracks desired sideslip angle and yaw rate trajectories to derive the stability-critical corrective yaw moment, while the lower layer utilizes a quadratic programming (QP) algorithm for precise four-wheel torque distribution. The proposed control strategy was verified through co-simulation using Simulink and Carsim, with results demonstrating that, compared to conventional MPC and PID algorithms, it significantly improves both the driving stability and control responsiveness of distributed-drive electric pickup trucks under medium- to high-speed conditions.
Human-machine co-driving is an important stage in the development of automatic driving, and accurate recognition of driver behavior is the basis for realizing human-machine co-driving. However, traditional detection methods exhibit limitations in driver behavior detection, including low accuracy and slow processing efficiency. Aiming at these challenges, this paper proposes a driver behavior detection method that improves the Swin transformer model. First, the efficient channel attention (ECA) module is added after the self-attention mechanism of the Swin transformer model so that the channel features can be dynamically adjusted according to their importance, thus enhancing the model’s attention to the important channel features. Then, the image preprocessing of the public State Farm dataset and expansion of the original image dataset is carried out. Then, the parameters of the model are tuned. Finally, through the comparison test with other models, an ablation test is performed to verify the performance of the proposed model. The results show that the proposed model algorithm has a better performance in 10 classifications of driver behavior detection, with an accuracy of 99.42%, which is improved by 3.8% and 1.68% compared to Vgg16 and MobileNetV2, respectively. It can provide a theoretical reference for the development of an intelligent automobile human-machine co-driving system.
随着电动汽车行业的快速发展,全球对电动汽车人才的需求也越来越旺盛.当前,本科高校培养的电动汽车人才与电动汽车市场需求存在匹配度不高,无法很好地适应电动汽车行业发展需求.本文深入分析了本科高校电动汽车课程教学中存在的问题,并探索了虚拟仿真技术在电动汽车课程教学中的应用措施,以期为其他应用型本科高校的电动汽车人才培养提供一些借鉴.
In this paper, a multiagent-based control method is proposed to design a transit signal priority (TSP) scheme at urban traffic networks. One agent controls an intersection. Coordination among different intersection agents is deployed to guarantee the benefits from TSP at upstream intersections. At one intersection there are usually many bus routes, and thus multiple TSP requests and coordination requests can occur at one phase simultaneously. Therefore, the proposed method also aims at resolving conflicting TSP or coordination requests. A multilevel fuzzy controller consisting of a transit signal priority controller, Green Time Adjustment Controller 1, fuzzy negotiation controller, and Green Time Adjustment Controller 2 is then introduced to realize these objectives. Following that, fuzzy inference decisions and algorithms are provided for describing the control process of the proposed multiagent method. An urban traffic network with 25 intersections is selected to conduct a case study to evaluate the proposed method by comparison and sensitivity analysis. The proposed method performed better than the other three methods under different scenarios in accordance with the simulation results.
The primary objective of an energy management strategy is to achieve optimal fuel economy through proper energy distribution. The adoption of a fuzzy energy management strategy is hindered due to different reasons, such as uncertainties surrounding its adaptability and sustainability compared to conventional energy control methods. To address this issue, a fuzzy energy management strategy based on long short-term memory neural network driving pattern recognition is proposed. The time-frequency characteristics of vehicle speed are obtained using the Hilbert–Huang transform method. The multi-dimensional features are composed of the time-frequency features of vehicle speed and the time-domain signals of the accelerator pedal and brake pedal. A novel driving pattern recognition approach is designed using a long short-term memory neural network. A dual-input and single-output fuzzy controller is proposed, which takes the required power of the vehicle and the state of charge of the battery as the input, and the comprehensive power of the range extender as the output. The parameters of the fuzzy controller are selected according to the category of driving pattern. The results show that the fuel consumption of the method proposed in this paper is 5.8% lower than that of the traditional fuzzy strategy, and 4.2% lower than the fuzzy strategy of the two-dimensional feature recognition model. In general, the proposed EMS can effectively improve the fuel consumption of extended-range electric vehicles.
将专业教育和思政育人深度融合是落实“立德树人”教育根本任务的重要举措。本文根据电动汽车结构与原理课程的特点和教学内容,重新修订了课程教学目标,深入挖掘本课程的思政元素,分析了思政育人点与专业知识点的衔接、切入总体思路,并探究了不同思政育人点采用的教学组织形式,为新能源汽车专业课程的“课程思政”建设提供参考。
将专业知识教育与思政教育有机融合是当前高等教育的一项重要要求.本文根据汽车试验学的课程教学目标,分析了思政育人点与课程知识点的衔接及切入思路,探讨了汽车试验学课程思政的教学组织与实施方法,为提高汽车试验学的课程思政育人效果提供思路.
1 引言 自2019年年末新冠肺炎疫情暴发以来,全球疫情一直没有得到彻底控制,这也给人们的工作生活带来了巨大影响.目前,国内疫情也是反复不断地出现,根据我国国情,新冠肺炎疫情常态化防疫仍然是全国人民的重要工作.防疫工作中的核酸检测已成为人们生活当中的日常工作.然而,在核酸检测过程中经常存在排队人员间距过小、未佩戴口罩或佩戴不规范以及体温异常等问题,主要由疫情防控工作人员来对核酸检测排队中出现的问题进行协调解决 [1].但是疫情工作人员现场管理的实际效果并不理想,还可能增加了工作人员感染病毒的几率,另外,在炎热、下雨等恶劣天气情况下无形中加重了工作人员的负担.
As a popular research field, autonomous driving may offer great benefits for human society. To achieve that, current studies often applied machine learning methods like reinforcement learning to enable an agent to interact and learn in a stimulating environment. However, most simulators lack realistic traffic which may cause a deficiency in realistic interaction. The present study adopted the SMARTS platform to create a simulator in which the trajectories of the vehicles in the NGSIM I-80 dataset were extracted as the background traffic. The built NGSIM simulator was used to train a model using the proximal policy optimization method. The actor-critic neural network was applied, and the model takes inputs including 38 features that encode the information of the host vehicle and the nearest surrounding vehicles in the current lane and adjacent lane. A2C was selected as a comparative method. The results revealed that the PPO model outperformed the A2C model in the current task by collecting more rewards, traveling longer distances, and encountering less dangerous events during model training and testing. The PPO model achieved an 84% success rate in the test which is comparable to the related studies. The present study proved that the public driving dataset and reinforcement learning can provide a useful tool to achieve autonomous driving.
Real-time driver behavior detection (DBD) is essential for developing driver-centered human-vehicle co-driving systems. This paper proposes an accurate and easily implementable DBD method based on driver behavior images and deep learning. Convolutional neural networks have difficulty dealing with changes in driver behavior images (rotation, scale, and translation). They suffer from a tradeoff between accuracy and number of trainable parameters, noise interference in images of driver behavior, and the black-box problem of neural networks, limiting the effectiveness of current DBD methods. Therefore, we design a novel deep learning-based DBD method consisting of a deep deformable inverted residual network with an attention mechanism. Deformable convolution is used to deal with image rotation and translation. An inverted residual block and linear bottleneck approach based on depthwise separable convolution is used to reduce the number of trainable parameters while maintaining high accuracy. An attention mechanism based on soft thresholding is incorporated into the nonlinear transformation layers to extract driver behavior-relevant features. We also innovatively propose a visualization method to improve the interpretability of the proposed method. The experimental results show that the proposed method can accurately detect driver behavior (95.17% mAP on the Kaggle driving test dataset), significantly outperforming state-of-the-art methods regarding accuracy, real-time performance, and reliability.
The prediction of the driver's focus of attention (DFoA) is becoming essential research for the driver distraction detection and intelligent vehicle. Therefore, this work makes an attempt to predict DFoA. However, traffic driving environment is a complex and dynamic changing scene. The existing methods lack full utilization of driving scene information and ignore the importance of different objects or regions of the driving scene. To alleviate this, we propose a multimodal deep neural network based on anthropomorphic attention mechanism and prior knowledge (MDNN-AAM-PK). Specifically, a more comprehensive information of driving scene (RGB images, semantic images, optical flow images and depth images of successive frames) is as the input of MDNN-AAM-PK. An anthropomorphic attention mechanism is developed to calculate the importance of each pixel in the driving scene. A graph attention network is adopted to learn semantic context features. The convolutional long short-term memory network (ConvLSTM) is used to achieve the transition of fused features in successive frames. Furthermore, a training method based on prior knowledge is designed to improve the efficiency of training and the performance of DFoA prediction. These experiments, including experimental comparison with the state-of-the-art methods, the ablation study of the proposed method, the evaluation on different datasets and the vi-sual assessment experiment in vehicle simulation platform, show that the proposed method can accurately predict DFoA and is better than the state-of-the-art methods.
To achieve efficient shared autonomy, driver behavior detection (DBD) is undoubtedly required. This paper investigates a deep driver behavior detection (DDBD) model. To overcome the low accuracy of DBD due to a lack of driver behavior data, the similarity of some driver behavior characteristics, and the ignorance of multi-scale structure and texture information, a DDBD model based on human brain consolidated learning (HBCL) is proposed. First, multiple DBD models with input information of different scales are trained based on transfer learning. Then, a new model called consolidation training (CT) using the Mish is trained based on the weight data from the first step. Finally, a novel method for the visualization of the attention area is proposed. The experimental results demonstrate that the proposed model achieved the highest accuracy (94.72% on the Kaggle-driving test dataset), generalization and real-time performance, the attention area is more anthropomorphic as compared with existing state-of-the-art models.
Numerous traffic crashes occur every year on zebra crossings in China. Pedestrians are vulnerable road users who are usually injured severely or fatally during human-vehicle collisions. The development of an effective pedestrian street-crossing decision-making model is essential to improving pedestrian street-crossing safety. For this purpose, this paper carried out a naturalistic field experiment to collect a large number of vehicle and pedestrian motion data. Through interviewed with many pedestrians, it is found that they pay more attention to whether the driver can safely brake the vehicle before reaching the zebra crossing. Therefore, this work established a novel decision-making model based on the vehicle deceleration-safety gap (VD-SGM). The deceleration threshold of VD-SGM was determined based on signal detection theory (SDT). To verify the performance of VD-SGM proposed in this work, the model was compared with the Raff model. The results show that the VD-SGM performs better and the false alarm rate is lower. The VD-SGM proposed in this work is of great significance to improve pedestrians’ safety. Meanwhile, the model can also increase the efficiency of autonomous vehicles.
为开展城市道路环境中驾驶人的工作负荷评价试验,研究了驾驶人视觉行为与工作负荷之间的相关性.在西安选取了1条典型城市道路路线,包括普通城市道路和城市快速路,招募了25名驾驶人开展实车试验,采用KF2生理检测仪和Eyelink Ⅱ型眼动仪,分别采集试验中每位驾驶人的心电指标(心率、心率变异性)和视觉行为数据,运用统计学方法分析驾驶人视觉行为参数与心率增长率和心率变异性(低频和相邻2个R-R间期差值的均方根)指标之间的相关性.研究结果表明:城市道路环境中,驾驶人扫视行为参数中的扫视幅度、扫视峰值速度、扫视平均速度和扫视持续时间与心率增长率和心率变异性之间分别存在正相关性(相关系数r≥0.5)和负相关性(r≤-0.5),扫视幅度、扫视峰值速度与工作负荷之间的相关性表现更强,而扫视平均速度、扫视持续时间与工作负荷之间的相关性表现较弱,说明城市道路环境中可以采用扫视行为参数来评价驾驶人的工作负荷.该研究丰富了驾驶人工作负荷的评价指标,为驾驶过程中驾驶人状态监测提供了理论借鉴,有助于提高城市道路环境中驾驶人的行车安全性.
Mobile phone use while driving has become one of the leading causes of traffic accidents and poses a significant threat to public health. This study investigated the impact of speech-based texting and handheld texting (two difficulty levels in each task) on car-following performance in terms of time headway and collision avoidance capability; and further examined the relationship between time headway increase strategy and the corresponding accident frequency. Fifty-three participants completed the car-following experiment in a driving simulator. A Generalized Estimating Equation method was applied to develop the linear regression model for time headway and the binary logistic regression model for accident probability. The results of the model for time headway indicated that drivers adopted compensation behavior to offset the increased workload by increasing their time headway by 0.41 and 0.59 s while conducting speech-based texting and handheld texting, respectively. The model results for the rear-end accident probability showed that the accident probability increased by 2.34 and 3.56 times, respectively, during the use of speech-based texting and handheld texting tasks. Additionally, the greater the deceleration of the lead vehicle, the higher the probability of a rear-end accident. Further, the relationship between time headway increase patterns and the corresponding accident frequencies showed that all drivers' compensation behaviors were different, and only a few drivers increased their time headway by 60% or more, which could completely offset the increased accident risk associated with mobile phone distraction. The findings provide a theoretical reference for the formulation of traffic regulations related to mobile phone use, driver safety education programs, and road safety public awareness campaigns. Moreover, the developed accident risk models may contribute to the development of a driving safety warning system.
The use of mobile phones while driving is a very common phenomenon that has become one of the main causes of traffic accidents. Many studies on the effects of mobile phone use on accident risk have focused on conversation and texting; however, few studies have directly compared the impacts of speech-based texting and handheld texting on accident risk, especially during sudden braking events. This study aims to statistically model and quantify the effects of potential factors on accident risk associated with a sudden braking event in terms of the driving behavior characteristics of young drivers, the behavior of the lead vehicle (LV), and mobile phone distraction tasks (i.e., both speech-based and handheld texting). For this purpose, a total of fifty-five licensed young drivers completed a driving simulator experiment in a Chinese urban road environment under five driving conditions: baseline (no phone use), simple speech-based texting, complex speech-based texting, simple handheld texting, and complex handheld texting. Generalized linear mixed models were developed for the brake reaction time and rear-end accident probability during the sudden braking events. The results showed that handheld texting tasks led to a delayed response to the sudden braking events as compared to the baseline. However, speech-based texting tasks did not slow down the response. Moreover, drivers responded faster when the initial time headway was shorter, when the initial speed was higher, or when the LV deceleration rate was greater. The rear-end accident probability respectively increased by 2.41 and 2.77 times in the presence of simple and complex handheld texting while driving. Surprisingly, the effects of speech-based texting tasks were not significant, but the accident risk increased if drivers drove the vehicle with a shorter initial time headway or a higher LV deceleration rate. In summary, these findings suggest that the effects of mobile phone distraction tasks, driving behavior characteristics, and the behavior of the LV should be taken into consideration when developing algorithms for forward collision warning systems.