Real-time assessment of driver trust is critical for safe human–automation collaboration, yet most existing approaches depend on post-hoc questionnaires or intrusive physiological sensing. This study proposes Perceived Safety Distance (PSD), a behavioral metric derived from vehicle dynamics. We conducted a driving simulator experiment (N = 36) using a within-subjects design to manipulate environmental uncertainty (sunny vs. foggy rainy) and safety urgency (low vs. high). Results revealed a robust inverse relationship between subjective trust and PSD (r = −0.72 in high-urgency scenarios). Contrary to physical expectations, drivers in the most critical condition (foggy/high-urgency) maintained larger safety margins (M = 8.11 m) than in clear conditions (M = 4.80 m). We interpret this pattern as a trust-compensatory mechanism: higher trust can foster complacency and delayed intervention, whereas under degraded visibility, trust erosion elicits protective hyper-vigilance, leading to earlier defensive actions despite delayed visual detection. Physiological metrics (heart rate, pupil diameter) further indicated that this conservative behavior was accompanied by elevated cognitive workload. These findings validate PSD as a sensitive real-time indicator for trust calibration, offering a promising foundation for developing trust-aware adaptive automation systems.
The proliferation of in-vehicle screens enhances human-machine interaction (HMI) in intelligent cockpits but intensifies visual distraction risks, threatening driving safety. Screens are one of the primary sources of visual distraction, and driver distraction majorly contributes to traffic accidents. Conducting empirical testing and user testing during the design and development stages of In-Vehicle Information Systems (IVIS) are effective means to reduce visual distractions caused by touchscreen systems. However, these methods are costly and time-consuming. With the accelerated pace of over-the-air (OTA) updates of IVIS, traditional testing methods struggle to keep up with the development cycle. Therefore, IVIS visual demand simulation algorithms offer a cost-effective alternative. Considering the difficulty of balancing model accuracy and interpretability with existing methods, we propose modeling driver visual demand based on LightGBM (LGBM) and Random Forest (RF) algorithms. Furthermore, building upon the previously proposed SHapley Additive exPlanations (SHAP) interpretation method, we adopted Partial Dependence Plots (PDP) to more thoroughly analyze the nonlinear relationships between interaction features and visual demand metrics. Our method assists in predicting visual demand and, combined with interpretative methods, identifies high visual distraction issues in interface design.
Automotive seat pressure sensing provides a non-invasive modality for occupant state recognition and adaptive seat functions in intelligent cockpits. However, creep-induced temporal drift after seating may reduce the reliability of short-term occupant weight classification. This study analyzed 90 cushion pressure records from 30 participants, each obtained from a 20 s controlled seated trial. A single-exponential model characterized the early pressure evolution, and a reference-state mapping method compensated for temporal drift. A random forest classifier using cumulative cushion pressure features from sliding windows was adopted to compare raw, filtered, and compensated signals. A total of 76 records met the fitting quality criteria. Compared with raw signals, compensated signals increased accuracy, Macro-F1, and balanced accuracy by 13.1%, 22.7%, and 17.9%, respectively, with improved prediction consistency across windows. These results suggest that drift compensation improves temporal feature comparability and supports more stable short-term occupant weight classification under controlled seated conditions.
Driver-centered Human–AI interaction is essential for SAE Level 3 automated driving, where drivers alternate between supervisory control and non-driving-related tasks (NDRTs). This study aims to investigate how layered transparency and explainability via a Surroundings Reality (SR) interface influence driver trust and experience under varying safety urgency and weather. A 2 (urgency: low vs. high) × 2 (weather: sunny vs. foggy & rainy) × 3 (interface tier: descriptive (T1), predictive (T2), and explanatory/uncertainty-aware (T3)) mixed factorial driving simulator experiment was conducted with 36 participants, who performed the 2048 game as an NDRT. The outcomes were evaluated using subjective trust and NASA-TLX scores, behavioral measures (voluntary manual interventions and NDRT performance), and eye-tracking metrics. Results showed an increase in trust from T1 to T3, suggesting that richer transparency and explainability may support reliance formation. NASA-TLX scores did not differ significantly across tiers. Behaviorally, higher transparency tiers tended to be associated with fewer voluntary manual interventions, while intermediate predictive transparency (T2) appeared to be associated with lower NDRT performance relative to T1 and T3, highlighting attentional competition when interpretive support is insufficient. Eye-tracking further revealed scenario-dependent information priorities: attention shifted toward perception & prediction cues under higher urgency, and toward explanatory & uncertainty cues under degraded visibility. Our findings suggest that transparency and explainability should be layered yet scenario-adaptive, dynamically adjusting the depth, priority, and timing of information to balance trust calibration with workload and attention management.
Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning and open-world generalization. However, the excessive computational overhead and high inference latency of these massive models severely hinder their deployment in resource-constrained AD systems. To address this challenge, we propose a novel decision-making framework utilizing a lightweight confidence-aware language model, which bridges the gap between complex multimodal intention reasoning and efficient inference. Specifically, we design a multi-agent collaborative workflow, comprising action voting, confidence assessment, and summarization agents, to generate high-quality, confidence-annotated decision demonstrations via explicit Chain-of-Thought (CoT) reasoning. These demonstrations are then distilled into a lightweight language model featuring a dual-head architecture, enabling the joint prediction of decision probabilities and the generation of textual rationales. The distillation is realized via a confidence-aware fine-tuning strategy coupled with Retrieval Augmented Generation (RAG) to enhance the model's adaptability and data efficiency. Comprehensive closed-loop experiments on the nuPlan benchmark demonstrate that our approach achieves state-of-the-art (SOTA) success rates in both regular and long-tail scenarios while maintaining low inference latency.
Touch screen has become the main interface for drivers to complete secondary tasks in in-vehicle information systems (IVIS), and the interactive input form of clicking touch buttons is the most commonly used interaction behavior in IVIS. However, the recognition and operation of touch buttons will increase the driver’s workload and cause driving distraction, which will affect driving safety. This study aims to reduce driving distraction and improve driving safety and driving experience by designing various touch buttons to improve visual search efficiency and interaction performance. First, we designed 15 schemes for touch buttons, based on a previous theoretical summary and effect screening. Then, using simulated driving, eye-tracking measurement, and user questionnaires, we obtained the data for four evaluation indicators: task, physiological, driving performance, and subjective questionnaire. Finally, the entropy weight method was adopted to evaluate the design comprehensively. The results indicate that touch buttons with dynamic effects of color change, color projection, circle shape, negative polarity, and boundary exhibit better visual usability in the secondary tasks. The proposed scheme in this paper provides suggestions on the visual usability of the touch button design of an automotive intelligent cabin, which is conducive to improving driving safety, task efficiency, and user experience.
Prolonged autonomous driving may lead to boredom in human drivers, which can be alleviated by entertaining non-driving related tasks (NDRTs), such as reading, watching videos, and playing games. Among these NDRTs, playing games is one of the most engaging interactive tasks. However, there are limited previous studies on its impact, which present conflicting findings. To address this gap, we conducted a mixed-design experiment to investigate the effects of gaming devices (mobile phones vs. in-vehicle infotainment system (IVIS)) and game types (action vs. logic games) on drivers’ cognitive load, situation awareness, and takeover safety. Specifically, cognitive load before takeover was assessed using pupil diameter and the NASA-TLX scale. Situation awareness during takeover was evaluated via attention recovery time and the situation awareness rating technique (SART). Moreover, takeover safety was measured by takeover time (TOT), maximum acceleration, and minimum time to collision (Min TTC). The results indicated that playing different types of games on various devices generally increased cognitive load and impaired situation awareness to some extent. Surprisingly, playing action games on IVIS showed potential benefits for takeover safety, as evidenced by longer Min TTC. Nevertheless, our study has a limitation: it was conducted in a driving simulator for safety reasons, which may not fully replicate real-world conditions. Overall, this study clarifies the combined effects of gaming devices and game types on multiple variables during the takeover process, with an innovative classification of games based on their physical and mental attributes. In the future, it is recommended to adapt some action games with low cognitive load from mobile phones to IVIS when designing autonomous vehicles, so as to alleviate boredom and improve takeover safety.
Users around the world exhibit diverse expectations for user experience (UX), making it essential to bridge culture and design through a structured approach. This paper proposes CXP-9, a systematic and quantitative framework for cross-cultural UX design, building upon Hofstede’s 6D cultural dimensions and Garrett’s Elements of User Experience. The model includes three layers—Appearance, Logic, and Value—each comprising three dimensions, for a total of nine. Each dimension uses a 0–100 scale to reflect cultural preferences. Based on a survey of 1,000 users across China, the U.S., Germany, Thailand, and Saudi Arabia. The results reveal significant cultural differences in UX expectations across the nine dimensions. CXP-9 offers a practical, data-driven tool for designers, product managers, and researchers to tailor UX strategies to specific cultural contexts.
In driver fatigue monitoring, camera-based methods lack direct medical interpretability, while EEG-integrated solutions are often intrusive and costly, posing a critical challenge in balancing accuracy, usability, and cost-effectiveness. To address this, a Multi-modal Fusion-based Transformer Knowledge Distillation (MFTKD) method for fatigue classification was proposed in this paper. Specifically, the teacher model leverages EEG to guide the learning process while integrating EOG signals via an ECA-Transformer network to capture cross-modal correlations. Knowledge distillation, implemented via soft label guidance and cross-modal feature distribution alignment via Maximum Mean Discrepancy (MMD), transfers EEG-based fatigue patterns to a lightweight, EEG-free student model. On the SEED-VIG dataset, the teacher model achieves 98.01% accuracy, while the student model reaches 95.14% without EEG input. To further validate engineering feasibility and robustness, a comprehensive engineering validation experiment was conducted, which involves 20 subjects using a consumer-grade camera and an embedded NVIDIA Jetson edge computing platform. Experimental results in this real-world simulation scenario demonstrate that the student model robustly maintains a high accuracy of 94.27% at an ultra-fast inference throughput of 118 frames per second. These findings confirm its exceptional advantages in balancing high accuracy, real-time inference speed, and practical deployability, thereby facilitating improved safety monitoring with better usability and significantly lower costs.
With the increasing adoption of regenerative braking technology in electric vehicles (EVs), one-pedal driving (OPD) mode has become a prevalent feature. While OPD offers technical advantages in energy efficiency, its implications for driver behavior and traffic safety remain unclear. To address the lack of human factors research in this domain, this study utilized a driving simulator to systematically compare driving performance between OPD and two-pedal driving (TPD) modes. Twenty-six participants engaged in car-following tasks under varying traffic densities (uncongested vs. congested) and cognitive load levels (normal vs. 1-back). Driving performance and safety were quantified using the absolute speed difference, distance headway, braking frequency, and Time-to-Collision at brake onset (TTCbrake). The results revealed a significant trade-off: while OPD simplified operation, it led to compromised driving performance compared to TPD in specific contexts. Specifically, OPD resulted in larger speed variations and reduced safety margins during the approach stage. Conversely, under high cognitive load, OPD demonstrated a protective effect by mitigating performance degradation. These findings suggest that while OPD can benefit drivers under mental pressure, its deployment requires adaptive safety strategies, such as the integration of Headway Monitoring Warning (HMW) and Forward Collision Warning (FCW), to compensate for performance deficits in complex traffic environments.
Electric vehicles equipped with regenerative braking systems provide drivers a new driving mode, the one-pedal mode, which enables drivers to accelerate and decelerate with the throttle alone. However, there is a lack of systematic research on driving behavior in one-pedal mode, and whether it actually enhances or reduces safety remains to be validated. A driving simulator was used to analyze driving behavior and safety in the one-pedal mode in situations with different urgency level, with the two-pedal mode (the traditional driving mode in internal combustion engine vehicles) serving as a comparative group. The driver's perception times, initial and final throttle release times, throttle to brake transition times, maximum brake pedal forces, collision ratios, and time-to-collision (TTC) were measured under the lead vehicle decelerating at 0.1 g, 0.2 g, 0.5 g, 0.75 g, as well as uncertainty (decelerating at 0.2 g to 25 km/h, then decelerating at 0.75 g to 0), and under headways of 1.5 s and 2.5 s. Results showed: 1) The regenerative braking system did not affect driver perception and reaction of the lead vehicle braking event and drivers extended throttle release to avoid rapid speed drops when the lead vehicle braked slowly; 2) the one-pedal mode exhibited a longer throttle to brake transition time and increased uncertainty in timing of brake pedal application; 3) the one-pedal mode was safer than the two-pedal mode in low urgency situations but became unsafe in high urgency or uncertain situations due to delayed braking. The implications of this research include enhancing regenerative braking systems and developing forward collision warning systems.
As vehicles are expected to shift from manual to automated driving, the role of human drivers changes from active control to passive supervision. However, prolonged disengagement from the driving task may induce passive task-related (PTR) fatigue, which may negatively impact takeover performance. The present study sought to investigate whether video entertainment during prolonged automated driving can mitigate these risks. A driving simulation experiment was conducted with 32 participants who experienced varying durations of automated driving time (ADT) ranging from 15 to 60 min, both with and without video entertainment. Cognitive load, fatigue, and takeover performance were assessed using objective and subjective measures. Linear mixed effects models were employed to analyze the effects of video entertainment, ADT, and their interaction on the dependent variables. The results indicated that video entertainment not only increased cognitive load, but also effectively suppressed the growth of fatigue with longer ADT, as reflected by blink frequency and KSS scores. Furthermore, video entertainment reduced safety risks in prolonged automation scenarios by shortening takeover time and increasing the minimum time to collision. These findings suggest that video entertainment can improve comfort and takeover safety under prolonged automated driving scenarios by mitigating PTR fatigue.
Despite advancements in In-Vehicle Information Systems (IVIS) and extensive research on screen layouts, the influence of drivers' peripheral vision on interactions with evolving multi-screen and large display technologies remains poorly understood. This study examines drivers' responses to in-vehicle interactive information through peripheral vision, aiming to optimize visual interaction efficiency and enhance driving safety. Analyzing data from 216 participants in a driving simulator, we explored how horizontal eccentricity, screen type, cognitive load, visual crowding, and stimulus type affect perception rates and reaction times. Our findings highlight the significance of these factors and the need for driver-centered design. The results suggest designing IVIS that align with natural visual tendencies to improve interaction efficiency and driving safety.
Rapid advances in automated driving technology and the widespread adoption of in-vehicle information systems (IVIS) have led to an increasing prevalence of drivers engaging in non-driving-related tasks (NDRTs) during autonomous operation, thereby introducing potential safety hazards. In this study, we conducted a driving simulator experiment with 30 participants to examine the effects of IVIS NDRTs (i.e., navigation, video, audio, and reading tasks) and takeover time budgets on takeover timing, takeover quality, and visual behavior. Results from linear mixed-effects models indicate that IVIS touchscreen interactions significantly prolonged takeover time and lane change time, increased maximum lateral acceleration, and reduced minimum time-to-collision (TTC), suggesting that drivers adopted aggressive control behaviors during takeovers, which in turn elevated collision risk. Moreover, visual behavior analysis revealed an increased proportion of long glances directed away from the forward roadway and a delayed reallocation of visual attention to key regions (such as mirrors, the road, and the malfunctioning vehicle) following the takeover request. These findings enhance our understanding of human factors in automated driving and provide empirical evidence for optimizing driver-vehicle interaction protocols and improving the safety of riding in conditionally automated driving systems.
With the development of the smart cockpit, in-vehicle list selection tasks have become increasingly common in IVIS (In-Vehicle Infotainment System). This study investigated three steering wheel gestures-Arrow Buttons (AB), Direct Manipulation (DM), and Direct Manipulation with Scroll-bar (DMS)-for controlling the IVIS. The effects on efficiency, driving performance, visual load / cognitive load, safety, and subjective task load were examined in a simulated driving environment using common driving and eye movement metrics. Results indicated that the DMS gesture enhanced efficiency but deteriorated driving performance in certain scenarios. AB gesture exhibited the highest visual load, while DM gesture imposed the lowest cognitive load. Any gesture used on the central display caused the driver to spend an excessive maximum duration of glances to display. DMS gesture held the least subjective task load. This study comprehensively evaluated these gestures' effects, offering insights into steering wheel button design for future automotive interfaces.
Text-to-image (T2I) models are emerging as a powerful tool for designers to create user interface (UI) prototypes from natural language inputs (i.e., prompts). However, the discrepancy between designer inputs and model-preferred prompts makes it challenging for designers to consistently deliver effective results to end users. To bridge this gap, we introduce a novel hybrid method that assists designers in crafting user- centric prompts for T2I models, ensuring that the generated UIs align with end-user expectations. First, this method merges text mining and Kansei Engineering (KE) to analyze online user reviews and construct a Knowledge Graph (KG), mapping the intricate relationships between diverse affective requirements of users, design features, and corresponding text prompts for UI generation. Then, our approach automatically transforms designer inputs into model-preferred prompts through entity mention recognition and entity linking during the human-AI collaborative design process. Finally, we validate the proposed approach with a case study on automotive human-machine interface design. Experimental results demonstrate that our approach achieves high performance in perceived efficiency, satisfaction, and expectation disconfirmation. Overall, this study represents a step forward in integrating human and AI contributions in design and innovation within engineering disciplines, enabling AI to inspire, develop, and reinforce human creativity from a human factors perspective.
Human-computer interaction (HCI) of in-vehicle information systems (IVISs) is crucial to the safety and experience of drivers. However, the operation of secondary tasks of IVIS will cause distraction when driving. This paper aims to reduce driving distractions by multimodal interaction design, to enhance the user’s driving safety and interaction experience. Firstly, we analyze the secondary task and the associated interaction modes and develop a theory tool for multimodal interaction design modes. Then, three multimodal interaction schemes of typical secondary tasks are designed. Finally, the data of three evaluation indexes, namely expert evaluation, lane position, and user score, are obtained through expert interviews, simulated driving, and user questionnaires. The results demonstrate that the proposed multimodal interaction design of the secondary task is generally superior to the traditional voice and screen interaction, which is conducive to improving driving safety, performance, and interaction experience. The theoretical framework presented in this work provides a potential opportunity for the expansion of theoretical methods and applications of multimodal interaction design for secondary tasks.
With the application trend of large language models (LLMs) in the industry, the individuality of vertical domain LLMs has become a new direction worthy of study. Therefore, we focus on exploring the personality traits of the in-vehicle LLMs in the intelligent cockpit, which is regarded as the "third living space" and an important application scenario of LLMs. This study uses the psychological scale Big Five Inventory (BFI) to evaluate the personality traits of the in-vehicle LLM quantitatively and proposes methods for shaping and fine-tuning the personality traits of the in-vehicle LLM. The research shows that: 1) The in-vehicle LLM has measurable, stable, and consistent personality traits by the psychological scale, showing a high level of extraversion, agreeableness, conscientiousness, a lower level of neuroticism and a medium level of openness; 2) The personality traits of the in-vehicle LLM can be shaped and fine-tuned through the prompt of trait marker words + Likert qualifiers, and it can affect the subsequent behavioral tendencies and generated content. This study migrates the methods applied in the personality research of general LLMs to the in-vehicle LLM, providing new ideas for establishing a psychological and personality evaluation system in intelligent cockpits and contributing to the further exploration of the emotional and personalized design of the intelligent cockpit in the future to enhance user experience.
In-vehicle navigation systems play a crucial role in the automotive intelligent cockpit and provide users with various functions. However, current systems make users suffer from passive experiences, which result in many inconveniences, such as low error prevention, limited user control, and insufficient user freedom. To address the existing limitations, we propose an interactive navigation paradigm based on conversations to enable users to actively engage in navigation, express their needs, and manage existing limitations. To explore the feasibility of this concept, we develop a simulated conversational navigation assistant based on prompting a large language model (LLM). The prompt designed for the LLM leveraged for conversational navigation is curated with tailored instructions and navigation functions. We conduct a subjective experiment to assess the impact on user perceptions of the conversational navigation system. Findings indicate that conversational navigation is deemed satisfactory, efficient, and more natural. Furthermore, users express that involvement with navigation during conversational navigation can reduce driver distraction, thus enhancing driving safety. This study highlights the potential to transform traditional navigation experiences into more engaging and efficient processes, paving the way for future research to further enhance user-centered navigation through additional external functions and optimized dialogue systems.
Alarm sounds significantly influence a user’s sensory perception while driving, directly affecting driving judgement and safety. Personal experience and the environment play an important role in information cognition, but they are rarely considered in the current warning design. We propose a methodology enabling engineers and designers to locally optimize the advanced driver-assistance system (ADAS) functions and applied it to the Shanghainese ecosystem to improve performance. The alarm sound content is studied and sorted out to conduct user research and spatial sound collection evaluation. Local optimization and the subdivision of data are carried out to generate a user perception set on which the experimental tests and evaluation analysis are implemented. The framework increases the overall efficiency of auditory warning systems and minimizes Human–Machine Interface misunderstandings, thus providing the optimal security scheme for users.