Situational awareness—the ability to comprehend the surrounding environment and anticipate future events—is necessary for safely avoiding hazards while performing a task (e.g., driving a car without crashing). Good situational awareness relies on sufficient visual search of the environment, known as visual search strategy. Insufficient visual search can result in a failure to detect hazards, such as a pedestrian entering the roadway from an obstructed crosswalk. To date, the study of visual search strategy remains limited to small cohorts and short experimental time periods due to the conventional labor-intensive approach of characterizing visual search strategy: video coding. To address this limitation, this study presents an efficient, scalable, machine learning approach to characterize visual search strategy: time series clustering. To demonstrate the feasibility of this novel approach, time series clustering was applied to eye-tracking during driving in a simulated environment. Eye-tracking data were collected from a cohort of young (16–24 years) drivers (n = 36) during a virtual driving assessment of performance. Raw time-series eye-tracking data were compared using dynamic localized coordinate aligned warping (DCLAW), an extension of dynamic time warping. Visual search strategy clusters were identified using k-medoids unsupervised clustering. Characteristics of the visual search strategy clusters were defined by review of the cluster medoids. Time series clustering successfully identified generalizable visual search strategies during a curved roadway driving scenario. This methodology can reduce the need for manual preprocessing of raw eye-tracking data, allowing for analysis of larger, more generalizable datasets to study visual search strategy.
BACKGROUND AND OBJECTIVES Young drivers are overrepresented in crashes, and newly licensed drivers are at high risk, particularly in the months immediately post-licensure. Using a virtual driving assessment (VDA) implemented in the licensing workflow in Ohio, this study examined how driving skills measured at the time of licensure contribute to crash risk post-licensure in newly licensed young drivers. METHODS This study examined 16 914 young drivers (<25 years of age) in Ohio who completed the VDA at the time of licensure and their subsequent police-reported crash records. By using the outcome of time to first crash, a Cox proportional hazard model was used to estimate the risk of a crash during the follow-up period as a function of VDA Driving Class (and Skill Cluster) membership. RESULTS The best performing No Issues Driving Class had a crash risk 10% lower than average (95% confidence interval [CI] 13% to 6%), whereas the Major Issues with Dangerous Behavior Class had a crash risk 11% higher than average (95% CI 1% to 22%). These results withstood adjusting for covariates (age, sex, and tract-level socioeconomic status indicators). At the same time, drivers licensed at age 18 had a crash risk 16% higher than average (95% CI 6% to 27%). CONCLUSIONS This population-level study reveals that driving skills measured at the time of licensure are a predictor of crashes early in licensure, paving the way for better prediction models and targeted, personalized interventions. The authors of future studies should explore time- and exposure-varying risks.
In this work we focus on the problem of identifying drivers with neurocognitive impairment (NCI), specifically an NCI specific to people with HIV (PWH) called HIV-associated neurocognitive disorders (HAND) directly from driving simulator data. Since NCI-screening is typically only effective for more progressed forms of HAND, there is a critical need to identify individuals that should be referred to specialists in order to mitigate potentially dangerous driving behaviors and improve their quality of life. Data collected from (n = 81) study participants that used the virtual driving test (VDT) platform were analyzed in order to predict which drivers had NCI. Of the (n = 62) PWH participants recruited, (n = 35) had HAND; of the remaining (n = 19) HIV negative participants, (n = 7) had non-HAND NCI (e.g., Parkinson’s Disease, Alzheimer’s, etc.). In three separate experiments, subsets of VDT data were first selected via Kruskal-Wallis feature ranking and then used as ensemble inputs to classify whether or not drivers had NCI. Within the PWH population, HAND could be classified with 69.4% accuracy and a risk ratio of 2.09 (95% CI 1.52, 2.65); within the HIV negative population, non-HAND NCI could be classified with 84.2% accuracy, risk ratio of 8.25 (6.34, 10.16); and within the combined population, NCI (regardless of causation) could be classified with 63.0% accuracy, risk ratio of 1.67 (1.22, 2.11).
Motor vehicle crash rates are highest immediately after licensure, and driver error is one of the leading causes. Yet, few studies have quantified driving skills at the time of licensure, making it difficult to identify at-risk drivers before independent driving. Using data from a virtual driving assessment implemented into the licensing workflow in Ohio, this study presents the first population-level study classifying degree of skill at the time of licensure and validating these against a measure of on-road performance: license exam outcomes. Principal component and cluster analysis of 33,249 virtual driving assessments identified 20 Skill Clusters that were then grouped into 4 major summary "Driving Classes"; i) No Issues (i.e. careful and skilled drivers); ii) Minor Issues (i.e. an average new driver with minor vehicle control skill deficits); iii) Major Issues (i.e. drivers with more control issues and who take more risks); and iv) Major Issues with Aggression (i.e. drivers with even more control issues and more reckless and risk-taking behavior). Category labels were determined based on patterns of VDA skill deficits alone (i.e. agnostic of the license examination outcome). These Skill Clusters and Driving Classes had different distributions by sex and age, reflecting age-related licensing policies (i.e. those under 18 and subject to GDL and driver education and training), and were differentially associated with subsequent performance on the on-road licensing examination (showing criterion validity). The No Issues and Minor Issues classes had lower than average odds of failing, and the other two more problematic Driving Classes had higher odds of failing. Thus, this study showed that license applicants can be classified based on their driving skills at the time of licensure. Future studies will validate these Skill Cluster classes in relation to their prediction of post-licensure crash outcomes.
Significance Existing screening tools for HIV-associated neurocognitive disorders (HAND) are often clinically impractical for detecting milder forms of impairment. The formal diagnosis of HAND requires an assessment of both cognition and impairment in activities of daily living (ADL). To address the critical need for identifying patients who may have disability associated with HAND, we implemented a low-cost screening tool, the Virtual Driving Test (VDT) platform, in a vulnerable cohort of people with HIV (PWH). The VDT presents an opportunity to cost-effectively screen for milder forms of impairment while providing practical guidance for a cognitively demanding ADL. Objectives We aimed to: (1) evaluate whether VDT performance variables were associated with a HAND diagnosis and if so; (2) systematically identify a manageable subset of variables for use in a future screening model for HAND. As a secondary objective, we examined the relative associations of identified variables with impairment within the individual domains used to diagnose HAND. Methods In a cross-sectional design, 62 PWH were recruited from an established HIV cohort and completed a comprehensive neuropsychological assessment (CNPA), followed by a self-directed VDT. Dichotomized diagnoses of HAND-specific impairment and impairment within each of the seven CNPA domains were ascertained. A systematic variable selection process was used to reduce the large amount of VDT data generated, to a smaller subset of VDT variables, estimated to be associated with HAND. In addition, we examined associations between the identified variables and impairment within each of the CNPA domains. Results More than half of the participants (N = 35) had a confirmed presence of HAND. A subset of twenty VDT performance variables was isolated and then ranked by the strength of its estimated associations with HAND. In addition, several variables within the final subset had statistically significant associations with impairment in motor function, executive function, and attention and working memory, consistent with previous research. Conclusion We identified a subset of VDT performance variables that are associated with HAND and assess relevant functional abilities among individuals with HAND. Additional research is required to develop and validate a predictive HAND screening model incorporating this subset.
In this paper we introduce a novel algorithm called Iterative Section Reduction (ISR) to automatically identify spatial regions wherein time series were recorded that are predictive of a target classification task. Specifically, using data collected from a driving simulator study, we identify which spatial regions (dubbed sections) along the simulated routes tend to manifest driving behaviors that are predictive of the presence of Attention Deficit Hyperactivity Disorder (ADHD). Identifying these sections is important for two main reasons: (1) to improve predictive accuracy of the trained ADHD screening models by filtering out non-predictive time series data, and (2) to gain insights into which on-road scenarios (dubbed events) elicit distinctly different driving behaviors from patients undergoing treatment for ADHD versus those that are not. Our experimental results show both improved classification performance over prior efforts and good alignment between the predictive sections identified and scripted on-road events in the simulator (negotiating turns and curves).
Young drivers are overrepresented in crashes and newly licensed drivers are at highest risk with average crash rates peaking in the months immediately post-licensure. Using a virtual driving assessment (VDA) implemented in the licensing workflow in Ohio, our prior work identified distinct classifications of skills (4 summary Driving Classes consisting of specific Skill Clusters) in license applicants that were differentially associated with license examination outcomes. Building on this, the current study examined how these skill classifications from the time of licensure predict post-licensure crash-risk, among 16,914 newly licensed young drivers (under age 25). Using the outcome of time-to-first crash, a Cox proportional hazard model was used to estimate the risk of crash during the follow-up period as a function of VDA Driving Class (and Skill Cluster) membership. The results confirmed that VDA driving skills measured at the point of licensure could predict crash risk early in licensure: specifically, the best performing No Issues Driving Class had a crash risk 10% lower than average risk (95% CI 13% - 6%), while the Major Issues with Dangerous Behavior Class had a crash risk 11% higher than average (95% CI 1% - 22%). However, the more moderate Minor Issues and the Major Issues Driving Classes were not significantly different from average crash risk. These results withstood adjusting for covariates (age, sex, and tract-level SES indicators), which had little impact on hazard ratios for crash risk. However, drivers licensed at age 18 had a crash risk 16% higher than average (95% CI 6% -27%). Thus, this is the first population-level study showing that driving skills measured at the time of licensure can predict crash-risk early in licensure, paving the way for targeted and personalized interventions. Future studies should explore time-varying risk and the effect of varying exposure on crash risk.
In this paper we introduce a novel algorithm called Iterative Section Reduction (ISR) to automatically identify sub-intervals of spatiotemporal time series that are predictive of a target classification task. Specifically, using data collected from a driving simulator study, we identify which spatial regions (dubbed sections) along the simulated routes tend to manifest driving behaviors that are predictive of the presence of Attention Deficit Hyperactivity Disorder (ADHD). Identifying these sections is important for two main reasons: (1) to improve predictive accuracy of the trained models by filtering out non-predictive time series sub-intervals, and (2) to gain insights into which on-road scenarios (dubbed events) elicit distinctly different driving behaviors from patients undergoing treatment for ADHD versus those that are not. Our experimental results show both improved performance over prior efforts (+10% accuracy) and good alignment between the predictive sections identified and scripted on-road events in the simulator (negotiating turns and curves).
In this paper, we identify the on-road scenarios within a simulated driving environment where a group of clinical trial participants (n= 30) with and without Attention Deficit Hyper-activity Disorder (ADHD) drive perceivably different fromone another. We partition the simulated routes into smaller non-overlapping sections in order to determine which sections elicit behaviors that are predictive of ADHD. Then, we develop section-specific classifiers, which are used as voters in bagging ensemble classifiers. Our results show gains in classifying ADHD (increase in 5-fold average evaluation accuracy) over our previous efforts, as well as providing explainable evidence that driving behaviors indicative of ADHD tend to be exhibited in turns and curves.
Background A large Midwestern state commissioned a virtual driving test (VDT) to assess driving skills preparedness before the on-road examination (ORE). Since July 2017, a pilot deployment of the VDT in state licensing centers (VDT pilot) has collected both VDT and ORE data from new license applicants with the aim of creating a scoring algorithm that could predict those who were underprepared. Objective Leveraging data collected from the VDT pilot, this study aimed to develop and conduct an initial evaluation of a novel machine learning (ML)–based classifier using limited domain knowledge and minimal feature engineering to reliably predict applicant pass/fail on the ORE. Such methods, if proven useful, could be applicable to the classification of other time series data collected within medical and other settings. Methods We analyzed an initial dataset that comprised 4308 drivers who completed both the VDT and the ORE, in which 1096 (25.4%) drivers went on to fail the ORE. We studied 2 different approaches to constructing feature sets to use as input to ML algorithms: the standard method of reducing the time series data to a set of manually defined variables that summarize driving behavior and a novel approach using time series clustering. We then fed these representations into different ML algorithms to compare their ability to predict a driver’s ORE outcome (pass/fail). Results The new method using time series clustering performed similarly compared with the standard method in terms of overall accuracy for predicting pass or fail outcome (76.1% vs 76.2%) and area under the curve (0.656 vs 0.682). However, the time series clustering slightly outperformed the standard method in differentially predicting failure on the ORE. The novel clustering method yielded a risk ratio for failure of 3.07 (95% CI 2.75-3.43), whereas the standard variables method yielded a risk ratio for failure of 2.68 (95% CI 2.41-2.99). In addition, the time series clustering method with logistic regression produced the lowest ratio of false alarms (those who were predicted to fail but went on to pass the ORE; 27.2%). Conclusions Our results provide initial evidence that the clustering method is useful for feature construction in classification tasks involving time series data when resources are limited to create multiple, domain-relevant variables.
Voice User Interfaces (VUIs) are becoming increasingly popular. However, how VUIs can adapt to user differences remains insufficiently understood. We analyze usage data from a user study (n=50) where participants interacted with an unfamiliar VUI. Through automated clustering and statistical analysis, we present user models of their behavior patterns. We found user behavior can be grouped into three clusters: people who become proficient with the system and typically stay proficient while completing different tasks, people who exhibit an exploratory approach to completing tasks, and people who struggled to complete tasks. We discuss design implications based on these behavior clusters.