This study is about smallest acceptable sample size determination in experimental design studies involving a driving simulator. The smallest acceptable sample size should be specified so researchers can make accurate inferences about their studied populations. However, the number of samples typically collected is largely subject to the expense of data collection. Working out the methodology of estimating the required number of subjects based on an initially small number is a better way for researchers to determine the smallest acceptable sample size in the experiment. Predictor estimate precision and prediction accuracy are major factors for conducting experiments. Accordingly, this study estimates the smallest acceptable sample size, with emphasis on coefficient estimation and prediction accuracy for selected significant variables. The smallest acceptable sample size is chosen to be the maximum value returned by both coefficient estimation calculation and accuracy prediction calculation approaches. This methodology is flexible and scalable, and can be tailored to other experimental situations. To validate the appropriateness of this procedure, a more than sufficient sample of 50 drivers was recruited. The smallest acceptable sample size was determined backwardly, based on the variable coefficient convergence trends of the mean squared error (MSE) curves of the significant variables. Both the clear converging trends of the MSE curves and the proposed method indicated that 30 was an acceptable sample size.
Short-window crash prediction is a fundamental step in proactive traffic safety management that can monitor traffic conditions in real time, identify unsafe traffic dynamics, and implement suitable interventions for traffic conflicts. Short-window (e.g., hourly) traffic-collision count data, however, exhibits excessive zeros and serial autocorrelation. Most of the commonly used regression-based models fail to address excessive zeros and temporal structure simultaneously in hourly traffic-collision prediction. For example, Hurdle models and zero-inflated (ZI) models can address the overdispersion issue caused by excessive zeros but lack power to control the significant spatiotemporal characteristics inherent in the time-based collision data. To overcome these issues simultaneously, this paper develops a novel statistical model termed Zero-Inflated Logarithmic link for count Time series (ZILT) which is based on the framework of ZI models. Covariates (e.g., speed, vehicle type, and traffic volume) were extracted through deep-learning computer vision methods in vehicle detection and tracking on the image space. This new statistical model (i.e., ZILT) performs better at solving the issues of excessive zeros and serial dependencies. The prediction accuracy of the ZILT model improved by around 5% in relation to zero-inflated Poisson (ZIP) and Hurdle models. Results show that traffic crashes happening in the previous hour and other covariates such as truck-to-car ratio, holiday effect, traffic flow, and speed have significant influence on collision occurrence. Findings from this study could be utilized by relevant transport agencies in developing engineering interventions and countermeasures to proactively manage road safety.
Traditionally, traffic characteristics such as speed, volume, and travel time are obtained from a range of sensors and systems such as inductive loop detectors (ILDs), automatic number plate recognition cameras (ANPR), and GPS-equipped floating cars. However, many issues associated with these data have been identified in the existing literature. Although roadside surveillance cameras cover most road segments, especially on freeways, existing techniques to extract traffic data (e.g., speed measurements of individual vehicles) from video are not accurate enough to be employed in a proactive traffic management system. Therefore, this paper aims to develop a technique for estimating traffic data from video captured by surveillance cameras. This paper then develops a deep learning-based video processing algorithm for detecting, tracking, and predicting highly disaggregated vehicle-based data, such as trajectories and speed, and transforms such data into aggregated traffic characteristics such as speed variance, average speed, and flow. By taking traffic observations from a high-quality LiDAR sensor as ‘ground truth’, the results indicate that the developed technique estimates lane-based traffic volume with an accuracy of 97%. With the application of the deep learning model, the computer vision technique can estimate individual vehicle-based speed calculations with an accuracy of 90–95% for different angles when the objects are within 50 m of the camera. The developed algorithm was then utilised to obtain dynamic traffic characteristics from a freeway in southern China and employed in a statistical model to predict monthly crashes.
Predicting short-term traffic crashes is challenging due to an imbalanced data set characterized by excessive zeros in noncrash counts, random crash occurrences, spatiotemporal correlation in crash counts, and inherent heterogeneity. Existing models struggle to effectively address these distinct characteristics in crash data. This paper proposes a new joint model by combining the time-series generalized regression neural network (TGRNN) model and the binomially weighted convolutional neural network (BWCNN) model. The joint model aims to capture all these characteristics in short-term crash prediction. The model was trained and tested using real-world, highly disaggregated traffic data collected with inductive loop detectors on the M1 motorway in the UK in 2019, along with crash data extracted from the UK National Accident Database for the same year. The short-term is defined as a 30-min interval, providing sufficient time for a traffic control center to implement interventions and mitigate potential hazards. The year was segmented into 30-min intervals, resulting in a highly imbalanced data set with over 99.99% noncrash samples. The joint model was applied to predict the probability of a crash occurrence by updating both the crash and traffic data every 30 min. The findings revealed that 75.3% of crashes and 81.6% of noncrash events were correctly predicted in the southbound direction. In the northbound direction, 78.1% of crashes and 80.2% of noncrash events were accurately captured. Causal analysis and model-based interpretation were used to analyze the relative importance of explanatory variables regarding their contribution to crashes. The results reveal that speed variance and speed are the most influential factors contributing to crash occurrence.
Tunnels on mountainous freeways are affected by abrupt changes in brightness, complex geometric alignments, heavy traffic flow, bad weather, and other factors, some of which contribute toward tunnels having more traffic crashes than other sections of the freeway. Previous research, however, has given limited attention to tunnel length and heterogeneity in the parts of the tunnels, such as at the tunnels’ entrance and exit zones. Focusing on 36 tunnels on the Guidu Freeway in China’s Guizhou Province, this study collects data on crashes and their influencing factors over 2 years (2020–2021), constructs a negative binomial panel data random effects model, and analyzes single-vehicle crashes, multi-vehicle crashes, and total crashes. The results show that: 1) multi-vehicle crashes occur throughout the tunnel sections, 2) crashes are more likely to occur in long tunnel sections, 3) the crash frequency from the tunnel entrance zone to the mid zone is higher than in other areas of the tunnel, 4) the crash frequency is higher for circular curve/easy curve tunnel sections than for straight tunnel sections, 5) the crash frequency is higher for downhill and concave curve sections than for flat sections, 6) the crash frequency increases with heavy traffic flow and adverse weather conditions, and 7) the crash frequency increases as road surface skidding resistance and ride quality decrease. These findings can provide theoretical support for engineering improvement and the formulation and revision of specifications for designing freeway tunnel sections, especially in mountainous areas.
Driving Simulator, a powerful simulation tool, has already been used in safety evaluation of roadway geometric design during the pre-construction design stage. Conventional ways of estimating proper sample size include the empirical method, the resource equation, power analysis, and the Bayesian method. However, significant boundaries and prior distributions of operational indices are hard to identify in simulator studies, which makes it difficult to use conventional ways in choosing the acceptable sample size. This study proposes an empirical method to infer proper sample size. The Tongji University eight-degree-of-freedom driving simulator was utilized to collect continuous driving behavior data from a simulated mountainous freeway. Vehicle speed and lane departure events were selected as the indices to measure the influence of geometric design features on operational efficiency and safety. A mixed linear model and a mixed logistic regression model were used to assess the relationships between geometric design features and vehicle speed and lane departure. Random sampling was used to choose 10 samples of 5 to 50 drivers from a total of 55 drivers. Acceptable sample size was determined based on the parameter coefficient convergence elbow points of the mean squared error (MSE) curves of significant variables. The clear elbow points of the MSE curves indicate that 30 is an acceptable sample size.
A traffic crash is becoming one of the major factors that leads to unexpected death in the world. Short window traffic crash prediction in the near future is becoming more pragmatic with the advancements in the fields of artificial intelligence and traffic sensor technology. Short window traffic prediction can monitor traffic in real time, identify unsafe traffic dynamics, and implement suitable interventions for traffic conflicts. Crash prediction being an important component of intelligent traffic systems, it plays a crucial role in the development of proactive road safety management systems. Some near future crash prediction models were put forward in recent years; further improvements need to be implemented for actual applications. This paper utilizes traffic accident data from the study Freeway in China to build a time series-based count data model for daily crash prediction. Lane traffic flow, weather information, vehicle speed, and truck to car ratio were extracted from the deployment of non-intrusive detection systems with support of the Bridge Management Administration study and were input into the model as independent variables. Different types of prediction models in machine learning and time series forecasting methods such as boosting, ARIMA, time-series count data model, etc. are compared within the paper. Results show that integrating time series with a count data model can capture traffic accident features and account for the temporal structure for variable serial correlation. A prediction error of 0.7 was achieved according to Root Mean Squared Deviation.
In February 2020, the Stockholm Declaration was announced, urging states toward the United Nation’s target of a 50% reduction in traffic deaths and injuries by 2030, with the potential to achieve Vision Zero by 2050. The aim of this research is threefold namely i) to assess if selected developed countries are likely to achieve the 2030 target, ii) to use the Gompertz model to predict future road trauma trends based on historical data and iii) to understand how number of road traffic fatality vary between the developed countries. After identifying potential reasons behind the patterns, time series models were applied to identify the effects of exposure variables on traffic fatalities. To assess the likelihood of meeting the U.N. target, autoregressive integrated moving average (ARIMA) models were used for obtaining trustworthy forecasts of road traffic fatalities using data from the last five decades from seven high-income countries. The total number of fatalities, vehicle-km travelled, vehicle ownership, GDP, GDP per capita, urbanization, population density and country population were used to develop the ARIMA models. The predictive performance of the models was validated for each country, and all were found to be within the 95% confidence interval. Estimated forecasts in all seven countries appear to be realistic, with Greece and the U.K., the only countries falling short of achieving the U.N.’s 2030 target. With these results, both developed and developing countries can review and reconsider the effects of safety interventions and other socioeconomic influences on achieving reductions in road fatalities. Interventions can be added to the existing model to ascertain their effect on the predicted number of fatalities.
Abstract The authors have requested that this preprint be removed from Research Square.
In order to study the influence of the traffic characteristics on traffic accidents in extra long tunnel, the main measurement indicators of traffic flow during the time of traffic accidents are matched with the accident information to form a data set of the number of traffic accidents and the hourly traffic flow of the accident. Vehicle ratio and the number of accidents are mainly used as the characteristic indicators of traffic flow. At the same time, the longitudinal distribution law of the average speed of traffic flow and the number of traffic accidents in the extra long tunnel is studied. Based on the superposition principle, the extra long tunnel is divided into 5 traffic safety zones. This paper analyzes the distribution of time, morphology, cause of accident, and other characteristics in different traffic safety zones, finding that the shape of traffic accidents in extra long tunnel is mainly rear-end collisions. Improper operation and illegal lane changes are the main causes of accidents.
Excessive longitudinal deceleration and acceleration reduce driving comfort and increase safety risk. There is a dearth of research, however, on how geometric design characteristics, especially complex alignments and their adjacent segments, affect deceleration and acceleration. The Tongji University driving simulator was used to collect vehicle operation data. Deceleration and acceleration values were measured and classified as deceleration, near-cruising and acceleration. A random parameter multinomial logistic regression model was used to reflect the relationships between drivers' deceleration and acceleration and the geometric alignments of a mountainous freeway. Vehicle speed was treated as a random slope parameter to account for the heterogeneity among subjects. Model results showed deceleration and acceleration were significantly influenced by the fixed effects of maximum curvature in the following 50 m, difference between maximum and minimum slope in the preceding 100 m, maximum slope change in the following 400 m, and maximum curvature in the preceding 50 m. These findings can assist in developing safer design and evaluation standards for mountainous freeways.