The crash-frequency model is to use statistical methods for identifying factors that are significantly associated with the consequence of a traffic crash, and their relationships. The response variable is the number of crashes, vehicle occupants, road-users, or similar entities. The models can be used for establishing relationships between the response variable and explanatory variables, examining variables, evaluating the sensitivity of the variables, and prediction. A variety of models have been developed to account for prevailing data issues and methodological limitations such as over- and underdispersion. The common modeling approaches include the negative binomial or Poisson-gamma, Poisson-lognormal, Conway–Maxwell, Random Parameters, and Dirichlet process models. The selection of the models should be based on the study objectives, goodness-of-fit, and goodness-of-logic. More complex models are necessarily the best models.
Before–after studies are frequently used for assessing the safety impacts of various interventions. Although these studies have a better control over the within-observation variance, they can be affected by the regression-to-the-mean and site selection biases. Several types of methods have been proposed. They include the simple or naïve method, the simple method with the control group, the empirical Bayes, and full Bayes methods. More recently, a naïve adjustment method and a before–after method using survival analysis have been proposed to overcome or minimize the biases listed earlier. Equations are provided to properly estimate the sample size requirements for conducting before–after studies.
The study examines U.S. port disruptions under tropical cyclones (TCs) by assessing key factors associated with port operational impact, recovery duration, and freight network-level impact, where generalizable empirical insights across TCs remain limited. We construct the CyPort dataset to capture these impacts for 145 U.S. principal ports across 90 tropical cyclones (2015-2023), integrating TC exposure, port characteristics, and freight network centrality. We employ the Random Parameter Negative Binomial-Lindley model to obtain more reliable resilience insights, accounting for data imbalance and unobserved heterogeneity across port-TC interactions. Results show that port-level and freight network-level impacts are governed by distinct drivers. Factors associated with reduced disruption are identified, including hinterland characteristics at the port level (e.g., railway connection and workforce factors) and betweenness centrality at the network level. Heterogeneous effects show wind, surge, rainfall, and railway connection affect ports differently, highlighting context-dependent resilience. Our insights support port preparedness, coordination, and long-term infrastructure investment.
The standard negative binomial (NB) and negative binomial-Lindley (NBL) models may not completely capture variations associated with time and space. To address these limitations, this study derives an enhanced version of the NBL model that simultaneously incorporates spatiotemporal random parameters to account for three key factors: temporal variations, spatial variations, and datasets with a large amount of zero observations. The proposed framework extends the standard NBL model by allowing both model coefficients and Lindley parameters to vary over time and across space together. This approach captures the heterogeneity in crash data caused by variations in spatiotemporal factors associated with traffic patterns, environmental conditions, or geographic differences, while also addressing the high proportion of zero observations typically found in such datasets. The results demonstrate that the proposed model reasonably recovers the true parameters. Findings indicate that the proposed model outperforms those that solely account for temporal or spatial variations.
Bicyclist injury severity varies notably across age groups due to differences in riding experience, risk-taking behavior, physical vulnerability, and response times. This study explores the demographic and temporal heterogeneity of bicyclist injury outcomes by examining crashes involving young, middle-aged, and older cyclists before, during, and after the COVID-19 lockdowns. Using detailed single-vehicle/single-bicycle crash data from Florida spanning 2019 to 2021, random parameters logit models with heterogeneity in means and variances of the random parameters were estimated, considering severe injury, minor injury, and no visible injury as possible injury outcomes. Likelihood ratio tests indicated significant temporal shifts of the model parameters across the three yearly time periods as well as age-related differences in model parameters. Considering a wide range of explanatory variables related to bicyclists, drivers, vehicles, roadways, and environmental conditions, the findings underscore the importance of proper model specification and reveal statistically significant variations in injury severity patterns across age groups and time periods. The findings also show that, in some instances, the commonality of parameters over time periods and age groups is statistically justified, giving additional insights into the determinants of injury severity and underscoring the importance of proper model specification.
The crash-severity model is to use the statistical methods for identifying factors that are significantly associated with the consequence of a traffic crash, and their relationships. The response variable is the person who sustains the most severe injury in a crash in the KABCO scale (i.e., killed, incapacitating injury, nonincapacitating injury, possible injury, and no injury). A variety of models have been developed to account for prevailing data issues and methodological limitations such as the ordinal structure of KABCO, unobserved data heterogeneity, omitted variable bias, and imbalanced data between injury severity levels. The common modeling approaches include logistic, probit, and their variations. The impact of a factor on the injury severity levels can be estimated through its marginal effect or odds ratio.
This paper presents a digital-twin platform for active safety analysis in mixed traffic environments. The platform is built using a multi-modal data-enabled traffic environment constructed from drone-based aerial LiDAR, OpenStreetMap, and vehicle sensor data (e.g., GPS and inclinometer readings). High-resolution 3D road geometries are generated through AI-powered semantic segmentation and georeferencing of aerial LiDAR data. To simulate real-world driving scenarios, the platform integrates the CAR Learning to Act (CARLA) simulator, Simulation of Urban MObility (SUMO) traffic model, and NVIDIA PhysX vehicle dynamics engine. CARLA provides detailed micro-level sensor and perception data, while SUMO manages macro-level traffic flow. NVIDIA PhysX enables accurate modeling of vehicle behaviors under diverse conditions, accounting for mass distribution, tire friction, and center of mass. This integrated system supports high-fidelity simulations that capture the complex interactions between autonomous and conventional vehicles. Experimental results demonstrate the platform's ability to reproduce realistic vehicle dynamics and traffic scenarios, enhancing the analysis of active safety measures. Overall, the proposed framework advances traffic safety research by enabling in-depth, physics-informed evaluation of vehicle behavior in dynamic and heterogeneous traffic environments.
A crash occurs as a result of driver behavior in the traffic stream. Studying the driver's car-following behavior, modeling the relationships between crashes and traffic volume, and mapping crash typologies to a variety of traffic conditions are instrumental in discovering the underlying process of crash occurrence. When the relationships between crash occurrence and traffic flow variables are unraveled, real-time crash risk prediction models can be used to predict crash risk during a short time period given the observed or simulated traffic conditions.
As the world's population ages, ensuring the safety of older adult pedestrians has become an urgent priority in transportation planning. However, most existing studies rely on global models that overlook spatial heterogeneity and fail to capture nonlinear, location-specific interactions between the environmental factors and crash outcomes. Moreover, subjective perceptions (e.g., how safe or walkable an area feels) may influence pedestrian behavior and crash exposure but are underexplored in traffic safety research. This study addresses these gaps by integrating subjective perception indicators extracted from Street View Images (SVI) with machine learning models to examine the severity of older adult pedestrian crashes at intersections in Taipei City. Three modeling frameworks are evaluated and compared: global Negative Binomial Regression (NBR), Geographically Weighted Negative Binomial Regression (GWNBR), and GeoShapley, a spatially interpretable extension of the SHAP framework for XGBoost. A total of 36 environmental and perceptual variables are evaluated in relation to injury and fatal crash frequencies. Among these models, GeoShapley achieved the best performance and revealed that spatial location (GEO) and its interactions with environmental factors and subjective perceptions were among the most influential predictors. In some areas, higher walkability was associated with reduced injury crash frequencies, especially in older and urban districts. In addition, the effect of convenience stores and nursing homes on the frequency of fatal crashes varied significantly across locations, reflecting the spatial clustering of pedestrian activity and older adults. Overall, the findings demonstrate the value of spatially explicit machine learning tools and subjective perceptions in understanding localized crash dynamics in aging urban populations.
The exploratory analysis of data is the first important step in any research study. The main objective of conducting exploratory analyses is to achieve maximum insights into the data by employing a variety of techniques. Typically, these techniques are employed before building a statistical model or conducting more advanced analyses. Exploratory data analyses are mainly performed to discover variable patterns, to identify outliers, to test hypothesis, and to examine data assumptions using summary statistics and graphical representations. Through the visualization of data, exploratory analyses can reveal the underlying structure of the data and help select one or more models for analyzing the data.
Crashes are very complex and multidimensional events. They can be analyzed from the perspective of driver, roadway environment, and vehicle. They can also be analyzed from the perspective of mathematical equations. Various types of data can be collected for analyzing crash data. They include crash, traffic flow, roadway characteristics, vehicle occupants, site visits, Google Earth, and Streetview. A four-stage modeling framework is proposed for analyzing highway crash data. Models can be estimated using the maximum likelihood estimation or the Bayesian methods. Several goodness-of-fit methods are described for evaluating crash-frequency and crash severity models. The goodness-of-fit methods to evaluate the performance of models can be grouped as either likelihood-based or model error-based.
The concept of a multivariate distribution is essential in statistics, allowing the simultaneous analysis of related variables. In transportation safety, such models are effective for studying crash data across multiple categories, improving predictions and evaluations of safety measures. This paper extends the negative binomial weighted Lindley (NB-WLindley) distribution, known for handling highly dispersed or sparse data, into a multivariate framework. The proposed multivariate NB-WLindley generalized linear model treats each crash category as a random variable dependent on other categories within a joint framework while preserving the marginal distributional properties. It is hierarchically defined as a mixture of NB and multivariate weighted Lindley distributions and incorporates a dependence structure to explain correlations among categories. Applications to two crash datasets show that the multivariate NB-WLindley model can simultaneously capture different crash types and severities, identifying dependencies that univariate models cannot. The study also develops a random parameters version of the model to address unobserved heterogeneity, which consistently outperforms the fixed-parameter version and yields stronger predictive performance. Overall, this work demonstrates that the proposed multivariate model offers a more flexible and accurate approach to crash analysis, enhancing the ability to capture variability and interdependence across crash categories. It provides a practical tool for improving safety assessments and supporting data-driven decision-making in transportation safety research.
Spatial analysis of crash data is to study the distribution of crash locations to identify the spatial patterns and their underlying causes. Spatial association indicators such as Getis G and Moran's I can measure the clustering of crash attributes of a set of geographic features at a global or a local scale. Kernel density estimation, Ripley's k-function and cross-k function analyze crash points by calculating crash intensity or the correlation between two distinct sets of points. Spatial regression methods explicitly consider spatial dependency of crash observations and spatial heterogeneity in the relationship between crashes and their contributing factors.
For any transportation agency, safety is considered a priority and, thus, transportation agencies should routinely evaluate the safety risk of the transport network to identify sites for further investigation and implementing potential treatments. The important assumption is that the adverse roadway design or operational characteristics play a vital role in the occurrence of roadway crashes. Although treating all locations that experience safety issues could be an option, budget constraints and the scarcity of resources force agencies to identify a subsample of locations, called as hazardous sites, to implement countermeasures where the maximum benefit can be achieved from targeted and cost-effective treatments.
The application of data mining and machine learning techniques in the highway safety analysis has boomed, resulting from the new and emerging data sources, powerful algorithms, handy software applications, and comparable or superior performance in crash prediction. The broad selection of techniques ranging from exploratory data analysis such as association rules, clustering analysis, decision tree models, Bayesian networks to more sophisticated neural network models, and support vector machines presents great opportunities to consider a large multitude of factors and explore intricate relationships among them. Most techniques are readily implemented through commercial or free statistical software packages such as R.
The challenges associated with crash-based analyses offer the opportunities for surrogate safety analysis. Safety measures have been developed from and around the surrogate events that are related to and are more frequently observed than crashes, such as traffic conflicts, near misses. Most measures are based on the proximal of time or location between road users. Common indicators are time to collision, post encroachment time, and proportion of stopping distance. The theoretical breakthrough from data-driven approach to extreme value models, such as block maxima and peak over threshold, is vital to model the relationships between surrogate measures and other factors, and understand the collision mechanisms.
Real-time corridor-wide crash-occurrence risk (COR) prediction is challenging, since existing near-miss EVT models oversimplify collision geometry, neglect vehicle-infrastructure (V-I) interactions, and fail to adequately account for spatial heterogeneity in traffic and roadway conditions. To do so, this study develops a geometry-aware 2D-TTC near-miss extraction and integrates it with a hierarchical Bayesian structure grouped random parameters (HBSGRP-UGEV) to estimate short-term COR in urban corridors. Building on prior grouped EVT formulations while explicitly accommodating both V-V and V-I near-miss processes within a unified corridor-wide modeling framework. High-resolution trajectories from the Argoverse-2 dataset were analyzed across 28 sites on Miami's Biscayne Boulevard to extract extreme near-miss events. The model incorporates vehicle dynamics and roadway features as covariates, with partial pooling across segments and intersections to capture corridor-wide heterogeneity. Results show that the HBSGRP-UGEV framework outperforms fixed-parameter HBSFP-UGEV models, reducing DIC by up to 7.5% (V-V) and 3.1% (V-I). Predictive validation using ROC-AUC confirms strong accuracy (0.89 for V-V segments, 0.82 for intersections, 0.79 for V-I segments, and 0.75 for intersections). Grouped random-parameters (HBSGRP) framework indicate that relative (speed, distance, and deceleration) dominate V-V near-miss risk on segments, whereas V-I segment risk is primarily associated with relative distance; at intersections, V-V risk is driven by relative (speed and distance), while V-I dynamics exhibit no statistically significant effects. These findings demonstrate the value of a geometry-aware, spatially adaptive framework for proactive corridor safety management, supporting both real-time interventions and long-term Vision Zero goals.
Introduction: Pedestrian safety has become a critical concern with the rising global population of older adults. Older pedestrians face higher crash risks due to age-related physical limitations, yet road infrastructure often fails to address their specific needs. Most studies treat older adults as a single group, overlooking variations in mobility and behavior. Also, few studies have included micro-level environmental variables (such as points of interest) and spatial models. Method: This study categorizes older adults into three age groups and analyzes 783 pedestrian crashes at intersections in Taipei City. We identify key environmental factors influencing crash frequency using a Geographically Weighted Negative Binomial Regression (GWNBR) model. Results: Results show that intersections with more motorcycles or a higher number of nearby restaurants are associated with increased pedestrian crashes across all older adults. However, risk factors vary by age group: subway stations and supermarkets contribute to crashes among the young-old (65-74) and middle-old (75-84), while nursing homes are linked to higher crash rates for the oldest-old (85+). Additionally, shorter signal cycles pose a greater risk for the oldest old due to slower walking speeds. Practical applications: These findings highlight the need for age-specific safety measures, such as adjusted signal timing and improved pedestrian facilities near high-risk locations.
Pedestrian safety remains a pressing concern near bus stops along urban transit, where frequent pedestrian-vehicle interactions occur. While prior research has primarily focused on intersections and midblock locations, bus stops have often been treated as secondary contributors rather than as distinct sites requiring targeted safety assessments. This has left a critical gap in understanding how traffic exposure, roadway characteristics, and bus stop design features specifically influence pedestrian crash risks around bus stop locations. To address these gaps, this study develops a comprehensive framework focused on pedestrian safety in the vicinity of bus stops. The proposed approach employs a Random Parameters Negative Binomial-Lindley (RPNB-L) model to account for unobserved heterogeneity and site-specific variability. Using data from 596 bus stops in Fort Worth, Texas (2018-2022), the model identifies that higher pedestrian crash frequencies are significantly associated with increased AADT, elevated boarding activity, and the absence of key safety elements such as crosswalks, medians, and lighting. Conversely, far-side bus stop placement, signalized intersections, sidewalks, and mixed-use development are associated with lower crash risks. Roads near schools and those with speed limits of ≤35 mph show elevated crash risk. To support proactive safety management, the study integrates a Full Bayes-based Potential for Safety Improvement (PSI) metric, enabling the identification of hazardous stops and high-risk corridors. By unifying advanced count-based modeling with strategic risk prioritization, this research offers actionable, data-driven insights for improving pedestrian safety near bus stops.