Utility infrastructure assets in the United States continue to grow as millions of utility features were installed within the properties of state and local agencies. With this growth, the management of the utility data records is becoming a complex problem in terms of large amounts of data. On one hand, management of data for utility infrastructures is extremely valuable to state and local agencies because the timely access to utility-related information is a significant requirement for the delivery of construction and renovation projects on time and within budget. On the other hand, many challenges arise, such as difficulties in effective data storage of complex and messy datasets, data analysis, and data visualization. Utility owners face challenges in collecting utility data in standardized formats, data storage, and providing easy access to all stakeholders. Using a case study in Nevada, this paper demonstrates how tools and a strategic workflow process can be harnessed to develop an end-to-end management solution for large and complex data of a utility infrastructure. This end-to-end utility data management solution builds upon existing systems which are not adequate for large utility data management because they are non-scalable, do not allow for access by multiple users, involve manual data uploads, do not control consistency of data attributes, and lack visualization tools for non-GIS experts. In addition, they do not provide an end-to-end data management pipeline from data acquisition, through data integration, quality control, storage and finally to data access. The developed system in this case study was used for an end-to-end management test of large data during the testing phase and proved to perform seamlessly. Our approach could be adopted by other utility jurisdictions to manage their utility data. Such a data management system allows for automated and proper management of utility data thereby helping state and local agencies reduce utility conflicts and offset construction costs due to utility damages. This data could be combined with other rich data sources, such as financial data, and mined for valuable, hidden insights.
Scientific fields such as insider-threat detection and highway-safety planning often lack sufficient amounts of time-series data to estimate statistical models for the purpose of scientific discovery. Moreover, the available limited data are quite noisy. This presents a major challenge when estimating time-series models that are robust to overfitting and have well-calibrated uncertainty estimates. Most of the current literature in these fields involve visualizing the time-series for noticeable structure and hard coding them into pre-specified parametric functions. This approach is associated with two limitations. First, given that such trends may not be easily noticeable in small data, it is difficult to explicitly incorporate expressive structure into the models during formulation. Second, it is difficult to know $\textit{a priori}$ the most appropriate functional form to use. To address these limitations, a nonparametric Bayesian approach was proposed to implicitly capture hidden structure from time series having limited data. The proposed model, a Gaussian process with a spectral mixture kernel, precludes the need to pre-specify a functional form and hard code trends, is robust to overfitting and has well-calibrated uncertainty estimates.
Pedestrian and bicycle safety challenges are becoming more apparent as those modes increase in popularity. Policy-relevant safety analysis methods for these modes are rare, particularly related to ...
Estimation of flexible-statistical models of travel demand involves tuning varying parameters, hyperparameters, manually and iteratively. Proper tuning of hyperparameters results in superior models. However, considerable expertise, including technical knowledge of statistics, data mining or machine learning, and experience are required to tune hyperparameters and consequently generate appropriate models. Moreover, tuning hyperparameters is prone to subjective error and consequently produces travel demand models that are difficult to reproduce and extend, and makes the development more an art than a science. There is a need for methods to reduce or eliminate subjectivity during the tuning process. This study proposed a framework to reduce subjectivity during the tuning of hyperparameters required for the estimation of nonparametric models of activity-duration. That is, a flexible-statistical framework, which leverages state-of-the-art innovations in Bayesian optimization (BO), was proposed to estimate Gaussian process models of activity duration and associated hyperparameters. The framework was applied to estimate duration models for five types of out-of-home non-mandatory activity episodes for household individuals in the greater Los Angeles area. Experiments demonstrate that the accuracy of results from the proposed framework are superior to those from the current tuning process, and are obtained in a fraction of the time. The proposed framework could potentially increase the productivity of modelers by reducing time required to tune hyperparameters.
Pymc-learn is a Python package providing a variety of state-of-the-art probabilistic models for supervised and unsupervised machine learning. It is inspired by scikit-learn and focuses on bringing probabilistic machine learning to non-specialists. It uses a general-purpose high-level language that mimics scikit-learn. Emphasis is put on ease of use, productivity, flexibility, performance, documentation, and an API consistent with scikit-learn. It depends on scikit-learn and pymc3 and is distributed under the new BSD-3 license, encouraging its use in both academia and industry. Source code, binaries, and documentation are available on http://github.com/pymc-learn/pymc-learn.
Traditional methods for stochastic user equilibrium are either probit-based or logit-based; both have weaknesses that limit their usefulness. This paper proposes a mixed logit algorithm for Stochastic Dynamic User Equilibrium (SDUE) which is capable of capturing required correlations at a reasonable cost. That is, the flexibility in specification of error terms provided by the mixed-logit allows the SDUE model to capture spatial and temporal correlations in unobserved factors. In addition, it does not require route enumeration. The capabilities of a mixed-logit model enable the development of a SDUE traffic assignment model with correct flow propagation. A mathematical programming problem is formulated for the SDUE whose solution is found by an iterative mixed-logit-based network-loading procedure which in theory is expected to calculate link flows at a more efficient rate compared to probit-based models.
The objective of this study was to investigate factors influencing occurrence of pedestrian and bicycle crashes in Tennessee. Of interest were demographic and socio-economic, roadway geometry, traffic, and land use factors that could influence pedestrian crash rates on specific infrastructure. Geographic Information System (GIS) and statistical modeling were applied to study the crash patterns with respect to these factors. GIS was used to geo-locate and cluster the crash locations onto the roadway network, joined with background data of the crash locations. Negative Binomial (NB) regression was used to model the relationship between contributing factors and the crashes to detect any positive or negative correlations with the crashes. The following factors were found to have significant correlation with pedestrian and bicycle crash occurrences; percentage distribution of population by race, age groups, mean household income, percentage in the labor force, poverty level, and vehicle ownership. Land use, number of lanes crossed by pedestrians or bicyclists, posted speed limit and the presence of special speed zones, all were found to influence the occurrence of these crashes significantly. The findings were used to identify patterns of pedestrian and bicycle high crash locations in Tennessee and flagged combination of demographic, socioeconomic and geometry variables which if present are good indicators to Tennesee Department of Transportation (TDOT) as areas likely to experience pedestrian and bicycle crashes.
This study evaluated the impact of roadway cross-sectional and geometric features, traffic characteristics, and median cable barrier placement to the frequency of median-related crashes through statistical modeling using multiyear data. A unique aspect of the model specificationwas the inclusion of median cable barrier placement data, horizontal curve data, and differential elevation of opposite travel lanes. Negative binomial model was used in linearizing and quantifying these factors with respect to median crossover crash frequency. The variables that were found to significantly influence the frequency of median crashes include the number of lanes, differential elevation, and cable barrier offset from the inside shoulder. Increasing traffic volume was found to increase the frequency of median barrier crashes as well as the presence of curves on a median barrier section was found to increase the frequency of crashes. Higher differential elevation between opposite travel lanes was also found to increase the frequency of median barrier crashes. Increasing the median barrier offset from the inside shoulder of the travel way decreases median barrier injury and fatal crash frequency. Segments with a higher number of lanes and a wider median width were associated with low crash frequencies.
In order to identify high crash locations, the Tennessee Department of Transportation (TDOT) has an extensive road safety audit program which uses criteria based on the ratio of crashes to average daily traffic but does not target locations with a high number of pedestrian crashes since there are no pedestrian counts. Apart from ratio approach, a robust methodology is not currently available to identify pedestrian high-crash locations in Tennessee. The objective of this study is to develop a different methodology based on Anselin’s Local Moran I index in Geographic Information System (GIS) to detect high crash clusters and investigate the factors that influence the concentration of pedestrian crashes. Using pedestrian crash data from Shelby County in Tennessee, the study found that spatial dependence plays a strong role during the analyses of pedestrian crashes. These spatial dependencies, accounted through spatial autocorrelation, helped to detect statistically significant clusters of crashes in a GIS framework. These clusters were then overlaid with selected socio-economic and population demographic data in order to identify their association with high crash clusters. The study found the following factors to be associated with high crash clusters: when more than 25% percentage of the population is 18 years of age and younger, when the population of seniors is greater than 13%, when there’s a high population density of low income people, and when the percentage of families below poverty level is greater than 10%. The cluster maps may help transportation agencies to understand issues of pedestrian crashes for safety enhancements.
In most cases, abandoned and disabled vehicles are left within the roadway right of ways. It is common to find a vehicle left on the shoulder, median, gore area or on the travel lane for certain period of time. Experience from the state of Tennessee has shown that 78% of the freeway traffic related incidents are due to disabled and abandoned vehicles. It is hypothesized that the longer the vehicle is left unattended within the right of way, the higher the probability of new incidents and secondary crashes. This paper utilized 2004 to 2010 freeway incident data in Tennessee to evaluate the impact of the length of incident durations caused by disabled and abandoned vehicles. Analysis evaluated the impact of these incidents with respect to roadway location, queue lengths, weather conditions, towing times, lane closure, and the source of incident notification. Temporal factors, including the spectra of the time of the day, the day of the week, and the seasons of the year were evaluated with respect to the number of incidents and incident durations. It was found that vehicles left on the left and right shoulders generated more incidents compared to other locations followed by gore areas and the ramps. Parametric hazard based log-logistic survival model was applied to determine the factors affecting the abandoned and disabled vehicles incident duration. Number of closed lanes, length of the queue formed, construction zones, trucks and towing involvement were found to be significantly associated with longer incident duration.
Performances and safety effectiveness evaluation results of median cable barrier systems in Tennessee are presented in this paper. Twenty seven segments with at least three years of complete crash data before and after cable installations were analyzed. The segments were evaluated in terms of descriptive statistics of factors associated with median crashes whose occurrences were influenced by the presence or absence of the median cable barriers. The cable systems were also evaluated in terms of percentage safety effectiveness and confidence levels comparing before and after cable conditions. The study involved review of crash report hard copies where only 24% were found to be relevant for median cable barriers evaluation, 76% were not related. Descriptive statistics compared the percentage of a certain type of crashes, crash attributes and other elements to the total crashes before and after the barriers were installed. To evaluate the safety effectiveness, the research applied crash modeling in the form of safety performance models, and observational Empirical Bayes (EB) before and after analysis. Safety effectiveness of the installed median cable barrier systems was found to be 93% for fatal crashes, 85% for fatal and incapacitating injury crashes combined and 51% for the combination of fatal and all injury crashes all above 95% confidence level. The study also found that combined fatal and injury crashes were reduced by 21% after median cable installations while fatal crashes only were reduced by 80%. Total number of people killed or injured was reduced by 29% after installation.
The paper analyses integrating origin-destination (O-D) survey results with stochastic user equilibrium (SUE) in traffic assignment. The two methods are widely used in transportation planning but their applications have not yet fully integrated. While O-D gives a generalized trip patterns, purpose and characteristics, SUE provides optimal trip distributions using the characteristics found in O-D survey. The paper utilized O-D and SUE in route relocation study for the town of Coamo in Puerto Rico. The O-D survey was used initially in studying possible trip distribution and assignment for the new route. Initial distribution and assignment of traffic to the existing roadway networks and the proposed route were allocated utilizing the O-D survey findings. The SUE was then used to optimize the assignments considering roadway characteristics such as number of lanes, capacity limits, free flow speed, signal spacing density, travel time and gasoline cost. The travel time was optimized through the Bureau of Public Roads (BPR) equation found in 2000 HCM. The optimal trips found from the SUE were then used to propose the final alignment of the new route. Traffic assignment from the SUE was slightly different from those initially assigned using O-D, indicating there was optimization. The assignment on new route was increased by 13.8% from the one assigned using O-D while assignment on the existing link was reduced by 22%.
This paper evaluates different factors and parameters contributing to likelihood of bicycle crash injury severity levels. Multinomial Logit (MNL) model was used to analyze impact of different roadway features, traffic characteristics and environmental conditions associated with bicycle crash injury severities. The multinomial model was used due to its flexibility in quantifying the effect of the independent variables for each injury severity categories. Model results showed that, severity of bicycle crashes increases with increase in vehicles per lane, number of lanes, bicyclist alcohol or drug use, routes with 35-45 mph posted speed limits, riding along curved or sloped road sections, when bicyclists approach or cross a signalized intersection, and at driveways. In addition, routes with a high percentage of trucks, roadway sections with curb and gutter, cloudy or foggy weather and obstructed vision were found to have high probability of severe injury. Segments with wider lanes, wide median and wide shoulders were found to have low likelihood of severe bicycle injury severities. Limited lighting locations was found to be associated with incapacitating injury and fatal crashes, indicating that insufficient visibility can potentially lead to severe crashes. Other findings are also presented in the paper.