Applications of discrete choice models in personalization are becoming increasingly popular among researchers and practitioners. However, in such systems, when users are presented with successive menus (or choice situations), the alternatives and attributes in each menu depend on the choices made by the user in the previous menus. This gives rise to endogeneity which can result in inconsistent estimates. Our companion paper, Danaf et al. (2020), showed that the estimates are only consistent when the entire choice history of each user is included in estimation. However, this might not be feasible because of computational constraints or data availability. In this paper, we present a control-function (CF) correction for the cases where the choice history cannot be included in estimation. Our method uses the attributes of non-personalized attributes as instruments, and applies the CF correction by including interactions between the explanatory variables and the first stage residuals. Estimation can be done either sequentially or simultaneously, however, the latter is more efficient (if the model reflects the true data generating process). This method is able to recover the population means of the distributed coefficients, especially with a long choice history. The variances are underestimated, because part of the inter-consumer variability is explained by the residuals, which are included in the systematic utility. However, the population variances can be computed from the estimation results. The modified utility equations (which include the residuals) can be used in forecasting and model application, and provide superior fit and predictions.
Through the vast adoption and application of emerging technologies, the intelligence and autonomy of smart mobility can be substantially elevated to address more diversified demands and supplies. Along with this trend, a systematic collaboration among three essential elements of smart mobility services, namely devices, data and functions, is being studied to comprehensively break down the intrinsic barriers that existed in current solutions, to support the integration of connectable devices, the fusion of heterogeneous data, the composability of reusable functions, and the flexibility in their cooperations. To enable such a collaboration, this paper proposes a federated platform, called Future Mobility Sensing Advisor (FMSA), which can 1) manage the three elements through standardized interfaces separately and uniformly; 2) create a fully connected knowledge graph to orchestrate the three elements efficiently and effectively; 3) support the client-server interaction in centralized and federated modes to handle service requests and edge resources with various availability and accessibilities jointly and adaptively; and 4) accommodate various mobility services to foster harmonious and sustainable mobility tenderly and invisibly. Moreover, the efficiency and effectiveness of the platform are also tested through a performance evaluation, and a pilot supported at the Great Boston Area, respectively. As a result, it shows that FMSA can 1) achieve high performance by using the two interaction modes selectively, and 2) renovate smart mobility towards sustainability through personalized services that can measure user preferences and system objectives mutually.
This paper discusses capabilities that are essential to models applied in policy analysis settings and the limitations of direct applications of off-the-shelf machine learning methodologies to such settings. Traditional econometric methodologies for building discrete choice models for policy analysis involve combining data with modeling assumptions guided by subject-matter considerations. Such considerations are typically most useful in specifying the systematic component of random utility discrete choice models but are typically of limited aid in determining the form of the random component. We identify an area where machine learning paradigms can be leveraged, namely in specifying and systematically selecting the best specification of the random component of the utility equations. We review two recent novel applications where mixed-integer optimization and cross-validation are used to algorithmically select optimal specifications for the random utility components of nested logit and logit mixture models subject to interpretability constraints.
This paper introduces a new data-driven methodology for estimating sparse covariance matrices of the random coefficients in logit mixture models. Researchers typically specify covariance matrices in logit mixture models under one of two extreme assumptions: either an unrestricted full covariance matrix (allowing correlations between all random coefficients), or a restricted diagonal matrix (allowing no correlations at all). Our objective is to find optimal subsets of correlated coefficients for which we estimate covariances. We propose a new estimator, called MISC, that uses a mixed-integer optimization (MIO) program to find an optimal block diagonal structure specification for the covariance matrix, corresponding to subsets of correlated coefficients, for any desired sparsity level using Markov Chain Monte Carlo (MCMC) posterior draws from the unrestricted full covariance matrix. The optimal sparsity level of the covariance matrix is determined using out-of-sample validation. We demonstrate the ability of MISC to correctly recover the true covariance structure from synthetic data. In an empirical illustration using a stated preference survey on modes of transportation, we use MISC to obtain a sparse covariance matrix indicating how preferences for attributes are related to one another.
We develop a methodology to analyze pedestrian-vehicular interactions in urban streets in a mixed traffic environment, and then apply it to Bliss Street, an urban street in Beirut. Data on the street was collected before and after a crosswalk was installed using videography, radar speed guns, and manual counts. A pedestrian gap acceptance model indicated that installing the crosswalk did not have any significant effect on the pedestrians’ sensitivity to waiting time, gap size, or the speed of the approaching vehicles. However, it caused reductions in the speed of approaching vehicles which in turn encouraged pedestrians to accept shorter gaps. A micro-simulation model indicated that the crosswalk would reduce the speed on the street slightly, with significant reductions observed if more pedestrians who currently cross at midblock locations shift to use the crosswalk. The results of this study can be used to test interventions for enhancing pedestrian safety in Lebanon, and are generalizable to similar contexts in developing countries.
This paper presents a framework for estimating and updating user preferences in the context of app-based recommender systems. We specifically consider recommender systems which provide personalized menus of options to users. A Hierarchical Bayes procedure is applied in order to account for inter- and intra-consumer heterogeneity, representing random taste variations among individuals and among choice situations (menus) for a given individual, respectively. Three levels of preference parameters are estimated: population-level, individual-level and menu-specific. In the context of a recommender system, the estimation of these parameters is repeated periodically in an offline process in order to account for trends, such as changing market conditions. Furthermore, the individual-level parameters are updated in real-time as users make choices in order to incorporate the latest information from the users. This online update is computationally efficient which makes it feasible to embed it in a real-time recommender system. The estimated individual-level preferences are stored for each user and retrieved as inputs to a menu optimization model in order to provide recommendations. The proposed methodology is applied to both Monte-Carlo and real data. It is observed that the online update of the parameters is successful in improving the parameter estimates in real-time. This framework is relevant to various recommender systems that generate personalized recommendations ranging from transportation to e-commerce and online marketing, but is particularly useful when the attributes of the alternatives vary over time.
Logit mixture models have gained increasing interest among researchers and practitioners because of their ability to capture unobserved taste heterogeneity. Becker et al. (2018) proposed a Hierarchical Bayes (HB) estimator for logit mixtures with inter- and intra-consumer heterogeneity (defined as taste variations among different individuals and among different choices made by the same individual respectively). However, the underlying model relies on strong assumptions on the inter- and intra-consumer mixing distributions; these distributions are assumed to be normal (or log-normal), and the intra-consumer covariance matrix is assumed to be the same for all individuals. This paper presents a latent class extension to the model and the estimator proposed by Becker et al. (2018) to account for flexible, semi-parametric mixing distributions. This relaxes the normality assumptions and allows different individuals to have different intra-consumer covariance matrices. The proposed model and the HB estimator are validated using real and synthetic data sets, and the models are evaluated using goodness-of-fit statistics and out-of-sample validation. Our results show that when the data comes from two or more distinct classes (with different population means and inter- and intra-consumer covariance matrices), this model results in a better fit and predictions compared to the single class model.
Stated preferences surveys are most commonly used to provide behavioral insights on hypothetical travel scenarios such as new transportation services or attribute ranges beyond those observed in existing conditions. When designing SP surveys, considerable care is needed to balance the statistical objectives with the realism of the experiment. This paper presents an innovative method for smartphone-based stated preferences (SP) surveys leveraging state-of-the-art smartphone-based survey platforms and their revealed preferences sensing capabilities. A random experimental design generates context-aware SP profiles using user specific socioeconomic characteristics and past travel data along with relevant web data for scenario generation. The generated choice tasks are automatically validated to reduce the number of dominant or inferior alternatives in real-time, then validated using Monte-Carlo simulations offline. In this paper we focus our attention on mode choice and design an experiment that considers a wide range of possible existing mode alternatives along with a new alternative on-demand mobility service that does not exist in real life. This experiment is then used to collect SP data or a sample of 224 respondents in the Greater Boston Area. A discrete mode choice model is estimated to illustrate the benefit of the proposed method in capturing current context-specific preferences in response to the new scenario.
Endogeneity arises in discrete choice models due to several factors and results in inconsistent estimates of the model parameters. In adaptive choice contexts such as choice-based recommender systems and adaptive stated preferences (ASP) surveys, endogeneity is expected because the attributes presented to an individual in a specific menu (or choice situation) depend on the previous choices of the same individual (as well as the alternative attributes in the previous menus). Nevertheless, the literature is indecisive on whether the parameter estimates in such cases are consistent or not. In this paper, we discuss cases where the estimates are consistent and those where they are not. We provide a theoretical explanation for this discrepancy and discuss the implications on the design of these systems and on model estimation. We conclude that endogeneity is not a concern when the likelihood function properly accounts for the data generation process. This can be achieved when the system is initialized exogenously and all the data are used in the estimation. In line with previous literature, Monte Carlo results suggest that, even when exogenous initialization is missing, empirical bias decreases with the number of choices per individual. We conclude by discussing the practical implications and extensions of this research.
This paper presents a systematic way of understanding and modeling traveler behavior in response to on-demand mobility services. We explicitly consider the sequential and yet inter-connected decision-making stages specific to on-demand service usage. The framework includes a hybrid choice model for service subscription, and three logit mixture models with inter-consumer heterogeneity for the service access, menu product choice and opt-out choice. Different models are connected by feeding logsums. The proposed modeling framework is essential for accounting the impacts of real-time on-demand system’s dynamics on traveler behaviors and capturing consumer heterogeneity, thus being greatly relevant for integrations in multi-modal dynamic simulators. The methodology is applied to a case study of an innovative personalized on-demand real-time system which incentivizes travelers to select more sustainable travel options. The data for model estimation is collected through a smartphone-based context-aware stated preference survey. Through model estimation, lower values of time are observed when the respondents opt to use the reward system. The perception of incentives and schedule delay by different population segments are quantified. These results are fundamental in setting the ground for different behavioral scenarios of such a new on-demand system. The proposed methodology is flexible to be applied to model other on-demand mobility services such as ride-hailing services and the emerging mobility as a service.
The objective of this article is to investigate the differences in driving behavior and red-light violations between drivers in two countries: Lebanon and the United States of America. To realize the stated objective, two driving simulators were utilized. The first simulator is located at the American University of Beirut (AUB), Lebanon. The second simulator is located at the George Washington University (GWU), United States of America. An elaborate experimental scheme involving the occurrence of frustrating events at signalized intersections was designed, and 35 students from GWU and 81 students from AUB participated in the experiments. Detailed trajectory data was collected, and students were compared based on three surrogate measures: number of red-light violations, time-to-junction, average and maximum velocities. The results indicated that frustrating events occurring at intersections elicit red-light violations and speeding for both samples. In addition, speeding and red-light violations follow increasing trends as students drive through successive signalized intersections and experience frustrating events. In a postdriving survey, AUB students indicated that they engage in risky driving behavior more than GWU students. On the other hand, GWU students indicated that they are more likely to violate traffic rules and also committed more red-light violations in the simulator. The results of this study have implications on driving education and enforcement, as they indicate that driving violations might not be necessarily related to risky or aggressive driving, but to an individual's tendency to violate traffic rules in general.
This paper presents a personalized menu optimization model with preference updater in the context of an innovative Smart Mobility system that offers a personalized menu of travel options with incentives for each incoming traveler in real time. This Smart Mobility system can serve as a major travel demand management system that encourages energy-efficient travel options. The personalized menu optimization is built on a logit mixture model that captures each individual traveler's choice behavior. The personalized menu optimization model is enhanced with a preference updater that can update the estimates of individual traveler's preference parameters when new choice data is received. To illustrate the advantages of the proposed methodology, a case study is presented based on real travelers and trips in the greater Boston area from the Massachusetts Travel Survey data. The case study consists of two parts. In the first part, the personalized menu optimization with preference updater is tested in a setting where the travelers are new to the system and their preferences are updated through preference updater. A comparative analysis of the performance of the proposed method with preference updater is presented against the method without preference updater. In the second part, the benefit of using individual level preference parameters instead of population level preference parameters in the personalized menu optimization model is analyzed. The case study shows that the proposed method can outperform the hit rates of its two counterparts.
This paper presents a utility-maximizing approach to agent-based modeling with an application to the Greater Boston Area (GBA). It leverages day activity schedules (DAS) to create a framework for representing travel demand in an individual's day. DAS are composed of a sequence of stops that make up home-based tours with activity purposes, intermediate stops, and subtours. The framework introduced in this paper includes three levels: (1) the Day Pattern Level, which determines if an individual will travel and, if so, what types of primary activities and intermediate stops they will do; (2) the Tour Level, which models the mode, destination, and time-of-day of the different primary activities; and (3) the Intermediate Stop Level, which generates intermediate stops. The models are estimated for the GBA using the 2010 Massachusetts Travel Survey (MTS). They are then implemented in SimMobility, the agent-based, activity-based, multimodal simulator. It run in a microsimulation using a Synthetic Population. Produced results are consistent with the MTS. Compared with similar activity-based approaches, the proposed framework allows for more flexibility in modeling a wide range of activity and travel patterns.
Estimating discrete choice models on panel data allows for the estimation of preference heterogeneity in the sample. While the Logit Mixture model with random parameters is mostly used to account for variation across individuals, preferences may also vary across different choice situations of the same individual. Up to this point. Logit Mixtures incorporating both inter- and intra-consumer heterogeneity are estimated with the classical Maximum Simulated Likelihood (MSL) procedure. The MSL procedure becomes computationally expensive with an increasing sample size and can be burdensome in the presence of a multi-modal likelihood function. We therefore propose a Hierarchical Bayes estimator for Logit Mixtures with both levels of heterogeneity. It builds on the Allenby-Train procedure, which considers only inter-consumer heterogeneity. To test the proposed procedures, we analyze how well the true patterns of heterogeneity are recovered in a simulation environment. Results from the Monte Carlo simulation suggest that falsely ignoring intra-consumer heterogeneity despite its presence in the data leads to biased estimates and a decreased goodness of fit. The latter is confirmed by a real-world example of explaining mode choices for GPS traces. We further show that the runtime of the proposed estimator is substantially faster than for the corresponding MSL estimator. (C) 2018 Elsevier Ltd. All rights reserved.
This study investigates differences between the mode choice patterns of students of the American University of Beirut (AUB) and the general population of the Greater Beirut Area. Discrete choice models are developed to model the choice among car, bus, and shared taxi (or jitney). It is found that travel time, cost, income, auto ownership, gender, and residence location (whether within Municipal Beirut or not) are the main factors affecting mode choice, and that AUB students who come from wealthier families have a significantly higher value of time than the general population. The models are used to forecast students' commute mode shares under alternative scenarios to support the development of policies that would encourage students to switch toward more sustainable modes. It is found that increasing parking fees and decreasing bus travel time through the provision of shuttle services or taxi sharing could be promising strategies for mode switching from car to public transport for AUB students. The study contributes to the emerging literature on students' travel patterns and its findings are particularly relevant in travel contexts characterized by high congestion levels, high auto ownership rates, and low quality public transport system. (C) 2014 World Conference on Transport Research Society. Published by Elsevier Ltd. All rights reserved.
This paper develops a hybrid choice-latent variable model combined with a Hidden Markov model in order to analyze the causes of aggressive driving and forecast its manifestations accordingly. The model is grounded in the state-trait anger theory; it treats trait driving anger as a latent variable that is expressed as a function of individual characteristics, or as an agent effect, and state anger as a dynamic latent variable that evolves over time and affects driving behavior, and that is expressed as a function of trait anger, frustrating events, and contextual variables (e.g., geometric roadway features, flow conditions, etc.). This model may be used in order to test measures aimed at reducing aggressive driving behavior and improving road safety, and can be incorporated into micro-simulation packages to represent aggressive driving. The paper also presents an application of this model to data obtained from a driving simulator experiment performed at the American University of Beirut. The results derived from this application indicate that state anger at a specific time period is significantly affected by the occurrence of frustrating events, trait anger, and the anger experienced at the previous time period. The proposed model exhibited a better goodness of fit compared to a similar simple joint model where driving behavior and decisions are expressed as a function of the experienced events explicitly and not the dynamic latent variable. (C) 2014 Elsevier Ltd. All rights reserved.
The objective of this paper is to investigate the differences in drivers' aggressiveness at signalized intersections among drivers in two countries: Lebanon and the United States of America. To realize the stated objective, two driving simulators were utilized. The first simulator is located at the American University of Beirut, Lebanon. The second simulator is located at the George Washington University, U.S.A. An elaborate experimental scheme was designed taking into consideration the driving simulators' capabilities and specifications. Twenty-six subjects from the George Washington University – GWU – and 81 subjects from the American University of Beirut – AUB – participated in the driving experiments. Detailed trajectory data was collected and processed to investigate any statistically significant differences in behaviors. Specific surrogate measures for aggressiveness were used including number of violations, time-to-junction, average and maximum velocities. These measures were linked to signalized intersections' situational characteristics. The results indicate that anger and aggressive behavior are incrementally intensified for both the AUB sample and the GWU sample. AUB students considered themselves as more aggressive than GWU students. However, GWU students showed more red-light violations; frustrating events instigated more aggressive behavior for GWU students, apparently because of the rare occurrence of such events in the driving context where these students belong. AUB students were not as sensitive to other drivers' violations or blocked intersections possibly due to the frequent occurrence of such events on the roads of Beirut.