We study consumer demand in large-scale retail settings with many products, multiple categories and repeated purchase behavior. While inertia and brand loyalty are well documented, existing discrete choice models typically focus on single categories or become computationally infeasible in high-dimensional environments. We propose a dynamic product-level factor model that captures heterogeneity in baseline preferences, price sensitivity and inertia through a shared latent factor structure. By factorizing individual-product coefficients, the model pools information across individuals and categories and allows for correlated heterogeneity. We estimate the model using Bayesian variational inference, enabling scalable estimation with tens of thousands of parameters. In a simulation study calibrated to realistic retail data, we show that the dynamic factor model substantially improves predictive performance relative to static factor models and mixed logit benchmarks, particularly when individual purchase histories are sparse. Accounting for inertia also leads to more elastic demand estimates, underscoring the importance of dynamics for measuring consumer responsiveness. Our results highlight dynamic factor models as a scalable and flexible approach for demand estimation in modern, high-dimensional retail markets.
This study uses double/debiased machine learning (DML) to evaluate the impact of transitioning from lecture-based blended teaching to a flipped classroom concept. Our findings indicate effects on students' self-conception, procrastination, and enjoyment. We do not find significant positive effects on exam scores, passing rates, or knowledge retention. This can be explained by the insufficient use of the instructional approach that we can identify with uniquely detailed usage data and highlights the need for additional teaching strategies. Methodologically, we propose a powerful DML approach that acknowledges the latent structure inherent in Likert scale variables and, hence, aligns with psychometric principles.
The nonlinear GMM-IV estimator of Berry, Levinsohn and Pakes (1995) can suffer from numerical instability resulting in a wide range of parameter estimates and economic implications. This has been reported to depend on technical details such as the choice of the optimization algorithm, starting values, and convergence criteria. We show that numerical approximation errors in the estimator's moment function are the main driver of this instability. With accurate approximation, the estimation approach is well-behaved. We provide a simple method to determine the required number of simulation draws.
In the United States, individuals with disabilities and those aged >= 65 can supplement their Medicare with so-called stand-alone Medicare Part D prescription drug plans. Beneficiaries can switch their stand-alone prescription drug plans annually, but most do not. Indirect evidence has raised concerns that non-switchers do not even make plan comparisons (labeled "inattention"), but direct evidence is scarce. Therefore, we surveyed 439 beneficiaries of Medicare Part D plans from a nationally representative adult sample after the 2024 open-enrollment period. Overall, 53% self-reported making no comparisons. Of those who did not compare, 98% did not switch (vs 67% of those who did compare). Multinomial regressions revealed that beneficiaries who neither compared nor switched were more likely than switchers to report difficulties with comparing and switching, experiencing no plan-related discontinuation, changes, or dissatisfaction, not using advisors or the plan-finder website, and receiving potentially confusing mailings. Non-switchers who did compare were similar to switchers in reporting few difficulties and relying on advisors and the plan-finder website, but they were less likely than switchers to report plan-related changes, discontinuation, or dissatisfaction, while being more likely to report receiving mailings and having no college degree. We discuss insights for policy-making.
This paper studies consumers’ choice between two different water tariffs. We document a large inaction in a novel setting where customers face a binary decision and receive simple, detailed, and personalized information about the financial savings they would obtain if they were to switch water tariff. Our empirical framework separates two sources of inertia: inattention and switching costs. The model estimates that half of the customers that would benefit from changing tariff are not aware of the opportunity they are offered. Conditional on paying attention, we estimate median switching costs to be around £100. A model where all customers are assumed to pay attention delivers instead implausibly high switching costs, with a median of £400. This shows the importance of inattention in explaining consumers’ inaction. Looking at the characteristics of the households, our results confirm previous findings that areas where households have higher levels of education or the proportion of minorities is lower display a higher responsiveness to potential savings. The new insight offered by our analysis is that this is entirely driven by attention, whereas switching costs actually increase with education and ethnic homogeneity. Our findings suggest that policies aimed at increasing attention can play a central role in fostering competition among suppliers and reducing inequalities.
This paper investigates and extends the computationally attractive nonparametric random coefficients estimator of Fox, Kim, Ryan, and Bajari (2011). We show that their estimator is a special case of the nonnegative LASSO, explaining its sparse nature observed in many applications. Recognizing this link, we extend the estimator, transforming it to a special case of the nonnegative elastic net. The extension improves the estimator's recovery of the true support and allows for more accurate estimates of the random coefficients' distribution. Our estimator is a generalization of the original estimator and therefore, is guaranteed to have a model fit at least as good as the original one. A theoretical analysis of both estimators' properties shows that, under conditions, our generalized estimator approximates the true distribution more accurately. Two Monte Carlo experiments and an application to a travel mode data set illustrate the improved performance of the generalized estimator.
Consumers' health plan choices are highly persistent even though optimal plans change over time. This paper separates two sources of inertia, inattention to plan choice and switching costs. We develop a panel data model with separate attention and choice stages, linked by heterogeneity in acuity, i.e., the ability and willingness to make diligent choices. Using data from Medicare Part D, we find that inattention is an important source of inertia but switching costs also play a role, particularly for-low-acuity individuals. Separating the two stages and allowing for heterogeneity is crucial for counterfactual simulations of interventions that reduce inertia.
In recent years, consumer choice has become an important element of public policy. One reason is that consumers differ in their tastes and needs, which they can express most easily through their own choices. Elements that strengthen consumer choice feature prominently in the design of public insurance markets, for instance in the United States in the recent introduction of prescription drug coverage for older individuals via Medicare Part D. For policy makers who design such a market, an important practical question in the design phase of such a new program is how to deduce enrollment and plan selection preferences prior to its introduction. In this paper, we investigate whether hypothetical choice experiments can serve as a tool in this process. We combine data from hypothetical and real plan choices, elicited around the time of the introduction of Medicare Part D. We first analyze how well the hypothetical choice data predict willingness to pay and market shares at the aggregate level. We then analyze predictions at the individual level, in particular how insurance demand varies with observable characteristics. We also explore whether the extent of adverse selection can be predicted using hypothetical choice data alone.
Empirical economic research frequently applies maximum likelihood estimation in cases where the likelihood function is analytically intractable. Most of the theoretical literature focuses on maximum simulated likelihood (MSL) estimators, while empirical and simulation analyzes often find that alternative approximation methods such as quasi-Monte Carlo simulation, Gaussian quadrature, and integration on sparse grids behave considerably better numerically. This paper generalizes the theoretical results widely known for MSL estimators to a general set of maximum approximated likelihood (MAL) estimators. We provide general conditions for both the model and the approximation approach to ensure consistency and asymptotic normality. We also show specific examples and finite-sample simulation results.
Between 2004 and 2016, we elicited individuals’ subjective expectations of stock market returns in a Dutch internet panel at bi-annual intervals. In this paper, we develop a panel data model with a finite mixture of expectation types who differ in how they use past stock market returns to form current stock market expectations. The model allows for rounding in the probabilistic responses and for observed and unobserved heterogeneity at several levels. We estimate the type distribution in the population and find evidence for considerable heterogeneity in expectation types and meaningful variation over time, in particular during the financial crisis of 2008/09.
Tuesday, October 4 Session, 9:00–10:10 Session, 9:00–10:10 Suboptimal feedback control of PDEs by solving HJB equations on adaptive sparse grids Jochen Garcke, Ilja Kalmikov and Axel Kröner 1University of Bonn; garcke@ins.uni-bonn.de 2Fraunhofer SCAI; ilja.kalmykov@scai.fraunhofer.de 3INRIA Saclay; axel.kroener@inria.fr An approach to solve finite time horizon suboptimal feedback control problems for partial differential equations is proposed by solving dynamic programming equations on adaptive sparse grids. A semi-discrete optimal control problem is introduced and the feedback control is derived from the corresponding value function. The value function can be characterized as the solution of an evolutionary Hamilton-Jacobi Bellman (HJB) equation which is defined over a state space whose dimension is equal to the dimension of the underlying semi-discrete system. Besides a low dimensional semi-discretization it is important to solve the HJB equation efficiently to address the curse of dimensionality. We propose to apply a semi-Lagrangian scheme using spatially adaptive sparse grids. Sparse grids allow the discretization of the value functions in (higher) space dimensions since the curse of dimensionality of full grid methods arises to a much smaller extent. For additional efficiency an adaptive grid refinement procedure is explored. The approach is illustrated for the wave equation and an extension to equations of Schrödinger type is indicated. We present several numerical examples studying the effect the parameters characterizing the sparse grid have on the accuracy of the value function and the optimal trajectory. closed-loop suboptimal control of PDEs and HJB equations and sparse grids and curse of dimensionality Sparse grid techniques for particle-in-cell schemes Lee Ricketson and Antoine Cerfon 1Courant Institute; ricketson@cims.nyu.edu 2Courant Institute; cerfon@cims.nyu.edu The kinetic equations governing plasma dynamics are six-dimensional PDEs for which the dominant computational concern is the curse of dimensionality. Motivated by this, the dominant scheme in many applications is particle-in-cell (PIC), which has the major advantage of reducing the dimensionality of the grid from six to three. The price to pay for this dimensionality reduction is the statistical error inherent to any particle method. Moreover, the statistical figure of merit is the number of particles per grid cell, meaning the curse of dimensionality persists. Since the number of grid cells can be very large, good statistical resolution requires overwhelmingly large particle populations. We propose to rectify this situation using the combination technique for sparse grids. In addition to accelerating grid computations, the cells in the relevant combination grids are much larger than for a comparable regular grid. In this way, we can achieve many more particles per cell for any given number of total particles. This results in dramatic reduction of statistical error. We present results from test cases that demonstrate the accuracy and efficiency of the new scheme. 8 Sparse Grids and Applications 2016 Session, 10:45–12:30 Tuesday, October 4 Session, 10:45–12:30 A sparse grid collocation method based on LaVallée Poussin kernel Moulay Abdellah Chkifa 1Oak Ridge National Lab; bchkifa@gmail.com In this talk, we present a new approach to polynomial approximation of functions in high dimension. We present a new non-intrusive collocation method that is based on Smolyack formula applied to polynomial approximations in one dimension that are based on Fejer and LaValle Poussin kernel. Smolyak formula allows us to transform a hierarchical collocation strategy in one dimension to a hierarchical collocation strategy in high dimension with possibly a straightforward computational formula and similar stability properties. We discuss the accuracy and stability of the introduced method versus the number of collocations needed and compare it to other existing strategies such as sparse hierarchical interpolation and sparse least squares. Numerical examples are provided to support the theoretical results and demonstrate the computational efficiency of the introduced method. Optimal Integration in Reproducing Kernel Hilbert spaces Jens Oettershagen and Michael Griebel 1Institute for Numerical Simulation , University of Bonn; oettersh@ins.uni-bonn.de 2Institute for Numerical Simulation , University of Bonn; griebel@ins.uni-bonn.de We present a black box algorithm for the construction of integration algorithms in tensor products of reproducing kernel Hilbert spaces. Here, in a first step an optimization procedure is employed to compute both stable and optimally weighted sequences of nested quadrature rules. In a second step, quasi-optimal index sets are build to obtain (generalized) sparse grids that inherit certain optimality properties of the univariate quadrature rules. Preserving Positivity of Sparse Grid Surrogates Fabian Franzelin and Dirk Pflüger 1University of Stuttgart; fabian.franzelin@ipvs.uni-stuttgart.de 2University of Stuttgart; Dirk.Pflueger@ipvs.uni-stuttgart.de It is a challenging task to limit the range of values of a locally refined sparse grid surrogate uSG(x) to the range of some function u(x) we want to approximate. If u(x) is a probability density function, for example, we want to preserve positivity and unit integrand. A common approach is to approximate log(u(x)), which has been applied to the interpolation of functions with Gaussian shape in astrophysics [Griebel10] and in the context of density estimation [Pflueger10]. In this talk we present a different approach based on extending the grid of uSG such that we enforce a non-negative range of values everywhere on the input domain without enumerating the whole full grid. In our approach we search the smallest number of additional full grid points we need to add to the grid in order to achieve a non-negative range. We will theoretically derive which points we have to consider and present an algorithm that computes these points efficiently by evaluating Sparse Grids and Applications 2016 9 Tuesday, October 4 Session, 10:45–12:30 intersections of grid points with negative coefficients. We will present algorithms that compute the coefficients of the new grid points in linear time with respect to the number grid points. This approach has two main advantages over the log-approach: (1) the surrogate is still just a linear combination of basis functions, and (2) computing its integral is easier than integrating exponentials. This makes it interesting for a large variety of applications. We will present results in the context of density estimation with locally adaptive sparse grids [Peherstorfer14] and uncertainty quantification [Franzelin16]. 10 Sparse Grids and Applications 2016 Invited Talk, 13:45–14:45 Tuesday, October 4 Invited Talk, 13:45–14:45 Sparse Grids and Computational Chemistry
The trend towards giving consumers choice about their health plans has invited research on how good they actually are at making these decisions.The introduction of Medicare Part D is an important example.Initial plan choices in this market were generally far from optimal.In this paper, we focus on plan choice in the years after initial enrollment.Due to changes in plan supply, consumer health status, and prescription drug needs, consumers' optimal plans change over time.However, in Medicare Part D only about 10% of consumers switch plans every year, and on average, plan choices worsen for those who do not switch.We develop a two-stage panel data model of plan choice whose stages correspond to two separate reasons for inertia: inattention and switching costs.The model allows for unobserved heterogeneity that is correlated across the two decision stages.We estimate the model using administrative data on Medicare Part D claims from 2007 to 2010.We find that consumers are more likely to pay attention to plan choice if overspending in the last year is more salient and if their old plan gets worse, for instance due to premium increases.Moreover, conditional on attention there are significant switching costs.Separating the two stages of the switching decision is thus important when designing interventions that improve consumers' plan choice.
The trend towards giving consumers choice about their health plans has invited research on how good they actually are at making these decisions. The introduction of Medicare Part D is an important example. Initial plan choices in this market were generally far from optimal. In this paper, we focus on plan choice in the years after initial enrollment. Due to changes in plan supply, consumer health status, and prescription drug needs, consumers' optimal plans change over time. However, in Medicare Part D only about 10% of consumers switch plans every year, and on average, plan choices worsen for those who do not switch. We develop a two-stage panel data model of plan choice whose stages correspond to two separate reasons for inertia: inattention and switching costs. The model allows for unobserved heterogeneity that is correlated across the two decision stages. We estimate the model using administrative data on Medicare Part D claims from 2007 to 2010. We find that consumers are more likely to pay attention to plan choice if overspending in the last year is more salient and if their old plan gets worse, for instance due to premium increases. Moreover, conditional on attention there are significant switching costs. Separating the two stages of the switching decision is thus important when designing interventions that improve consumers' plan choice. Florian Heiss University of Duesseldorf LS Statistics and Econometrics Universitaetsstrasse 1, Geb. 24.31 40225 Düsseldorf Germany florian.heiss@hhu.de Daniel McFadden University of California, Berkeley Department of Economics 508-1 Evans Hall #3880 Berkeley, CA 94720-3880 and NBER mcfadden@econ.berkeley.edu Joachim Winter Department of Economics LMU Munich Ludwigstr. 33 D-80539 Munich Germany joachim.winter@lrz.uni-muenchen.de Amelie Wuppermann Department of Economics LMU Munich Ludwigstr. 33 D-80539 Munich Germany amelie.wuppermann@lrz.uni-muenchen.de Bo Zhou University of Southern California 635 Downey Way Los Angeles, CA 90089-3331 zhoub@usc.edu
Discrete Choice Methods with Simulation by Kenneth Train has been available in the second edition since 2009. The book is published by Cambridge University Press and is also available for download ...
For the parametric estimation of logit models with individual time-invariant effects the conditional and unconditional fixed effects maximum likelihood estimators exist. The conditional fixed effects logit (CL) estimator is consistent but it has the drawback that it does not deliver estimates of the fixed effects or marginal effects. It is also computationally costly if the number of observations per individual T is large. The unconditional fixed effects logit estimator (UCL) can be estimated by including a dummy variable for each individual (DVL). It suffers from the incidental parameters problem which causes severe biases for small T. Another problem is that with a large number of individuals N, the computational costs of the DVL estimator can be prohibitive. We suggest a pseudo-demeaning algorithm in spirit of Greene (2004) and Chamberlain (1980) that delivers the identical results as the DVL estimator without its computational burden for large N. We also discuss how to correct for the incidental parameters bias of parameters and marginal effects. Monte-Carlo evidence suggests that the bias-corrected estimator has similar properties as the CL estimator in terms of parameter estimation. Its computational burden is much lower than the CL or the DVL estimators, especially with large N and/or T.
Individuals' socioeconomic status (SES) is positively correlated with their health status. While the existence of this gradient may be uncontroversial, the same cannot be said about its explanation. In this paper, we extend the approach of testing for the absence of causal channels developed by Adams et al. (2003), which in a Granger causality sense promises insights on the causal structure of the health-SES nexus. We introduce some methodological refinements and integrate retrospective survey data on early childhood circumstances into this framework. We confirm that childhood health has lasting predictive power for adult health. We also uncover strong gender differences in the intertemporal transmission of SES and health: While the link between SES and functional as well as mental health among men appears to be established rather late in life, the gradient among women seems to originate from childhood circumstances.
We consider how age-health profiles differ by demographic characteristics such as education, race, and ethnicity. A key feature of the analysis is the joint estimation of health and mortality to correct for the effect of mortality selection on observed age-health profiles. The model also allows for heterogeneity in individual health at a point in time and the persistence of the unobserved component of health over time. The observed component of health is based on a multidimensional index based on 27 indicators of health. Most of the key results are shown by simulations that illustrate the range of issues that can be addressed using the model. Differences in health by education and racial-ethnic group at age 50 persist throughout the remainder of life. Based on observed profiles, the health of whites is about 8 percentile points greater than the health of blacks at age 50 but by age 90 the gap is only 5 percentile points. However, when corrected for mortality selection, the health of blacks is actually declining more rapidly with age than the health of whites; the true gap widens with age. We also find that much of the difference in age-health profiles by racial-ethnic group is accounted for by differences in the levels of education between race-ethnic groups--from two-thirds to 85 percent for men and about half for women. We also simulate differences in survival probabilities by level of education and health and use these probabilities to calculate the expected present discounted value (EPDV) of an immediate annuity with first payout at age 66 for persons by gender, level of education, and health decile. The range of EPDVs is over two-fold for both men and women suggesting enormous potential for adverse selection.