
Most users of educational technology disengage within weeks of joining, and whether temporary incentives can durably change engagement remains unclear. We study contest-based gamification in a randomized controlled trial with 10,000 students on an English-learning app in India. Treated students entered a reading contest with leaderboard rankings and prizes for top performers. During the contest, treated students completed 45% more stories than control (0.96, SE 0.32). Twelve weeks after all incentives were removed, treated students completed 76% more stories (0.30, SE 0.14), relative to a control mean of 0.40, and the share of active users rose from 6.5% to 9.6%. Effects are concentrated in the upper tail of the engagement distribution, where nearly all platform activity is generated. Each story completed during the contest generates more lasting engagement for treated than for control students, and persistence extends beyond students who received prizes. Temporary contest-based incentives can durably expand the small segment of engaged users on which educational technology platforms depend.
Why do consumers search some products more than once (revisit) before making a purchase decision? How are these revisits related to search outcomes, such as consumers’ consideration sets and choices? In this paper, we build an online shopping website and run an incentive-aligned research study to find out. We show that most consumers revisit with the goal of comparing products, followed second by an expressed desire to obtain more information, and third by forgetting. Also, we document the ways in which search behavior differs across revisit motivations. Products revisited with the goal of obtaining additional information and comparing are more likely to be purchased than those revisited due to forgetting. Interestingly, revisits also reveal information about consumers’ consideration sets, which are typically unobserved: most consumers have eliminated a product from consideration if they don’t revisit it. This behavior implies consideration sets are dynamic, which we illustrate. Finally, managerial implications and possible extensions of the results are discussed.
We study players’ engagement decisions with different game modes in a popular multi-generational franchise video game. Some game modes are competitive where players’ utility might derive from performance, whereas other game modes are casual and feature the aspect of social interactions with other players. Model-free data patterns suggest that three potential mechanisms could jointly determine players’ decisions to engage in competitive versus casual game mode: unobserved preference heterogeneity, learning about their preferences for different modes, and skill accumulation. We develop a novel model that explains players’ decisions along both the extensive and the intensive margins with respect to different modes, while incorporating all three mechanisms. These mechanisms can be separately identified in our data set due to the panel feature of the multi-generational game. Our estimation results show several key insights about each mechanism. First, playing the competitive mode accumulates skills much faster than the casual mode. In addition, we find that 27 ∼ 32 ∼ 28
This study draws implications for targeting by incorporating switching reasons into a factor-analytic choice model for conducting benefit segmentation. In contrast to survey or conjoint-based studies, our segmentation task relies on a time series of brand choice decisions in real practice to improve external validity. Using physician-level panel data, we estimate a factor-analytic choice model to identify the primary product benefits each physician seeks. The key empirical challenge is that standard prescription data are entirely agnostic about the set of product benefits that underlie drug preferences. To address this issue, we utilize self-reported switching reasons in addition to the observed prescription choices. Accordingly, we extend the standard factor-analytic choice model to incorporate this augmented data and develop a Markov chain Monte Carlo (MCMC) procedure for estimation. Our proposed model enables us to directly identify which physicians are more efficacy, side effects, and/or cost saving oriented, an essential input to conducting benefit segmentation and fine-tuning subsequent targeted marketing promotion activities. We also investigate how misleading statistical inferences from standard factor-analytic choice models can be without the aid of augmented switching reasons data.
We study the effect of a procurement market affirmative action program where buyers can provide bid discounts to preferred (small, women-owned, and minority-owned) vendors, thereby enabling them to win contracts even if they are not the cheapest option. These programs are typically framed as an exercise in social responsibility, since buyers will pay slightly more in order to support traditionally disadvantaged businesses. Using data from business-to-government procurement auctions in Virginia and a structural model of vendors’ bidding behavior, we demonstrate that preferred vendors generally have higher costs than their non-preferred counterparts. As a consequence, the affirmative action program can reduce buyers’ procurement expenditures in this context through intensifying competition and forcing large, low-cost vendors to significantly reduce their prices. The magnitude of these savings is over four times larger when buyers strategically use a variable bid discount policy rather than committing to a fixed bid discount ahead of time. Our findings demonstrate that these affirmative action programs need not be a financial burden for buyers. Instead, affirmative action programs that improve societal outcomes can also improve other key metrics like procurement spending.
Using two decades of household purchase records (2004–2023) and retail scanner data, we document political polarization in everyday consumption. We examine products with health or environmental claims–organic produce, cage-free eggs, plant-based milk, eco-friendly household products–across ten categories. Ideological differences are negligible through 2012 but emerge around 2013 and widen thereafter. Controlling for demographics, we find that by 2023, liberal households purchase 1.8 percentage points more mindful products by volume than conservative households–roughly 40
A sequence of three randomized controlled trials (RCTs) is conducted to support a small online business’ decision of whether and how to implement AI in the creation of email marketing content. Recent developments in frequentist statistical decision theory are used to accommodate small samples available for testing in the small-business setting. The RCTs comprise three test policy cells with email content created by (i) salaried writers (“human”), (ii) a large-language model (“LLM”), and (iii) a “hybrid” combination of a human editing the content created by the LLM, respectively. When a “no email” control policy is included, all three test cells approximately double the gross profits from orders relative to the control cell. The RCTs vary whether the hybrid cell is edited by a salaried writer or the marketing team. The LLM cells vary whether the AI is pre-trained using historic emails or uses a prompt-based generative pre-trained transformer (“GPT”). Decision theory always selects one of the AI cells over the standard human policy on the basis of total annual profit net of related labor and software overhead.
We study how TikTok affects demand for music on paid streaming platforms in the context of Universal Music Group’s (UMG) global withdrawal of its catalog from TikTok. Recent studies using this quasi-natural experiment have reached different conclusions about whether TikTok promotes or cannibalizes streaming demand. We show that these differences arise in part because the estimand changes with the specification. With heavy-tailed outcomes, common difference-in-differences implementations in levels, logs, and Poisson place weight on different parts of the distribution and therefore answer different economic questions. In our data, the top 10 DiDestimands.app .
This paper studies the expansion of dollar store chains in the U.S. since 2008, which has generated public interest in their impact on retail markets and food accessibility. We show evidence that entry of dollar store chains significantly reduces fresh produce consumption for nearby households with low income and high travel costs. The impact of dollar store entry increases in the number of entries and we find no significant changes in spending in other product categories. These effects are driven in part by the indirect effects of dollar store entry on local market structure. We find that dollar stores expansion has led to a large decline in the number of grocery stores and that this decline in grocery store access causes reductions in produce purchases.
We argue that reputation mechanisms used by platform markets suffer from two problems.First, buyers may draw conclusions about the quality of the platform from single transactions, causing a reputational externality across sellers.Second, for a variety of reasons we discuss, reputations will be biased.We document these problems using eBay data and claim that platforms can benefit from identifying and promoting higher quality sellers.We create an unobservable measure of seller quality and demonstrate the benefits of our approach through a controlled experiment that prioritizes better quality sellers.We highlight the importance of reputational externalities and chart an agenda that aims to create more realistic models of platform markets.
Employee burnout has long plagued firms. The prevalence of burnout shows that work-related effort is not only costly in the present but has carryover effects into the future. We incorporate this ‘effort cost spillover’ into a dynamic, two-period principal-agent model, where the worker’s effort cost in the second period increases in both their second-period and first-period efforts. We use this model to explore optimal compensation design and the connection between incentives, burnout, and turnover. Naturally, turnover may occur if it is easy to replace workers, or if firms fail to account for burnout when designing contracts. However, we show that even when turnover is very costly, and firms and workers properly understand effort cost spillover, the firm’s equilibrium strategy may be to offer high-powered incentives that induce workers to work so hard that they exit (i.e. reject any contract that the firm would offer) in the next period. Workplace measures that reduce spillover, such as flexible work arrangements, can limit turnover and improve profits dramatically. Committing to contracts for both periods in advance can also limit turnover (at the cost of reduced flexibility).
This study presents a modeling framework for predicting rare events in relational data settings. Focusing on the rare disease market, it introduces a factor graph model within a Bayesian classifier that jointly models physician and patient features through their complex visit relationships. The framework is applied to an empirical case focused on identifying physicians treating hereditary angioedema patients, using extensive prescription and medical claims data. Our analysis demonstrates the model’s effectiveness, showing it surpasses various benchmark models in identifying rare disease physicians, including those currently recognized in healthcare databases and those likely to emerge in the future. This research contributes to the existing literature by addressing the challenge of predicting rare disease physicians and highlighting the benefits of leveraging relational dependencies among distinct entities to forecast rare events.
Accurately predicting consumer visits is critical for businesses to plan resources and serve their customers better. We examine the extent to which geo-tracking data about consumers' locations can improve the ability of businesses to predict total visits to their store. Given the sensitive nature of geo-tracking data, we also examine how potential privacy regulations that restrict such data may limit their usefulness for prediction. Using proprietary data from a safe-driving app with over 120 million driving instances across 38,980 app users and aggregate data on the total number of visits to over 400 restaurants in Texas, we quantify the value of geo-tracking data by training machine learning models to predict total visits to a restaurant one week ahead. Our results show that geo-tracking data improve the performance of prediction models by 10.76% relative to that of models that use demographic and behavioral data only. Simulation exercises that limit what data are tracked, where and how frequently these data are tracked show a decrease in the predictive performance of models that use geo-tracking data. However, the decrease varies by the type of restriction; regulations that restrict what data are geo-tracked (i.e., summaries of driving behaviors) result in the largest decreases in predictive performance (13.78%), while regulations that restrict where users are geo-tracked (i.e., within a few miles of a business location) and how frequently result in smaller decreases (5.52% and 3.15-3.41%, depending on the frequency). Importantly, models with geo-tracking generally outperform models that do not use any geo-tracking data by reducing the extent of overpredicting and underpredicting visits. These findings can assist managers and policymakers in assessing the risks and benefits associated with the use of geo-tracking data.
We measure the causal effects of within-household, over time, changes in household income and wealth on several household spending decisions - overall spending, CPG vs. non-CPG spending, and CPG spending across different store types, brand types, and their combinations - using a comprehensive, long-term, and representative panel of U.S. households. We compare the baseline allocations of household expenditures prior to the economic shock with the allocations of the incremental (reduced) household expenditure dollars due to the shock. Since brand managers and retail category managers make decisions based on the flow of consumer spending, we use this to derive managerial implications. We find that, compared to baseline allocations, households allocate their incremental household expenditures due to an increase in income to (a) non-CPG versus CPG categories; (b) warehouse clubs and discount stores versus grocery stores; and (c) national brands versus private labels. National brands account for a larger portion of the loss in grocery store allocation of incremental household expenditures which benefits warehouse clubs more than discount stores. Also, discount stores benefit more than warehouse clubs from the reallocation of private label spending away from grocery stores. Our results are robust across households with different levels of income but differ in magnitude across income groups.
In this research, we leverage an exogenous shock that temporally deprives consumers of accessing stores of a major retail chain due to a strike resulting in a one-week closure. This temporary store unavailability gives insights into how persistent grocery shopping behavior and retailer choice are after consumers have switched to competing grocery retail chains. Using household purchase data and a difference-in-differences approach, we show that consumers are more likely to visit and spend more at competing grocery chains during the temporary store closure. However, these changes in purchase behavior are only temporary. Consumers quickly revert to their pre-strike shopping behavior once the stores reopen. In a set of calibrated simulations, we illustrate that this persistence of shopping behavior is consistent with low levels of state dependence. The results of persistence in grocery shopping behavior are robust across perishable vs. non-perishable categories and across households with varying levels of share of wallet to the affected retail chain.
We utilize a novel dataset that merges household administrative income and socio-demographic information from tax records with scanner data on CPG consumption. Our analysis reveals significant variation in household per-capita expenditures. However, we find only a modest economic relationship between CPG spending and income, even amidst substantial within-household income fluctuations during the Dutch "double-dip recession" from 2011 to 2018. This relationship remains small for households with low income and low liquidity and holds across both food and non-food expenditures.
This research explores heterogeneity in customers' reference-dependent sensitivities using rich, individual level CRM data from a large casino in the U.S. and discusses implications for targeting decisions. We use a unique panel dataset of over 12,000 slot machine gamblers over 14 years and model heterogeneity in reference-dependent sensitivities at the individual level using a hierarchical Bayesian model. This analysis focuses on gains and losses relative to three reference points unique to the casino industry but conceptually extends to many other settings such as the financial services industry and hospital- ity: gambling outcomes relative to 1) zero, 2) prior trip outcomes, and 3) expected losses based on the house advantage of the slot machines. Firms can use heterogeneous reference-dependent sensitivities to improve their targeting decisions by considering the sequences of gambler outcomes in tandem with gamblers' individual sensitivities to marketing promotions. In our empirical application, we estimate that incorporating individual-level reference-dependent sensitivities improves targeted offer profitability by at least 19.8% relative to a comparable RFM model, depending on the offer type.
A randomized experiment with almost 35 million Pandora listeners enables us to measure the sensitivity of consumers to advertising, an important topic of study in the era of ad-supported digital content provision. The experiment randomized listeners into nine treatment groups, each of which received a different level of audio advertising interrupting their music listening, with the highest treatment group receiving more than twice as many ads as the lowest treatment group. By keeping consistent treatment assignment for 21 months, we are able to measure long-run demand effects, with three times as much ad-load sensitivity as we would have obtained if we had run a month-long experiment. We estimate a demand curve that is strikingly linear, with the number of hours listened decreasing linearly in the number of ads per hour (also known as the price of ad-supported listening). We also show the negative impact on the number of days listened and on the probability of listening at all in the final month. Using an experimental design that separately varies the number of commercial interruptions per hour and the number of ads per commercial interruption, we find that neither makes much difference to listeners beyond their impact on the total number of ads per hour. Lastly, we find that increased ad load causes a significant increase in the number of paid ad-free subscriptions to Pandora, particularly among older listeners.
We estimate the persuasive and dissuasive effects of media-reported campaign-trail speech on candidate electoral performance using a model in which, during each bipartisan race, candidates choose whether to address partisan supporters or swing voters, newspapers choose whether to cover each remark, and reported speech moves turnout and voter support. The model captures a core tension: while candidates seek selective media amplification of their messages, media outlets chase the most polarizing content. Using text analysis to label more than 200,000 newspaper stories from 1980-2012 as partisan or swing voter-oriented, and linking those labels to high-frequency polls of U.S. Senate races, we estimate the model parameters that allow us to recover persuasive and dissuasive effects of campaign speech. We rely on instruments that exploit sports-driven crowding-out of political coverage. The estimates reveal that Democratic appeals to their base are approximately four times more effective at mobilizing partisan turnout than Republican appeals, but also provoke twice the level of backlash from swing voters. These offsetting forces moderate Democratic rhetoric and, because media outlets prefer partisan content, yield near-symmetric coverage of the two parties despite asymmetric mobilization returns.