Biomedical research may uncover insights regarding the interaction of the treatments of a disease with patient covariates. We show how to use such insights to improve the efficiency of adaptive clinical trials for precision medicine by extending ideas from optimal Bayesian learning. We present a model for response-adaptive multi-arm clinical trials that leverages knowledge about the predictive-prognostic covariate structure to accelerate the learning of personalized treatment strategies that obtain the best expected outcomes for post-trial patients. Our base model is a contextual linear bandit for best-arm identification, and outcomes may be observed with delay. We characterize the optimal policy for sequentially allocating treatments to in-trial patients and, because it is hard to compute, propose several computable heuristics based on Bayesian one-step look-ahead techniques. We prove that several of our proposed heuristics are asymptotically optimal in learning treatment strategies. Numerical results based on two case studies motivated by sepsis management show that our heuristics can significantly improve clinical trial efficiency to learn a treatment strategy for precision medicine. We provide extensions that allow for rewards from outcomes of in-trial patients (resolving the exploration-exploitation tradeoff) and for inferring covariate structure using Lasso when biomedical insights on covariate structure are lacking. Our proposed trial design is of interest to funders, designers, and managers of clinical trials. It may also apply to other contextual bandit problems in settings where insights about covariate-treatment interactions are available.
We propose an approach for determining the sample size required when using an experiment to train and certify a targeting policy. Calculating the rate at which the performance of a targeting model improves with additional training data is a complex problem. We address this challenge by assuming that customers are grouped into segments that capture relevant information about their responsiveness to the firm’s marketing actions. We consider two problem formulations. The first formulation identifies the sample size required to train a targeting policy and certify that its expected performance exceeds a predefined threshold. The second formulation identifies the sample size required to train a targeting policy and certify that it outperforms a baseline in an out-of-sample statistical test. We establish theoretical properties of these problems, based on which we propose computationally efficient algorithms for optimal sample size calculations. We illustrate our algorithms and analysis using data from a luxury fashion retailer. This paper was accepted by David Simchi-Levi, marketing. Supplemental Material: The online appendix and data files are available at https://doi.org/10.1287/mnsc.2022.02947 .
Contextual ranking & selection is attracting increasing attention in simulation and other fields. A successful approach to addressing related challenges uses arm allocation indices that compute the Bayesian expected value of information of one-step look-ahead policies. We recall recent work on such indices for linear contextual bandits that take advantage of structural information about the nature of the covariates that describe contexts. Such indices can be computed exactly with a finite number of contexts and no delay in observing outcomes, but may require Monte Carlo simulation otherwise. Our contribution is to describe and quantify the benefits of two variance reduction techniques (conditional Monte Carlo and common random numbers) to estimate such allocation indices for contextual ranking & selection problems when some covariates are continuous or outcomes are observed with delay. We find that both techniques significantly improve estimates and the speed of inference, but conditioning is particularly useful.
In the context of subscription-based services, many technologies improve over time and service providers can provide increasingly powerful service upgrades to their customers, but at a launching cost, and the expense of the sales of existing products. We propose a model of technology upgrades and characterize the optimal pricing and timing of technology introductions for a service provider who price-discriminates among customers based on their upgrade experience, in the face of customers who are averse to switching to improved offerings. We first characterize optimal discriminatory pricing for the infinite horizon pricing problem with fixed introduction times. We reduce the optimal pricing problem to a tractable optimization problem and propose an efficient algorithm for solving it. Our algorithm computes optimal discriminatory prices within a fraction of a second, even for large problem instances. We then show that periodic introduction times, combined with optimal pricing, enjoy optimality guarantees. In particular, we first show that as long as the introduction intervals are constrained to be non-increasing, it is optimal to have periodic introductions after an initial warm-up phase. When allowing general introduction intervals, we show that periodic introduction intervals after some time are optimal in a more restricted sense. Numerical experiments suggest that it is generally optimal to have periodic introductions after an initial warm-up phase. Finally, we focus on a setting in which the firm does not price-discriminate based on customers' experience. We show both analytically and numerically that in the non-discriminatory setting, a simple policy of Myerson (i.e., myopic) pricing and periodic introductions enjoys good performance guarantees.
Abstract Although covert warfare does not readily lend itself to scientific inquiry, new technologies are increasingly providing scholars with tools that enable such research. In this note, we examine the effects of drone strikes on patterns of communication in Yemen using big data and anomaly detection methods. The combination of these analytic tools allows us to not only quantify some of the effects of drone strikes, but also to compare them to other shocks. We find that on average drone strikes leave a footprint in their aftermath, spurring significant but localized spikes in communication. This suggests that drone strikes are not a purely surgical intervention, but rather have a disruptive impact on the local population.
We consider the problem of sequentially allocating sample observations to learn personalized treatment strategies, motivated by the design of adaptive clinical trials that aim to learn the best treatment as a function of patient covariates. In such settings there may be clinical knowledge of which covariates are predictive (they may interact with the treatment choice) and which are prognostic (they may influence the outcome independent of treatment choice). We extend the expected value of information (EVI)/knowledge gradient framework to develop useful heuristics for a context with predictive and prognostic covariates and a delay in observing outcomes. We also propose and analyze closely related Monte Carlo-based allocation policies to enhance our proposal's computational efficiency and applicability for adaptive contextual learning. We show that several of our proposed allocation policies are asymptotically optimal in learning treatment strategies. We run simulation experiments motivated by an application for clinical trial design to assess potential treatments of sepsis. We illustrate that the proposed EVI-based allocation policies, with knowledge about which covariates are predictive and prognostic, can improve the rate of inference relative to some existing approaches to adaptive contextual learning.
We consider the contextual ranking and selection problem that aims to learn the best treatment as a function of covariates when expected outcomes are unknown but can be learned from noisy observations. We develop a sequential allocation policy based on Bayesian expected value of information methods, called fEVI, to learn the best treatment for a finite set of covariates. We observe good performance of the $f$ EVI allocation policy in simulation experiments and find that prior distributions which accurately reflect correlations across treatments and patient types can improve sampling effectiveness with limited sample sizes. We compare the performance between the case when covariates are random arrivals from a population, and the case when the allocation policy chooses covariates. In experiments, the benefit of $f$ EVI over allocation policies that sample randomly is much larger than the benefit from being able to choose covariates, or from using a prior that accurately reflects correlations.
To respond to pandemics such as COVID-19, policy makers have relied on interventions that target specific population groups or activities. Because targeting is operationally challenging and contentious, rigorously quantifying its benefits and designing practically implementable policies that achieve some of these benefits is critical for effective and equitable pandemic control. We propose a flexible framework that leverages publicly available data and a novel optimization algorithm based on model predictive control and trust region methods to compute optimized interventions that can target two dimensions of heterogeneity: age groups and the specific activities that individuals normally engage in. We showcase a complete implementation focused on the Île-de-France region of France and use this case study to quantify the benefits of dual targeting and to propose practically implementable policies. We find that dual targeting can lead to Pareto improvements, reducing the number of deaths and the economic losses. Additionally, dual targeting allows maintaining higher activity levels for most age groups and, importantly, for those groups that are most confined, thus leading to confinements that are arguably more equitable. We then fit decision trees to explain the decisions and gains of dual-targeted policies and find that they prioritize confinements intuitively, by allowing increased activity levels for group-activity pairs with high marginal economic value prorated by social contacts, which generates important complementarities. Because dual targeting can face significant implementation challenges, we introduce two practical proposals inspired by real-world interventions — based on curfews and recommendations — that achieve a significant portion of the benefits without explicitly discriminating based on age.
We investigate how firms can use the results of field experiments to optimize the targeting of promotions when prospecting for new customers. We evaluate seven widely used machine-learning methods using a series of two large-scale field experiments. The first field experiment generates a common pool of training data for each of the seven methods. We then validate the seven optimized policies provided by each method together with uniform benchmark policies in a second field experiment. The findings not only compare the performance of the targeting methods, but also demonstrate how well the methods address common data challenges. Our results reveal that when the training data are ideal, model-driven methods perform better than distance-driven methods and classification methods. However, the performance advantage vanishes in the presence of challenges that affect the quality of the training data, including the extent to which the training data captures details of the implementation setting. The challenges we study are covariate shift, concept shift, information loss through aggregation, and imbalanced data. Intuitively, the model-driven methods make better use of the information available in the training data, but the performance of these methods is more sensitive to deterioration in the quality of this information. The classification methods we tested performed relatively poorly. We explain the poor performance of the classification methods in our setting and describe how the performance of these methods could be improved. This paper was accepted by Matthew Shum, marketing.
As the quality of computer hardware increases over time, cloud service providers have the ability to offer more powerful virtual machines (VMs) and other resources to their customers. But providers face several trade-offs as they seek to make the best use of improved technology. On one hand, more powerful machines are more valuable to customers and command a higher price. On the other hand, there is a cost to develop and launch a new product. Further, the new product competes with existing products. Thus, the provider faces two questions. First, when should new classes of VMs be introduced? Second, how should they be priced, taking into account both the VM classes that currently exist and the ones that will be introduced in the future? This decision problem, combining scheduling and pricing new product introductions, is common in a variety of settings. One aspect that is more specific to the cloud setting is that VMs are rented rather than sold. Thus existing customers can switch to a new offering, albeit with some inconvenience. There is indeed evidence of customers' aversion to upgrades in the cloud computing services market. Based on a study of Microsoft Azure, we estimate that customers who arrive after a new VM class is launched are 50% more likely to use it than existing customers, indicating that these switching costs may be substantial. This opens up a wide range of possible policies for the cloud service provider. Our main result is that a surprisingly simple policy is close to optimal in many situations: new VM classes are introduced on a periodic schedule and each is priced as if it were the only product being offered. (Periodic introductions have been noticeable in the practice of cloud computing. For example, Amazon Elastic Compute Cloud (EC2) launched new classes of the m.xlarge series in October 2007, February 2010, October 2012, and June 2015, i.e., in intervals of 28 months, 32 months, and 32 months.) We refer to this pricing policy as Myerson pricing, as these prices can be computed as in Myerson's classic paper (1981). This policy produces a marketplace where new customers always select the newest and best offering, while existing customers may stick with older VMs due to switching costs.
Many technologies improve over time and technology providers can provide increasingly powerful service upgrades to their customers, but at a launching cost, and the expense of the sales of existing products. We propose a model of technology introduction in the context of subscription-based services and characterize the optimal pricing and timing of technology introductions for a service provider, in the face of customers who are averse to switching to improved offerings. Overall, we show that a simple policy of Myerson (i.e., myopic) pricing and periodic introductions is approximately optimal. We first show that under a linear pricing rule, which subsumes Myerson pricing, there is no loss of optimality with a periodic schedule of introductions, and that under periodic introductions, the potential additional revenue of any pricing policy over Myerson pricing decays to zero after sufficiently many introductions. We then argue that Myerson pricing is approximately optimal under arbitrary introduction times. To do so, we first characterize prices that achieve optimal revenue in a single period, given arbitrary fixed introduction times. We then establish that Myerson pricing achieves a bounded bicriteria approximation ratio to both revenue and cost, for the infinite-horizon problem. Third, we provide analytical bounds for the approximation ratio in terms of the customer type distribution. Our bounds show that Myerson pricing is approximately optimal when switching costs for the customers who upgrade are small or large. Following our analysis, we examine our analytical bounds for Myerson pricing with simulations and show that, after sufficiently many introductions, they are tight for all values of the switching cost, for several natural distributions for the customer type. Furthermore, when we numerically compute optimal prices for fixed introduction times, rather than using our analytical bounds, we find that Myerson pricing is often several orders of magnitude closer to optimal revenue than our analytical bounds suggest. Our conclusions on the quality of Myerson pricing can robustly be carried over to realistic settings for the switching costs, the customer lifetime, and the frequency of technology introductions.
Firms must often decide how to target households that did not respond to past promotions. For example, when prospecting for new customers, households that purchase are no longer eligible, and so the remaining households are an increasingly pure pool of non-responders. Past response data reflects the behavior of households that responded, and these households generally differ from the rest in unobservable ways. We show that despite this, the decisions of the past responders in a group can help firms target the remaining non-responders. In particular, the timing of past responses can help to reveal whether a geographic region is (a) exhausted of future responders, or (b) the future responders just require additional exposures. We use this insight to develop three different timing measures. The measures are calculated and validated using a sequence of mailings in a large-scale field experiment. The measures perform well, even when some regions have only a handful of past responses to calibrate them. They confirm that the decisions of past responders can help firms target non-responders.
Champion versus challenger field experiments are widely used to compare the performance of different targeting policies. These experiments randomly assign customers to receive marketing actions recommended by either the existing (champion) policy or the new (challenger) policy, and then compare the aggregate outcomes. We recommend an alternative experimental design and propose an alternative estimation approach to improve the evaluation of targeting policies. The recommended experimental design randomly assigns customers to marketing actions. This allows evaluation of any targeting policy without requiring an additional experiment, including policies designed after the experiment is implemented. The proposed estimation approach identifies customers for whom different policies recommend the same action and recognizes that for these customers there is no difference in performance. This allows for a more precise comparison of the policies. We illustrate the advantages of the experimental design and estimation approach using data from an actual field experiment. We also demonstrate that the grouping of customers, which is the foundation of our estimation approach, can help to improve the training of new targeting policies. This paper was accepted by Matthew Shum, marketing.
We study the role of local information channels in enabling coordination among strategic agents. Building on the standard finite-player global games framework, we show that the set of equilibria of a coordination game is highly sensitive to how information is locally shared among different agents. In particular, we show that the coordination game has multiple equilibria if there exists a collection of agents such that (i) they do not share a common signal with any agent outside of that collection and (ii) their information sets form an increasing sequence of nested sets. Our results thus extend the results on the uniqueness and multiplicity of equilibria beyond the well-known cases in which agents have access to purely private or public signals. We then provide a characterization of the set of equilibria as a function of the penetration of local information channels. We show that the set of equilibria shrinks as information becomes more decentralized.
The feasibility of using field experiments to optimize marketing decisions remains relatively unstudied. We investigate category pricing decisions that require estimating a large matrix of cross-product demand elasticities and ask the following question: How many experiments are required as the number of products in the category grows? Our main result demonstrates that if the categories have a favorable structure, we can learn faster and reduce the number of experiments that are required: the number of experiments required may grow just logarithmically with the number of products. These findings potentially have important implications for the application of field experiments. Firms may be able to obtain meaningful estimates using a practically feasible number of experiments, even in categories with a large number of products. We also provide a relatively simple mechanism that firms can use to evaluate whether a category has a structure that makes it feasible to use field experiments to set prices. We illustrate how to accomplish this using either a sample of historical data or a pilot set of experiments. We also discuss how to evaluate whether field experiments can help optimize other marketing decisions, such as selecting which products to advertise or promote.Data, as supplemental material, are available at http://dx.doi.org/10.1287/mnsc.2014.2066.This paper was accepted by Pradeep Chintagunta, marketing.
We infer local influence relations between networked entities from data on outcomes and assess the value of temporal data by formulating relevant binary hypothesis testing problems and characterizing the speed of learning of the correct hypothesis via the Kullback-Leibler divergence, under three different types of available data: knowing the set of entities who take a particular action; knowing the order that the entities take an action; and knowing the times of the actions.
Protection of one's intellectual property is a topic with important technological and legal facets. We provide mechanisms for establishing the ownership of a dataset consisting of multiple objects. The algorithms also preserve important properties of the dataset, which are important for mining operations, and so guarantee both right protection and utility preservation. We consider a right-protection scheme based on watermarking. Watermarking may distort the original distance graph. Our watermarking methodology preserves important distance relationships, such as: the Nearest Neighbors (NN) of each object and the Minimum Spanning Tree (MST) of the original dataset. This leads to preservation of any mining operation that depends on the ordering of distances between objects, such as NN-search and classification, as well as many visualization techniques. We prove fundamental lower and upper bounds on the distance between objects post-watermarking. In particular, we establish a restricted isometry property, i.e., tight bounds on the contraction/expansion of the original distances. We use this analysis to design fast algorithms for NN-preserving and MST-preserving watermarking that drastically prune the vast search space. We observe two orders of magnitude speedup over the exhaustive schemes, without any sacrifice in NN or MST preservation.
How can the details of who observes what affect the outcomes of economic, social, and political interactions? Our thesis is that outcomes do not depend merely on the status quo and the available (noisy) information on it; they also crucially depend on how the available pieces of information are allocated among strategic agents. We study the dependence of coordination outcomes on local information sharing. In the economic literature, common knowledge of the fundamentals leads to the standard case of multiple equilibria due to the self-fullling nature of agents' beliefs. The global-games framework has been used extensively as a toolkit for arriving at a unique equilibrium selection in the context of coordination games, which can model bank runs, currency attacks, and social uprisings, among others. Yet, there is a natural mechanism through which multiplicity can reemerge, while keeping information solely exogenous: the (exogenous) information structure per se, namely, the details of who observes what noisy observation. The aim of this paper is to understand the role of the exogenous information structure in the determination and characterization of equilibria in the coordination game. We answer the question of how the equilibria of the coordination game depend on the details of local information sharing. Our main contribution is to provide conditions for uniqueness and multiplicity that pertain solely to the details of information sharing. The findings in the present paper give an immediate answer as to the determinacy of equilibria using only the characterization of what agent observes what pieces of information. We build on the standard global game framework for coordination games with incomplete and asymmetric information and consider a coordination game in which each of a collection of agents decides whether to take a risky action (whose payoff depends on how many agents made the same decision, and the fundamentals) or a safe action, based on their noisy observations regarding the fundamentals. Generalizing away from the standard practice of considering only private and public signals, we allow for signals that are observed by arbitrary subsets of the agents. We refer to signals that are neither private nor public as local signals. We pose the following question: how do the equilibria of the coordination game depend on the information locality, i.e., on the details of local information sharing. Our key finding is that the number of equilibria is highly sensitive to the details of information locality. As a result, a new dimension of indeterminacy regarding the outcomes is being introduced: not only may the same fundamentals well lead to different outcomes in different societies, due to different realizations of the noisy observations; the novel message of this work is that the same realization of the noisy observations is compatible with different equilibrium outcomes in societies with different structures of local information sharing. In particular, we show that as long as a collection of agents share the same observations, and no other agent's observations overlap with their common observations, multiple equilibria arise. Identical observations is not, nevertheless, a necessary condition for multiplicity: we show that as long as the observations of a collection of agents form a cascade of containments, and no other agent's observations overlap with the observations of the collection, then multiplicity emerges. This is not to say however that common knowledge of information at the local level necessarily implies multiplicity: in particular, in the absence of identical observations or cascade of containments of observations, or if the condition of no overlap of information is violated, then, despite the presence of some signals that are common knowledge between agents, a unique equilibrium may be selected. In the case where each agent observes exactly one signal, we characterize the set of equilibria as a function of the details of the information structure. We show how the distance between the largest and smallest equilibria depends on how information is locally shared among the agents. In particular, the more equalized the sizes of the sets of agents who observe the same signal, the more diverse the information of each group becomes, heightening inter-group strategic uncertainty, and leading to a more refined set of equilibria. We use our characterization to study the set of equilibria in large coordination games. We show that as the number of agents grows, the game exhibits a unique equilibrium if and only if the largest set of agents with access to a common signal grows sublinearly in the number of agents, thus identifying a sharp threshold for uniqueness versus multiplicity.
Consumers adopting a new product; an epidemic spreading across a population; a sovereign debt crisis hitting several countries; a cellular process during which the expression of a gene affects the expression of other genes; an article trending in the blogosphere, a topic trending on an online social network, computer malware spreading across a network; all of these are temporal processes governed by local interactions of networked entities, which influence one another. Due to the increasing capability of data acquisition technologies, rich data on the outcomes of such processes are oftentimes available (possibly with time stamps), yet the underlying network of local interactions is hidden. In this work, we infer who influences whom in a network of interacting entities based on data of their actions/ decisions, and quantify the gain of learning based on sequences of actions versus sets of actions. We answer the following question: how much faster can we learn influences with access to increasingly informative temporal data (sets versus sequences)?