Many interventions, such as vaccines in clinical trials or coupons in online marketplaces, must be assigned sequentially without full knowledge of their effects. Multi-armed bandit algorithms have proven successful in such settings. However, standard independence assumptions fail when the treatment status of one individual impacts the outcomes of others, a phenomenon known as interference. We study optimal-policy learning under interference on a dynamic network. Existing approaches to this problem require repeated observations of the same fixed network and struggle to scale in sample size beyond as few as fifteen connected units – both limit applications. We show that under common assumptions on the structure of interference, rewards become linear. This enables us to develop a scalable Thompson sampling algorithm that maximizes policy impact when a new n-node network is observed each round. We prove a Bayesian regret bound that is sublinear in n and the number of rounds. Simulation experiments show that our algorithm learns quickly and outperforms existing methods. The results close a key scalability gap between causal inference methods for interference and practical bandit algorithms, enabling policy optimization in large-scale networked systems.
The contextual bandit framework is widely used to solve sequential optimization problems where the reward of each decision depends on auxiliary context variables. In settings such as medicine, business, and engineering, the decision maker often possesses additional structural information on the generative model that can potentially be used to improve the efficiency of bandit algorithms. We consider settings in which the mean reward is known to be a concave function of the action for each fixed context. Examples include patient-specific dose-response curves in medicine and expected profit in online advertising auctions. We propose a contextual bandit algorithm that accelerates optimization by conditioning the posterior of a Bayesian Gaussian Process model on this concavity information. We design a novel shape-constrained reward function estimator using a specially chosen regression spline basis and constrained Gaussian Process posterior. Using this model, we propose a UCB algorithm and derive corresponding regret bounds. We evaluate our algorithm on numerical examples and test functions used to study optimal dosing of Anti-Clotting medication.
BACKGROUND/AIMS:Pain is common in cancer patients and results in lower quality of life, depression, poor physical functioning, financial difficulty, and decreased survival time. Behavioral pain interventions are effective and nonpharmacologic. Traditional randomized controlled trials (RCT) test interventions of fixed time and dose, which poorly represent successive treatment decisions in clinical practice. We utilize a novel approach to conduct a RCT, the sequential multiple assignment randomized trial (SMART) design, to provide comparative evidence of: 1) response to differing initial doses of a pain coping skills training (PCST) intervention and 2) intervention dose sequences adjusted based on patient response. We also examine: 3) participant characteristics moderating intervention responses and 4) cost-effectiveness and practicality. METHODS/DESIGN:Breast cancer patients (N=327) having pain (ratings≥5) are recruited and randomly assigned to: 1) PCST-Full or 2) PCST-Brief. PCST-Full consists of 5 PCST sessions. PCST-Brief consists of one 60-min PCST session. Five weeks post-randomization, participants re-rate their pain and are re-randomized, based on intervention response, to receive additional PCST sessions, maintenance calls, or no further intervention. Participants complete measures of pain intensity, interference and catastrophizing. CONCLUSIONS:Novel RCT designs may provide information that can be used to optimize behavioral pain interventions to be adaptive, better meet patients' needs, reduce barriers, and match with clinical practice. This is one of the first trials to use a novel design to evaluate symptom management in cancer patients and in chronic illness; if successful, it could serve as a model for future work with a wide range of chronic illnesses.
The center of pressure (COP) position reflects a combination of proprioceptive, motor and mechanical function. As such, it can be used to quantify and characterize neurologic dysfunction. The aim of this study was to describe and quantify the movement of COP and its variability in healthy chondrodystrophoid dogs while walking to provide a baseline for comparison to dogs with spinal cord injury due to acute intervertebral disc herniations. Fifteen healthy adult chondrodystrophoid dogs were walked on an instrumented treadmill that recorded the location of each dog's COP as it walked. Center of pressure (COP) was referenced from an anatomical marker on the dogs' back. The root mean squared (RMS) values of changes in COP location in the sagittal (y) and horizontal (x) directions were calculated to determine the range of COP variability. Three dogs would not walk on the treadmill. One dog was too small to collect interpretable data. From the remaining 11 dogs, 206 trials were analyzed. Mean RMS for change in COPx per trial was 0.0138 (standard deviation, SD 0.0047) and for COPy was 0.0185 (SD 0.0071). Walking speed but not limb length had a significant effect on COP RMS. Repeat measurements in six dogs had high test retest consistency in the x and fair consistency in the y direction. In conclusion, COP variability can be measured consistently in dogs, and a range of COP variability for normal chondrodystrophoid dogs has been determined to provide a baseline for future studies on dogs with spinal cord injury.
Double-blind randomized placebo-controlled trials, in which a new approval-seeking drug is compared to placebo, are generally considered to be the gold standard for evaluating the efficacy of a psychopharmacological intervention in the treatment of affective disorders. However, this type of study has substantial limitations regarding the external validity of its results. Therefore, alternatives within randomized controlled trials, such as non-inferiority trials or study designs in which the new drug or placebo is given simultaneously on top of standard antidepressant or anti-manic treatment are of substantial interest. In addition, innovative statistical approaches such as marginal structural models or Q-learning should be employed to evaluate potential causal effects associated with different fixed treatment regimes that are realized in actual clinical practice and help to establish treatment recommendations that are tailored to individual patient characteristics.
9.1 ▪ IntroductionThe area of personalized medicine is based on the premise that a patient's individual characteristics are implicated in which treatments are likely to benefit him/her. The most popular perspective on personalized medicine emphasizes identification of subgroups of patients who share certain characteristics, most often genetic/genomic features, and who are likely to benefit from a specific treatment. This goal understandably is of considerable interest in pharmaceutical research and in a regulatory context and has centered around development of biomarkers that can be used to identify such patients and targeting of new products to biomarker-defined subgroups. This point of view can be summarized as focusing on "the right patient for the treatment."
• Sequential, multiple assignment, randomized trials are an effective way for gathering data to learn personalized treatment strategies [1] • Many analytic techniques have been developed to learn the best individualized treatment strategies from data, including iterative minimization of regrets [2], G-estimation [3], reinforcement learning methods [4, 5, 6], regret-regression [7] and more parametric approaches [8, 9, 10]. • Many practical challenges arise when applying these methods to data collected from clinical trials [11]. • Missing data: drop out and missed exams can lead to missing information in clinical trials • Imputation replaces missing data with predicted values to obtain complete data sets • Having complete data is especially important when outcome is a function of the whole trial
Kang, Janes and Huang propose an interesting boosting method to combine biomarkers for treatment selection. The method requires modeling the treatment effects using markers. We discuss an alternative method, outcome weighted learning. This method sidesteps the need for modeling the outcomes, and thus can be more robust to model misspecification.
The Big Data Research and Development Initiative is now in its third year and making great strides to address the challenges of Big Data. To further advance this initiative, we describe how statistical thinking can help tackle the many Big Data challenges, emphasizing that often the most productive approach will involve multidisciplinary teams with statistical, computational, mathematical, and scientific domain expertise.
Inspired by the non-regular framework studied in Laber and Murphy (2011), we propose a family of adaptive classifiers. We discuss briefly their asymptotic properties and show that under the non-regular framework these classifiers have an "oracle property," and consequently have smaller asymptotic variance and smaller asymptotic test error variance than those of the original classifier. We also show that confidence intervals for the test error of the adaptive classifiers, based on either normal approximation or centered percentile bootstrap, are consistent.