Supplementary Materials from Cancer-Specific High-Throughput Annotation of Somatic Mutations: Computational Prediction of Driver Missense Mutations
We describe a method for evaluating an ensemble of predictive models given a sample of observations comprising the model predictions and the outcome event measured with error. Our formulation allows us to simultaneously estimate measurement error parameters, true outcome — aka the gold standard — and a relative weighting of the predictive scores. We describe conditions necessary to estimate the gold standard and for these estimates to be calibrated and detail how our approach is related to, but distinct from, standard model combination techniques. We apply our approach to ∗This work was funded in part by the Cancer Genetics Network at Duke, U24 CA78157, and at The Johns Hopkins University, U24 CA78148. In addition, G.P. and S.C. were supported by the Johns Hopkins SPORE in Breast Cancer (P50CA88843) and NCI grants R01CA105090-01A1 and P50CA62924-05; E.I. also received support from NCI through R01CA105090-01A1. The authors wish to thank this manuscript’s editors and referees for their comments and suggestions and the following individuals for their invaluable contributions to the success of the CGN BRCA1/2 Models Validation Study: Tara Friebel, Center for Clinical Epidemiology and Biostatistics, University of Pennsylvania; Dianne Finkelstein, The Massachusetts General Hospital; Hoda Anton-Culver, Department of Medicine, University of California, Irvine; Argyrios Ziogas, Department of Medicine, University of California, Irvine; Barbara L. Weber, Abramson Cancer Center, University of Pennsylvania; Andrea Eisen, Hamilton Regional Cancer Centre; Kathleen E. Malone, Program in Epidemiology, Fred Hutchinson Cancer Research Center; Li Hsu, Public Health Sciences Division, Fred Hutchinson Cancer Research Center; Leif E. Peterson, Departments of Medicine and Molecular and Human Genetics Baylor College of Medicine; Joellen M. Schildkraut, Department of Community and Family Medicine, Duke University; Claudine Isaacs, Georgetown University Lombardi Cancer Center; Beth N. Peshkin, Georgetown University Lombardi Cancer Center; Camille Corio, Georgetown University Lombardi Cancer Center; Leoni Leondaridis, Georgetown University Lombardi Cancer Center; Gail Tomlinson, Department of Pediatrics, University of Texas Southwestern; Christopher I. Amos, Department of Epidemiology, University of Texas M. D. Anderson Cancer Center; Louise C. Strong, Department of Pediatrics, University of Texas M. D. Anderson Cancer Center; Donald A. Berry, Department of Biostatistics, University of Texas M. D. Anderson Cancer Center; Jeffrey Weitzel, City of Hope National Medical Center; Sharon Sand, City of Hope National Medical Center; Debra Dutson, Huntsman Cancer Center, University of Utah; Rich Kerber, Huntsman Cancer Center, University of Utah; David M. Euhus, Department of Surgery, University of Texas Southwestern.
We study mobile user behavior from cell-level location trace (CLLT). Since CLLT contains no GPS coordinates of mobile users, we infer approximate user locations from the cell locations they visit. We build upon the vast literature on user behavior analysis and demonstrate the ability to extract user behavior in the absence of the more precise GPS information. We focus on the “leisure time” behavior, i.e. activities outside home and office. In particular, we compare pairs of users and study their similarity or the lack of it. Such similarity comparison can be done directly using the users' actual cell-level locations and the times of their visits. We observe that a user's behavior on different days tends to be more similar to oneself than to others. We then compare users in terms of their activities irrespective of physical locations. We develop the notion of semantic cell type which classifies the cells according to the consistency of points of interest within the cells. In this way, we can compare two users based on the type of cells they visit and extract similarity from there. As a result, we gain understanding of the general profiles of the cells and the users. We are able to differentiate user behavior and cluster them in a meaningful way.
Data analytics in smart grids can be leveraged to channel the data downpour from individual meters into knowledge valuable to electric power utilities and end-consumers. Short-term load forecasting (STLF) can address issues vital to a utility but it has traditionally been done mostly at system (city or country) level. In this case study, we exploit rich, multi-year, and high-frequency annotated data collected via a metering infrastructure to perform STLF on aggregates of power meters in a mid-sized city. For smart meter aggregates complemented with geo-specific weather data, we benchmark several state-of-the-art forecasting algorithms, including kernel methods for nonlinear regression, seasonal and temperature-adjusted auto-regressive models, exponential smoothing and state-space models. We show how STLF accuracy improves at larger meter aggregation (at feeder, substation, and system-wide level). We provide an overview of our algorithms for load prediction and discuss system performance issues that impact real time STLF. © 2014 Alcatel-Lucent.
Today, data collected by service providers can track an individual user's experience in detail, at flow or packet level in real time. However, we still lack analytics methods that can translate this information into a comprehensive and ever-evolving representation of the user experience. In this paper, we provide a layered dynamic model that addresses the problem of how to relate low-level network performance metrics to a user's perception of network service and their subsequent actions. Using time-stamped observations from networks, devices, and customer care, we build probabilistic models to link network performance to an inferred state of customer satisfaction, and then to explicit and implicit customer disengagement events. We provide inference algorithms for the model parameters, and report test results on synthesized datasets based on real, but incomplete, observations. We discuss how popular anonymization techniques such as data masking, encryption, k-anonymization, and differential privacy can be used to protect sensitive and private user data without impacting the user experience inference. (c) 2014 Alcatel-Lucent.
In retrospective studies, odds ratio is often used as the measure of association. Under independent beta prior assumption, the exact posterior distribution of odds ratio given a single 2 × 2 table has been derived in the literature. However, independence between risks within the same study may be an oversimplified assumption because cases and controls in the same study are likely to share some common factors and thus to be correlated. Furthermore, in a meta-analysis of case–control studies, investigators usually have multiple 2 × 2 tables. In this article, we first extend the published results on a single 2 × 2 table to allow within study prior correlation while retaining the advantage of closed-form posterior formula, and then extend the results to multiple 2 × 2 tables and regression setting. The hyperparameters, including within study correlation, are estimated via an empirical Bayes approach. The overall odds ratio and the exact posterior distribution of the study-specific odds ratio are inferred based on the estimated hyperparameters. We conduct simulation studies to verify our exact posterior distribution formulas and investigate the finite sample properties of the inference for the overall odds ratio. The results are illustrated through a twin study for genetic heritability and a meta-analysis for the association between the N-acetyltransferase 2 (NAT2) acetylation status and colorectal cancer.
Major efforts to sequence cancer genomes are now occurring throughout the world. Though the emerging data from these studies are illuminating, their reconciliation with epidemiologic and clinical observations poses a major challenge. In the current study, we provide a mathematical model that begins to address this challenge. We model tumors as a discrete time branching process that starts with a single driver mutation and proceeds as each new driver mutation leads to a slightly increased rate of clonal expansion. Using the model, we observe tremendous variation in the rate of tumor development—providing an understanding of the heterogeneity in tumor sizes and development times that have been observed by epidemiologists and clinicians. Furthermore, the model provides a simple formula for the number of driver mutations as a function of the total number of mutations in the tumor. Finally, when applied to recent experimental data, the model allows us to calculate the actual selective advantage provided by typical somatic mutations in human tumors in situ. This selective advantage is surprisingly small—0.004 ± 0.0004—and has major implications for experimental cancer research.
Voxel‐based morphometry (VBM) is widely used as a high‐resolution approach to understanding the relationship between anatomical structures and variables of interest. Controlling for the false discovery rate (FDR) is an attractive choice for thresholding the resulting statistical maps and has been commonly used in fMRI studies. However, we caution against the use of nonadaptive FDR control procedures, such as the most commonly used Benjamini–Hochberg procedure (B‐H), in VBM analyses. This is because, in VBM analyses, specific risk factors may be associated with volume change in a global, rather than local, manner, which means the proportion of truly associated voxels among all voxels is large. In such a case, the achieved FDR obtained by nonadaptive procedures can be substantially lower than the nominal, or controlled, level. Such conservatism deprives researchers of power for detecting true associations. In this article, we advocate for the use of adaptive FDR control in VBM‐type analyses. Specifically, we examine two representative adaptive procedures: the two‐stage step‐up procedure by Benjamini, Krieger and Yekutieli ( 2006 : Biometrika 93:491–507) and the procedure of Storey and Tibshirani ([2003]: Proc Natl Acad Sci USA 100:9440–9445). We demonstrate mathematically, with simulations, and with a data example that these procedures provide improved performance over the B‐H procedure. Hum Brain Mapp, 2009. © 2008 Wiley‐Liss, Inc.
In studies of the accuracy of diagnostic tests, it is common that both the diagnostic test itself and the reference test are imperfect. This is the case for the microsatellite instability test, which is routinely used as a prescreening procedure to identify individuals with Lynch syndrome, the most common hereditary colorectal cancer syndrome. The microsatellite instability test is known to have imperfect sensitivity and specificity. Meanwhile, the reference test, mutation analysis, is also imperfect. We evaluate this test via a random effects meta-analysis of 17 studies. Study-specific random effects account for between-study heterogeneity in mutation prevalence, test sensitivities and specificities under a nonlinear mixed effects model and a Bayesian hierarchical model. Using model selection techniques, we explore a range of random effects models to identify a best-fitting, model. We also evaluate sensitivity to the conditional independence assumption between the microsatellite instability test and the Mutation analysis by allowing for correlation between them. Finally. we use simulations to illustrate the importance of including appropriate random effects and the impact of overfitting. underfitting and misfitting on model performance. Our approach can be used to estimate the accuracy of two imperfect diagnostic tests from a meta-analysis of multiple studies or a multicenter study when the prevalence of disease, test sensitivities and/or specificities may be heterogeneous among studies or centers.
AbstractLarge-scale sequencing of cancer genomes has uncovered thousands of DNA alterations, but the functional relevance of the majority of these mutations to tumorigenesis is unknown. We have developed a computational method, called Cancer-specific High-throughput Annotation of Somatic Mutations (CHASM), to identify and prioritize those missense mutations most likely to generate functional changes that enhance tumor cell proliferation. The method has high sensitivity and specificity when discriminating between known driver missense mutations and randomly generated missense mutations (area under receiver operating characteristic curve, >0.91; area under Precision-Recall curve, >0.79). CHASM substantially outperformed previously described missense mutation function prediction methods at discriminating known oncogenic mutations in P53 and the tyrosine kinase epidermal growth factor receptor. We applied the method to 607 missense mutations found in a recent glioblastoma multiforme sequencing study. Based on a model that assumed the glioblastoma multiforme mutations are a mixture of drivers and passengers, we estimate that 8% of these mutations are drivers, causally contributing to tumorigenesis. [Cancer Res 2009;69(16):OF6660–8]
We describe a method for evaluating an ensemble of predictive models given a sample of observations comprising the model predictions and the outcome event measured with error. Our formulation allows us to simultaneously estimate measurement error parameters, true outcome—the “gold standard”—and a relative weighting of the predictive scores. We describe conditions necessary to estimate the gold standard and to calibrate these estimates and detail how our approach is related to, but distinct from, standard model combination techniques. We apply our approach to data from a study to evaluate a collection of BRCA1/BRCA2 gene mutation prediction scores. In this example, genotype is measured with error by one or more genetic assays. We estimate true genotype for each individual in the data set, operating characteristics of the commonly used genotyping procedures, and a relative weighting of the scores. Finally, we compare the scores against the gold standard genotype and find that Mendelian scores are, on average, the more refined and better calibrated of those considered and that the comparison is sensitive to measurement error in the gold standard.
Mendelian models can predict who carries an inherited deleterious mutation of known disease genes based on family history. For example, the BRCAPRO model is commonly used to identify families who carry mutations of BRCA1 and BRCA2, based on familial breast and ovarian cancers. These models incorporate the age of diagnosis of diseases in relatives and current age or age of death. We develop a rigorous foundation for handling multiple diseases with censoring. We prove that any disease unrelated to mutations can be excluded from the model, unless it is sufficiently common and dependent on a mutation-related disease time. Furthermore, if a family member has a disease with higher probability density among mutation carriers, but the model does not account for it, then the carrier probability is deflated. However, even if a family only has diseases the model accounts for, if the model excludes a mutation-related disease, then the carrier probability will be inflated. In light of these results, we extend BRCAPRO to account for surviving all non-breast/ovary cancers as a sin le outcome. The extension also enables BRCAPRO to extract more useful information from male relatives. Using 1500 families from the Cancer Genetics Network, accounting for surviving other cancers improves BRCAPRO's concordance index from 0.758 to 0.762 (p=0.046), improves its positive predictive value from 35 to 39 per cent (P<10(-6)) without impacting its negative predictive value, and improves its overall calibration, although calibration slightly worsens for those with carrier probability <10 per cent. Copyright (c) 2008 John Wiley & Sons, Ltd.
Recent studies have demonstrated that histopathologic features of inherited breast cancers due to BRCA1 gene mutations often differ from those of sporadic breast cancer and from breast cancers caused by germline BRCA2 mutations. Invasive breast carcinomas in individuals with germline BRCA1 gene mutations tend to be of higher grade, either basal-type or basal-like, estrogen receptor (ER) negative, progesterone receptor (PR) negative, HER2-Neu (ERBB-2) negative, cytokeratin 5/6 (CK5/6) positive, CK14 positive. Incorporating this information into our Mendelian risk prediction model, BRCAPRO may allow for improved estimation of BRCA1 and BRCA2 carrier risk.
To rigorously determine whether a gene or a set of genes have alterations that are involved in carcinogenesis requires a comparison of the prevalence of identified changes to a control mutation frequency present in tumor DNA. To facilitate this task, we develop a testing approach and the associated R library, called TRAB, that evaluates whether the frequency of somatic mutation in a given gene is higher than that observed in a control group of genes. Specifically, we test the null hypothesis that the frequency belongs to a control population of frequencies, against the alternative hypothesis that the frequency is higher. Mutation frequencies in the control group are themselves allowed to be variable. TRAB computes the a posteriori probability and the Bayes factor for the hypothesis using a hierarchical Bayesian approach.
We previously showed that lifetime cumulative lead dose, measured as lead concentration in the tibia bone by X-ray fluorescence, was associated with persistent and progressive declines in cognitive function and with decreases in MRI-based brain volumes in former lead workers. Moreover, larger region-specific brain volumes were associated with better cognitive function. These findings motivated us to explore a novel application of path analysis to evaluate effect mediation, specifically, whether the association of lead dose with cognitive function is mediated through brain volumes, on a voxel-wise basis. Voxel-wise path analysis, at face value, represents the natural evolution of voxel-based morphometry methods to answer questions of mediation. However, application of these methods to the former lead worker data demonstrated potential limitations in this approach. In particular, there was a tendency for results to be strongly biased towards the null hypothesis (lack of mediation). Moreover, a complimentary analysis using anatomically-derived regions of interest (ROI) volumes yielded opposing results, suggesting evidence of mediation. Specifically, in the ROI-based approach, there was evidence that the association of tibia lead with function in three cognitive domains (e.g., visuo-construction, executive functioning, eye-hand coordination) was mediated through the volumes of total brain, frontal gray matter, and/or possibly cingulate. A simulation study was conducted to investigate whether the voxel-wise results arose from an absence of localized mediation, or more subtle defects in the methodology. The simulation results showed the same null bias evidenced as seen in the lead workers data. Both the lead worker data results and the simulation study suggest that a null-bias in voxel-wise path analysis limits its inferential utility for producing confirmatory results.
Oxidative stress-mediated destruction of normal parenchymal cells during hepatic inflammatory responses contributes to the pathogenesis of immune-mediated hepatitis and is implicated in the progression of acute inflammatory liver injury to chronic inflammatory liver disease. The transcription factor NF-E2-related factor 2 (Nrf2) regulates the expression of a battery of antioxidative enzymes and Nrf2 signaling can be activated by small-molecule drugs that disrupt Keap1-mediated repression of Nrf2 signaling. Therefore, genetic and pharmacologic approaches were used to activate Nrf2 signaling to assess protection against inflammatory liver injury. Profound increases in indicators of cell death were observed in both Nrf2 wild-type (Nrf2-WT) mice and Nrf2-disrupted (Nrf2-KO) mice 24 h following intravenous injection of concanavalin A (12.5 mg/kg, ConA), a model for T cell-mediated acute inflammatory liver injury. However, hepatocyte-specific conditional Keap1 null (Alb-Cre:Keap1(flox/-), cKeap1-KO) mice with constitutively enhanced expression of Nrf2-regulated antioxidative genes as well as Nrf2-WT mice but not Nrf2-KO mice pretreated with three daily doses of a triterpenoid that potently activates Nrf2 (30 mu mol/kg, cyano-3,12-dioxooleana-1,9(11)-dien-28-oyl-imidazolide [CDDO-Im]) were highly resistant to ConA-mediated inflammatory liver injury. CDDO-Im pretreatment of both Nrf2-WT and Nrf2-KO mice resulted in equivalent suppression of serum proinflammatory soluble proteins suggesting that the hepatoprotection afforded by CDDO-Im pretreatment of Nrf2-WT mice but not Nrf2-KO mice was not due to suppression of systemic proinflammatory signaling, but instead was due to activation of Nrf2 signaling in the liver. Enhanced hepatic expression of Nrf2-regulated antioxidative genes inhibited inflammation-mediated oxidative stress, thereby preventing hepatocyte necrosis. Attenuation of hepatocyte death in cKeap1-KO mice and CDDO-Im pretreated Nrf2-WT mice resulted in decreased late-phase proinflammatory gene expression in the liver thereby diminishing the sustained influx of inflammatory cells initially stimulated by the ConA challenge. Taken together, these results clearly illustrate that targeted cytoprotection of hepatocytes through Nrf2 signaling during inflammation prevents the amplification of inflammatory responses in the liver.
same regimens. 8,9 The way to minimize risks from treatment is to treat at the earliest stages of disease.We agree that it is difficult to summarize adequately the findings of a study from a media press release or news article.This is not a novel concept in reporting the findings of a study, and we hope no clinician gathers evidence for practice from reading only the headline of a layman's press article.We believe that it is the responsibility of clinicians to advise patients on all clinical studies, because even randomized studies do not often match individual patient needs and scenarios.The investigators hope that this study encourages such discussions between the patient and physicians.In summary, as more American women live longer and as the means of treating their comorbid conditions continue to improve, how to screen and manage elderly women for breast cancer will become an even more pressing problem.It is clear that more data are needed in these areas.Ideally, such data would be obtained through randomized controlled trials and would address not only mortality but local control and quality of life.Until such questions are answered, controversy will abound, and clinicians and patients will need to consider all facets of the debate in making appropriate individualized decisions.