OBJECTIVE To identify genetic risk factors for incident cardiovascular disease (CVD) among people with type 2 diabetes (T2D). RESEARCH DESIGN AND METHODS We conducted a multi-ancestry time-to-event genome-wide association study for incident CVD among people with T2D. We also tested 204 known coronary artery disease (CAD) variants for association with incident CVD. RESULTS Among 49,230 participants with T2D, 8,956 had incident CVD events (event rate 18.2%). We identified three novel genetic loci for incident CVD: rs147138607 (near CACNA1E/ZNF648, HR 1.23, P=3.6×10-9), rs11444867 (near HS3ST1, HR 1.89, P=9.9×10-9), and rs335407 (near TFB1M/NOX3, HR 1.25, P=1.5×10-8). Among 204 known CAD loci, 5 were associated with incident CVD in T2D (multiple comparison-adjusted P < 0.00024, 0.05/204). A standardized polygenic score of these 204 variants was associated with incident CVD with HR 1.14 (P=1.0×10-16). CONCLUSIONS The data point to novel and known genomic regions associated with incident CVD among individuals with T2D.
Background Clinical use of genotype data requires high positive predictive value (PPV) and thorough understanding of the genotyping platform characteristics. BeadChip arrays, such as the Global Screening Array (GSA), potentially offer a high-throughput, low-cost clinical screen for known variants. We hypothesize that quality assessment and comparison to whole-genome sequence and benchmark data establish the analytical validity of GSA genotyping. Methods To test this hypothesis, we selected 263 samples from Coriell, generated GSA genotypes in triplicate, generated whole genome sequence (rWGS) genotypes, assessed the quality of each set of genotypes, and compared each set of genotypes to each other and to the 1000 Genomes Phase 3 (1KG) genotypes, a performance benchmark. For 59 genes (MAP59), we also performed theoretical and empirical evaluation of variants deemed medically actionable predispositions. Results Quality analyses detected sample contamination and increased assay failure along the chip margins. Comparison to benchmark data demonstrated that > 82% of the GSA assays had a PPV of 1. GSA assays targeting transitions, genomic regions of high complexity, and common variants performed better than those targeting transversions, regions of low complexity, and rare variants. Comparison of GSA data to rWGS and 1KG data showed > 99% performance across all measured parameters. Consistent with predictions from prior studies, the GSA detection of variation within the MAP59 genes was 3/261. Conclusion We establish the analytical validity of GSA assays using quality analytics and comparison to benchmark and rWGS data. GSA assays meet the standards of a clinical screen although assays interrogating rare variants, transversions, and variants within low-complexity regions require careful evaluation.
Traditionally, the use of genomic information for personalized medical decisions relies on prior discovery and validation of genotype-phenotype associations. This approach constrains care for patients presenting with undescribed problems. The National Institutes of Health (NIH) Undiagnosed Diseases Program (UDP) hypothesized that defining disease as maladaptation to an ecological niche allows delineation of a logical framework to diagnose and evaluate such patients. Herein, we present the philosophical bases, methodologies, and processes implemented by the NIH UDP. The NIH UDP incorporated use of the Human Phenotype Ontology, developed a genomic alignment strategy cognizant of parental genotypes, pursued agnostic biochemical analyses, implemented functional validation, and established virtual villages of global experts. This systematic approach provided a foundation for the diagnostic or non-diagnostic answers provided to patients and serves as a paradigm for scalable translational research.
In this work, we present the Genome Modeling System (GMS), an analysis information management system capable of executing automated genome analysis pipelines at a massive scale. The GMS framework provides detailed tracking of samples and data coupled with reliable and repeatable analysis pipelines. The GMS also serves as a platform for bioinformatics development, allowing a large team to collaborate on data analysis, or an individual researcher to leverage the work of others effectively within its data management system. Rather than separating ad-hoc analysis from rigorous, reproducible pipelines, the GMS promotes systematic integration between the two. As a demonstration of the GMS, we performed an integrated analysis of whole genome, exome and transcriptome sequencing data from a breast cancer cell line (HCC1395) and matched lymphoblastoid line (HCC1395BL). These data are available for users to test the software, complete tutorials and develop novel GMS pipeline configurations. The GMS is available at https://github.com/genome/gms.
We have designed an integrated RNA-seq pipeline that determines both sequence variants and the differential expression of the transcripts from the same alignment in cancer samples. Comparison with other widely used protocols for RNA-seq data analysis showed that our method was more sensitive for determining subtle phenotypes in studies with large sample size, which frequently occur in compound screening approaches for cancer therapy. This whole transcriptome sequencing technology provides a cost-effective approach to gain understanding into the RNA relative biological events by deciphering RNA sequences. Benefiting from the fast developing technology, the cost of deep RNA-seq is no longer a barrier for current applications. Furthermore, an RNA-seq dataset provides more insight into disease processes than DNA-seq, since single nucleotide variant (SNV) detection in addition to gene and exon level expression and new isoform detection can be garnered from the same information. Based on the recent public data for mixed cancer cell lines and pooled normal human samples in MicroArray Quality Control (MAQC) phase III, also referred to as Sequencing Quality Control (SEQC), we evaluated our method with previous published results by means of different alignment strategies, sequencing depths and replicates numbers. Receiver operating characteristic (ROC) analysis showed our gene expression results agreed with the ones from the SEQC teams. In most cases, our solution has similar performance with other widely used strategies, while our STAR 2 pass based strategy has better true positive and true negative rates in the experiments with large sample size and subtle difference between groups. At the same time, the Genome Analysis Toolkit (GATK) 3 in our pipeline gained a reasonable balance between accuracy and performance. In summary, our new RNA-seq analysis pipeline has good performance and balance between precision and speed. Furthermore, this work also implicitly indicates that our gene expression method generates more accurate results in a large sample size experiment among similar groups. Citation Format: Gang Feng, Jamie Osman, Sean Leighton, Lynn Carmichael, Yuchen Bai, John Begemann, Richard Mazzarella. Validation of a procedure for simultaneous variant calling and differential expression in cancer study. [abstract]. In: Proceedings of the 106th Annual Meeting of the American Association for Cancer Research; 2015 Apr 18-22; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2015;75(15 Suppl):Abstract nr 4869. doi:10.1158/1538-7445.AM2015-4869
BACKGROUND:The full complement of DNA mutations that are responsible for the pathogenesis of acute myeloid leukemia (AML) is not yet known.METHODS:We used massively parallel DNA sequencing to obtain a very high level of coverage (approximately 98%) of a primary, cytogenetically normal, de novo genome for AML with minimal maturation (AML-M1) and a matched normal skin genome.RESULTS:We identified 12 acquired (somatic) mutations within the coding sequences of genes and 52 somatic point mutations in conserved or regulatory portions of the genome. All mutations appeared to be heterozygous and present in nearly all cells in the tumor sample. Four of the 64 mutations occurred in at least 1 additional AML sample in 188 samples that were tested. Mutations in NRAS and NPM1 had been identified previously in patients with AML, but two other mutations had not been identified. One of these mutations, in the IDH1 gene, was present in 15 of 187 additional AML genomes tested and was strongly associated with normal cytogenetic status; it was present in 13 of 80 cytogenetically normal samples (16%). The other was a nongenic mutation in a genomic region with regulatory potential and conservation in higher mammals; we detected it in one additional AML tumor. The AML genome that we sequenced contains approximately 750 point mutations, of which only a small fraction are likely to be relevant to pathogenesis.CONCLUSIONS:By comparing the sequences of tumor and skin genomes of a patient with AML-M1, we have identified recurring mutations that may be relevant for pathogenesis.
Background: Investigators in the biological sciences continue to exploit laboratory automation methods and have dramatically increased the rates at which they can generate data. In many environments, the methods themselves also evolve in a rapid and fluid manner. These observations point to the importance of robust information management systems in the modern laboratory. Designing and implementing such systems is non-trivial and it appears that in many cases a database project ultimately proves unserviceable.Results: We describe a general modeling framework for laboratory data and its implementation as an information management system. The model utilizes several abstraction techniques, focusing especially on the concepts of inheritance and meta-data. Traditional approaches commingle event-oriented data with regular entity data in ad hoc ways. Instead, we define distinct regular entity and event schemas, but fully integrate these via a standardized interface. The design allows straightforward definition of a "processing pipeline" as a sequence of events, obviating the need for separate workflow management systems. A layer above the event-oriented schema integrates events into a workflow by defining "processing directives", which act as automated project managers of items in the system. Directives can be added or modified in an almost trivial fashion, i.e., without the need for schema modification or re-certification of applications. Association between regular entities and events is managed via simple "many-to-many" relationships. We describe the programming interface, as well as techniques for handling input/output, process control, and state transitions.Conclusion: The implementation described here has served as the Washington University Genome Sequencing Center's primary information system for several years. It handles all transactions underlying a throughput rate of about 9 million sequencing reactions of various kinds per month and has handily weathered a number of major pipeline reconfigurations. The basic data model can be readily adapted to other high-volume processing environments.
Background The ever-expanding population of gene expression profiles (EPs) from specified cells and tissues under a variety of experimental conditions is an important but difficult resource for investigators to utilize effectively. Software tools have been recently developed to use the distribution of gene ontology (GO) terms associated with the genes in an EP to identify specific biological functions or processes that are over- or under-represented in that EP relative to other EPs. Additionally, it is possible to use the distribution of GO terms inherent to each EP to relate that EP as a whole to other EPs. Because GO term annotation is organized in a tree-like cascade of variable granularity, this approach allows the user to relate ( e.g ., by hierarchical clustering) EPs of varying length and from different platforms ( e.g ., GeneChip, SAGE, EST library). Results Here we present GOurmet, a software package that calculates the distribution of GO terms represented by the genes in an individual expression profile (EP), clusters multiple EPs based on these integrated GO term distributions, and provides users several tools to visualize and compare EPs. GOurmet is particularly useful in meta-analysis to examine EPs of specified cell types ( e.g ., tissue-specific stem cells) that are obtained through different experimental procedures. GOurmet also introduces a new tool, the Targetoid plot, which allows users to dynamically render the multi-dimensional relationships among individual elements in any clustering analysis. The Targetoid plotting tool allows users to select any element as the center of the plot, and the program will then represent all other elements in the cluster as a function of similarity to the selected central element. Conclusion GOurmet is a user-friendly, GUI-based software package that greatly facilitates analysis of results generated by multiple EPs. The clustering analysis features a dynamic targetoid plot that is generalizable for use with any clustering application.
The human gut is colonized with a vast community of indigenous microorganisms that help shape our biology. Here, we present the complete genome sequence of the Gram-negative anaerobe Bacteroides thetaiotaomicron , a dominant member of our normal distal intestinal microbiota. Its 4779-member proteome includes an elaborate apparatus for acquiring and hydrolyzing otherwise indigestible dietary polysaccharides and an associated environment-sensing system consisting of a large repertoire of extracytoplasmic function sigma factors and one- and two-component signal transduction systems. These and other expanded paralogous groups shed light on the molecular mechanisms underlying symbiotic host-bacterial relationships in our intestine.