3’ untranslated region (3’UTR) variants are strongly associated with human traits and diseases, yet few have been causally identified. We developed the Massively Parallel Reporter Assay for 3’UTRs (MPRAu) to sensitively assay 12,173 3’UTR variants. We applied MPRAu to six human cell lines, focusing on genetic variants associated with genome-wide association studies (GWAS) and human evolutionary adaptation. MPRAu expands our understanding of 3’UTR function, suggesting that low-complexity sequences predominately explain 3’UTR regulatory activity. We adapt MPRAu to uncover diverse molecular mechanisms at base-pair resolution, including an AU-rich element of LEPR linked to potential metabolic evolutionary adaptations in East Asians. We nominate hundreds of 3’UTR causal variants with genetically fine-mapped phenotype associations. Using endogenous allelic replacements, we characterize one variant that disrupts a miRNA site regulating the viral defense gene TRIM14, and one that alters PILRB abundance, nominating a causal variant underlying transcriptional changes in age-related macular degeneration.
Exploring the Phenotypic Consequences of Tissue Specific Gene Expression Variation Inferred GWAS Summary Statistics. Scalable, integrative methods to understand mechanisms that link genetic variants with phenotypes are needed. Here we derive a mathematical expression to compute PrediXcan (a gene mapping approach) results using summary data (S-PrediXcan) and show its accuracy and general robustness to misspeci fi ed reference sets. We apply this framework to 44 GTEx tissues and 100 + phenotypes from GWAS and meta-analysis studies, creating a growing public catalog of associations that seeks to capture the effects of gene expression variation on human phenotypes. Replication in an independent cohort is shown. Most of the associations are tissue speci fi c, suggesting context speci fi city of the trait etiology. Colocalized signi fi cant associations in unexpected tissues underscore the need for an agnostic scanning of multiple contexts to improve our ability to detect causal regulatory mechanisms. Monogenic disease genes are enriched among signi fi cant associations for related traits, suggesting that smaller alterations of these genes may cause a spectrum of milder phenotypes. in concordance with COLOC. A few genes fall in the “ colocalized ” region, in disagreement with COLOC classi fi cation. Unlike COLOC results, HEIDI does not partition the genes into distinct clusters and an arbitrary cutoff p -value has to be chosen. e Shows genes with large HEIDI p -value (no evidence of heterogeneity) which fall in large part in the “ colocalized ” region. However a substantial number fall in “ independent signal ” region, disagreeing with COLOC ’ s classi fi cation signals) from S-TWAS signi fi cant vs. S-PrediXcan signi fi cant results. For most phenotypes, TWAS has lower proportion of colocalized signals compared to S-PrediXcan, as estimated via COLOC. Phenotype abbreviations are as follows: FNBD Femoral Neck Bone Density, LSBD Lumbar Spine Bone Density, BMI Body Mass Index, HEIGHT Height, LDL Low-Density Lipoprotein Cholesterol, HDL High-Density Lipoprotein Cholesterol, TRYG Tryglicerides, CROHN Crohn ’ s Disease, INFBOWEL In fl ammatory Bowel ’ s Disease, ULCERC Ulcerative Colitis, HBA1C Hemogoblin Levels, HOMA-IR HOMA Insulin Response, SCZ Schizophrenia, RA Rheumatoid Arthritis, COLLEGE College Completion, EDUCYEARS Education Years and genes with low some Bone Low-Density Cholesterol, HDL High-Density Lipoprotein Tryglicerides, In Bowel HBA1C Hemogoblin Levels, HOMA-IR HOMA
Rare genetic variants are abundant across the human genome, and identifying their function and phenotypic impact is a major challenge. Measuring aberrant gene expression has aided in identifying functional, large-effect rare variants (RVs). Here, we expanded detection of genetically driven transcriptome abnormalities by analyzing gene expression, allele-specific expression, and alternative splicing from multitissue RNA-sequencing data, and demonstrate that each signal informs unique classes of RVs. We developed Watershed, a probabilistic model that integrates multiple genomic and transcriptomic signals to predict variant function, validated these predictions in additional cohorts and through experimental assays, and used them to assess RVs in the UK Biobank, the Million Veterans Program, and the Jackson Heart Study. Our results link thousands of RVs to diverse molecular effects and provide evidence to associate RVs affecting the transcriptome with human traits.
It is estimated that 350 million individuals worldwide suffer from rare diseases, which are predominantly caused by mutation in a single gene1. The current molecular diagnostic rate is estimated at 50%, with whole-exome sequencing (WES) among the most successful approaches2–5. For patients in whom WES is uninformative, RNA sequencing (RNA-seq) has shown diagnostic utility in specific tissues and diseases6–8. This includes muscle biopsies from patients with undiagnosed rare muscle disorders6,9, and cultured fibroblasts from patients with mitochondrial disorders7. However, for many individuals, biopsies are not performed for clinical care, and tissues are difficult to access. We sought to assess the utility of RNA-seq from blood as a diagnostic tool for rare diseases of different pathophysiologies. We generated whole-blood RNA-seq from 94 individuals with undiagnosed rare diseases spanning 16 diverse disease categories. We developed a robust approach to compare data from these individuals with large sets of RNA-seq data for controls (n = 1,594 unrelated controls and n = 49 family members) and demonstrated the impacts of expression, splicing, gene and variant filtering strategies on disease gene identification. Across our cohort, we observed that RNA-seq yields a 7.5% diagnostic rate, and an additional 16.7% with improved candidate gene resolution. A diagnostic tool based on blood RNA-seq is shown to identify causal genes and variants linked to clinical phenotypes in individuals with rare diseases for which whole-exome genetic sequencing was uninformative.
Characterization of the molecular function of the human genome and its variation across individuals is essential for identifying the cellular mechanisms that underlie human genetic traits and diseases. The Genotype-Tissue Expression (GTEx) project aims to characterize variation in gene expression levels across individuals and diverse tissues of the human body, many of which are not easily accessible. Here we describe genetic effects on gene expression levels across 44 human tissues. We find that local genetic variation affects gene expression levels for the majority of genes, and we further identify inter-chromosomal genetic effects for 93 genes and 112 loci. On the basis of the identified genetic effects, we characterize patterns of tissue specificity, compare local and distal effects, and evaluate the functional properties of the genetic effects. We also 34Department of Operations Research and Financial Engineering, Princeton University, Princeton, New Jersey 84540, USA. 35Biomedical Informatics Program, Stanford University, Stanford, California 94305, USA. 36Department of Convergence Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, 05505, Korea. 37Department of Human Genetics, University of California, Los Angeles, California 90095, USA. 38Department of Biology, Stanford University, Stanford, California 94305, USA. 39Department of Computer Science, University of California, Los Angeles, California 90095, USA. 40Computational Biology and Bioinformatics Graduate Program, Duke University, Durham, North Carolina 27708, USA. 41Department of Statistics and Operations Research, University of North Carolina, Chapel Hill, North Carolina 27599, USA. 42Department of Biostatistics, The University of Texas MD Anderson Cancer Center, Houston, Texas 77030, USA. 43Computer Science and Artificial Intelligence Laboratory, MIT, Cambridge, Massachusetts 02139, USA. 44Stanley Center for Psychiatric Research, Broad Institute of MIT and Harvard, Cambridge, Massachusetts 02142, USA. 45Center for Biomarker Research and Personalized Medicine, Virginia Commonwealth University, Richmond, Virginia 23298, USA. 46Department of Psychiatry and Biobehavioral Sciences, University of California, Los Angeles, California 90095, USA. 47Bioinformatics Research Center and Department of Biological Sciences, North Carolina State University, Raleigh, North Carolina 27695, USA. 48Department of Biomedical Data Science, Stanford University, Stanford, California 94305, USA. 49Wellcome Trust Centre for Human Genetics, Nuffield Department of Medicine, University of Oxford, Oxford OX3 7BN, UK. 50Oxford Centre for Diabetes, Endocrinology and Metabolism, Radcliffe Department of Medicine, University of Oxford, Oxford OX3 9DU, UK. 51Oxford NIHR Biomedical Research Centre, Oxford University Hospitals Trust, Oxford OX3 7LE, UK. 52Department of Pathology & Immunology, Washington University School of Medicine, St. Louis, Missouri 63110, USA. 53Department of Genetics, Washington University School of Medicine, St. Louis, Missouri 63110, USA. 54Department of Biostatistics, Mailman School of Public Health, Columbia University, New York, New York 10032, USA. 55Department of Statistics, Stanford University, Stanford, California 94305, USA. 56Section of Genetic Medicine, Department of Medicine, Institute for Genomics and Systems Biology, Center for Data Intensive Science, University of Chicago, Chicago, Illinois 60637, USA. 57Department of Biostatistics, University of Michigan, Ann Arbor, Michigan 48109, USA. 58Bioinformatics Research Center, Departments of Statistics and Biological Sciences, North Carolina State University, Raleigh, North Carolina 27695, USA. 59Department of Computer Science and Center for Statistics and Machine Learning, Princeton University, Princeton, New Jersey 08540, USA. *These authors contributed equally to this work. §These authors jointly supervised this work. Online Content Methods, along with any additional Extended Data display items and Source Data, are available in the online version of the paper; references unique to these sections appear only in the online paper. Supplementary Information is available in the online version of the paper. Author Contributions All authors reviewed and revised the manuscript. Detailed author contributions are available in the Supplementary Information. The authors declare competing financial interests: details are available in the online version of the paper. Readers are welcome to comment on the online version of the paper. Page 2 Nature. Author manuscript; available in PMC 2018 January 22. A uhor M anscript
RNA sequencing (RNA-seq) is a complementary approach for Mendelian disease diagnosis for patients in whom exome-sequencing is not informative. For both rare neuromuscular and mitochondrial disorders, its application has improved diagnostic rates. However, the generalizability of this approach to diverse Mendelian diseases has yet to be evaluated. We sequenced whole blood RNA from 56 cases with undiagnosed rare diseases spanning 11 diverse disease categories to evaluate the general application of RNA-seq to Mendelian disease diagnosis. We developed a robust approach to compare rare disease cases to existing large sets of RNA-seq controls (N=1,594 external and N=31 family-based controls) and demonstrated the substantial impacts of gene and variant filtering strategies on disease gene identification when combined with RNA-seq. Across our cohort, we observed that RNA-seq yields a 8.5% diagnostic rate. These diagnoses included diseases where blood would not intuitively reflect evidence of disease. We identified RARS2 as an under-expression outlier containing compound heterozygous pathogenic variants for an individual exhibiting profound global developmental delay, seizures, microcephaly, hypotonia, and progressive scoliosis. We also identified a new splicing junction in KCTD7 for an individual with global developmental delay, loss of milestones, tremors and seizures. Our study provides a broad evaluation of blood RNA-seq for the diagnosis of rare disease.
Living organisms constantly maintain their structural and biochemical integrity by the critical means of response, healing, and regeneration. Inanimate objects, on the other hand, are axiomatically considered incapable of responding to damage and healing it, leading to the profound negative environmental impact of their continuous manufacturing and trashing. Objects with such biological properties would be a significant step towards sustainable technology. In this work we present a feasible strategy for driving regeneration in fabric by means of integration with a bacterial biofilm to obtain a symbiotic-like hybrid - the fabric provides structural framework to the biofilm and supports its growth, whereas the biofilm responds to mechanical tear by synthesizing a silk protein engineered to self-assemble upon secretion from the cells. We propose the term crossbiosis to describe this and other hybrid systems combining organism and object. Our strategy could be implemented in other systems and drive sensing of integrity and response by regeneration in other materials as well.
Genetic studies of complex traits have mainly identified associations with noncoding variants. To further determine the contribution of regulatory variation, we combined whole-genome and transcriptome data for 624 individuals from Sardinia to identify common and rare variants that influence gene expression and splicing. We identified 21,183 expression quantitative trait loci (eQTLs) and 6,768 splicing quantitative trait loci (sQTLs), including 619 new QTLs. We identified high-frequency QTLs and found evidence of selection near genes involved in malarial resistance and increased multiple sclerosis risk, reflecting the epidemiological history of Sardinia. Using family relationships, we identified 809 segregating expression outliers (median z score of 2.97), averaging 13.3 genes per individual. Outlier genes were enriched for proximal rare variants, providing a new approach to study large-effect regulatory variants and their relevance to traits. Our results provide insight into the effects of regulatory variants and their relationship to population history and individual genetic risk.
Rare genetic variants are abundant in humans and are expected to contribute to individual disease risk. While genetic association studies have successfully identified common genetic variants associated with susceptibility, these studies are not practical for identifying rare variants. Efforts to distinguish pathogenic variants from benign rare variants have leveraged the genetic code to identify deleterious protein-coding alleles, but no analogous code exists for non-coding variants. Therefore, ascertaining which rare variants have phenotypic effects remains a major challenge. Rare non-coding variants have been associated with extreme gene expression in studies using single tissues, but their effects across tissues are unknown. Here we identify gene expression outliers, or individuals showing extreme expression levels for a particular gene, across 44 human tissues by using combined analyses of whole genomes and multi-tissue RNA-sequencing data from the Genotype-Tissue Expression (GTEx) project v6p release. We find that 58% of underexpression and 28% of overexpression outliers have nearby conserved rare variants compared to 8% of non-outliers. Additionally, we developed RIVER (RNA-informed variant effect on regulation), a Bayesian statistical model that incorporates expression data to predict a regulatory effect for rare variants with higher accuracy than models using genomic annotations alone. Overall, we demonstrate that rare variants contribute to large gene expression changes across tissues and provide an integrative method for interpretation of rare variants in individual genomes.
17 Despite the abundance of rare genetic variants—variants carried by less than one percent of the population—in human genomes, the impact of these variants on specific tissues has been largely uncharacterized. Population-level test statistics, while effective in understanding the impact of common variants—variants carried by at least five percent of the population, have had limited success in characterizing the effect of rare variants mainly due to limited statistical power. In addition, the effect of each rare variant can vary greatly between specific tissues. This heterogeneity coupled with limited sample sizes and a lack of known disease-causing rare variants makes predicting tissue-specific cellular consequences of rare variants a difficult task. To make these predictions, we propose a new method called SPEER (SPecific tissuE variant Effect predictoR): a hierarchical Bayesian model that uses transfer learning, allowing separate predictions in each tissue while flexibly sharing signal across tissues to improve power. Our probabilistic model capitalizes on a growing body of rich epigenetic annotations to inform the consequences of a variant in specific tissues. These annotations are integrated with tissue-specific RNA expression levels and common variants. We show our method improves prediction accuracy in simulations and in genomic data from the Genotype-Tissue Expression (GTEx) project. 18
Structural variants (SVs) are an important source of human genetic diversity, but their contribution to traits, disease and gene regulation remains unclear. We mapped cis expression quantitative trait loci (eQTLs) in 13 tissues via joint analysis of SVs, single-nucleotide variants (SNVs) and short insertion/deletion (indel) variants from deep whole-genome sequencing (WGS). We estimated that SVs are causal at 3.5-6.8% of eQTLs-a substantially higher fraction than prior estimates-and that expression-altering SVs have larger effect sizes than do SNVs and indels. We identified 789 putative causal SVs predicted to directly alter gene expression: most (88.3%) were noncoding variants enriched at enhancers and other regulatory elements, and 52 were linked to genome-wide association study loci. We observed a notable abundance of rare high-impact SVs associated with aberrant expression of nearby genes. These results suggest that comprehensive WGS-based SV analyses will increase the power of common-and rare-variant association studies.
Motivation: Interpreting genetic variation in noncoding regions of the genome is an important challenge for personal genome analysis. One mechanism by which noncoding single nucleotide variants (SNVs) influence downstream phenotypes is through the regulation of gene expression. Methods to predict whether or not individual SNVs are likely to regulate gene expression would aid interpretation of variants of unknown significance identified in whole‐genome sequencing studies. Results: We developed FIRE (Functional Inference of Regulators of Expression), a tool to score both noncoding and coding SNVs based on their potential to regulate the expression levels of nearby genes. FIRE consists of 23 random forests trained to recognize SNVs in cis‐expression quantitative trait loci (cis‐eQTLs) using a set of 92 genomic annotations as predictive features. FIRE scores discriminate cis‐eQTL SNVs from non‐eQTL SNVs in the training set with a cross‐validated area under the receiver operating characteristic curve (AUC) of 0.807, and discriminate cis‐eQTL SNVs shared across six populations of different ancestry from non‐eQTL SNVs with an AUC of 0.939. FIRE scores are also predictive of cis‐eQTL SNVs across a variety of tissue types. Availability and implementation: FIRE scores for genome‐wide SNVs in hg19/GRCh37 are available for download at https://sites.google.com/site/fireregulatoryvariation/. Contact: nilah@stanford.edu Supplementary information: Supplementary data are available at Bioinformatics online.
Identifying interactions between genetics and the environment (GxE) remains challenging. We have developed EAGLE, a hierarchical Bayesian model for identifying GxE interactions based on associations between environmental variables and allele-specific expression. Combining whole-blood RNA-seq with extensive environmental annotations collected from 922 human individuals, we identified 35 GxE interactions, compared with only four using standard GxE interaction testing. EAGLE provides new opportunities for researchers to identify GxE interactions using functional genomic data.
A growing knowledge base of genetic and environmental information has greatly enabled the study of disease risk factors. However, the computational complexity and statistical burden of testing all variants by all environments has required novel study designs and hypothesis-driven approaches. We discuss how incorporating biological knowledge from model organisms, functional genomics, and integrative approaches can empower the discovery of novel gene-environment interactions and discuss specific methodological considerations with each approach. We consider specific examples where the application of these approaches has uncovered effects of gene-environment interactions relevant to drug response and immunity, and we highlight how such improvements enable a greater understanding of the pathogenesis of disease and the realization of precision medicine.
Although mice are the most widely used mammalian model organism, genetic studies have suffered from limited mapping resolution due to extensive linkage disequilibrium ( LD) that is characteristic of crosses among inbred strains. Carworth Farms White ( CFW) mice are a commercially available outbred mouse population that exhibit rapid LD decay in comparison to other available mouse populations. We performed a genome-wide association study ( GWAS) of behavioral, physiological and gene expression phenotypes using 1,200 male CFW mice. We used genotyping by sequencing ( GBS) to obtain genotypes at 92,734 SNPs. We also measured gene expression using RNA sequencing in three brain regions. Our study identified numerous behavioral, physiological and expression quantitative trait loci ( QTLs). We integrated the behavioral QTL and eQTL results to implicate specific genes, including Azi2 in sensitivity to methamphetamine and Zmynd11 in anxiety-like behavior.The combination of CFW mice, GBS and RNA sequencing constitutes a powerful approach to GWAS in mice.
1 Department of Pathology, Stanford University, Stanford, CA. 2 Department of Computer Science, Johns Hopkins 5 University, Baltimore, MD. 3 Biomedical Informatics Program, Stanford University, Stanford, CA. 4 Department of 6 Genetics, Stanford University, Stanford, CA. 5 McDonnell Genome Institute, Washington University School of 7 Medicine, St. Louis, MO. 6 Department of Biomedical Engineering, Johns Hopkins University, Baltimore, MD. 7 8 Analytic and Translational Genetics Unit, Massachusetts General Hospital, Boston, MA. 8 Program in Medical and 9 Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA. 9 Stanley Center for Psychiatric 10 Research, Broad Institute of MIT and Harvard, Cambridge, MA. 10 Department of Medicine, Washington University 11 School of Medicine, St. Louis, MO. 11 Department of Genetics, Washington University School of Medicine, St. 12 Louis, MO. 13
Identifying functional non-coding variants can enhance genome interpretation and inform novel genetic risk factors. We used whole genomes and peripheral white blood cell transcriptomes from 624 Sardinian individuals to identify non-coding variants that contribute to population, family, and individual differences in transcript abundance. We identified 21,183 independent expression quantitative trait loci (eQTLs) and 6,768 independent splicing quantitative trait loci (sQTLs) influencing 73 and 41% of all tested genes. When we compared Sardinian eQTLs to those previously identified in Europe, we identified differentiated eQTLs at genes involved in malarial resistance and multiple sclerosis, reflecting the long-term epidemiological history of the island’s population. Taking advantage of pedigree data for the population sample, we identify segregating patterns of outlier gene expression and allelic imbalance in 61 Sardinian trios. We identified 809 expression outliers (median z-score of 2.97) averaging 13.3 genes with outlier expression per individual. We then connected these outlier expression events to rare non-coding variants. Our results provide new insight into the effects of non-coding variants and their relationship to population history, traits and individual genetic risk.
The X Chromosome, with its unique mode of inheritance, contributes to differences between the sexes at a molecular level, including sex-specific gene expression and sex-specific impact of genetic variation. Improving our understanding of these differences offers to elucidate the molecular mechanisms underlying sex-specific traits and diseases. However, to date, most studies have either ignored the X Chromosome or had insufficient power to test for the sex-specific impact of genetic variation. By analyzing whole blood transcriptomes of 922 individuals, we have conducted the first large-scale, genome-wide analysis of the impact of both sex and genetic variation on patterns of gene expression, including comparison between the X Chromosome and autosomes. We identified a depletion of expression quantitative trait loci (eQTL) on the X Chromosome, especially among genes under high selective constraint. In contrast, we discovered an enrichment of sex-specific regulatory variants on the X Chromosome. To resolve the molecular mechanisms underlying such effects, we generated chromatin accessibility data through ATAC-sequencing to connect sex-specific chromatin accessibility to sex-specific patterns of expression and regulatory variation. As sex-specific regulatory variants discovered in our study can inform sex differences in heritable disease prevalence, we integrated our data with genome-wide association study data for multiple immune traits identifying several traits with significant sex biases in genetic susceptibilities. Together, our study provides genome-wide insight into how genetic variation, the X Chromosome, and sex shape human gene regulation and disease.
Expression quantitative trait locus (eQTL) mapping provides a powerful means to identify functional variants influencing gene expression and disease pathogenesis. We report the identification of cis-eQTLs from 7,051 post-mortem samples representing 44 tissues and 449 individuals as part of the Genotype-Tissue Expression (GTEx) project. We find a cis-eQTL for 88% of all annotated protein-coding genes, with one-third having multiple independent effects. We identify numerous tissue-specific cis-eQTLs, highlighting the unique functional impact of regulatory variation in diverse tissues. By integrating large-scale functional genomics data and state-of-the-art fine-mapping algorithms, we identify multiple features predictive of tissue-specific and shared regulatory effects. We improve estimates of cis-eQTL sharing and effect sizes using allele specific expression across tissues. Finally, we demonstrate the utility of this large compendium of cis-eQTLs for understanding the tissue-specific etiology of complex traits, including coronary artery disease. The GTEx project provides an exceptional resource that has improved our understanding of gene regulation across tissues and the role of regulatory variation in human genetic diseases.
Indigenous peoples globally have relied on customary practices for safeguarding and optimizing harvest of wildlife populations, including seabirds. Increasingly, there have been efforts to engage these indigenous and local knowledge systems to inform responses to the global biodiversity crisis. We considered how customary harvest management as practiced historically by New Zealand Maori might have contributed to sustaining a burrowing seabird population. We used a deterministic age-structured model of a population of grey-faced petrels (Pterodroma gouldi). By harvesting pre-fledging chicks, rather than adult birds, Maori harvesters had less of an impact on population growth rates; annual population growth rate dropped below 1.0 if 2% or more of adult birds were harvested compared with threshold harvests of 25% or more of either eggs or chicks. Restrictions such as prohibition of the digging out of burrows or the use of a snaring pole to catch chicks would lead to an annual harvest of less than our estimated deterministic threshold for a sustainable harvest of 25% of chicks even if all the chicks within arm's reach down the burrows were caught. The cultural practice of rotating harvests or resting populations was also effective; harvesting every 3 years allowed up to 75% of chicks to be taken without causing the theoretical population to decline. The development and implementation of ecologically meaningful but also culturally appropriate approaches to managing wildlife is an important part of efforts to de-colonialize conservation management. We recommend adaptive or multi-evidence-based approaches where scientific information and methods complement indigenous knowledge and practices in the management of wildlife. (c) 2015 The Wildlife Society.