Measures of nucleotide sequence conservation across species are useful for identifying functional genomic loci, but can fail when regulatory function is maintained, often in a cell type-specific manner, even when sequence is not. We introduce CTACIT, the Cell Type-Aware Conservation Inference Toolkit, to identify trait-associated regulatory variants. CTACIT integrates sequence conservation scores with cell type-specific open chromatin data collected from a few mammalian species to impute function for hundreds more. Applying CTACIT to neuropsychiatric trait loci identifies higher heritability enrichment and more fine-mapped variants than nucleotide conservation and human chromatin data alone. Our in vivo reporter assays validate predictions for enhancers with risk variants near the DRD2 schizophrenia risk locus. By integrating genome conservation and multi-species open chromatin data, CTACIT prioritizes variants within regions of conserved regulatory function for in vivo characterization and addresses a major challenge in translating disease associations to mechanistic understanding.
Vocal production learning ("vocal learning") is a convergently evolved trait in vertebrates. To identify brain genomic elements associated with mammalian vocal learning, we integrated genomic, anatomical, and neurophysiological data from the Egyptian fruit bat (Rousettus aegyptiacus) with analyses of the genomes of 215 placental mammals. First, we identified a set of proteins evolving more slowly in vocal learners. Then, we discovered a vocal motor cortical region in the Egyptian fruit bat, an emergent vocal learner, and leveraged that knowledge to identify active cis-regulatory elements in the motor cortex of vocal learners. Machine learning methods applied to motor cortex open chromatin revealed 50 enhancers robustly associated with vocal learning whose activity tended to be lower in vocal learners. Our research implicates convergent losses of motor cortex regulatory elements in mammalian vocal learning evolution.
Epstein-Barr virus (EBV) is an aetiologic risk factor for the development of multiple sclerosis (MS). However, the role of EBV-infected B cells in the immunopathology of MS is not well understood. Here we characterized spontaneous lymphoblastoid cell lines (SLCLs) isolated from MS patients and healthy controls (HC) ex vivo to study EBV and host gene expression in the context of an individual's endogenous EBV. SLCLs derived from MS patient B cells during active disease had higher EBV lytic gene expression than SLCLs from MS patients with stable disease or HCs. Host gene expression analysis revealed activation of pathways associated with hypercytokinemia and interferon signalling in MS SLCLs and upregulation of forkhead box protein 1 (FOXP1), which contributes to EBV lytic gene expression. We demonstrate that antiviral approaches targeting EBV replication decreased cytokine production and autologous CD4+ T cell responses in this ex vivo model. These data suggest that dysregulation of intrinsic B cell control of EBV gene expression drives a pro-inflammatory, pathogenic B cell phenotype that can be attenuated by suppressing EBV lytic gene expression. Characterizing EBV and host gene expression profiles of spontaneous lymphoblastoid cell lines isolated from multiple sclerosis (MS) patients with acute or stable disease, as well as healthy donors, suggests antivirals as a potential road to treat MS.
ABSTRACT Shotgun metagenomic sequencing can determine both taxonomic and functional content of microbiomes. However, current functional classification methods for metagenomic reads require substantial computational resources and yield ambiguous classifications, limiting downstream quantitative analyses. Existing k -mer based methods to classify microbial sequences into species-level groups have immensely improved taxonomic classification, but this concept has not been extended to functional classification. Here we introduce k Mermaid, for classifying metagenomic reads into functional clusters of proteins. Using protein k -mers, k Mermaid allows for highly accurate and ultrafast functional classification, with a fixed memory usage, and can easily be employed on a typical computer.
Protein-coding differences between species often fail to explain phenotypic diversity, suggesting the involvement of genomic elements that regulate gene expression such as enhancers. Identifying associations between enhancers and phenotypes is challenging because enhancer activity can be tissue-dependent and functionally conserved despite low sequence conservation. We developed the Tissue-Aware Conservation Inference Toolkit (TACIT) to associate candidate enhancers with species' phenotypes using predictions from machine learning models trained on specific tissues. Applying TACIT to associate motor cortex and parvalbumin-positive interneuron enhancers with neurological phenotypes revealed dozens of enhancer-phenotype associations, including brain size-associated enhancers that interact with genes implicated in microcephaly or macrocephaly. TACIT provides a foundation for identifying enhancers associated with the evolution of any convergently evolved phenotype in any large group of species with aligned genomes.
Local microbiome shifts are implicated in the development and progression of gastrointestinal cancers, and in particular, esophageal carcinoma (ESCA), which is among the most aggressive malignancies. Short-read RNA sequencing (RNAseq) is currently the leading technology to study gene expression changes in cancer. However, using RNAseq to study microbial gene expression is challenging. Here, we establish a new tool to efficiently detect viral and bacterial expression in human tissues through RNAseq. This approach employs a neural network to predict reads of likely microbial origin, which are targeted for assembly into longer contigs, improving identification of microbial species and genes. This approach is applied to perform a systematic comparison of bacterial expression in ESCA and healthy esophagi. We uncover bacterial genera that are over or underabundant in ESCA vs healthy esophagi both before and after correction for possible covariates, including patient metadata. However, we find that bacterial taxonomies are not significantly associated with clinical outcomes. Strikingly, in contrast, dozens of microbial proteins were significantly associated with poor patient outcomes and in particular, proteins that perform mitochondrial functions and iron-sulfur coordination. We further demonstrate associations between these microbial proteins and dysregulated host pathways in ESCA patients. Overall, these results suggest possible influences of bacteria on the development of ESCA and uncover new prognostic biomarkers based on microbial genes. In addition, this study provides a framework for the analysis of other human malignancies whose development may be driven by pathogens.
Evolutionary constraint and acceleration are powerful, cell-type agnostic measures of functional importance. Previous studies in mammals were limited by species number and reliance on human-referenced alignments. We explore the evolution of placental mammals, including humans, through reference-free whole-genome alignment of 240 species and protein-coding alignments for 428 species. We estimate 10.7% of the human genome is evolutionarily constrained. We resolve constraint to single nucleotides, pinpointing functional positions, and refine and expand by over seven-fold the catalog of ultraconserved elements. Overall, 48.5% of constrained bases are as yet unannotated, suggesting yet-to-be-discovered functional importance. Using species-level phenotypes and an updated phylogeny, we associate coding and regulatory variation with olfaction and hibernation. Focusing on biodiversity conservation, we identify genomic metrics that predict species at risk of extinction.
About 15% of human cancer cases are attributed to viral infections. To date, virus expression in tumor tissues has been mostly studied by aligning tumor RNA sequencing reads to databases of known viruses. To allow identification of divergent viruses and rapid characterization of the tumor virome, we develop viRNAtrap, an alignment-free pipeline to identify viral reads and assemble viral contigs. We utilize viRNAtrap, which is based on a deep learning model trained to discriminate viral RNAseq reads, to explore viral expression in cancers and apply it to 14 cancer types from The Cancer Genome Atlas (TCGA). Using viRNAtrap, we uncover expression of unexpected and divergent viruses that have not previously been implicated in cancer and disclose human endogenous viruses whose expression is associated with poor overall survival. The viRNAtrap pipeline provides a way forward to study viral infections associated with different clinical conditions.
Epidemiological studies have demonstrated that Epstein-Barr virus (EBV) is a known etiologic risk factor, and perhaps prerequisite, for the development of MS. EBV establishes life-long latent infection in a subpopulation of memory B cells. Although the role of memory B cells in the pathobiology of MS is well established, studies characterizing EBV-associated mechanisms of B cell inflammation and disease pathogenesis in EBV (+) B cells from MS patients are limited. Accordingly, we analyzed spontaneous lymphoblastoid cell lines (SLCLs) from multiple sclerosis patients and healthy controls to study host-virus interactions in B cells, in the context of an individual's endogenous EBV. We identify differences in EBV gene expression and regulation of both viral and cellular genes in SLCLs. Our data suggest that EBV latency is dysregulated in MS SLCLs with increased lytic gene expression observed in MS patient B cells, especially those generated from samples obtained during "active" disease. Moreover, we show increased inflammatory gene expression and cytokine production in MS patient SLCLs and demonstrate that tenofovir alafenamide, an antiviral that targets EBV replication, decreases EBV viral loads, EBV lytic gene expression, and EBV-mediated inflammation in both SLCLs and in a mixed lymphocyte assay. Collectively, these data suggest that dysregulation of EBV latency in MS drives a pro-inflammatory, pathogenic phenotype in memory B cells and that this response can be attenuated by suppressing EBV lytic activation. This study provides further support for the development of antiviral agents that target EBV-infection for use in MS.
Zoonomia is the largest comparative genomics resource for mammals produced to date. By aligning genomes for 240 species, we identify bases that, when mutated, are likely to affect fitness and alter disease risk. At least 332 million bases (~10.7%) in the human genome are unusually conserved across species (evolutionarily constrained) relative to neutrally evolving repeats, and 4552 ultraconserved elements are nearly perfectly conserved. Of 101 million significantly constrained single bases, 80% are outside protein-coding exons and half have no functional annotations in the Encyclopedia of DNA Elements (ENCODE) resource. Changes in genes and regulatory elements are associated with exceptional mammalian traits, such as hibernation, that could inform therapeutic development. Earth’s vast and imperiled biodiversity offers distinctive power for identifying genetic variants that affect genome function and organismal phenotypes.
Background Evolutionary conservation is an invaluable tool for inferring functional significance in the genome, including regions that are crucial across many species and those that have undergone convergent evolution. Computational methods to test for sequence conservation are dominated by algorithms that examine the ability of one or more nucleotides to align across large evolutionary distances. While these nucleotide alignment-based approaches have proven powerful for protein-coding genes and some non-coding elements, they fail to capture conservation of many enhancers, distal regulatory elements that control spatial and temporal patterns of gene expression. The function of enhancers is governed by a complex, often tissue- and cell type-specific code that links combinations of transcription factor binding sites and other regulation-related sequence patterns to regulatory activity. Thus, function of orthologous enhancer regions can be conserved across large evolutionary distances, even when nucleotide turnover is high. Results We present a new machine learning-based approach for evaluating enhancer conservation that leverages the combinatorial sequence code of enhancer activity rather than relying on the alignment of individual nucleotides. We first train a convolutional neural network model that can predict tissue-specific open chromatin, a proxy for enhancer activity, across mammals. Next, we apply that model to distinguish instances where the genome sequence would predict conserved function versus a loss of regulatory activity in that tissue. We present criteria for systematically evaluating model performance for this task and use them to demonstrate that our models accurately predict tissue-specific conservation and divergence in open chromatin between primate and rodent species, vastly out-performing leading nucleotide alignment-based approaches. We then apply our models to predict open chromatin at orthologs of brain and liver open chromatin regions across hundreds of mammals and find that brain enhancers associated with neuron activity have a stronger tendency than the general population to have predicted lineage-specific open chromatin. Conclusion The framework presented here provides a mechanism to annotate tissue-specific regulatory function across hundreds of genomes and to study enhancer evolution using predicted regulatory differences rather than nucleotide-level conservation measurements.
About 15% of human cancer cases are attributed to viral infections. To date, virus expression in tumor tissues has been mostly studied by aligning tumor RNA sequencing reads to databases of known viruses. To allow identification of divergent viruses and rapid characterization of the tumor virome, we developed viRNAtrap, an alignment-free pipeline to identify viral reads and assemble viral contigs. We apply viRNAtrap, which is based on a deep learning model trained to discriminate viral RNAseq reads, to 14 cancer types from The Cancer Genome Atlas (TCGA). We find that expression of exogenous cancer viruses is associated with better overall survival. In contrast, expression of human endogenous viruses is associated with worse overall survival. Using viRNAtrap, we uncover expression of unexpected and divergent viruses that have not previously been implicated in cancer. The viRNAtrap pipeline provides a way forward to study viral infections associated with different clinical conditions.
The origin of eukaryotes was marked by the emergence of several novel subcellular systems. One such is the calcium (Ca 2+ )-stores system of the endoplasmic reticulum, which profoundly influences diverse aspects of cellular function including signal transduction, motility, division, and biomineralization. We use comparative genomics and sensitive sequence and structure analyses to investigate the evolution of this system. Our findings reconstruct the core form of the Ca 2+ - stores system in the last eukaryotic common ancestor as having at least 15 proteins that constituted a basic system for facilitating both Ca 2+ flux across endomembranes and Ca 2+ -dependent signaling. We present evidence that the key EF-hand Ca 2+ -binding components had their origins in a likely bacterial symbiont other than the mitochondrial progenitor, whereas the protein phosphatase subunit of the ancestral calcineurin complex was likely inherited from the asgardarchaeal progenitor of the stem eukaryote. This further points to the potential origin of the eukaryotes in a Ca 2+ -rich biomineralized environment such as stromatolites. We further show that throughout eukaryotic evolution there were several acquisitions from bacteria of key components of the Ca 2+ -stores system, even though no prokaryotic lineage possesses a comparable system. Further, using quantitative measures derived from comparative genomics we show that there were several rounds of lineage-specific gene expansions, innovations of novel gene families, and gene losses correlated with biological innovation such as the biomineralized molluscan shells, coccolithophores, and animal motility. The burst of innovation of new genes in animals included the wolframin protein associated with Wolfram syndrome in humans. We show for the first time that it contains previously unidentified Sel1, EF-hand, and OB-fold domains, which might have key roles in its biochemistry.
Cyclic di- and linear oligo-nucleotide signals activate defenses against invasive nucleic acids in animal immunity; however, their evolutionary antecedents are poorly understood. Using comparative genomics, sequence and structure analysis, we uncovered a vast network of systems defined by conserved prokaryotic gene-neighborhoods, which encode enzymes generating such nucleotides or alternatively processing them to yield potential signaling molecules. The nucleotide-generating enzymes include several clades of the DNA-polymerase β-like superfamily (including Vibrio cholerae DncV), a minimal version of the CRISPR polymerase and DisA-like cyclic-di-AMP synthetases. Nucleotide-binding/processing domains include TIR domains and members of a superfamily prototyped by Smf/DprA proteins and base (cytokinin)-releasing LOG enzymes. They are combined in conserved gene-neighborhoods with genes for a plethora of protein superfamilies, which we predict to function as nucleotide-sensors and effectors targeting nucleic acids, proteins or membranes (pore-forming agents). These systems are sometimes combined with other biological conflict-systems such as restriction-modification and CRISPR/Cas. Interestingly, several are coupled in mutually exclusive neighborhoods with either a prokaryotic ubiquitin-system or a HORMA domain-PCH2-like AAA+ ATPase dyad. The latter are potential precursors of equivalent proteins in eukaryotic chromosome dynamics. Further, components from these nucleotide-centric systems have been utilized in several other systems including a novel diversity-generating system with a reverse transcriptase. We also found the Smf/DprA/LOG domain from these systems to be recruited as a predicted nucleotide-binding domain in eukaryotic TRPM channels. These findings point to evolutionary and mechanistic links, which bring together CRISPR/Cas, animal interferon-induced immunity, and several other systems that combine nucleic-acid-sensing and nucleotide-dependent signaling.