The determination of a protein's functional residues is key for understanding its function at the molecular level and devising ways of modifying it to our benefit. A plethora of computational methods exist for predicting these sites, and most of them make use of the information contained in a multiple sequence alignment of homologous proteins. Nowadays, these approaches are mature enough, both in terms of reliability and ease of use, to become part of the toolbox of any biologist. This chapter describes a simple protocol for generating a multiple sequence alignment for a protein of interest and detecting different types of functional positions on it, all using freely available web and desktop tools with graphical user interfaces.
Maximum lifespan varies widely among mammals, but the genomic basis of this variation remains incompletely understood. We used body-mass-corrected longevity and a phylogenetically informed phenotype-shift framework to compare 126 primate species, selecting seven long-lived and six short-lived species for convergent amino-acid substitution analysis. Stringent pooling across species identified 1,068 recurrent phenotype-associated positions in 934 genes. These genes showed significant agreement with mammal-wide longevity candidates, with 226 genes shared compared with 120 expected by chance (1.88-fold enrichment; P = 1.16 x 10exp-22). Functional analysis detected 17 significant terms, with the strongest coherent signal involving male infertility and decreased male fertility. Our results show that focused primate sampling recovers a non-random component of the broader mammalian longevity signal and point to male reproductive biology as a promising direction for investigating the evolution of lifespan.
Abstract Complex diseases are influenced by both genetic and environmental factors. Immune cells are key mediating interactions with the environment, but the impact of genetic variation on the immune system and how it influences complex diseases is not fully understood. Moreover, most genome-wide analyses (GWAS) variants associated with complex diseases are non-coding and difficult to interpret. Here, we investigated the association of non-coding variants with immune cell enhancers. As part of BLUEPRINT and the International Human Epigenome (IHEC) consortia, we generated and analysed a comprehensive set of epigenomes for human primary immune cells, including 107 epigenomes derived from 749 ChIP-Seq experiments across 24 cell types. We identified multicell enhancer activity patterns across the genome and examined their links with non-coding variants from 518 GWAS traits. This analysis revealed 117 significant associations, including novel links between cardiovascular disease variants and macrophage-specific enhancers that regulate genes involved in lipid metabolism and immunity, such as the gene encoding for the nuclear receptor LXR-alpha ( NR1H3) and many of its known target genes. Together, these data will help to better understand the influence of genetic variability in immune function and related diseases.
Ancient tooth enamel, and to some extent dentin and bone, contain characteristic peptides that persist for long periods of time. In particular, peptides from the enamel proteome (enamelome) have been used to reconstruct the phylogenetic relationships of fossil taxa. However, the enamelome is based on only about 10 genes, whose protein products undergo fragmentation in vivo and post mortem. This raises the question as to whether the enamelome alone provides enough information for reliable phylogenetic inference. We address these considerations on a selection of enamel-associated proteins that has been computationally predicted from genomic data from 232 primate species. We created multiple sequence alignments for each protein and estimated the evolutionary rate for each site. We examined which sites overlap with the parts of the protein sequences that are typically isolated from fossils. Based on this, we simulated ancient data with different degrees of sequence fragmentation, followed by phylogenetic analysis. We compared these trees to a reference species tree. Up to a degree of fragmentation that is similar to that of fossil samples from 1 to 2 million years ago, the phylogenetic placements of most nodes at family level are consistent with the reference species tree. We tested phylogenetic analysis on combinations of different enamel proteins and found that the composition of the proteome can influence deep splits in the phylogeny. With our methods, we provide guidance for researchers on how to evaluate the potential of paleoproteomics for phylogenetic studies before sampling valuable ancient specimens.
Environmental DNA (eDNA) is a powerful, non-invasive tool for biodiversity monitoring, but traditional approaches face challenges including limited sensitivity, taxonomic restrictions, lengthy processing times, and reliance on specialized lab equipment and computational resources. To address these limitations, we present an integrated eDNA workflow from extraction to real-time analysis. Our approach combines the Tagsteady protocol for multiplexed eDNA metabarcoding with Oxford Nanopore library preparation and sequencing, enabling simultaneous detection of animals (Metazoa) and plants (Viridiplantae) using universal markers (COI and ITS2). By integrating metabarcoding and barcoding in a single protocol, this workflow can assess species presence in environmental samples and barcode individual specimens to expand reference databases, providing a streamlined solution for broad, field-based biodiversity assessment.
Topologically associated domains (TADs) are interaction subnetworks of chromosomal regions in 3D genomes. TAD boundaries frequently coincide with genome breaks while boundary deletion is under negative selection, suggesting that TADs may facilitate genome rearrangements and evolution. We show that genes co-localize by evolutionary age in humans and mice, resulting in TADs having different proportions of younger and older genes. We observe a major transition in the age co-localization patterns between the genes born during vertebrate whole-genome duplications (WGDs) or before and those born afterward. We also find that genes recently duplicated in primates and rodents are more frequently essential when they are located in old-enriched TADs and interact with genes that last duplicated during the WGD. Therefore, the evolutionary relevance of recent genes may increase when located in TADs with established regulatory networks. Our data suggest that TADs could play a role in organizing ancestral functions and evolutionary novelty.
When somatic cells acquire complex karyotypes, they often are removed by the immune system. Mutant somatic cells that evade immune surveillance can lead to cancer. Neurons with complex karyotypes arise during neurotypical brain development, but neurons are almost never the origin of brain cancers. Instead, somatic mutations in neurons can bring about neurodevelopmental disorders, and contribute to the polygenic landscape of neuropsychiatric and neurodegenerative disease. A subset of human neurons harbors idiosyncratic copy number variants (CNVs, "CNV neurons"), but previous analyses of CNV neurons are limited by relatively small sample sizes. Here, we develop an allele-based validation approach, SCOVAL, to corroborate or reject read-depth based CNV calls in single human neurons. We apply this approach to 2,125 frontal cortical neurons from a neurotypical human brain. SCOVAL identifies 226 CNV neurons, which include a subclass of 65 CNV neurons with highly aberrant karyotypes containing whole or substantial losses on multiple chromosomes. Moreover, we find that CNV location appears to be nonrandom. Recurrent regions of neuronal genome rearrangement contain fewer, but longer, genes.
The rich diversity of morphology and behavior displayed across primate species provides an informative context in which to study the impact of genomic diversity on fundamental biological processes. Analysis of that diversity provides insight into long-standing questions in evolutionary and conservation biology and is urgent given severe threats these species are facing. Here, we present high-coverage whole-genome data from 233 primate species representing 86% of genera and all 16 families. This dataset was used, together with fossil calibration, to create a nuclear DNA phylogeny and to reassess evolutionary divergence times among primate clades. We found within-species genetic diversity across families and geographic regions to be associated with climate and sociality, but not with extinction risk. Furthermore, mutation rates differ across species, potentially influenced by effective population sizes. Lastly, we identified extensive recurrence of missense mutations previously thought to be human specific. This study will open a wide range of research avenues for future primate genomic research.
Archaic admixture has had a substantial impact on human evolution with multiple events across different clades, including from extinct hominins such as Neanderthals and Denisovans into modern humans. In great apes, archaic admixture has been identified in chimpanzees and bonobos but the possibility of such events has not been explored in other species. Here, we address this question using high-coverage whole-genome sequences from all four extant gorilla subspecies, including six newly sequenced eastern gorillas from previously unsampled geographic regions. Using approximate Bayesian computation with neural networks to model the demographic history of gorillas, we find a signature of admixture from an archaic ‘ghost’ lineage into the common ancestor of eastern gorillas but not western gorillas. We infer that up to 3% of the genome of these individuals is introgressed from an archaic lineage that diverged more than 3 million years ago from the common ancestor of all extant gorillas. This introgression event took place before the split of mountain and eastern lowland gorillas, probably more than 40 thousand years ago and may have influenced perception of bitter taste in eastern gorillas. When comparing the introgression landscapes of gorillas, humans and bonobos, we find a consistent depletion of introgressed fragments on the X chromosome across these species. However, depletion in protein-coding content is not detectable in eastern gorillas, possibly as a consequence of stronger genetic drift in this species.
Noncoding DNA is central to our understanding of human gene regulation and complex diseases 1 , 2 , and measuring the evolutionary sequence constraint can establish the functional relevance of putative regulatory elements in the human genome 3 – 9 . Identifying the genomic elements that have become constrained specifically in primates has been hampered by the faster evolution of noncoding DNA compared to protein-coding DNA 10 , the relatively short timescales separating primate species 11 , and the previously limited availability of whole-genome sequences 12 . Here we construct a whole-genome alignment of 239 species, representing nearly half of all extant species in the primate order. Using this resource, we identified human regulatory elements that are under selective constraint across primates and other mammals at a 5% false discovery rate. We detected 111,318 DNase I hypersensitivity sites and 267,410 transcription factor binding sites that are constrained specifically in primates but not across other placental mammals and validate their cis -regulatory effects on gene expression. These regulatory elements are enriched for human genetic variants that affect gene expression and complex traits and diseases. Our results highlight the important role of recent evolution in regulatory sequence elements differentiating primates, including humans, from other placental mammals.
Primate genomics holds the key to understanding fundamental aspects of human evolution and disease. However, genetic diversity and functional genomics data sets are currently available for only a few of the more than 500 extant primate species. Concerted efforts are under way to characterize primate genomes, genetic polymorphism and divergence, and functional landscapes across the primate phylogeny. The resulting data sets will enable the connection of genotypes to phenotypes and provide new insight into aspects of the genetics of primate traits, including human diseases. In this Review, we describe the existing genome assemblies as well as genetic variation and functional genomic data sets. We highlight some of the challenges with sample acquisition. Finally, we explore how technological advances in single-cell functional genomics and induced pluripotent stem cell-derived organoids will facilitate our understanding of the molecular foundations of primate biology.
Recent advances in long-read sequencing technologies have allowed the generation and curation of more complete genome assemblies, enabling the analysis of traditionally neglected chromosomes, such as the human Y chromosome (chrY). Native DNA was sequenced on a MinION Oxford Nanopore Technologies sequencing device to generate genome assemblies for seven major chrY human haplogroups. We analyzed and compared the chrY enrichment of sequencing data obtained using two different selective sequencing approaches: adaptive sampling and flow cytometry chromosome sorting. We show that adaptive sampling can produce data to create assemblies comparable to chromosome sorting while being a less expensive and time-consuming technique. We also assessed haplogroup-specific structural variants, which would be otherwise difficult to study using short-read sequencing data only. Finally, we took advantage of this technology to detect and profile epigenetic modifications among the considered haplogroups. Altogether, we provide a framework to study complex genomic regions with a simple, fast, and affordable methodology that could be applied to larger population genomics datasets.
MOTIVATION:Coincidence of Convergent Amino Acid Substitutions (CAAS) with phenotypic convergences allow pinpointing genes and even individual mutations that are likely to be associated with trait variation within their phylogenetic context. Such findings can provide useful insights into the genetic architecture of complex phenotypes.RESULTS:Here we introduce CAAStools, a set of bioinformatics tools to identify and validate CAAS in orthologous protein alignments for predefined groups of species representing the phenotypic values targeted by the user.AVAILABILITY AND IMPLEMENTATION:CAAStools source code is available at http://github.com/linudz/caastools, along with documentation and examples.
Personalized genome sequencing has revealed millions of genetic differences between individuals, but our understanding of their clinical relevance remains largely incomplete. To systematically decipher the effects of human genetic variants, we obtained whole-genome sequencing data for 809 individuals from 233 primate species and identified 4.3 million common protein-altering variants with orthologs in humans. We show that these variants can be inferred to have nondeleterious effects in humans based on their presence at high allele frequencies in other primate populations. We use this resource to classify 6% of all possible human protein-altering variants as likely benign and impute the pathogenicity of the remaining 94% of variants with deep learning, achieving state-of-the-art accuracy for diagnosing pathogenic variants in patients with genetic diseases.
Human populations have been shaped by catastrophes that may have left long-lasting signatures in their ge-nomes. One notable example is the second plague pandemic that entered Europe in ca. 1,347 CE and repeat-edly returned for over 300 years, with typical village and town mortality estimated at 10%-40%.1 It is assumed that this high mortality affected the gene pools of these populations. First, local population crashes reduced genetic diversity. Second, a change in frequency is expected for sequence variants that may have affected survival or susceptibility to the etiologic agent (Yersinia pestis).2 Third, mass mortality might alter the local gene pools through its impact on subsequent migration patterns. We explored these factors using the Nor-wegian city of Trondheim as a model, by sequencing 54 genomes spanning three time periods: (1) prior to the plague striking Trondheim in 1,349 CE, (2) the 17th-19th century, and (3) the present. We find that the pandemic period shaped the gene pool by reducing long distance immigration, in particular from the British Isles, and inducing a bottleneck that reduced genetic diversity. Although we also observe an excess of large FST values at multiple loci in the genome, these are shaped by reference biases introduced by mapping our relatively low genome coverage degraded DNA to the reference genome. This implies that attempts to detect selection using ancient DNA (aDNA) datasets that vary by read length and depth of sequencing coverage may be particularly challenging until methods have been developed to account for the impact of differential refer-ence bias on test statistics.
The role of somatic mutations in complex diseases, including neurodevelopmental and neurodegenerative disorders, is becoming increasingly clear. However, to date, no study has shown their relation to Parkinson disease's phenotype. To explore the relevance of embryonic somatic mutations in sporadic Parkinson disease, we performed whole-exome sequencing in blood and four brain regions of ten patients. We identified 59 candidate somatic single nucleotide variants (sSNVs) through sensitive calling and a careful filtering strategy (COSMOS). We validated 27 of them with amplicon-based ultra-deep sequencing, with a 70% validation rate for the highest-confidence variants. The identified sSNVs are in genes with synaptic functions that are co-expressed with genes previously associated with Parkinson disease. Most of the sSNVs were only called in blood but were also found in the brain tissues with ultra-deep amplicon sequencing, demonstrating the strength of multi-tissue sampling designs.
Transcriptomic diversity greatly contributes to the fundamentals of disease, lineage-specific biology, and environmental adaptation. However, much of the actual isoform repertoire contributing to shaping primate evolution remains unknown. Here, we combined deep long- and short-read sequencing complemented with mass spectrometry proteomics in a panel of lymphoblastoid cell lines (LCLs) from human, three other great apes, and rhesus macaque, producing the largest full-length isoform catalog in primates to date. Around half of the captured isoforms are not annotated in their reference genomes, significantly expanding the gene models in primates. Furthermore, our comparative analyses unveil hundreds of transcriptomic innovations and isoform usage changes related to immune function and immunological disorders. The confluence of these evolutionary innovations with signals of positive selection and their limited impact in the proteome points to changes in alternative splicing in genes involved in immune response as an important target of recent regulatory divergence in primates.
Extreme phenotypic diversity, a history of artificial selection, and socioeconomic value make domestic dog breeds a compelling subject for genomic research. Copy number variation (CNV) is known to account for a significant part of inter-individual genomic diversity in other systems. However, a comprehensive genome-wide study of structural variation as it relates to breed-specific phenotypes is lacking. We have generated whole genome CNV maps for more than 300 canids. Our data set extends the canine structural variation landscape to more than 100 dog breeds, including novel variants that cannot be assessed using microarray technologies. We have taken advantage of this data set to perform the first CNV-based genome-wide association study (GWAS) in canids. We identify 96 loci that display copy number differences across breeds, which are statistically associated with a previously compiled set of breed-specific morphometrics and disease susceptibilities. Among these, we highlight the discovery of a long-range interaction involving a CNV near MED13L and TBX3, which could influence breed standard height. Integration of the CNVs with chromatin interactions, long noncoding RNA expression, and single nucleotide variation highlights a subset of specific loci and genes with potential functional relevance and the prospect to explain trait variation between dog breeds.
Changes in the epigenetic regulation of gene expression have a central role in evolution. Here, we extensively profiled a panel of human, chimpanzee, gorilla, orangutan, and macaque lymphoblastoid cell lines (LCLs), using ChIP-seq for five histone marks, ATAC-seq and RNA-seq, further complemented with whole genome sequencing (WGS) and whole genome bisulfite sequencing (WGBS). We annotated regulatory elements (RE) and integrated chromatin contact maps to define gene regulatory architectures, creating the largest catalog of RE in primates to date. We report that epigenetic conservation and its correlation with sequence conservation in primates depends on the activity state of the regulatory element. Our gene regulatory architectures reveal the coordination of different types of components and highlight the role of promoters and intragenic enhancers (gE) in the regulation of gene expression. We observe that most regulatory changes occur in weakly active gE. Remarkably, novel human-specific gE with weak activities are enriched in human-specific nucleotide changes. These elements appear in genes with signals of positive selection and human acceleration, tissue-specific expression, and particular functional enrichments, suggesting that the regulatory evolution of these genes may have contributed to human adaptation.
The authors wish to make the following correction to this paper [...].