Sex biases in admixture and other demographic processes are recurrent features throughout human evolution. For admixture between Neanderthals and anatomically modern humans (AMHs), sex bias has been proposed as an explanation for the relative lack of Neanderthal ancestry in modern human X chromosomes compared with that in modern human autosomes. By observing a 62% relative excess of AMH ancestry in Neanderthal X chromosomes, we characterized the interbreeding between the two groups as predominantly male Neanderthals with female AMHs. Analytic and numerical modeling presents mate preference as a more parsimonious cause of the sex bias than purely demographic processes with differential patterns of male and female migration.
Modern sequencing instruments bring unprecedented opportunity to study within-host viral evolution in conjunction with viral transmissions between hosts. However, no computational simulators are available to assist the characterization of within-host dynamics. This limits our ability to interpret epidemiological predictions incorporating within-host evolution and to validate computational inference tools. To fill this need we developed Apollo, a GPU-accelerated, out-of-core tool for within-host simulation of viral evolution and infection dynamics across population, tissue, and cellular levels. Apollo is scalable to hundreds of millions of viral genomes and can handle complex demographic and population genetic models. Apollo can replicate real within-host viral evolution; accurately recapturing observed viral sequences from HIV and SARS-CoV-2 cohorts derived from initial population-genetic configurations. For practical applications, using Apollo-simulated viral genomes and transmission networks, we validated and uncovered the limitations of a widely used viral transmission inference tool.
The advent of high-throughput sequencing technologies coupled with breakthroughs in third-generation sequencing have allowed new exploration into within-host viral dynamics. However, a simulation platform to analyze this new world of within host/tissue/cell viral populations does not exist. We present a solution. Apollo is a state-of-the-art within-host viral simulator developed to comprehensively model viral transmission, replication dynamics, natural selection, and host behaviors such as Lost to Follow Up, across population, host, tissue, and cell levels. Leveraging CATE (https://doi.org/10.1111/2041-210X.14168), our proven large-scale GPU CUDA-powered parallel processing architecture, Apollo achieves unprecedented speeds and hardware efficiency. Apollo is built on the standard Wright Fisher (WF) evolutionary model, but, thanks to its scriptable parameter structure, users are able to design simulations that mimic real world dynamics that expand beyond the WF model. Through rigorous testing we have been able to demonstrate Apollo’s accuracy and resource efficiency. We present a complete simulation of an HIV epidemic with within-tissue factors, recombination, and mutation mechanisms that characterizes HIV viral evolution and dynamics both the within hosts and across host levels. The simulations correspond/align with clinical findings and enhances the real-world data by providing further insight into the pedigree of viral variants and within-host quasispecies dynamics. Apollo represents a significant advancement in structured viral evolution and offers a powerful new tool for studying complex viral dynamics aimed to inform individual therapeutics and public health interventions.
Comparisons of Neanderthal genomes to anatomically modern human (AMH) genomes show a history of Neanderthal-to-AMH introgression stemming from interbreeding after the migration of AMHs from Africa to Eurasia. All non-sub-Saharan African AMHs have genomic regions genetically similar to Neanderthals that descend from this introgression. Regions of the genome with Neanderthal similarities have also been identified in sub-Saharan African populations, but their origins have been unclear. To better understand how these regions are distributed across sub-Saharan Africa, the source of their origin, and what their distribution within the genome tells us about early AMH and Neanderthal evolution, we analyzed a dataset of high-coverage, whole-genome sequences from 180 individuals from 12 diverse sub-Saharan African populations. In sub-Saharan African populations with non-sub-Saharan African ancestry, as much as 1% of their genomes can be attributed to Neanderthal sequence introduced by recent migration, and subsequent admixture, of AMH populations originating from the Levant and North Africa. However, most Neanderthal homologous regions in sub-Saharan African populations originate from migration of AMH populations from Africa to Eurasia -250 kya, and subsequent admixture with Neanderthals, resulting in -6% AMH ancestry in Neanderthals. These results indicate that there have been multiple migration events of AMHs out of Africa and that Neanderthal and AMH gene flow has been bi-directional. Observing that genomic regions where AMHs show a depletion of Neanderthal introgression are also regions where Neanderthal genomes show a depletion of AMH introgression points to deleterious interactions between introgressed variants and background genomes in both groups-a hallmark of incipient speciation.
Linkage disequilibrium (LD) is a fundamental concept in genetics; critical for studying genetic associations and molecular evolution. However, LD measurements are only reliable for common genetic variants, leaving low-frequency variants unanalyzed. In this work, we introduce cumulative LD (cLD), a stable statistic that captures the rare-variant LD between genetic regions, which reflects more biological interactions between variants, in addition to lack of recombination. We derived the theoretical variance of cLD using delta methods to demonstrate its higher stability than LD for rare variants. This property is also verified by bootstrapped simulations using real data. In application, we find cLD reveals an increased genetic association between genes in 3D chromatin interactions, a phenomenon recently reported negatively by calculating standard LD between common variants. Additionally, we show that cLD is higher between gene pairs reported in interaction databases, identifies unreported protein-protein interactions, and reveals interacting genes distinguishing case/control samples in association studies.
A classic population genetic prediction is that alleles experiencing directional selection should swiftly traverse allele frequency space, leaving detectable reductions in genetic variation in linked regions. However, despite this expectation, identifying clear footprints of beneficial allele passage has proven to be surprisingly challenging. We addressed the basic premise underlying this expectation by estimating the ages of large numbers of beneficial and deleterious alleles in a human population genomic data set. Deleterious alleles were found to be young, on average, given their allele frequency. However, beneficial alleles were older on average than non-coding, non-regulatory alleles of the same frequency. This finding is not consistent with directional selection and instead indicates some type of balancing selection. Among derived beneficial alleles, those fixed in the population show higher local recombination rates than those still segregating, consistent with a model in which new beneficial alleles experience an initial period of balancing selection due to linkage disequilibrium with deleterious recessive alleles. Alleles that ultimately fix following a period of balancing selection will leave a modest ‘soft’ sweep impact on the local variation, consistent with the overall paucity of species-wide ‘hard’ sweeps in human genomes.Analyses of allele age and evolutionary impact reveal that beneficial alleles in a human population are often older than neutral controls, suggesting a large role for balancing selection in adaptation.
SUMMARYColor variation is a frequent evolutionary substrate for camouflage in small mammals but the underlying genetics and evolutionary forces that drive color variation in natural populations of large mammals are mostly unexplained. The American black bear, Ursus americanus, exhibits a range of colors including the cinnamon morph which has a similar color to the brown bear, U. arctos, and is found at high frequency in the American southwest. Reflectance and chemical melanin measurements showed little distinction between U. arctos and cinnamon U. americanus individuals. We used a genome-wide association for hair color as a quantitative trait in 151 U. americanus individuals and identified a single major locus (P < 10−13). Additional genomic and functional studies identified a missense alteration (R153C) in Tyrosinase-related protein 1 (TYRP1) that impaired protein localization and decreased pigment production. Population genetic analyses and demographic modeling indicated that the R153C variant arose 9.36kya in a southwestern population where it likely provided a selective advantage, spreading both northwards and eastwards by gene flow. A different TYRP1 allele, R114C, contributes to the characteristic brown color of U. arctos, but is not fixed across the range.HIGHLIGHTSThe cinnamon morph of American black bears and brown bears have different missense mutations in TYRP1 that account for their similar colorationTYRP1 variants in American black bears and brown bears are loss-of-function alleles associated with impaired protein localization to melanosomesIn American black bears, the variant causing the cinnamon morph arose 9,360 years ago in the western lineage where it provides an adaptive advantage, and has spread northwards and eastwards by migration
The Carnegie Council for Ethics in International Affairs is an independent, nonpartisan, nonsectarian, tax-exempt organization founded in by Andrew Carnegie.Since its beginnings, the Carnegie Council has asserted its strong belief that ethics, as informed by the world's principal moral and religious traditions, is an inevitable and integral component of all policy decisions, whether in the realm of economics
An abstract is not available for this content so a preview has been provided. As you have access to this content, a full PDF is available via the ‘Save PDF’ action button.
The Carnegie Council for Ethics in International Affairs is an independent, nonpartisan, nonsectarian, tax-exempt organization founded in by Andrew Carnegie.Since its beginnings, the Carnegie Council has asserted its strong belief that ethics, as informed by the world's principal moral and religious traditions, is an inevitable and integral component of all policy decisions, whether in the realm of economics
For many traits, human variation is less a matter of categorical differences than quantitative variation, such as height, where individuals fall along a continuum from short to tall. Most recent studies utilize large population-based samples with whole-genome sequences to study the evolution of these traits and have made significant progress implementing a broad spectrum of techniques. However, relatively few studies of quantitative trait evolution include ethnically diverse populations, which often harbor the highest levels of genetic and phenotypic diversity. Thus, our ability to draw inferences about quantitative trait adaptation has been limited. Here, we review recent studies examining human quantitative trait adaptation, and argue that including ethnically diverse populations, particularly from Africa, will be especially informative for our understanding of how humans adapt to the world around them.
AbstractThe observation that even a tiny sample of genome sequences from a natural population contains a plethora of information about the history of the population has enticed researchers to use these data to fit complex demographic histories and make detailed inference about the changes a population has experienced through time. Unfortunately, the standard assumptions required to make these inferences are often violated by natural populations in such ways as to produce specious results. This paper examines two phenomena of particular concern: when a sample is drawn from a single sub-population of a larger meta-population these models infer a spurious recent population decline, and when a genome contains loci under weak or recessive purifying selection these models infer a spurious recent population expansion.
For many traits, human variation is less a matter of categorical differences than quantitative variation, such as height, where individuals fall along a continuum from short to tall. Most recent studies utilize large population-based samples with whole-genome sequences to study the evolution of these traits and have made significant progress implementing a broad spectrum of techniques. However, relatively few studies of quantitative trait evolution include ethnically diverse populations, which often harbor the highest levels of genetic and phenotypic diversity. Thus, our ability to draw inferences about quantitative trait adaptation has been limited. Here, we review recent studies examining human quantitative trait adaptation, and argue that including ethnically diverse populations, particularly from Africa, will be especially informative for our understanding of how humans adapt to the world around them.
An abstract is not available for this content so a preview has been provided. As you have access to this content, a full PDF is available via the ‘Save PDF’ action button.
Allele age has long been a focus of population genetic research, primarily because it can be an important clue to the fitness effects of an allele. By virtue of their effects on fitness, alleles under directional selection are expected to be younger than neutral alleles of the same frequency. We developed a new coalescent-based estimator of a close proxy for allele age, the time when a copy of an allele first shares common ancestry with other chromosomes in a sample not carrying that allele. The estimator performs well, including for the very rarest of alleles that occur just once in a sample, with a bias that is typically negative. The estimator is mostly insensitive to population demography and to factors that can arise in population genomic pipelines, including the statistical phasing of chromosomes. Applications to 1000 Genomes Data and UK10K genome data confirm predictions that singleton alleles that alter proteins are significantly younger than those that do not, with a greater difference in the larger UK10K dataset, as expected. The 1000 Genomes populations varied markedly in their distributions for singleton allele ages, suggesting that these distributions can be used to inform models of demographic history, including recent events that are only revealed by their impacts on the ages of very rare alleles.
Matrices representing genetic relatedness among individuals (i.e., Genomic Relationship Matrices, GRMs) play a central role in genetic analysis. The eigen-decomposition of GRMs (or its alternative that generates fewer top singular values using genotype matrices) is a necessary step for many analyses including estimation of SNP-heritability, Principal Component Analysis (PCA), and genomic prediction. However, the GRMs and genotype matrices provided by modern biobanks are too large to be stored in active memory. To accommodate the current and future bigger-data, we develop a disk-based tool, Out-of-Core Matrices Analyzer (OCMA), using state-of-the-art computational techniques that can nimbly perform eigen and Singular Value Decomposition (SVD) analyses. By integrating memory mapping (mmap) and the latest matrix factorization libraries, our tool is fast and memory-efficient. To demonstrate the impressive performance of OCMA, we test it on a personal computer. For full eigen-decomposition, it solves an ordinary GRM (N = 10,000) in 55 sec. For SVD, a commonly used faster alternative of full eigen-decomposition in genomic analyses, OCMA solves the top 200 singular values (SVs) in half an hour, top 2,000 SVs in 0.95 hr, and all 5,000 SVs in 1.77 hr based on a very large genotype matrix (N = 1,000,000, M = 5,000) on the same personal computer. OCMA also supports multi-threading when running in a desktop or HPC cluster. Our OCMA tool can thus alleviate the computing bottleneck of classical analyses on large genomic matrices, and make it possible to scale up current and emerging analytical methods to big genomics data using lightweight computing resources.
The human genome contains hundreds of thousands of missense mutations. However, only a handful of these variants are known to be adaptive, which implies that adaptation through protein sequence change is an extremely rare phenomenon in human evolution. Alternatively, existing methods may lack the power to pinpoint adaptive variation. We have developed and applied an Evolutionary Probability Approach (EPA) to discover candidate adaptive polymorphisms (CAPs) through the discordance between allelic evolutionary probabilities and their observed frequencies in human populations. EPA reveals thousands of missense CAPs, which suggest that a large number of previously optimal alleles had experienced a reversal of fortune in the human lineage. We explored non-adaptive mechanisms to explain CAPs, including the effects of demography, mutation rate variability, and negative and positive selective pressures in modern humans. Our analyses suggest that a large proportion of CAP alleles have increased in frequency due to beneficial selection. This conclusion is supported by the facts that a vast majority of adaptive missense variants discovered previously in humans are CAPs, and that hundreds of CAP alleles are protective in genotype-phenotype association data. Our integrated phylogenomic and population genetic EPA approach predicts the existence of thousands of signatures of non-neutral evolution in the human proteome. We expect this collection to be enriched in beneficial variation. EPA approach can be applied to discover candidate adaptive variation in any protein, population, or species for which allele frequency data and reliable multispecies alignments are available.
That population size affects the fate of new mutations arising in genomes, modulating both how frequently they arise and how efficiently natural selection is able to filter them, is well established. It is therefore clear that these distinct roles for population size that characterize different processes should affect the evolution of proteins and need to be carefully defined. Empirical evidence is consistent with a role for demography in influencing protein evolution, supporting the idea that functional constraints alone do not determine the composition of coding sequences. Given that the relationship between population size, mutant fitness and fixation probability has been well characterized, estimating fitness from observed substitutions is well within reach with well-formulated models. Molecular evolution research has, therefore, increasingly begun to leverage concepts from population genetics to quantify the selective effects associated with different classes of mutation. However, in order for this type of analysis to provide meaningful information about the intra- and inter-specific evolution of coding sequences, a clear definition of concepts of population size, what they influence, and how they are best parameterized is essential. Here, we present an overview of the many distinct concepts that “population size” and “effective population size” may refer to, what they represent for studying proteins, and how this knowledge can be harnessed to produce better specified models of protein evolution.