We review three areas of human genetics that have been developed in the past few decades, in which statistical innovation has made a crucial contribution with recent important advances and the potential for further rapid progress. The first topic is the development of mathematical models for the genealogy underlying samples of genome-wide genetic data. Coalescent theory emerged in the 1980s, leaped ahead in the past decade and is now burgeoning into new application areas in population, evolutionary and medical genetics. The second is the development of statistical methods for genome-wide association studies which has made great strides over two decades, including exciting recent developments for association testing based on coalescent theory and improved methods for trait prediction. Finally, we review the statistical ideas that helped resolve the controversies surrounding the introduction of forensic DNA profiling in the early 1990s. Big advances in interpretation of the predominant autosomal DNA profiles have set a benchmark for other areas of forensic science, but the statistical assessment of uniparentally inherited profiles (derived from the mitochondrial DNA or the Y chromosome) remains unsatisfactory.
Inference of evolutionary and demographic parameters from a sample of genome sequences often proceeds by first inferring identical-by-descent (IBD) genome segments. By exploiting efficient data encoding based on the ancestral recombination graph (ARG), we obtain three major advantages over current approaches: (i) no need to impose a length threshold on IBD segments, (ii) IBD can be defined without the hard-to-verify requirement of no recombination, and (iii) computation time can be reduced with little loss of statistical efficiency using only the IBD segments from a set of sequence pairs that scales linearly with sample size. We first demonstrate powerful inferences when true IBD information is available from simulated data. For IBD inferred from real data, we propose an approximate Bayesian computation inference algorithm and use it to show that even poorly-inferred short IBD segments can improve estimation. Our mutation-rate estimator achieves precision similar to a previously-published method despite a 4 000-fold reduction in data used for inference, and we identify significant differences between human populations. Computational cost limits model complexity in our approach, but we are able to incorporate unknown nuisance parameters and model misspecification, still finding improved parameter inference.
Aims Prevention of human hypertension is an important challenge and has been achieved in experimental models. Brief treatment with renin-angiotensin system (RAS) inhibitors permanently reduces the genetic hypertension of the spontaneously hypertensive rat (SHR). The kidney is involved in this fascinating phenomenon, but relevant changes in gene expression are unknown. Methods and results In SHR, we studied the effect of treatment between 10 and 14 weeks of age with the angiotensin receptor blocker, losartan, or the angiotensin-converting enzyme inhibitor, perindopril [with controls for non-specific effects of lowering blood pressure (BP)], on differential RNA expression, DNA methylation, and renin immunolabelling in the kidney at 20 weeks of age. RNA sequencing revealed a six-fold increase in renin gene (Ren) expression during losartan treatment (P < 0.0001). Six weeks after losartan, arterial pressure remained lower (P = 0.006), yet kidney Ren showed reduced expression by 23% after losartan (P = 0.03) and by 43% after perindopril (P = 1.4 x 10(-6)) associated with increased DNA methylation (P = 0.04). Immunolabelling confirmed reduced cortical renin after earlier RAS blockade (P = 0.002). RNA sequencing identified differential expression of mRNAs, miRNAs, and lncRNAs with evidence of networking and co-regulation. These included 13 candidate genes (Grhl1, Ammecr1l, Hs6st1, Nfil3, Fam221a, Lmo4, Adamts1, Cish, Hif3a, Bcl6, Rad54l2, Adap1, Dok4), the miRNA miR-145-3p, and the lncRNA AC115371. Gene ontogeny analyses revealed that these networks were enriched with genes relevant to BP, RAS, and the kidneys. Conclusion Early RAS inhibition in SHR resets genetic pathways and networks resulting in a legacy of reduced Ren expression and BP persisting for a minimum of 6 weeks.
Population genomics has revolutionized our ability to study bacterial evolution by enabling data-driven discovery of the genetic architecture of trait variation. Genome-wide association studies (GWAS) have more recently become accompanied by genome-wide epistasis and co-selection (GWES) analysis, which offers a phenotype-free approach to generating hypotheses about selective processes that simultaneously impact multiple loci across the genome. However, existing GWES methods only consider associations between distant pairs of loci within the genome due to the strong impact of linkage-disequilibrium (LD) over short distances. Based on the general functional organisation of genomes it is nevertheless expected that majority of co-selection and epistasis will act within relatively short genomic proximity, on co-variation occurring within genes and their promoter regions, and within operons. Here, we introduce LDWeaver, which enables an exhaustive GWES across both short- and long-range LD, to disentangle likely neutral co-variation from selection. We demonstrate the ability of LDWeaver to efficiently generate hypotheses about co-selection using large genomic surveys of multiple major human bacterial pathogen species and validate several findings using functional annotation and phenotypic measurements. Our approach will facilitate the study of bacterial evolution in the light of rapidly expanding population genomic data.
Current methods for inferring historical population sizes from DNA sequences often impose a heavy computational burden, or relieve that burden by imposing a fixed parametric form. In addition, they can be marred by sequencing errors or uncertainty about recombination rates, and the quality of inference is often poor in the recent past. We propose "InferNo" for flexible, nonparametric inference of effective population sizes. It requires modest computing resources and little prior knowledge of the recombination and mutation maps, and is robust to sequencing error and gene conversion. We illustrate the statistical and computational advantages of InferNo over previous approaches using a range of simulation scenarios. In particular, we demonstrate the ability of InferNo to exploit biobank-scale datasets for accurate inference of rapid population size changes in the recent past. We also apply Inferno to worldwide human data, finding remarkable similarities in inferences from different populations in the same region. Unlike previous studies, we show two historic bottlenecks for most of the non-African populations. ### Competing Interest Statement The authors have declared no competing interest.
A GWAS in Latin Americans highlights the convergent evolution of lighter skin pigmentation in Eurasia
ABSTRACT BACKGROUND Prevention of human hypertension is an important challenge and has been achieved in experimental models. Brief treatment with renin-angiotensin system (RAS) inhibitors permanently reduces the genetic hypertension of the spontaneously hypertensive rat (SHR). The kidney is involved in this reprogramming, but relevant genetic changes are unknown. METHODS In SHR, we studied the effect of treatment between 10 and 14 weeks of age with the angiotensin receptor blocker, losartan, or the angiotensin-converting enzyme (ACE) inhibitor, perindopril (with controls for non-specific effects of lowering BP) on differential RNA expression, DNA methylation and renin immunolabelling in the kidney at 20 weeks of age. RESULTS RNA sequencing revealed a 6-fold increase in renin gene ( Ren ) expression during losartan treatment (P < 0.0001). At 20 weeks, six weeks after treatment cessation, mean arterial pressure remained lower in the treated SHR (P = 0.006), kidney Ren expression was reduced by 23% (P = 0.03) and DNA methylation within the Ren promoter region was increased (P = 0.04). Experiments with the ACE inhibitor perindopril confirmed a long-term reduction in kidney Ren expression of 43% (P = 1.4 x 10 -6 ). Renin immunolabelling was also lower after losartan or perindopril treatment (P = 0.002). RNA sequencing identified differential expression of 13 candidate genes ( Grhl1 , Ammecr1l , Hs6st1 , Nfil3 , Fam221a , Lmo4 , Adamts1 , Cish , Hif3a , Bcl6 , Rad54l2 , Adap1 , Dok4 ) and the miRNA miR-145-3p. We found correlations between expression of mRNAs, miRNAs and lncRNAs that we believe represent genetic networks underpinning the decreased Ren expression and lower BP. Gene ontogeny analyses revealed that these networks were enriched with genes relevant to BP, RAS and the kidneys. CONCLUSIONS Early RAS inhibition in SHR reprograms genetic pathways and networks resulting in a legacy of reduced Ren expression and the persistent reduction in BP.
Full text Figures and data Side by side Abstract Editor's evaluation Introduction Results Discussion Appendix 1 Appendix 2 Data availability References Decision letter Author response Article and author information Metrics Abstract Our understanding of population history in deep time has been assisted by fitting admixture graphs (AGs) to data: models that specify the ordering of population splits and mixtures, which along with the amount of genetic drift and the proportions of mixture, is the only information needed to predict the patterns of allele frequency correlation among populations. The space of possible AGs relating populations is vast, and thus most published studies have identified fitting AGs through a manual process driven by prior hypotheses, leaving the majority of alternative models unexplored. Here, we develop a method for systematically searching the space of all AGs that can incorporate non-genetic information in the form of topology constraints. We implement this findGraphs tool within a software package, ADMIXTOOLS 2, which is a reimplementation of the ADMIXTOOLS software with new features and large performance gains. We apply this methodology to identify alternative models to AGs that played key roles in eight publications and find that in nearly all cases many alternative models fit nominally or significantly better than the published one. Our results suggest that strong claims about population history from AGs should only be made when all well-fitting and temporally plausible models share common topological features. Our re-evaluation of published data also provides insight into the population histories of humans, dogs, and horses, identifying features that are stable across the models we explored, as well as scenarios of populations relationships that differ in important ways from models that have been highlighted in the literature. Editor's evaluation This is a rigorous and critical analysis of the performance of a popular suite of methods for inferring population history, accompanied by improvements. Should be of broad interest to anyone interested in human history. https://doi.org/10.7554/eLife.85492.sa0 Decision letter eLife's review process Introduction Admixture graph models provide a powerful intellectual framework for describing the relationships among populations that allows not only branching of populations from a common ancestor but also mixture events. An admixture graph (abbreviated below as AG), as fit in the widely used software packages ADMIXTOOLS (Patterson et al., 2012) and TreeMix (Pickrell and Pritchard, 2012; Molloy et al., 2021), is a directed acyclic bifurcating graph with two types of edges: those representing genetic drift, and those representing gene flow. Each admixture event is represented as a confluence of two gene flow edges. Nodes of such a graph represent unsampled intermediate populations, and terminal nodes (leaves) represent sampled present-day or ancient groups (see a mathematical definition in Soraggi and Wiuf, 2019). An attractive feature of AGs is that they can summarize important features of population history without requiring specification of all parameters such as population sizes, split times, mixture times, and distinguishing between sudden splits or drawn-out separations. All these parameters describe important features of demographic history and are fit by many methods for fitting demographic models (Gutenkunst et al., 2009; Gronau et al., 2011; Schiffels et al., 2016; Flegontov et al., 2019; Kamm et al., 2020, Rogers, 2019; Hubisz et al., 2020). However, the fact that it is possible to factor this difficult problem by first inferring important aspects of the topology (AGs fitted to allele frequency correlation statistics), and then fitting additional demographic parameters to data such as site frequency spectra, simplifies demographic inference (Patterson et al., 2012; Pickrell and Pritchard, 2012; Lipson et al., 2013; Leppälä et al., 2017; Lipson et al., 2020b; Molloy et al., 2021; Yan et al., 2021). AGs thus serve both as conceptual frameworks that allow us to think about the relationships of populations deep in time, and as mathematical models we can fit to genetic data. AGs are fitted to f-statistics (Reich et al., 2009; Patterson et al., 2012; Peter, 2016; Soraggi and Wiuf, 2019). For convenience, below we use a concise definition of f-statistics by Lipson, 2020a: 'The most general definition is that of the f4-statistic f4(A, B; C, D), which measures the average correlation in allele frequency differences between (1) populations A and B and (2) populations C and D that is, (pA – pB)∗ (pC – pD), for allele frequencies p, typically averaged over many biallelic single-nucleotide polymorphisms. This f4-statistic is the same as the D-statistic up to a normalization factor.' The other f-statistics (f2 and f3) can be defined as special cases of f4-statistics: f2(A, B) = f4(A, B; A, B) and f3(A; B, C) = f4(A, B; A, C). f4-Statistics can be written as linear combinations of f3- or f2-statistics, and f3-statistics can be written as linear combinations of f4- and f2-statistics. f2-, f3-, and f4-statistics have straightforward interpretations in terms of drift edges along the tree, see Figure 2 in Patterson et al., 2012 and Appendix 1—figure 1b. A challenge for fitting AG models is that they are often not uniquely constrained by the data, with many providing equally good fits to the f2-, f3-, and f4-statistics used to constrain them within the limits of statistical resolution. Previously published methods for finding fitting AGs (mainly qpGraph, Patterson et al., 2012 and TreeMix, Pickrell and Pritchard, 2012; Molloy et al., 2021) were not well equipped to handle the large range of equally well-fitting models for three reasons: (1) They did not reliably provide information on whether there is a uniquely fitting parsimonious model or alternatively whether there are many models that fit equally well to the limits of statistical resolution, (2) they did not provide formal goodness-of-fit tests, and related to this, (3) they did not provide tests for whether the difference between the fits of any two models is statistically significant. As a consequence, and as we demonstrate in what follows, many published AG models have been interpreted as providing more confidence than is merited about the extent to which genetic data allows us to disentangle ancestral relationships. To appreciate these problems, we first need to consider two main approaches that were utilized to study demographic history with AGs. The first approach is to identify AGs automatically, either without human intervention or with guidance. It is possible in theory to exhaustively test all possible graphs for a given set of populations and pre-specified number of admixture events, as implemented, for example, in the admixturegraph R package (Leppälä et al., 2017). An exhaustive approach can provide a complete view of the range of models that are consistent with the data for a specified level of parsimony (total number of admixture events allowed in the graph), which is not biased by the algorithm used to explore the space of possible AGs. However, this approach is limited to small graphs (typically up to six groups, two admixture events) due to the rapid increase in the number of possible AGs as the number of populations and admixture events grows. As we show in our discussion of case studies, the simple models explored with an exhaustive approach can lead to misleading conclusions about population history because not including additional populations can blind users to additional mixture events that occurred (and whose existence is revealed by examining data from additional populations). Furthermore, models with additional admixture events that are qualitatively different to the best-fitting parsimonious graph and that capture the true history, will sometimes be completely missed when constraining the number of gene flows. Alternatively, the programs TreeMix (Pickrell and Pritchard, 2012; Molloy et al., 2021), MixMapper (Lipson et al., 2013), miqoGraph (Yan et al., 2021), and AdmixtureBayes (Nielsen et al., 2023 preprint) all address the problem of how to rapidly explore the vast space of AGs relating a set of populations by applying algorithmic ideas or heuristics; all of these methods speed up model search by orders of magnitude. The second approach to fitting AGs is to manually build them up by grafting additional populations onto simpler smaller graphs that fit the data. This approach involves stepwise addition of populations in an order that is chosen based on the best judgment of the user, and for each newly added population involves adding admixture events or tweaks in the graph until a fit is obtained; the user then moves on to adding the next population (see Reich et al., 2009; Reich et al., 2011; Reich et al., 2012; Lazaridis et al., 2014; Seguin-Orlando et al., 2014; Fu et al., 2016; Skoglund et al., 2016; Yang et al., 2017; McColl et al., 2018; Moreno-Mayar et al., 2018; Tambets et al., 2018; van de Loosdrecht et al., 2018; Flegontov et al., 2019; Sikora et al., 2019; Wang et al., 2019; Lipson et al., 2020b; Shinde et al., 2019; Yang et al., 2020; Hajdinjak et al., 2021; Wang et al., 2021; Bergström et al., 2022 for examples). The program qpGraph in the ADMIXTOOLS package (Patterson et al., 2012) has been the most common computational method used for testing fits of individual AGs. Most AGs in the literature have been constructed manually in this way, often acknowledging the existence of alternative models by presenting plausible models side-by-side, and this approach has been the basis for many claims about population history (Lazaridis et al., 2014; Yang et al., 2017; Posth et al., 2018; Sikora et al., 2019; Shinde et al., 2019; Bergström et al., 2020; Lipson et al., 2020b; Hajdinjak et al., 2021; Wang et al., 2021; Bergström et al., 2022). A strength of this approach is that it takes advantage of human judgment and outside knowledge about what graphs best fit the history of the human or animal populations being analyzed. This external information is powerful as it can incorporate non-genetic evidence such as geographic plausibility and temporal ordering of populations or linguistic similarity, or other genetic data such as estimates of population split times, or shared Y chromosomes, or rejection of proposed scenarios based on joint analysis of much larger numbers of populations than can reasonably be analyzed within a single AG. Thus, while manual approaches explore many orders of magnitude fewer topologies than automatic approaches often do, they still may provide inferences about population history that are more useful than those provided by automatic approaches. These methods' strength is also their weakness: by relying on intuition, following a manual approach has the potential to validate the biases users have as to what types of histories are most plausible (these may be the only types of histories that will be carefully explored). This can blind users to surprises: to profoundly different topologies that may correspond more closely to the true history, and we discuss examples of this in the Results section. In this study, we introduce a new method, findGraphs, that belongs to the first class of algorithms (those for automated AG topology inference). Algorithmic innovations and speedups in findGraphs enable us to explore a much larger proportion of plausible AG space than many other methods reported to date. The findGraphs method combines the advantages of automated and manual topology exploration by allowing users to encode various sources of information as constraints on the space of AGs, which is then explored automatically. However, the main innovations in findGraphs are not computational, but instead conceptual. Instead of finding one or a few AGs fitting the data well, we use findGraphs for exploring AG spaces and assessing if any reliable information on population history can be extracted from a given AG space (defined by a population set and parsimony constraints) in the first place. Results Regardless of the approach used to search through the space of possibly fitting AGs, a challenge in the effort to find a uniquely well-fitting AG (or group of topologically similar AGs) is that it has been difficult to quantify the absolute goodness of fit of a model to date. We have not been entirely successful with this and are not aware of other work that has been successful. It is also difficult to assess the relative fits of multiple models, especially if they differ in complexity. Performance gains relative to the original implementation of qpGraph allow us to address this problem by obtaining bootstrap confidence intervals and p-values for estimated parameters of single models, as well as for the difference in fit quality of two models (see Appendix 1, Sections 1.B.3 and 2.E). In combination with the approach to automating the search of well-fitting AGs, this leads to a situation where we are able to find and test a large number of models, many of which fit equally well despite often having very different topological features. Published approaches to comparing the fits of AG models based on Akaike information criterion (AIC) or Bayesian information criterion (BIC), see Flegontov et al., 2019; Shinde et al., 2019 have the problem that it is often not clear what the effective number of degrees of freedom is in the two models being compared since in the case of AGs it depends not only on the number of graph edges, but also on graph topology. The methods for automated graph topology inference and model comparison relying on bootstrap resampling are implemented in ADMIXTOOLS 2, a comprehensive platform for learning about population history from f-statistics. It is built to provide a stand-alone workspace for research in this area and is implemented as an R package. For all computations, ADMIXTOOLS 2 exhibits large speedups relative to previously published platforms for f-statistic analysis (e.g., popstats and ADMIXTOOLS version 6.0 which we call 'Classic ADMIXTOOLS' in what follows to distinguish it from updated ADMIXTOOLS version 7.0.2 which implements some of the speedup ideas also implemented in ADMIXTOOLS 2). This is achieved by deploying a series of algorithmic improvements, most notably storage of precomputed f-statistics in random access memory, which avoids having to rely on reading in extremely large genotype matrices to perform most computations. In addition to the new algorithmic ideas allowing efficient searching through the space of AGs and comparing the fits of two AGs, ADMIXTOOLS 2 also provides a solution to the question of which parameters of an AG are identifiable in the limit of infinite data. Methodological details are presented in Appendix 1, and below we focus on documenting problems of AG inference on simulated data and revisiting AGs from the literature to understand the extent to which methodological challenges with AG fitting biased previous studies. Topological diversity of well-fitting models and effects of parsimony constraints on simulated data First, we explored the performance of the findGraphs method for automated topology inference on simulated AGs of random topology, focusing on the following questions: (1) among findGraphs results, how common are AGs fitting nominally or significantly better than the true one but different topologically; (2) what is the degree of topological diversity among these models fitting the data better than the true one? For this purpose, we simulated AGs of four complexity classes using msprime v.1.1.1: eight or nine non-outgroup populations, and four or five admixture events. Only simulations where pairwise FST for groups were in the range characteristic for anatomically modern and archaic humans were selected for further analysis, resulting in 20 random topologies per complexity class, each including a distant outgroup that facilitates automated exploration of the topology space. We ran findGraphs on each simulated dataset starting from random graphs and pre-specifying the true number of admixture events (n), or n − 1, or n + 1 events. For each of these graph complexity levels, we performed 100 independent findGraphs runs and recorded 5 AGs from each run having the best log-likelihood (LL) scores. Topologically redundant AGs were discarded, and for the remaining AGs we calculated worst f4-statistic residuals (WR) and tested if the newly found models fit significantly better than the true model, using the bootstrap model comparison method developed in this study (see Appendix 1, Sections 1.B.3 and 2.E). In Figure 1a–c, we show the following statistics for each simulated AG, summarized across simulated complexity classes and parsimony levels allowed at the stage of topology exploration: fraction of topologies found with findGraphs that fit better than the true AG (according to LL score), or that fit significantly better than the true AG, or those with plausible absolute fits (WR < 3 SE). It is clear that for the great majority of simulated datasets, even a shallow exploration of the topology space with findGraph (100 independent runs) uncovers AGs that fit nominally better than the true topology (Figure 1a and d) and are topologically diverse (see Figure 1e for examples). When allowing for n admixture events, at least one AG fitting significantly better than the true one was found for 60% of simulated datasets. When n + 1 admixture events were allowed, this grew to 100% (all 80 datasets). It should be noted that some admixture events are indistinguishable with f-statistics; for instance, successive gene flows between two lineages, with no other edges branching off between the gene flows. If such gene flows were included in the random topologies we simulated, AGs with n events were overly complex for representing the true history. Thus, if we are dealing with random histories, choosing an optimal complexity class for topology search is not straightforward. Figure 1 Download asset Open asset Computer simulations show that when the true admixture graph (AG) topology is complex, findGraphs frequently finds AGs fitting the data better than the true AG. (a) Fractions of distinct AGs found with findGraphs that fit the data nominally better than the true AG (according to log-likelihood [LL] scores). The simulated datasets are grouped by complexity class (eight or nine leaves, four or five admixture events) and by the number of admixture events allowed at the topology search stage (n − 1 on the left, n in the middle, and n + 1 on the right, where n is the true number of simulated admixture events). Each dot represents a simulated random history, and 20 such histories were simulated for each complexity class. (b) Fractions of distinct AGs found with findGraphs that fit the data significantly better than the true AG (two-tailed empirical p-value of the bootstrap model comparison method <0.05). (c) Fractions of distinct AGs found with findGraphs that fit the data well in absolute terms (WR < 3 SE). (d) Distinct AGs found for a particular simulated history (eight groups and four admixture events) in the LL and WR coordinates. Only best-fitting graphs with WR < 3 SE are shown. The fit of the true topology is shown in yellow, and topologies that fit the data significantly better than the true one are in purple. The true topology was not recovered by our findGraphs searches. (e) The true model from panel (d) and two alternative models found with findGraphs, both fitting significantly better than the true one (based on the bootstrap p-value) and very different topologically. This is presented as an example of very high topological diversity seen among well-fitting models. Model parameters (graph edges) that were inferred to be unidentifiable (see Appendix 1, Section 2.F) are plotted in red. These results on simulated data raise concerns about the extent to which fitting AG topologies provide reliable information about population history. Even for histories including eight or nine groups, an outgroup, and four or five pulse-like admixture events, perfect diploid data, and groups as differentiated as Neanderthals and anatomically modern humans—a complexity class that is simpler than many models fitted to real genetic data in published papers—models fitting the data as well as or better than the true one are common, and their topological diversity is in most cases so high that it precludes consensus inference of topology by analysis of multiple topologies. As we demonstrate on another set of simulated data in Appendix 1 (Section 1.B.1), the probability of finding a 'wrong' model that fits better than the true one grows with increasing graph complexity, and that effect is reproduced with both findGraphs and TreeMix (Appendix 1—figure 2). We expect this problem to be even more acute when researchers are dealing with realistic complex histories. However, geneticists often rely on external constraints on AG topologies (such as temporal plausibility of a topology, results from qpAdm modelling, f4-statistics, PCA, ADMIXTURE, geographical, archaeological, and linguistic considerations) that were not used for filtering the results of our topology searches on simulated data. Thus, it is possible in principle that published AG models are more robust than our results on simulated data, and we explore this issue in depth in the next section and in Appendix 2. Revisiting published AGs We studied AGs from eight publications (Lazaridis et al., 2014; Shinde et al., 2019; Sikora et al., 2019; Bergström et al., 2020; Lipson et al., 2020b; Hajdinjak et al., 2021Librado et al., 2021; Wang et al., 2021) with the goal of comparing published models to models identified by our algorithm for automatically inferring optimal (best-fitting) AGs (Table 1). In all but one study, qpGraph or its automated reimplementation (admixturegraph, Leppälä et al., 2017) was used for fitting topologies to genetic data, while Librado et al., 2021 relied on the automated OrientAGraph method (Molloy et al., 2021). The main question we were interested in is whether we can find alternative models which (1) fit as well as, or better than the published graph, (2) differ in important ways from the published graph, and (3) cannot immediately be rejected based on other evidence such as temporal plausibility. The studies were selected according to the criterion that an AG model inferred in the study is used as primary evidence for at least one statement about population history in the main text of the study. In other words, the AG method was used in the original studies to support new conclusions about population history, and not simply to show that there is a model that exists that does not contradict results of other genetic analyses, an approach that is a valid use of AGs and has been taken in some studies (e.g., Seguin-Orlando et al., 2014; Narasimhan et al., 2019; Wang et al., 2019). There are many published studies that could have been included in our re-evaluation exercise as they meet our key criterion (e.g., Yang et al., 2017; McColl et al., 2018; Posth et al., 2018; Flegontov et al., 2019; Carlhoff et al., 2021, Kutanan et al., 2021; Bergström et al., 2022; Lipson et al., 2022; Vallini et al., 2022). However, critical re-evaluation of each published graph is an intensive process, and the sample of studies we revisited is diverse enough to identify some general patterns. Table 1 Published graphs in the context of automatically found graphs. We compared graphs from eight publications to alternative graphs inferred on the same or very similar data (see Supplementary file 1 for details). PublicationFigure in the original publicationGroups (populations)Admixture eventsSNPs usedPubl. model: worst residual, SEDistinct alternative topologies foundSignificantly better fitting topologies, %Non-significantly better fitting topologies, %Non-significantly worse fitting topologies, %Significantly worse fitting topologies, %Bergström et al., 20201e73312,2822.12210.52.316.780.5Lazaridis et al., 2014374642,2472.23061.012.180.76.2Shinde et al., 201938*32,49,0092.61430.02.83.593.7Librado et al., 20213b10*31,767,41923.93246.815.724.153.4Ext 5d414.15350.00.04.595.5Ext 5e56.97840.00.328.471.3Hajdinjak et al., 20212d128263,6984.8198815.755.76.622.0Lipson et al., 2020bExt 41211†211,7382.320000.011.977.110.4Wang et al., 2021Ext 6128203,7533.8177812.684.33.10.0Sikora et al., 20193f (left)136†344,9033.88940.317.134.648.03f (right)146†613,5094.227850.10.99.889.2 Publication: Last name of the first author and year of the relevant publication. Figure in the original publication: Figure number in the original paper where the AG is presented. Groups (populations): The number of populations in each graph. Admixture events: The number of admixture events in each graph. SNPs used: The number of single-nucleotide polymorphisms (SNPs; with no missing data at the group level) used for fitting the AGs. For all case studies, we tested the original data (SNPs, population composition, and the published graph topology) and obtained model fits very similar to the published ones. However, for the purpose of efficient topology search, we in some cases adjusted settings for f3-statistic calculation, population composition, or graph complexity as noted in the footnotes, in Supplementary file 1, and discussed in the text. Publ. model: Worst residual, SE: The worst f-statistic residual of the published graph fitted to the SNP set shown in the 'SNPs used' column, measured in standard errors (SE). Distinct alternative topologies found: The number of distinct newly found topologies differing from the published one. Significantly better fitting topologies, %: The percentage of distinct alternative topologies that fit significantly better than the published graph according to the bootstrap model comparison test (two-tailed empirical p-value <0.05). If the number of distinct topologies was very large, a representative sample of models (1/20 to 1/3 of models evenly distributed along the log-likelihood spectrum) was compared to the published one instead, and the percentages in this and following columns were calculated on this sample. Non-significantly better fitting topologies, %: The percentage of distinct topologies that fit non-significantly (nominally) better than the published graph according to the bootstrap model comparison test (two-tailed empirical p-value ≥0.05). Non-significantly worse fitting topologies, %: The percentage of distinct topologies that fit non-significantly (nominally) worse than the published graph according to the bootstrap model comparison test (two-tailed empirical p-value ≥0.05). Significantly worse fitting topologies, %: The percentage of distinct topologies that fit significantly worse than the published graph according to the bootstrap model comparison test (two-tailed empirical p-value <0.05). * The population composition was modified, see Supplementary file 1 and the text. † Certain gene flows were removed from the published model for simplicity, see Supplementary file 1 and the text. Here we present a high-level summary of these analyses. Discussion of individual case studies follows below, and for details see the exposition in Appendix 2. For 19 out of 22 published graphs we examined, we were able to find at least one, but usually many, graphs of the same complexity (number of groups and admixture events), with an LL score that was nominally better than that of the published graph (see results for 11 selected graphs in Table 1 and full results for all 22 graphs in Supplementary file 1). The 22 graphs were drawn from the 8 publications as there were multiple final graphs presented in some of the publications (Shinde et al., 2019; Sikora et al., 2019; Librado et al., 2021), or we examined selected intermediates in the model construction process (Bergström et al., 2020; Lazaridis et al., 2014; Lipson et al., 2020b; Wang et al., 2021), or we introduced an outgroup not used in the original study (Hajdinjak et al., 2021; Sikora et al., 2019), or we tested additional graph complexity classes dropping 'unnecessary' admixture events (Lipson et al., 2020b; Sikora et al., 2019). These alternative graphs often fit not significantly better than the published one after taking into account variability across single-nucleotide polymorphisms (SNPs) via bootstrapping. In the following cases, at least one model that fits significantly better than the published one according to our bootstrap model comparison method was found: the Bergström et al. and Lazaridis et al. 7-population graphs; the Librado et al., 2021 graph with 3 admixture events; the Hajdinjak et al. graphs with or without adding a chimpanzee outgroup; the Lipson et al., 2020b. intermediate graphs with 7 groups and 4 admixture events and with 10 groups and 8 admixture events; the Wang et al. 12-population graph; and the Sikora et al. graphs for West Eurasians and for East Eurasians with 10 or 6 admixture events (Supplementary file 1). In nearly all cases (except for the Lazaridis et al. six-population graph, Shinde et al. graph with eight populations and three admixture events, and the Librado et al., 2021. graph with four admixture events), we also identified many additional graphs that fit the data not significantly worse than the published ones. In every example, some of these graphs have topologies that are qualitatively different in important ways from those of the published graphs. Features such as which populations are admixed or unadmixed, direction of gene flow, or the order of split events, if not constrained a priori, are generally not the same between alternative fit
We introduce a fast, new algorithm for inferring from allele count data the FST parameters describing genetic distances among a set of populations and/or unrelated diploid individuals, and a tree with branch lengths corresponding to FST values. The tree can reflect historical processes of splitting and divergence, but seeks to represent the actual genetic variance as accurately as possible with a tree structure. We generalise two major approaches to defining FST, via correlations and mismatch probabilities of sampled allele pairs, which measure shared and non-shared components of genetic variance. A diploid individual can be treated as a population of two gametes, which allows inference of coancestry coefficients for individuals as well as for populations, or a combination of the two. A simulation study illustrates that our fast method-of-moments estimation of FST values, simultaneously for multiple populations/individuals, gains statistical efficiency over pairwise approaches when the population structure is close to tree-like. We apply our approach to genome-wide genotypes from the 26 worldwide human populations of the 1000 Genomes Project. We first analyse at the population level, then a subset of individuals and in a final analysis we pool individuals from the more homogeneous populations. This flexible analysis approach gives advantages over traditional approaches to population structure/coancestry, including visual and quantitative assessments of long-standing questions about the relative magnitudes of within- and between-population genetic differences.
Association mapping using crop cultivars allows identification of genetic loci of direct relevance to breeding. Here, 150 U.K. wheat (Triticum aestivum L.) cultivars genotyped with 23,288 single nucleotide polymorphisms (SNPs) were used for genome-wide association studies (GWAS) using historical phenotypic data for grain protein content, Hagberg falling number (HFN), test weight, and grain yield. Power calculations indicated experimental design would enable detection of quantitative trait loci (QTL) explaining ≥20% of the variation (PVE) at a relatively high power of >80%, falling to 40% for detection of a SNP with an R2 ≥ .5 with the same QTL. Genome-wide association studies identified marker-trait associations for all four traits. For HFN (h 2 = .89), six QTL were identified, including a major locus on chromosome 7B explaining 49% PVE and reducing HFN by 44 s. For protein content (h 2 = 0.86), 10 QTL were found on chromosomes 1A, 2A, 2B, 3A, 3B, and 6B, together explaining 48.9% PVE. For test weight, five QTL were identified (one on 1B and four on 3B; 26.3% PVE). Finally, 14 loci were identified for grain yield (h 2 = 0.95) on eight chromosomes (1A, 2A, 2B, 2D, 3A, 5B, 6A, 6B; 68.1% PVE), of which five were located within 16 Mbp of genetic regions previously identified as under breeder selection in European wheat. Our study demonstrates the utility of exploiting historical crop datasets, identifying genomic targets for independent validation, and ultimately for wheat genetic improvement.
Genetic epilepsy with febrile seizures plus (GEFS+) is an autosomal dominant familial epilepsy syndrome characterized by distinctive phenotypic heterogeneity within families. The SCN1B c.363C>G (p.Cys121Trp) variant has been identified in independent, multi-generational families with GEFS+. Although the variant is present in population databases (at very low frequency), there is strong clinical, genetic, and functional evidence to support pathogenicity. Recurrent variants may be due to a founder event in which the variant has been inherited from a common ancestor. Here, we report evidence of a single founder event giving rise to the SCN1B c.363C>G variant in 14 independent families with epilepsy. A common haplotype was observed in all families, and the age of the most recent common ancestor was estimated to be approximately 800 years ago. Analysis of UK Biobank whole-exome-sequencing data identified 74 individuals with the same variant. All individuals carried haplotypes matching the epilepsy-affected families, suggesting all instances of the variant derive from a single mutational event. This unusual finding of a variant causing an autosomal dominant, early-onset disease in an outbred population that has persisted over many generations can be attributed to the relatively mild phenotype in most carriers and incomplete penetrance. Founder events are well established in autosomal recessive and late-onset disorders but are rarely observed in early-onset, autosomal dominant diseases. These findings suggest variants present in the population at low frequencies should be considered potentially pathogenic in mild phenotypes with incomplete penetrance and may be more important contributors to the genetic landscape than previously thought.
Complex‐trait genetics has advanced dramatically through methods to estimate the heritability tagged by SNPs, both genome‐wide and in genomic regions of interest such as those defined by functional annotations. The models underlying many of these analyses are inadequate, and consequently many SNP‐heritability results published to date are inaccurate. Here, we review the modelling issues, both for analyses based on individual genotype data and association test statistics, highlighting the role of a low‐dimensional model for the heritability of each SNP. We use state‐of‐art models to present updated results about how heritability is distributed with respect to functional annotations in the human genome, and how it varies with allele frequency, which can reflect purifying selection. Our results give finer detail to the picture that has emerged in recent years of complex trait heritability widely dispersed across the genome. Confounding due to population structure remains a problem that summary statistic analyses cannot reliably overcome. Also see the video abstract here: https://youtu.be/WC2u03V65MQ
The inclusion of ancestrally diverse participants in genetic studies can lead to new discoveries and is important to ensure equitable health care benefit from research advances. Here, members of the Ethical, Legal, Social, Implications (ELSI) committee of the International Genetic Epidemiology Society (IGES) offer perspectives on methods and analysis tools for the conduct of inclusive genetic epidemiology research, with a focus on admixed and ancestrally diverse populations in support of reproducible research practices. We emphasize the importance of distinguishing socially defined population categorizations from genetic ancestry in the design, analysis, reporting, and interpretation of genetic epidemiology research findings. Finally, we discuss the current state of genomic resources used in genetic association studies, functional interpretation, and clinical and public health translation of genomic findings with respect to diverse populations.
Clustering genetic variants based on their associations with different traits can provide insight into their underlying biological mechanisms. Existing clustering approaches typically group variants based on the similarity of their association estimates for various traits. We present a new procedure for clustering variants based on their proportional associations with different traits, which is more reflective of the underlying mechanisms to which they relate. The method is based on a mixture model approach for directional clustering and includes a noise cluster that provides robustness to outliers. The procedure performs well across a range of simulation scenarios. In an applied setting, clustering genetic variants associated with body mass index generates groups reflective of distinct biological pathways. Mendelian randomization analyses support that the clusters vary in their effect on coronary heart disease, including one cluster that represents elevated body mass index with a favourable metabolic profile and reduced coronary heart disease risk. Analysis of the biological pathways underlying this cluster identifies inflammation as potentially explaining differences in the effects of increased body mass index on coronary heart disease.
We present LDAK-GBAT, a tool for gene-based association testing using summary statistics from genome-wide association studies that is computationally efficient, produces well-calibrated p values, and is significantly more powerful than existing tools. LDAK-GBAT takes approximately 30 min to analyze imputed data (2.9M common, genic SNPs), requiring less than 10 Gb memory. It shows good control of type 1 error given an appropriate reference panel. Across 109 phenotypes (82 from the UK Biobank, 18 from the Million Veteran Pro-gram, and nine from the Psychiatric Genetics Consortium), LDAK-GBAT finds on average 19% (SE: 1%) more significant genes than the existing tool sumFREGAT-ACAT, with even greater gains in comparison with MAGMA, GCTA-fastBAT, sumFREGAT-SKAT-O, and sum-FREGAT-PCA.
We present a novel algorithm, implemented in the software ARGinfer, for probabilistic inference of the Ancestral Recombination Graph under the Coalescent with Recombination. Our Markov Chain Monte Carlo algorithm takes advantage of the Succinct Tree Sequence data structure that has allowed great advances in simulation and point estimation, but not yet probabilistic inference. Unlike previous methods, which employ the Sequentially Markov Coalescent approximation, ARGinfer uses the Coalescent with Recombination, allowing more accurate inference of key evolutionary parameters. We show using simulations that ARGinfer can accurately estimate many properties of the evolutionary history of the sample, including the topology and branch lengths of the genealogical tree at each sequence site, and the times and locations of mutation and recombination events. ARGinfer approximates posterior probability distributions for these and other quantities, providing interpretable assessments of uncertainty that we show to be well calibrated. ARGinfer is currently limited to tens of DNA sequences of several hundreds of kilobases, but has scope for further computational improvements to increase its applicability.
Whole-genome sequencing has facilitated genome-wide analyses of association, prediction and heritability in many organisms. However, such analyses in bacteria are still in their infancy, being limited by difficulties including genome plasticity and strong population structure. Here we propose a suite of methods including linear mixed models, elastic net and LD-score regression, adapted to bacterial traits using innovations such as frequency-based allele coding, both insertion/deletion and nucleotide testing and heritability partitioning. We compare and validate our methods against the current state-of-art using simulations, and analyse three phenotypes of the major human pathogen Streptococcus pneumoniae, including the first analyses of minimum inhibitory concentrations (MIC) for penicillin and ceftriaxone. We show that the MIC traits are highly heritable with high prediction accuracy, explained by many genetic associations under good population structure control. In ceftriaxone MIC, this is surprising because none of the isolates are resistant as per the inhibition zone criteria. We estimate that half of the heritability of penicillin MIC is explained by a known drug-resistance region, which also contributes a quarter of the ceftriaxone MIC heritability. For the within-host carriage duration phenotype, no associations were observed, but the moderate heritability and prediction accuracy indicate a moderately polygenic trait.