Early postzygotic mutations (PZMs) that arise after fertilization but prior to primordial germ cell specification may be present in both somatic and germ cells, causing mosaicism in a parent and constitutive inheritance in their offspring. In clinical family-trio whole-genome sequencing (WGS), such variants are systematically missed because their sub-heterozygous variant allele fraction (VAF) prevents heterozygous calling in the parent, while residual parental allele support disqualifies the variant as a candidate germline de novo mutation (DNM) in the child. Here, we developed a bioinformatic approach to ascertain parental PZMs from unfiltered DNM candidates in standard-depth (∼30×) trio WGS and applied it to 12,015 trios from the Genomics England 100,000 Genomes Project. We identified 1,015 high-confidence early autosomal parental PZMs, a large single-source catalog of this mutation class. These exhibited a monomodal VAF distribution centered around 5% in parental blood, consistent with empirically characterized ascertainment boundaries imposed by standard-depth sequencing and germline variant calling. PZMs showed no parental age or sex bias and displayed a mutational spectrum distinct from that of DNMs, with enrichment for C>A and T>A substitutions and depletion of T>C. Mutational signature analysis revealed that both mutation types are shaped by clock-like signatures SBS1 and SBS5 in similar proportions, suggesting that spectral differences reflect shifts within shared mutagenic processes. Exploratory genomic distribution analysis revealed a negative PZM association with GC content, in contrast to the positive association for DNMs. Among these, we found variants in DYNC1H1 and WT1 with potential clinical relevance that were missed by routine diagnostic pipelines.
Understanding the history of admixture events and population size changes leading to modern humans is central to human evolutionary genetics. Here we introduce a coalescence-based hidden Markov model, cobraa, that explicitly represents an ancestral population split and rejoin, and demonstrate its application on simulated and real data across multiple species. Using cobraa, we present evidence for an extended period of structure in the history of all modern humans, in which two ancestral populations that diverged ~1.5 million years ago came together in an admixture event ~300 thousand years ago, in a ratio of ~80:20%. Immediately after their divergence, we detect a strong bottleneck in the major ancestral population. We inferred regions of the present-day genome derived from each ancestral population, finding that material from the minority correlates strongly with distance to coding sequence, suggesting it was deleterious against the majority background. Moreover, we found a strong correlation between regions of majority ancestry and human–Neanderthal or human–Denisovan divergence, suggesting the majority population was also ancestral to those archaic humans.
The North Sea's historical migrations have impacted the genetic structure of its neighbouring populations. We analysed haplotype sharing among 858,635 modern individuals from Denmark and Britain to infer migration history from the Middle Ages to the modern day. We estimated the genetic relationship among 370,259 Danes using Danish healthcare registries and validated those with retrievable pedigree relationships from a national family registry. We also compared their IBD sharing with 488,376 British individuals from the UK Biobank. Our analysis revealed the fine-grained population genetic history of Denmark and identified distinct coastal and island communities with a history of genetic isolation and bottlenecks. We observed a significant population decline in Jutland compared to Zealand during the late medieval period to the start of early modern period, accompanied by migration from Jutland to Zealand, corresponding to historical evidence. We identified two major IBD sharing patterns between Denmark and Great Britain: early coast-to-coast connections between South Jutland and eastern England, likely driven by Viking settlements, continuous trade and people movements across the North Sea from the Middle Ages through early modern times, and later city-to-city connections such as those between London and København likely influenced by urbanisation and Industrial Revolution. Further comparisons using other North Sea countries showed both shared and unique histories of genetic exchange with Denmark and Britain. Our study provides novel genetic evidence of migration across the North Sea from the Middle Ages to the Industrial Revolution. It highlights the power of nationwide biobanks in reconstructing fine-scale historical population movements among closely related populations. ### Competing Interest Statement SB has ownerships in Intomics A/S, Hoba Therapeutics Aps, Novo Nordisk A/S, Lundbeck A/S, ALK abello A/S, Eli Lilly and Co. The other authors declare no competing interests.
Complex de novo structural variants (dnSVs) are crucial genetic factors in rare disorders, yet their prevalence and characteristics in rare disorders remain poorly understood. Here, we conduct a comprehensive analysis of whole-genome sequencing data of 12,568 families, including 13,698 offspring with rare diseases, obtained as part of the UK 100,000 Genomes Project. We identify 1,870 dnSVs, constituting the largest dnSV dataset reported to date. Complex dnSVs (n = 158; 8.4%) emerge as the third most common type of SV, following simple deletions and duplications. We classify 65% of these complex dnSVs into 11 subtypes. Among probands with dnSVs (n = 1,696), 9% exhibit exon-disrupting pathogenic dnSVs associated with the probands' phenotype. Notably, 12% of exon-disrupting pathogenic dnSVs and 22% of de novo deletions or duplications previously identified by array-based or whole-exome sequencing methods are found to be complex dnSVs. We also find distinct genomic properties of de novo deletions depending on the parent of origin. This study highlights the importance of complex dnSVs in the cause of rare disorders and demonstrates the necessity of specific genomic analysis to avoid overlooking these variants.
De novo germline mutation is an important factor in the evolution of allelic diversity and disease predisposition in a population. Here, we study the influence of genetically-inferred ancestry and environmental factors on de novo mutation rates and spectra. Using a genetically diverse sample of ~10 K whole-genome sequenced trios, one of the largest de novo mutation catalogues to date, we found that genetically-inferred ancestry is associated with modest but significant changes in both germline mutation rate and spectra across continental populations. These effects may be due to genetic or environmental factors correlated with ancestry. We find epidemiological evidence that cigarette smoking is significantly associated with increased de novo mutation rate, but it does not mediate the observed ancestry effects. Investigation of several other potential mutagenic factors using Mendelian randomisation showed no consistent effects, except for age at menopause, where factors increasing this corresponded to a reduction in de novo mutation rate. Overall, our study sheds light on factors influencing de novo mutation rates and spectra.
Population differences in cardiometabolic disease remain unexplained. Misleading assumptions over genetic explanations are partly due to terminology used to distinguish populations, specifically ancestry, race, and ethnicity. These terms differentially implicate environmental and biological causal pathways, which should inform their use. Genetic variation alone accounts for a limited fraction of population differences in cardiometabolic disease. Research effort should focus on societally driven, lifelong environmental determinants of population differences in disease. Rather than pursuing population stratifiers to personalize medicine, we advocate removing socioeconomic barriers to receipt of and adherence to healthcare interventions, which will have markedly greater impact on improving cardiometabolic outcomes. This requires multidisciplinary collaboration and public and policymaker engagement to address inequalities driven by society rather than biology per se.
De novo structural variants (dnSVs) have emerged as crucial genetic factors in the context of rare disorders. However, these variations often go undiagnosed in routine genetic screening practices. To shed light on their significance in rare disease, we conducted a comprehensive analysis of the largest cohort of parent-offspring whole-genome sequencing data from the UK 100,000 Genomes Project. Our study encompassed a vast cohort of 12,568 families, including 13,702 offspring affected by rare genetic diseases. We identified a total of 1,872 dnSVs, revealing that approximately 12% of the probands harboured at least one dnSV, of which 9% were identified as likely pathogenic in affected probands (151/1696). Advanced parents' age was found to be associated with an increased chance of having dnSVs in probands. We discovered 148 clustered breakpoints resulting from a single event. 60% of these complex dnSVs were classified into 9 major SV types, and could be observed in multiple individuals, while the remaining 40% were private events and had not been previously reported. We found 12% of pathogenic dnSVs are complex SVs, emphasising the critical importance of thoroughly examining and considering complex dnSVs in the context of rare disorders. Furthermore, we discovered an enrichment of maternal dnSVs at subtelometric, early-replicating regions of chromosome 16, suggesting possible sex-specific mechanisms in generation of dnSVs. This study sheds light on the extent of diversity of dnSVs in the germline and their contribution to rare genetic disorders.
Gene duplication events can drive evolution by providing genetic material for new gene functions, and they create opportunities for diverse developmental strategies to emerge between species. To study the contribution of duplicated genes to human early development, we examined the evolution and function of NANOGP1, a tandem duplicate of the transcription factor NANOG. We found that NANOGP1 and NANOG have overlapping but distinct expression profiles, with high NANOGP1 expression restricted to early epiblast cells and naïve-state pluripotent stem cells. Sequence analysis and epitope-tagging revealed that NANOGP1 is protein coding with an intact homeobox domain. The duplication that created NANOGP1 occurred earlier in primate evolution than previously thought and has been retained only in great apes, whereas Old World monkeys have disabled the gene in different ways, including homeodomain point mutations. NANOGP1 is a strong inducer of naïve pluripotency; however, unlike NANOG, it is not required to maintain the undifferentiated status of human naïve pluripotent cells. By retaining expression, sequence and partial functional conservation with its ancestral copy, NANOGP1 exemplifies how gene duplication and subfunctionalisation can contribute to transcription factor activity in human pluripotency and development.
De novo mutations (DNMs) in the germline have long been identified as a key element in the causes of developmental and other genetic disorders. Previous attempts to investigate genetic factors affecting DNMs have suffered from a lack of statistical power, due to the difficulty of obtaining a sufficient number of parent-offspring trios. Thus, the rare disease cohort of the UK’s 100k Genomes Project (100kGP), comprising more than 10,000 trios, represents an unprecedented opportunity to investigate the genetics of germline mutation. Here we estimate SNP heritability of DNM count in offspring, as a measure of the relative contribution of genetic factors to the variance of the trait, in a PCA-selected subset of the 100kGP cohort. We estimate separate SNP heritabilities for paternally and maternally transmitted mutations (based on parentally phased DNMs in offspring), computed using parental genetic variants at a range of minimum frequencies and a variety of methodologies. We estimate a heritability of 10-20% for paternal DNMs; by contrast, for maternal DNMs we find no significant evidence for non-zero heritability. We investigated the partitioning of heritability among genes with different expression profiles in different tissue or cell states, and found a relative heritability enrichment for genes expressed in gonadal tissues, particularly testis. Among germ cells in adult testes we observed relative enrichment of heritability in genes associated with the (undifferentiated) spermatogonial stem cell state.
Attempts to identify a 'homeland' for our species from genetic data are widespread in the academic literature. However, even when putting aside the question of whether a 'homeland' is a useful concept, there are a number of inferential pitfalls in attempting to identify the geographic origin of a species from contemporary patterns of genetic variation. These include making strong claims from weakly informative data, treating genetic lineages as representative of populations, assuming a high degree of regional population continuity over hundreds of thousands of years, and using circumstantial observations as corroborating evidence without considering alternative hypotheses on an equal footing, or formally evaluating any hypothesis. In this commentary we review the recent publication that claims to pinpoint the origins of 'modern humans' to a very specific region in Africa (Chan et al., 2019), demonstrate how it fell into these inferential pitfalls, and discuss how this can be avoided.
Many complex genomic rearrangements arise through template switch errors, which occur in DNA replication when there is a transient polymerase switch to an alternate template nearby in three-dimensional space. While typically investigated at kilobase-to-megabase scales, the genomic and evolutionary consequences of this mutational process are not well characterised at smaller scales, where they are often interpreted as clusters of independent substitutions, insertions and deletions. Here we present an improved statistical approach using pair hidden Markov models, and use it to detect and describe short-range template switches underlying clusters of mutations in the multi-way alignment of hominid genomes. Using robust statistics derived from evolutionary genomic simulations, we show that template switch events have been widespread in the evolution of the great apes’ genomes and provide a parsimonious explanation for the presence of many complex mutation clusters in their phylogenetic context. Larger-scale mechanisms of genome rearrangement are typically associated with structural features around breakpoints, and accordingly we show that atypical patterns of secondary structure formation and DNA bending are present at the initial template switch loci. Our methods improve on previous non-probabilistic approaches for computational detection of template switch mutations, allowing the statistical significance of events to be assessed. By specifying realistic evolutionary parameters based on the genomes and taxa involved, our methods can be readily adapted to other intra- or inter-species comparisons.
The language commonly used in human genetics can inadvertently pose problems for multiple reasons. Terms like ‘ancestry’, ‘ethnicity’, and other ways of grouping people can have complex, often poorly understood, or multiple meanings within the various fields of genetics, between different domains of biological sciences and medicine, and between scientists and the general public. Furthermore, some categories in frequently used datasets carry scientifically misleading, outmoded or even racist perspectives derived from the history of science. Here, we discuss examples of problematic lexicon in genetics, and how commonly used statistical practices to control for the non-genetic environment may exacerbate difficulties in our terminology, and therefore understanding. Our intention is to stimulate a much-needed discussion about the language of genetics, to begin a process to clarify existing terminology, and in some cases adopt a new lexicon that both serves scientific insight, and cuts us loose from various aspects of a pernicious past.
Genome sequences from diverse human groups are needed to understand the structure of genetic variation in our species and the history of, and relationships between, different populations. We present 929 high-coverage genome sequences from 54 diverse human populations, 26 of which are physically phased using linked-read sequencing. Analyses of these genomes reveal an excess of previously undocumented common genetic variation private to southern Africa, central Africa, Oceania, and the Americas, but an absence of such variants fixed between major geographical regions. We also find deep and gradual population separations within Africa, contrasting population size histories between hunter-gatherer and agriculturalist groups in the past 10,000 years, and a contrast between single Neanderthal but multiple Denisovan source populations contributing to present-day human populations.
Ancestry connects genetics and society in fundamental ways.For many people it has cultural, religious or even political significance, and can play a key role in shaping personal and public identities.People's desire to discover their own ancestry drives the multibillion-dollar genealogy industry, which has grown rapidly in the era of consumer genomics.Companies such as 23andMe and Ancestry now claim tens of millions of customers worldwide.In parallel, our scientific understanding of the human past is being transformed by studies of ancient and modern genetic data, which allow us to track changes in ancestry over space and time.Sophisticated methods have been developed to infer and visualise these relationships.Thus, it seems that both scientists and the wider public are learning more and more about ancestry, and there is an optimistic sense that genetic data provide an exhaustive repository of ancestral information.However, although frequently discussed, ancestry itself is rarely defined.We argue that this reflects widespread underlying confusion about what it means in different contexts and what genetic data can really tell us.This leads to miscommunication between researchers in different fields, and leaves customers open to spurious claims about consumer genomics products and overinterpretation of individual results.In wider usage, the terms ancestry and ancestors often indicate a general connection to people or things in the past.But in a genetic context they have a more specific meaning: your ancestors are the individuals from whom you are biologically descended and ancestry is information about them and their genetic relationship to you.Even here however, confusion arises from the way that ancestry is presented and discussed.Rather than emphasising its complex structure, results are often simplified in terms of discrete categories.While convenient and sometimes useful, ultimately this is misleading about the nature of ancestry.These labels can also impose contemporary political or cultural divisions which may be misrepresentative of ancestral relationships.Another source of confusion is that three distinct concepts-genealogical ancestry, genetic ancestry, and genetic similarity-are frequently conflated.We discuss them in turn, but note that only the first two are explicitly forms of ancestry, and that genetic data are surprisingly uninformative about either of them.Consequently, most statements about ancestry are really statements about genetic similarity, which has a complex relationship with ancestry, and can only be related to it by making assumptions about human demography whose validity is uncertain and difficult to test.Genealogical ancestry probably reflects the most common and intuitive understanding of the term ancestry.Consider your parents, grandparents, or even great-grandparents.You likely have a sense of these people as individuals, even if you have never met them.If one of them belonged to a particular group X, you might say that you have some "X" ancestry.You might even be able to claim ancestry in this way from more distant ancestors, based on
We challenge the view that our species, Homo sapiens, evolved within a single population and/or region of Africa. The chronology and physical diversity of Pleistocene human fossils suggest that morphologically varied populations pertaining to the H. sapiens clade lived throughout Africa. Similarly, the African archaeological record demonstrates the polycentric origin and persistence of regionally distinct Pleistocene material culture in a variety of paleoecological settings. Genetic studies also indicate that present-day population structure within Africa extends to deep times, paralleling a paleoenvironmental record of shifting and fractured habitable zones. We argue that these fields support an emerging view of a highly structured African prehistory that should be considered in human evolutionary inferences, prompting new interpretations, questions, and interdisciplinary research directions.