The theorised risk that confounded rare variant associations will emerge from population based genetic studies has not been investigated empirically. Here, we use 306,991 sequenced exomes from the UK Biobank to demonstrate that recent demography is poorly captured by common and rare variant principal components, and accounting for haplotype sharing does not eliminate false-positive rare variant associations with non-heritable spatially structured traits. Through re-analysis of 155 phenotypes in siblings, we show a trend of higher effect estimates bias for non-uniformly distributed traits, suggesting population stratification is most pervasive in these settings. Despite its spatial structure, bias of rare variant associations with height appeared most strongly influenced by assortative mating. We explore the risk of elevated false discovery rates for recent variants private to extended families sharing polygenic liability to extreme phenotypes, as well as through local linkage with common causal variants. Overall, we consider the complex confounding mechanisms that can impact rare variant studies and demonstrate family-based approaches can enable important sensitivity analyses.
High throughput metabolomic assays offer a huge opportunity to quantify the cellular processes underlying disease and intervention pathways. However, the multi-dimensional inter-relatedness between these processes coupled with the complex noisy measurement environment create a need for generation of new methods that move beyond simple pairwise associations. Here we develop a computationally simple, multivariate, relational comparison method called CLARITY to compare metabolomic data before and after an intervention. This generates a relational anomaly score that combines with traditional methods to increase classification performance of the underlying cause of changes to the levels of and covariances between metabolites. We demonstrate utility in the By-Band-Sleeve (BBS) clinical trial of bariatric surgery using NMR metabolomics data. On supplementing linear regression analysis with CLARITY, previously identified changes form two clusters that imply involvement in different underlying biological pathways. An additional cluster of metabolites are identified as undergoing a relational change which would not have been detected using traditional methods. Gathering insights about metabolites and the biomarkers they capture in the causal pathway between intervention and effect, from observations at scale, will inform the future design of modelling and laboratory experiments to capture the underlying biological process.
Roman writers found the relative empowerment of Celtic women remarkable1. In southern Britain, the Late Iron Age Durotriges tribe often buried women with substantial grave goods2. Here we analyse 57 ancient genomes from Durotrigian burial sites and find an extended kin group centred around a single maternal lineage, with unrelated (presumably inward migrating) burials being predominantly male. Such a matrilocal pattern is undescribed in European prehistory, but when we compare mitochondrial haplotype variation among European archaeological sites spanning six millennia, British Iron Age cemeteries stand out as having marked reductions in diversity driven by the presence of dominant matrilines. Patterns of haplotype sharing reveal that British Iron Age populations form fine-grained geographical clusters with southern links extending across the channel to the continent. Indeed, whereas most of Britain shows majority genomic continuity from the Early Bronze Age to the Iron Age, this is markedly reduced in a southern coastal core region with persistent cross-channel cultural exchange3. This southern core has evidence of population influx in the Middle Bronze Age but also during the Iron Age. This is asynchronous with the rest of the island and points towards a staged, geographically granular absorption of continental influence, possibly including the acquisition of Celtic languages.
Increasingly efficient methods for inferring the ancestral origin of genome regions are needed to gain insights into genetic function and history as biobanks grow in scale. Here we describe two near-linear time algorithms to learn ancestry harnessing the strengths of a Positional Burrows-Wheeler Transform. SparsePainter is a faster, sparse replacement of previous model-based 'chromosome painting' algorithms to identify recently shared haplotypes, whilst PBWTpaint uses further approximations to obtain lightning-fast estimation optimized for genome-wide relatedness estimation. The computational efficiency gains of these tools for fine-scale local ancestry inference offer the possibility to analyse large-scale genomic datasets using different approaches. Application to the UK Biobank shows that haplotypes better represent ancestries than principal components, whilst linkage-disequilibrium of ancestry identifies signals of recent changes to population-specific selection for many genomic regions associated with immune responses, suggesting avenues for understanding the pathogen-immune system interplay on a historical timescale.
Stability for dynamic network embeddings ensures that nodes behaving the same at different times receive the same embedding, allowing comparison of nodes in the network across time. We present attributed unfolded adjacency spectral embedding (AUASE), a stable unsupervised representation learning framework for dynamic networks in which nodes are attributed with time-varying covariate information. To establish stability, we prove uniform convergence to an associated latent position model. We quantify the benefits of our dynamic embedding by comparing with state-of-the-art network representation learning methods on four real attributed networks. To the best of our knowledge, AUASE is the only attributed dynamic embedding that satisfies stability guarantees without the need for ground truth labels, which we demonstrate provides significant improvements for link prediction and node classification.
Genome-wide association studies (GWAS) have revolutionized our understanding of the genetic basis of complex traits and diseases, but limitations in SNP-centric approaches to population stratification limit the resolution of fine-scale population structures. Here we consider the use of haplotypes to represent population structure, leveraging haplotype components (HCs) for an improved understanding of trait associations and adjustment for population stratification. Using data from the UK Biobank, we showed that HCs have stronger associations with a range of phenotypes than principal components (PCs) while containing more predictive power for birthplaces globally. In GWAS, HCs-correction identifies more genome-wide significant association signals for birthplace and lifestyle-related phenotypes, which are missed by PCs-corrected GWAS. Through thorough testing and simulation, we highlight challenges in performing ancestry-specific GWAS, underscoring the critical role of accurate local ancestry inference in studying admixed populations. We analyzed the haplotype structure of the UK Biobank in terms of 93 genetically-distinct populations, which enabled the computation of Ancestral Risk Scores (ARS) across 8 continental populations, providing insights into population-specific genetic risks for traits and diseases. By integrating haplotype information, this framework provides the potential to address challenges in population stratification, enhances GWAS resolution, and supports equitable health research by facilitating genetic studies in diverse populations.
Dynamic graphs provide a flexible data abstraction for modelling many sorts of real-world systems, such as transport, trade, and social networks. Graph neural networks (GNNs) are powerful tools allowing for different kinds of prediction and inference on these systems, but getting a handle on uncertainty, especially in dynamic settings, is a challenging problem.In this work we propose to use a dynamic graph representation known in the tensor literature as the unfolding, to achieve valid prediction sets via conformal prediction. This representation, a simple graph, can be input to any standard GNN and does not require any modification to existing GNN architectures or conformal prediction routines. One of our key contributions is a careful mathematical consideration of the different inference scenarios which can arise in a dynamic graph modelling context. For a range of practically relevant cases, we obtain valid prediction sets with almost no assumptions, even dispensing with exchangeability. In a more challenging scenario, which we call the semi-inductive regime, we achieve valid prediction under stronger assumptions, akin to stationarity. We provide real data examples demonstrating validity, showing improved accuracy over baselines, and sign-posting different failure modes which can occur when those assumptions are violated.
Multiple sclerosis (MS) is a neuro-inflammatory and neurodegenerative disease that is most prevalent in Northern Europe. Although it is known that inherited risk for MS is located within or in close proximity to immune-related genes, it is unknown when, where and how this genetic risk originated 1 . Here, by using a large ancient genome dataset from the Mesolithic period to the Bronze Age 2 , along with new Medieval and post-Medieval genomes, we show that the genetic risk for MS rose among pastoralists from the Pontic steppe and was brought into Europe by the Yamnaya-related migration approximately 5,000 years ago. We further show that these MS-associated immunogenetic variants underwent positive selection both within the steppe population and later in Europe, probably driven by pathogenic challenges coinciding with changes in diet, lifestyle and population density. This study highlights the critical importance of the Neolithic period and Bronze Age as determinants of modern immune responses and their subsequent effect on the risk of developing MS in a changing environment.
The Holocene (beginning around 12,000 years ago) encompassed some of the most significant changes in human evolution, with far-reaching consequences for the dietary, physical and mental health of present-day populations. Using a dataset of more than 1,600 imputed ancient genomes 1 , we modelled the selection landscape during the transition from hunting and gathering, to farming and pastoralism across West Eurasia. We identify key selection signals related to metabolism, including that selection at the FADS cluster began earlier than previously reported and that selection near the LCT locus predates the emergence of the lactase persistence allele by thousands of years. We also find strong selection in the HLA region, possibly due to increased exposure to pathogens during the Bronze Age. Using ancient individuals to infer local ancestry tracts in over 400,000 samples from the UK Biobank, we identify widespread differences in the distribution of Mesolithic, Neolithic and Bronze Age ancestries across Eurasia. By calculating ancestry-specific polygenic risk scores, we show that height differences between Northern and Southern Europe are associated with differential Steppe ancestry, rather than selection, and that risk alleles for mood-related phenotypes are enriched for Neolithic farmer ancestry, whereas risk alleles for diabetes and Alzheimer’s disease are enriched for Western hunter-gatherer ancestry. Our results indicate that ancient selection and migration were large contributors to the distribution of phenotypic diversity in present-day Europeans.
Populations that have experienced a bottleneck are regularly used in Genome Wide Association Studies (GWAS) to investigate variants associated with complex traits. It is generally understood that these isolated sub-populations may experience high frequency of otherwise rare variants with large effect size, and therefore provide a unique opportunity to study said trait. However, the demographic history of the population under investigation affects all SNPs that determine the complex trait genome-wide, changing its heritability and genetic architecture. We use a simulation based approach to identify the impact of the demographic processes of drift, expansion, and migration on the heritability of complex trait. We show that demography has considerable impact on complex traits. We then investigate the power to resolve heritability of complex traits in GWAS studies subjected to demographic effects. We find that demography is an important component for interpreting inference of complex traits and has a nuanced impact on the power of GWAS. We conclude that demographic histories need to be explicitly modelled to properly quantify the history of selection on a complex trait.
This paper asks the question: can genomic information be used to recover a species that is already on the pathway to extinction due to genetic swamping from a related and more numerous population? We show that a breeding strategy in a captive breeding program can use whole genome sequencing to identify and remove segments of DNA introgressed through hybridisation. The proposed policy uses a generalized measure of kinship or heterozygosity accounting for local ancestry, that is, whether a specific genetic location was inherited from the target of conservation. We then show that optimizing these measures would minimize undesired ancestry while also controlling kinship and/or heterozygosity, in a simulated breeding population. The process is applied to real data representing the hybridized Scottish wildcat breeding population, with the result that it should be possible to breed out domestic cat ancestry. The ability to reverse introgression is a powerful tool brought about through the combination of sequencing with computational advances in ancestry estimation. Since it works best when applied early in the process, important decisions need to be made about which genetically distinct populations should benefit from it and which should be left to reform into a single population.
Microhaplotypes (MHs) describe physically close genetic markers that are inherited together and are gaining prominence due to their efficiency in forensic, clinical, and population studies. They excel in kinship analysis, DNA mixture detection, and ancestry inference, offering advantages in precision over individual SNPs and STRs. In this study, a pipeline was developed to efficiently select highly informative MHs from large-scale genomic datasets. Over 120,000 MHs were identified from almost a million markers, which allow this non-independent information to be efficiently used for inference. The MHs were compared to SNPs in terms of their informativeness and performance of their subsets in ancestry inference and all the results consistently favored MHs. A method for ranking markers by specific population informativeness was also introduced, which showed improvement in the accuracy of Native American ancestry estimation, overcoming the challenges of its underrepresentation in datasets. In conclusion, this study presents a comprehensive way for selecting highly informative MHs for accurate ancestry inference. The proposed approach and the subsets selected by specific population informativeness offer valuable tools for improving ancestry inference accuracy, particularly for admixed populations as demonstrated for a Brazilian dataset.
Major migration events in Holocene Eurasia have been characterized genetically at broad regional scales 1 – 4 . However, insights into the population dynamics in the contact zones are hampered by a lack of ancient genomic data sampled at high spatiotemporal resolution 5 – 7 . Here, to address this, we analysed shotgun-sequenced genomes from 100 skeletons spanning 7,300 years of the Mesolithic period, Neolithic period and Early Bronze Age in Denmark and integrated these with proxies for diet ( 13 C and 15 N content), mobility ( 87 Sr/ 86 Sr ratio) and vegetation cover (pollen). We observe that Danish Mesolithic individuals of the Maglemose, Kongemose and Ertebølle cultures form a distinct genetic cluster related to other Western European hunter-gatherers. Despite shifts in material culture they displayed genetic homogeneity from around 10,500 to 5,900 calibrated years before present, when Neolithic farmers with Anatolian-derived ancestry arrived. Although the Neolithic transition was delayed by more than a millennium relative to Central Europe, it was very abrupt and resulted in a population turnover with limited genetic contribution from local hunter-gatherers. The succeeding Neolithic population, associated with the Funnel Beaker culture, persisted for only about 1,000 years before immigrants with eastern Steppe-derived ancestry arrived. This second and equally rapid population replacement gave rise to the Single Grave culture with an ancestry profile more similar to present-day Danes. In our multiproxy dataset, these major demographic events are manifested as parallel shifts in genotype, phenotype, diet and land use.
The rapid growth of earthquake catalogs, driven by machine learning-based phase picking and denser seismic networks, calls for the application of a broader range of models to determine whether the new data enhances earthquake forecasting capabilities. Additionally, this growth demands that existing forecasting models efficiently scale to handle the increased data volume. Approximate inference methods such as inlabru, which is based on the Integrated nested Laplace approximation, offer improved computational efficiencies and the ability to perform inference on more complex point-process models compared to traditional MCMC approaches. We present SB-ETAS: a simulation based inference procedure for the epidemic-type aftershock sequence (ETAS) model. This approximate Bayesian method uses sequential neural posterior estimation (SNPE) to learn posterior distributions from simulations, rather than typical MCMC sampling using the likelihood. On synthetic earthquake catalogs, SB-ETAS provides better coverage of ETAS posterior distributions compared with inlabru. Furthermore, we demonstrate that using a simulation based procedure for inference improves the scalability from 𝒪(n^2) to 𝒪(nlog n) . This makes it feasible to fit to very large earthquake catalogs, such as one for Southern California dating back to 1981. SB-ETAS can find Bayesian estimates of ETAS parameters for this catalog in less than 10 h on a standard laptop, a task that would have taken over 2 weeks using MCMC. Beyond the standard ETAS model, this simulation based framework allows earthquake modellers to define and infer parameters for much more complex models by removing the need to define a likelihood function.
Western Eurasia witnessed several large-scale human migrations during the Holocene 1 – 5 . Here, to investigate the cross-continental effects of these migrations, we shotgun-sequenced 317 genomes—mainly from the Mesolithic and Neolithic periods—from across northern and western Eurasia. These were imputed alongside published data to obtain diploid genotypes from more than 1,600 ancient humans. Our analyses revealed a ‘great divide’ genomic boundary extending from the Black Sea to the Baltic. Mesolithic hunter-gatherers were highly genetically differentiated east and west of this zone, and the effect of the neolithization was equally disparate. Large-scale ancestry shifts occurred in the west as farming was introduced, including near-total replacement of hunter-gatherers in many areas, whereas no substantial ancestry shifts happened east of the zone during the same period. Similarly, relatedness decreased in the west from the Neolithic transition onwards, whereas, east of the Urals, relatedness remained high until around 4,000 bp , consistent with the persistence of localized groups of hunter-gatherers. The boundary dissolved when Yamnaya-related ancestry spread across western Eurasia around 5,000 bp , resulting in a second major turnover that reached most parts of Europe within a 1,000-year span. The genetic origin and fate of the Yamnaya have remained elusive, but we show that hunter-gatherers from the Middle Don region contributed ancestry to them. Yamnaya groups later admixed with individuals associated with the Globular Amphora culture before expanding into Europe. Similar turnovers occurred in western Siberia, where we report new genomic data from a ‘Neolithic steppe’ cline spanning the Siberian forest steppe to Lake Baikal. These prehistoric migrations had profound and lasting effects on the genetic diversity of Eurasian populations.
The European wildcat population in Scotland is considered critically endangered as a result of hybridization with introduced domestic cats,1,2 though the time frame over which this gene flow has taken place is unknown. Here, using genome data from modern, museum, and ancient samples, we reconstructed the trajectory and dated the decline of the local wildcat population from viable to severely hybridized. We demonstrate that although domestic cats have been present in Britain for over 2,000 years,3 the onset of hybridization was only within the last 70 years. Our analyses reveal that the domestic ancestry present in modern wildcats is markedly over-represented in many parts of the genome, including the major histocompatibility complex (MHC). We hypothesize that introgression provides wildcats with protection against diseases harbored and introduced by domestic cats, and that this selection contributes to maladaptive genetic swamping through linkage drag. Using the case of the Scottish wildcat, we demonstrate the importance of local ancestry estimates to both understand the impacts of hybridization in wild populations and support conservation efforts to mitigate the consequences of anthropogenic and environmental change.
Domestic cats were derived from the Near Eastern wildcat (Felis lybica), after which they dispersed with people into Europe. As they did so, it is possible that they interbred with the indigenous population of European wildcats (Felis silvestris). Gene flow between incoming domestic animals and closely related indigenous wild species has been previously demonstrated in other taxa, including pigs, sheep, goats, bees, chickens, and cattle. In the case of cats, a lack of nuclear, genome-wide data, particularly from Near Eastern wildcats, has made it difficult to either detect or quantify this possibility. To address these issues, we generated 75 ancient mitochondrial genomes, 14 ancient nuclear genomes, and 31 modern nuclear genomes from European and Near Eastern wildcats. Our results demonstrate that despite cohabitating for at least 2,000 years on the European mainland and in Britain, most modern domestic cats possessed less than 10% of their ancestry from European wildcats, and ancient European wildcats possessed little to no ancestry from domestic cats. The antiquity and strength of this reproductive isolation between introduced domestic cats and local wildcats was likely the result of behavioral and ecological differences. Intriguingly, this long-lasting reproductive isolation is currently being eroded in parts of the species’ distribution as a result of anthropogenic activities.
In this paper, we address the problem of dynamic network embedding, that is, representing the nodes of a dynamic network as evolving vectors within a low-dimensional space. While the field of static network embedding is wide and established, the field of dynamic network embedding is comparatively in its infancy. We propose that a wide class of established static network embedding methods can be used to produce interpretable and powerful dynamic network embeddings when they are applied to the dilated unfolded adjacency matrix. We provide a theoretical guarantee that, regardless of embedding dimension, these unfolded methods will produce stable embeddings, meaning that nodes with identical latent behaviour will be exchangeable, regardless of their position in time or space. We additionally define a hypothesis testing framework which can be used to evaluate the quality of a dynamic network embedding by testing for planted structure in simulated networks. Using this, we demonstrate that, even in trivial cases, unstable methods are often either conservative or encode incorrect structure. In contrast, we demonstrate that our suite of stable unfolded methods are not only more interpretable but also more powerful in comparison to their unstable counterparts.