The leading cause of human pregnancy loss is aneuploidy, often tracing to errors in chromosome segregation during female meiosis1,2. Although abnormal crossover recombination is known to confer risk for aneuploidy3,4, limited data have hindered understanding of the potential shared genetic basis of these key molecular phenotypes. To address this gap, we performed retrospective analysis of pre-implantation genetic testing data from 139,416 in vitro fertilized embryos from 22,850 sets of biological parents. By tracing transmission of haplotypes, we identified 3,809,412 crossovers, as well as 92,485 aneuploid chromosomes. Counts of crossovers were lower in aneuploid versus euploid embryos, consistent with their role in chromosome pairing and segregation. Our analyses further revealed that a common haplotype spanning the meiotic cohesin SMC1B is associated significantly with both crossover count and maternal meiotic aneuploidy, with evidence supporting a non-coding cis-regulatory mechanism. Transcriptome- and phenome-wide association tests also implicated variation in the synaptonemal complex component C14orf39 and crossover-regulating ubiquitin ligases CCNB1IP1 and RNF212 in meiotic aneuploidy risk. More broadly, variants associated with aneuploidy often showed secondary associations with recombination, and several also exhibited associations with reproductive ageing traits. Our findings highlight the dual role of recombination in generating genetic diversity, while ensuring meiotic fidelity.
Admixture between modern humans and extinct hominins has shaped the genomes of present-day individuals, but reconstructing this history has been constrained by the scarcity of archaic samples and unadmixed outgroup populations. We introduce TRACE, a reference- and outgroup-free approach that uses features of ancestral recombination graphs to identify archaic ancestry. Simulations demonstrate that TRACE achieves high precision and low false discovery rates. Applied to 1000 genomes, TRACE recovers known Neanderthal and Denisovan introgression and uncovers ghost admixture from uncharacterized hominins in both Africans and non-Africans. Ghost ancestry persists in Neanderthal and Denisovan ancestry deserts, challenging their interpretation as Homo sapiens-specific regions. In Oceanians, TRACE finds that deep lineages are enriched in Denisovan compared with Neanderthal regions, supporting super-archaic introgression. TRACE enables mapping of archaic introgression without archaic reference genomes.
Ancestral recombination graphs (ARGs) are an increasingly important component of population and statistical genetics. The tskit library has become key infrastructure for the field, providing an expressive and general representation of ARGs together with a suite of efficient fundamental operations. In this note, we announce tskit version 1.0, describe its underlying rationale, and document its stability guarantees. These guarantees provide a foundation for durable computational artefacts and support long-term reproducibility of code and analyses.
Neanderthal and Denisovan genomes have reshaped our understanding of archaic introgression. Yet, the limited number of archaic genomes sequenced and the reliance on unadmixed outgroups have left much of this history unresolved. We introduce TRACE, a method to identify archaic ancestry using features of ancestral recombination graphs inferred from contemporary genomes alone. Simulations show that TRACE reliably detects archaic introgression without requiring archaic genomes or unadmixed outgroups. Applied to 1000 Genomes data, TRACE recovers known Neanderthal and Denisovan introgression and reveals signals of ghost admixture from previously uncharacterized hominins in both Africans and non-Africans. Strikingly, ghost ancestry persists in Neanderthal and Denisovan ancestry deserts, challenging their interpretation as Homo sapiens-specific regions. In Oceanians, TRACE finds deep lineages enriched in Denisovan--and not Neanderthal--regions, supporting a model of super-archaic gene flow. TRACE provides a scalable framework for mapping the legacy of archaic introgression in the absence of archaic genome sequences.
Conventional genome mapping-based approaches systematically overlook genetic variation, particularly in regions that substantially differ from the reference. To explore this hidden variation, here we examine unmapped and poorly mapped reads from the genomes of 640 human individuals from South Asian populations in the 1000 Genomes Project and the Simons Genome Diversity Project. We assemble tens of megabases of non-redundant sequence in tens of thousands of large contigs, a significant portion of which is present in both South Asian and other populations. We demonstrate that much of this sequence is not discovered by traditional variant discovery approaches even when using complete genomes and pangenomes. Across 20,000 placed contigs, we find 8215 intersections with 106 protein coding genes and over 15,000 placements within 1 kbp of a known GWAS hit. We use long-read data from a subset of samples to validate the majority of their assembled sequences, align RNA-seq data to identify hundreds of unplaced contigs with transcriptional potential, and query existing nucleotide databases to infer the origins of the remaining unplaced sequences. Our results highlight the limitations of even the most complete reference genomes and provide a model for understanding the distribution of hidden variation in any human population.
One key component of study design in population genetics is the "geographic breadth" of a sample (i.e., how broad a region across which individuals are sampled). How the geographic breadth of a sample impacts observations of rare, deleterious variants is unclear, even though such variants are of particular interest for biomedical and evolutionary applications. Here, in order to gain insight into the effects of sample design on ascertained genetic variants, we formulate a stochastic model of dispersal, genetic drift, selection, mutation, and geographically concentrated sampling. We use this model to understand the effects of the geographic breadth of sampling effort on the discovery of negatively selected variants. We find that samples which are more geographically broad will discover a greater number of variants as compared to geographically narrow samples (an effect we label "discovery"); though the variants will be detected at lower average frequency than in narrow samples (e.g., as singletons, an effect we label "dilution"). Importantly, these effects are amplified for larger sample sizes and fitness effects. We validate these results using both population genetic simulations and empirical analyses in the UK Biobank. Our results are particularly important in two contexts: the association of large-effect rare variants with particular phenotypes and the inference of negative selection from allele frequency data. Overall, our findings emphasize the importance of considering geographic breadth when designing and carrying out genetic studies, especially at biobank scale.
The leading cause of human pregnancy loss is aneuploidy, often tracing to errors in chromosome segregation during female meiosis. While abnormal crossover recombination is known to confer risk for aneuploidy, limited data have hindered understanding of the potential shared genetic basis of these key molecular phenotypes. To address this gap, we performed retrospective analysis of preimplantation genetic testing data from 139,416 in vitro fertilized embryos from 22,850 sets of biological parents. By tracing transmission of haplotypes, we identified 3,656,198 crossovers, as well as 92,485 aneuploid chromosomes. Counts of crossovers were lower in aneuploid versus euploid embryos, consistent with their role in chromosome pairing and segregation. Our analyses further revealed that a common haplotype spanning the meiotic cohesin SMC1B is significantly associated with both crossover count and maternal meiotic aneuploidy, with evidence supporting a non-coding cis-regulatory mechanism. Transcriptome- and phenome-wide association tests also implicated variation in the synaptonemal complex component C14orf39 and crossover-regulating ubiquitin ligases CCNB1IP1 and RNF212 in meiotic aneuploidy risk. More broadly, recombination and aneuploidy possess a partially shared genetic basis that also overlaps with reproductive aging traits. Our findings highlight the dual role of recombination in generating genetic diversity, while ensuring meiotic fidelity.
Sri Lanka has played a key role in the peopling of South Asia, with archaeological evidence for human presence on the island dating back to ⁓40,000 years ago. Present-day Indigenous peoples of the island, the Adivasi, are proposed to have descended from early inhabitants of the region, while urban populations like the Sinhalese, the major ethnic group on the island, migrated from India in historical times. Using whole genomes from 19 Adivasi individuals belonging to two clans and from 35 Sinhalese, we find that the Adivasi and Sinhalese share high genetic similarities with each other and with other Sri Lankan and Indian populations, especially those with greater genetic affinity to Ancestral South Indians (ASI). Admixture modeling of the Sri Lankan groups reveals that despite shared ancestral components, the Adivasi retain higher genetic contributions from ancient hunter-gatherers compared with the Sinhalese. Additionally, in contrast to the Sinhalese, the Adivasi have maintained low effective population size and undergone strong founder events, which is consistent with their hunter-gatherer lifestyle, historic relocations, and habitat fragmentation. While the two Adivasi clans are genetically more similar to each other than to any other populations, we observe differing demographic histories, with the Interior Adivasi experiencing a stronger bottleneck than the Coastal Adivasi since their split. This whole-genome-based study addresses gaps in our understanding of the demographic and migratory history of two key Sri Lankan groups and, consequently, of broader South Asia by illuminating complex population structure that has been shaped by both demographic and socio-cultural factors.
The most dynamic and repetitive regions of great ape genomes have traditionally been excluded from comparative studies 1–3 . Consequently, our understanding of the evolution of our species is incomplete. Here we present haplotype-resolved reference genomes and comparative analyses of six ape species: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan and siamang. We achieve chromosome-level contiguity with substantial sequence accuracy (<1 error in 2.7 megabases) and completely sequence 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, to provide in-depth evolutionary insights. Comparative analyses enabled investigations of the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference genome. Such regions include newly minted gene families in lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes and subterminal heterochromatin. This resource serves as a comprehensive baseline for future evolutionary studies of humans and our closest living ape relatives.
Genetic variation influencing gene expression and splicing is a key source of phenotypic diversity. Though invaluable, studies investigating these links in humans have been strongly biased toward participants of European ancestries, diminishing generalizability and hindering evolutionary research. To address these limitations, we developed MAGE, an open-access RNA-seq data set of lymphoblastoid cell lines from 731 individuals from the 1000 Genomes Project spread across 5 continental groups and 26 populations. Most variation in gene expression (92%) and splicing (95%) was distributed within versus between populations, mirroring variation in DNA sequence. We mapped associations between genetic variants and expression and splicing of nearby genes (cis-eQTLs and cis-sQTLs, respective), identifying >15,000 putatively causal eQTLs and >16,000 putatively causal sQTLs that are enriched for relevant epigenomic signatures. These include 1310 eQTLs and 1657 sQTLs that are largely private to previously underrepresented populations. Our data further indicate that the magnitude and direction of causal eQTL effects are highly consistent across populations and that apparent "population-specific" effects observed in previous studies were largely driven by low resolution or additional independent eQTLs of the same genes that were not detected. Together, our study expands understanding of gene expression diversity across human populations and provides an inclusive resource for studying the evolution and function of human genomes.
We present haplotype-resolved reference genomes and comparative analyses of six ape species, namely: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. We achieve chromosome-level contiguity with unparalleled sequence accuracy (<1 error in 500,000 base pairs), completely sequencing 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, providing more in-depth evolutionary insights. Comparative analyses, including human, allow us to investigate the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference. This includes newly minted gene families within lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes, and subterminal heterochromatin. This resource should serve as a definitive baseline for all future evolutionary studies of humans and our closest living ape relatives.
Apes possess two sex chromosomes-the male-specific Y and the X shared by males and females. The Y chromosome is crucial for male reproduction, with deletions linked to infertility. The X chromosome carries genes vital for reproduction and cognition. Variation in mating patterns and brain function among great apes suggests corresponding differences in their sex chromosome structure and evolution. However, due to their highly repetitive nature and incomplete reference assemblies, ape sex chromosomes have been challenging to study. Here, using the state-of-the-art experimental and computational methods developed for the telomere-to-telomere (T2T) human genome, we produced gapless, complete assemblies of the X and Y chromosomes for five great apes (chimpanzee, bonobo, gorilla, Bornean and Sumatran orangutans) and a lesser ape, the siamang gibbon. These assemblies completely resolved ampliconic, palindromic, and satellite sequences, including the entire centromeres, allowing us to untangle the intricacies of ape sex chromosome evolution. We found that, compared to the X, ape Y chromosomes vary greatly in size and have low alignability and high levels of structural rearrangements. This divergence on the Y arises from the accumulation of lineage-specific ampliconic regions and palindromes (which are shared more broadly among species on the X) and from the abundance of transposable elements and satellites (which have a lower representation on the X). Our analysis of Y chromosome genes revealed lineage-specific expansions of multi-copy gene families and signatures of purifying selection. In summary, the Y exhibits dynamic evolution, while the X is more stable. Finally, mapping short-read sequencing data from >100 great ape individuals revealed the patterns of diversity and selection on their sex chromosomes, demonstrating the utility of these reference assemblies for studies of great ape evolution. These complete sex chromosome assemblies are expected to further inform conservation genetics of nonhuman apes, all of which are endangered species.
Over the past decade, genomic data have contributed to several insights on global human population histories. These studies have been met both with interest and critically, particularly by populations with oral histories that are records of their past and often reference their origins. While several studies have reported concordance between oral and genetic histories, there is potential for tension that may stem from genetic histories being prioritized or used to confirm community -based knowledge and ethnography, especially if they differ. To investigate the interplay between oral and genetic histories, we focused on the southwestern region of India and analyzed wholegenome sequence data from 156 individuals identifying as Bunt, Kodava, Nair, and Kapla. We supplemented limited anthropological records on these populations with oral history accounts from community members and historical literature, focusing on references to non -local origins such as the ancient Scythians in the case of Bunt, Kodava, and Nair, members of Alexander the Great's army for the Kodava, and an African -related source for Kapla. We found these populations to be genetically most similar to other Indian populations, with the Kapla more similar to South Indian tribal populations that maximize a genetic ancestry related to Ancient Ancestral South Indians. We did not find evidence of additional genetic sources in the study populations than those known to have contributed to many other present-day South Asian populations. Our results demonstrate that oral and genetic histories may not always provide consistent accounts of population origins and motivate further community -engaged, multi -disciplinary investigations of non -local origin stories in these communities.
Genome-wide genealogies compactly represent the evolutionary history of a set of genomes and inferring them from genetic data has the potential to facilitate a wide range of analyses. We introduce a method, ARG-Needle, for accurately inferring biobank-scale genealogies from sequencing or genotyping array data, as well as strategies to utilize genealogies to perform association and other complex trait analyses. We use these methods to build genome-wide genealogies using genotyping data for 337,464 UK Biobank individuals and test for association across seven complex traits. Genealogy-based association detects more rare and ultra-rare signals ( N = 134, frequency range 0.0007−0.1%) than genotype imputation using ~65,000 sequenced haplotypes ( N = 64). In a subset of 138,039 exome sequencing samples, these associations strongly tag (average r = 0.72) underlying sequencing variants enriched (4.8×) for loss-of-function variation. These results demonstrate that inferred genome-wide genealogies may be leveraged in the analysis of complex traits, complementing approaches that require the availability of large, population-specific sequencing panels.
African populations have been drastically underrepresented in genomics research, and failure to capture the genetic diversity across the numerous ethnolinguistic groups (ELGs) found on the continent has hindered the equity of precision medicine initiatives globally. Here, we describe the whole-genome sequencing of 449 Nigerian individuals across 47 unique self-reported ELGs. Population structure analysis reveals genetic differentiation among our ELGs, consistent with previous findings. From the 36 million SNPs and insertions or deletions (indels) discovered in our dataset, we provide a high-level catalog of both novel and medically relevant variation present across the ELGs. These results emphasize the value of this resource for genomics research, with added granularity by representing multiple ELGs from Nigeria. Our results also underscore the potential of using these cohorts with larger sample sizes to improve our understanding of human ancestry and health in Africa.
Simulation is a key tool in population genetics for both methods development and empirical research, but producing simulations that recapitulate the main features of genomic datasets remains a major obstacle. Today, more realistic simulations are possible thanks to large increases in the quantity and quality of available genetic data, and the sophistication of inference and simulation software. However, implementing these simulations still requires substantial time and specialized knowledge. These challenges are especially pronounced for simulating genomes for species that are not well-studied, since it is not always clear what information is required to produce simulations with a level of realism sufficient to confidently answer a given question. The community-developed framework stdpopsim seeks to lower this barrier by facilitating the simulation of complex population genetic models using up-to-date information. The initial version of stdpopsim focused on establishing this framework using six well-characterized model species (Adrion et al., 2020). Here, we report on major improvements made in the new release of stdpopsim (version 0.2), which includes a significant expansion of the species catalog and substantial additions to simulation capabilities. Features added to improve the realism of the simulated genomes include non-crossover recombination and provision of species-specific genomic annotations. Through community-driven efforts, we expanded the number of species in the catalog more than threefold and broadened coverage across the tree of life. During the process of expanding the catalog, we have identified common sticking points and developed the best practices for setting up genome-scale simulations. We describe the input data required for generating a realistic simulation, suggest good practices for obtaining the relevant information from the literature, and discuss common pitfalls and major considerations. These improvements to stdpopsim aim to further promote the use of realistic whole-genome population genetic simulations, especially in non-model organisms, making them available, transparent, and accessible to everyone.
Archaeogenetics has been revolutionary, revealing insights into demographic history and recent positive selection in many organisms. However, most studies to date have ignored the non-random association of genetic variants at different loci (i.e., linkage disequilibrium, LD). This may be in part because basic properties of LD in samples from different times are still not well understood. Here, we derive several results for summary statistics of haplotypic variation under a model with time-stratified sampling: 1) The correlation between the number of pairwise differences observed between time-staggered samples ( π Δ t ) in models with and without strict population continuity; 2) The product of the LD coefficient, D, between ancient and modern samples, which is a measure of haplotypic similarity between modern and ancient samples; and 3) The expected switch rate in the Li and Stephens haplotype copying model. The latter has implications for genotype imputation and phasing in ancient samples with modern reference panels. Overall, these results provide a characterization of how haplotype patterns are affected by sample age, recombination rates, and population sizes. We expect these results will help guide the interpretation and analysis of haplotype data from ancient and modern samples.
ABSTRACTProteomic variation between individuals has immense potential for identifying novel drug targets and disease mechanisms. However, with high-throughput proteomic technologies still in their infancy, they have largely been applied in large majority European ancestry cohorts (e.g. the UK Biobank). An open question is the degree to which proteomic signatures seen in European and other groups mirror those seen in diverse populations, such as cohorts from Africa. Coupled with genetic information, we can also gain a better understanding of the role of genetic variants in the regulation of the proteome and subsequent disease mechanisms.To address the gap in our understanding of proteomic variation in individuals of African ancestry, we collected proteomic data from 176 individuals across two ethnic groups (Igbo and Yoruba) in Nigeria. These individuals were also stratified into high BMI (BMI > 30 kg/m2) and normal BMI (20 kg/m2< BMI < 30 kg/m2) categories. We characterized differences in plasma protein abundance using the Olink Explore 1536 panel between high and normal BMI individuals, finding strong associations consistent with previously known signals in individuals of European descent. We additionally found 73 sentinel cis-pQTL in this dataset, with 21 lead cis-pQTL not observed in catalogs of variation from European-ancestry individuals. In summary, our study highlights the value of leveraging proteomic data in cohorts of diverse ancestry for investigating trait-specific mechanisms and discovering novel genetic regulators of the plasma proteome.
Several Indian populations have genetic contributions from ancient hunter-gatherer, Iranian, and Eurasian Steppe lineages. However, sparse sampling across finer-scale regions impedes our understanding of the emergence of present-day population structure. Moreover, oral histories maintained by many Indian populations are only partially documented. To address these gaps and investigate the interplay between oral and genetic histories, we analyzed whole-genome sequences from 158 Southwest Indian individuals identifying as Bunt, Kodava, Nair, and Kapla. From community interviews and historical literature, we found that all four populations have references to non-Indian origins. We observed that while the Kapla are genetically differentiated from the Bunt, Kodava, and Nair, all four populations are genetically most similar to other Indian populations. Our results demonstrate that oral and genetic histories may not always provide consistent accounts of population origins and suggest community-engaged, multi-disciplinary investigations to understand their relationship.
Background Asthma is the most common chronic disease in children, occurring at higher frequencies and with more severe disease in children with African ancestry. Methods We tested for association with haplotypes at the most replicated and significant childhood-onset asthma locus at 17q12-q21 and asthma in European American and African American children. Following this, we used whole-genome sequencing data from 1060 African American and 100 European American individuals to identify novel variants on a high-risk African American–specific haplotype. We characterized these variants in silico using gene expression and ATAC-seq data from airway epithelial cells, functional annotations from ENCODE, and promoter capture (pc)Hi-C maps in airway epithelial cells. Candidate causal variants were then assessed for correlation with asthma-associated phenotypes in African American children and adults. Results Our studies revealed nine novel African-specific common variants, enriched on a high-risk asthma haplotype, which regulated the expression of GSDMA in airway epithelial cells and were associated with features of severe asthma. Using ENCODE annotations, ATAC-seq, and pcHi-C, we narrowed the associations to two candidate causal variants that are associated with features of T2 low severe asthma. Conclusions Previously unknown genetic variation at the 17q12-21 childhood-onset asthma locus contributes to asthma severity in individuals with African ancestries. We suggest that many other population-specific variants that have not been discovered in GWAS contribute to the genetic risk for asthma and other common diseases.