Meiotic recombination is a fundamental process that generates genetic diversity by creating new combinations of existing alleles1. Whereas crossovers in humans are well characterized2, the more frequent non-crossovers that lead to gene conversion remain challenging to study. Here we show that single high-fidelity long sequencing reads from sperm can capture both crossovers and non-crossovers, which enables effectively arbitrary sample sizes for analysis from a single male. We analysed 2,382 candidate non-crossovers in 15 sperm samples from 13 donors, and identified a consistent component with properties distinct from PRDM9-induced recombination. This phenomenon was not associated with meiotic double-strand break sites identified by DMC1 binding, the crossover recombination map or GC-biased gene conversion, but was associated with genomic fragile sites. This component is also seen in paternal non-crossover gene conversions in pedigree data3. Applying the same analysis to 12 blood samples4, we observed non-crossover gene conversions with similar properties, but very few crossover events. Further, we demonstrate variation between donors for the different types of recombination, even when they share the same PRDM9 genotype. We suggest that a substantial fraction of the non-crossover gene conversion events seen in sperm arise prior to meiosis.
Abstract Diptera genome evolution is traditionally viewed through the lens of Drosophila melanogaster . To provide an unbiased picture of chromosome evolution in Diptera, we reconstructed six ancestral linkage groups (ALGs) and their rearrangements using 340 chromosome-level genome assemblies from 59 dipteran Families and two outgroups. Notably, Diptera ALGs are conserved in many Nematoceran lineages, while emergence of Brachycera coincided with four chromosomal fissions, which subsequently fused in various combinations. In Schizophora and relatives, a stable karyotype of five metacentric chromosomes and a chromosome homologous to the dot emerged that is remarkably stable with the notable exception of Drosophila, where the metacentric chromosomes became acrocentric Muller elements. Furthermore, our reconstruction reveals an ancient sex chromosome system associated with a small and gene-poor ALG that frequently fused to other chromosomes.
Background:Notothenioids are a well characterised species flock endemic to the Antarctic and an important model group for the study of genome adaptation to extreme cold. We used a new reference assembly and clade-wide comparative genomic analysis to investigate cryonotothenioid evolution and the appearance of novel functionalities linked to cold adaptation. Results:A new phased assembly of a model notothenioid, Harpagifer antarcticus, demonstrated low levels of haplotypic variability across the genome. Nevertheless, numerous insertions from multiple LINE-L2 clades were found, suggesting ongoing transposition with potential contribution to speciation. Contrary to expectations the afgp locus was highly similar between haplotypes, except for large length allelic variants of afgp genes. Analysis suggests a model for the afgp locus expansion in H. antarcticus through segmental tandem duplications involving two pairs of afgp genes at time. Syntenic reconstruction of genomes from across the clade demonstrates conserved macrosyntenic relationships and group specific chromosomal fusions of notothenioids. Quantification of genome gain and transposition rates during cryonotothenioid diversification showed a first ancestral slow genome expansion concurrent with historic temperature drops. This was followed by lineage-specific massive peaks of genomic gain and transposition activity. Finally, we identified a set of genes that underwent ancestral diversifying selection and acquired novel conserved non-coding elements during the cryonotothenioid emergence. These were related to antioxidants and proteostasis, which may have facilitated the notothenioid Antarctic radiation. Conclusion:Diversifying selection and genomic gain linked to transposon activity are primary contributors to lineage-specific evolutionary dynamics through the clade which facilitated adaptation to life in the cold.
De novo genome assembly is challenging in highly repetitive regions; however, reference-guided assemblers often suffer from bias. We propose a framework for pangenome-guided sequence assembly that can resolve short-read data in complex regions without bias towards a single reference genome. Our primary contribution is to frame the assembly as a graph traversal optimization problem, which can be implemented classically or on a quantum computer. The workflow involves first annotating pangenome graphs with estimated copy numbers for each node, then finding a path on the graph that best explains those copy numbers. On simulated data, our approach significantly reduces the number of contigs compared with de novo assemblers. While they introduce a small increase in inaccuracies, such as false joins, our optimization-based methods are competitive with current exhaustive search techniques. They are also designed to scale more efficiently as the problem size grows and will run effectively on future quantum computers; a small experiment on a real quantum device showcases this behaviour. Moreover, they are more resilient to noise in copy number estimation inherent in short-read-based assembly. We also develop novel tools for creating realistic synthetic pangenomes, aligning reads to pangenomes and for evaluating assembly quality.
Abstract How evolutionary and developmental processes interact to determine axes of neural variation that produce behavioural diversity has been debated for many decades, with alternative hypotheses giving differential emphasis to functional coupling, which favours co-evolution, and developmental constraint, which enforces it. A critical omission is data on the genetic architecture of brain size and structure, which more closely illuminates the shared developmental dependencies between components of an integrated system. Here, we exploit ecological divergence between Astatotilapia calliptera and Aulonocara stuartgranti , two closely related cichlid species from Lake Malawi, to explore the genetic architecture of brain evolution. Using computer vision and machine learning techniques to extract volumetric data from micro-tomographic images, we first demonstrate significant divergence in brain composition between these species. Genomic and micro-tomographic imaging data from a population of hybrids generated between the two species were used to investigate genetic factors shaping this differentiation. We show that the majority of brain components are integrated phenotypically in hybrids, but genetic correlations between them are generally weaker. We further show that variation in multiple brain components is associated with variation in largely structure-specific quantitative trait loci, rather than determined by genetic factors with broad effects across the entire brain. These results suggest a genetic architecture that can facilitate modular changes in brain structure, and imply that individual components are independently evolvable.
Understanding the genetic basis of widespread phenotypic convergence, particularly for complex morphological traits, remains a major challenge in evolutionary biology. The Mediterranean gravel beach clingfishes of the genus Gouania provide an excellent system to study this phenomenon. Within this genus, two distinct morphotypes, "slender" and "stout," have repeatedly evolved, adapting to different microhabitats. These morphotypes differ in multiple complex traits, including body elongation, head compression, vertebral number, eye size, and the structure of the adhesive disc. First, to scrutinize phylogenetic convergence, we combined 3D morphometrics of the pelvic girdle and skull, with molecular species delimitation based on >660 DNA barcodes, and a phylogenomic framework based on more than 3,400 single-copy orthologs. Second, by employing whole-genome resequencing and a novel "convergence score" statistic, we examined genomic convergence across multiple levels: nucleotides, sequences, genes, and functional pathways. While we found no evidence of large-scale genomic or protein-level convergence, we identified promising candidate regions at the level of single variants, genes, and biological pathways. Notably, a longer shared (but interrupted) haplotype around the candidate gene adam12 was associated with convergent traits. The lack of simple genomic patterns may reflect the radiation's age and the complex genetic basis of the underlying morphological traits (eg eye size, neurocranium shape). Altogether, our findings highlight the importance of assessing genomic convergence at multiple molecular levels to uncover diagnostic signals across varying evolutionary processes and timescales.
Abstract Genetics may help address the biodiversity crisis by providing information about genetic diversity and temporal changes in demography for species of interest. Advances in whole-genome sequencing create new opportunities for demographic analysis, even based on the two copies of a genome found in a single diploid individual. The Vertebrate Genomes Project (VGP) is generating high-quality, chromosome-level reference genomes across the full range of extant vertebrate species, with its first phase delivering assemblies spanning approximately 95% of vertebrate orders. Using 512 diploid VGP genomes, we quantified intra-species heterozygosity, runs of homozygosity (ROH), and inferred past effective population sizes ( N e ) with the Pairwise Sequentially Markovian Coalescent (PSMC). Threatened species are more likely to exhibit lower heterozygosity and longer ROH, though there is large variation in both measures across all IUCN categories. Interestingly, PSMC suggests that estimated historical N e several thousand generations ago is a better predictor of threatened status than the present day estimate. Co-analysing with life history traits, we found that marine species tend to have lower ROH content, while fossorial species show significantly higher inbreeding levels. Indeed, habitat and foraging strata are much stronger predictors of IUCN status than genetics, with estimated historical N e providing a small but significant amount of additional information. Together, these results suggest that, while measures of genetic diversity are correlated with IUCN status, much of that correlation may derive from ecological factors such as habitat, with only a relatively small direct contribution. Nevertheless, reference genomes like those generated by the VGP can yield valuable information, like historical N e , while facilitating population monitoring and management for species of interest.
Recombination is central to genetics and to evolution of sexually reproducing organisms. However, obtaining accurate estimates of recombination rates, and of how they vary along chromosomes, continues to be challenging. To advance our ability to estimate recombination rates, we present Hi-reComb, a new method and software for estimation of recombination maps from bulk gamete chromosome conformation capture sequencing (Hi-C). Simulations show that Hi-reComb produces robust, accurate recombination landscapes. With empirical data from sperm of five fish species we show the advantages of this approach, including joint assessment of recombination maps and large structural variants, map comparisons using bootstrap, and workflows with trio phasing vs. Hi-C phasing. With off-the-shelf library construction and a straightforward rapid workflow, our approach will facilitate routine recombination landscape estimation for a broad range of studies and model organisms in genetics and evolutionary biology. Hi-reComb is open-source and freely available at https://github.com/millanek/Hi-reComb.
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
The black soldier fly (Hermetia illucens) is the main species in the developing global industry of insects as food and feed, but little is known about its natural diversity or the genetic basis of its domestication. We obtained whole-genome sequences for 54 individuals from both wild and captive populations. We identified two major genetic clusters at least 3 million years divergent, revealing cryptic diversity within the species. Our study indicates that the most common populations used for commercial and academic applications are primarily derived from just one of these sampled lineages, likely originating from a wild North American progenitor. We find that captive populations show strong reductions in genetic diversity, consistent with genome-wide effects of population bottlenecks and drift associated with rearing in captivity. Some limited evidence of gene flow between divergent lineages was observed, as well as evidence of hybridization from domesticated populations into the wild. Our study suggests that natural genetic diversity could provide important variation for industrial purposes in this novel agricultural species.
The Vertebrate Genomes Project (VGP) aims to produce complete and near-error-free reference genomes for all ~70,000 extant vertebrate species1. Organized in four phases, it progressively targets all vertebrate orders, families, genera, and eventually all species. Here we present the completion of VGP Phase I, delivering reference genomes for ~95% of vertebrate orders, along with additional lineages within those orders, totaling 816 species and 1.6 trillion base pairs of main haplotype sequence. These genomes were assembled and annotated over an 8-year period (2018-2026) of rapid advances in genome sequencing, assembly, and annotation methods2-4, alongside the growth of associated consortium initiatives and international collaborations5-9. They represent some of the highest-quality vertebrate genomes currently available, and most have become the primary reference for their respective species in public databases. Comparative analyses across a subset of 579 species when we reached a threshold of 85% of orders allowed us to reconstruct the genome of the last common ancestor of all vertebrates 500 million years ago, identify diverse modes of sex chromosome evolution, reveal clade-specific three-dimensional genome architecture, discover methylated epigenetic landscapes across vertebrates, and provide a framework for studying gene and pseudogene evolution, immune loci, cancer-associated genes, and other trait-associated loci. Approximately a quarter of this subset are listed as Vulnerable to Critically Endangered by the IUCN Red List of Threatened Species, and have enabled more advanced genomic investigations of extinction risk. VGP Phase I delivers a reference backbone for vertebrate genomics, enabling discoveries that would otherwise remain out of reach across evolution, conservation, and medicine.
Motivation:FastGA finds alignments between two genome sequences more than an order of magnitude faster than previous methods that have comparable sensitivity. Its speed is due to (i) a fully cache-local architecture involving only MSD radix sorts and merges, (ii) an algorithm for finding adaptive seed hits in a linear merge of sorted k-mer tables, and (iii) a variant of the Myers adaptive wave algorithm to find alignments around a chain of seed hits. It further stores alignments in a fraction of the space of a conventional CIGAR string using a trace-point encoding and our ONEcode data system introduced here. Results:For example, two 2 Gbp bat genomes are compared in 2.1 min with eight threads on an Apple laptop using 5.7 GB of memory and producing 1.05 million alignments covering 60% of each genome. Our ALN format file occupies 66 MB and in just 6 s can be converted to a standard 1.03 GB PAF file. Availability and implementation:FastGA is freely available at GitHub: http://www.github.com/thegenemyers/FASTGA along with utilities for viewing inputs, intermediates, and outputs and transforming ALN files to PSL or PAF with or without CIGAR strings and common formats. There is also a utility to chain FastGA's alignments and display them in a dot-plot view in PostScript files.
Anopheles funestus s.s. is a major human malaria vector across Africa. To study its evolution, especially under vector control pressure, we sequenced 656 modern specimens (collected 2014 to 2018) and 45 historic specimens (collected 1927 to 1967) from 16 African countries. Despite high genetic diversity, the species shows stable but considerable continental population structure. Although one population showed little differentiation over a century and 4000 kilometers, nearby, we found two genetically distinct ecotypes. Vector control has resulted in strong signals of selection, with some resistance alleles shared across populations through gene flow and others arising independently. Fortunately, we found that a promising gene drive target in Anopheles gambiae is highly conserved in An. funestus. These insights will enable more strategic insecticide usage and gene drive deployment, supporting malaria elimination.
Understanding the history of admixture events and population size changes leading to modern humans is central to human evolutionary genetics. Here we introduce a coalescence-based hidden Markov model, cobraa, that explicitly represents an ancestral population split and rejoin, and demonstrate its application on simulated and real data across multiple species. Using cobraa, we present evidence for an extended period of structure in the history of all modern humans, in which two ancestral populations that diverged ~1.5 million years ago came together in an admixture event ~300 thousand years ago, in a ratio of ~80:20%. Immediately after their divergence, we detect a strong bottleneck in the major ancestral population. We inferred regions of the present-day genome derived from each ancestral population, finding that material from the minority correlates strongly with distance to coding sequence, suggesting it was deleterious against the majority background. Moreover, we found a strong correlation between regions of majority ancestry and human–Neanderthal or human–Denisovan divergence, suggesting the majority population was also ancestral to those archaic humans.
Increasingly efficient methods for inferring the ancestral origin of genome regions are needed to gain insights into genetic function and history as biobanks grow in scale. Here we describe two near-linear time algorithms to learn ancestry harnessing the strengths of a Positional Burrows-Wheeler Transform. SparsePainter is a faster, sparse replacement of previous model-based 'chromosome painting' algorithms to identify recently shared haplotypes, whilst PBWTpaint uses further approximations to obtain lightning-fast estimation optimized for genome-wide relatedness estimation. The computational efficiency gains of these tools for fine-scale local ancestry inference offer the possibility to analyse large-scale genomic datasets using different approaches. Application to the UK Biobank shows that haplotypes better represent ancestries than principal components, whilst linkage-disequilibrium of ancestry identifies signals of recent changes to population-specific selection for many genomic regions associated with immune responses, suggesting avenues for understanding the pathogen-immune system interplay on a historical timescale.
Plant organelle genomes, particularly large mitochondrial genomes with complex repeats, present significant challenges for assembly. The advent of long-read sequencing enables the assembly of complete genomes, but problems of resolving alternative structures remain. Here we introduce a novel tool that employs a syncmer-based assembler for rapid assembly graph construction, integrates a profile-HMM database for robust organelle identification, and leverages a new search method to find the best supported path through the assembly graph. We describe high-quality organelle assemblies for 195 plant species, demonstrating improvements over other methods, and providing multiple insights into structural complexity, heteroplasmy, and DNA exchange between organelles.
The characterisation of mutational processes active in somatic and germline cells in vivo has predominantly focused on human cancers1,2, normal tissues3-5 and model organisms6-8. Beyond mammals9, little is known about mutational processes across the tree of life. We developed himut (high-fidelity mutation), an algorithm to identify somatic mutations from circular consensus long-read sequencing data, deploying it on 708 samples from 661 species from Britain and Ireland10. The spectra of somatic mutations, categorised by mutation type and local sequence context, showed considerable between-species divergence but within-species similarity. From normalised spectra across the dataset, we extracted 95 distinct patterns, or 'signatures', of somatic mutations, together with 18 signatures of germline mutational processes. Only two of the somatic signatures resembled mutational signatures extracted from human cancers1,2. Of the somatic signatures, 22 had significant clustering within the taxonomic classification, with some distributed across an entire kingdom or phylum, while others were restricted to a single family or species. Three somatic signatures were found only in water-dwelling species, with one distributed across multiple clades that might suggest it derives from a water-borne mutagen. Of germline signatures, eight were found in multiple species, showed significant clustering by taxonomic classification and had counterpart somatic signatures with matching mutational spectrum and species distribution. Thus, there is considerably greater diversity of mutational processes across the tree of life than found in human samples, likely reflecting either cell-intrinsic biological processes or environmental exposures (or both) operative within distinct clades. ### Competing Interest Statement S.L. is a shareholder and was previously an employee at Pacific Biosciences. P.J.C. is an employee and shareholder of Quotient Therapeutics Ltd. Wellcome Trust, https://ror.org/029chgv08, 218328, 206194, 220540/Z/20/A
BACKGROUND:East African cichlid fishes have diversified in an explosive fashion, but the (epi)genetic basis of the phenotypic diversity of these fishes remains largely unknown. Although transposable elements (TEs) have been associated with phenotypic variation in cichlids, little is known about their transcriptional activity and epigenetic silencing. We set out to bridge this gap and to understand the interactions between TEs and their cichlid hosts. RESULTS:Here, we describe dynamic patterns of TE expression in African cichlid gonads and during early development. Orthology inference revealed strong conservation of TE silencing factors in cichlids, and an expansion of piwil1 genes in Lake Malawi cichlids, likely driven by PiggyBac TEs. The expanded piwil1 copies have signatures of positive selection and retain amino acid residues essential for catalytic activity. Furthermore, the gonads of African cichlids express a Piwi-interacting RNA (piRNA) pathway that targets TEs. We define the genomic sites of piRNA production in African cichlids and find divergence in closely related species, in line with fast evolution of piRNA-producing loci. CONCLUSIONS:Our findings suggest dynamic co-evolution of TEs and host silencing pathways in the African cichlid radiations. We propose that this co-evolution has contributed to cichlid genomic diversity.
Pangenome methods have the potential to uncover hitherto undiscovered sequences missing from established reference genomes, making them useful to study evolutionary and speciation processes in diverse organisms. The cichlid fishes of the East African Rift Lakes represent one of nature's most phenotypically diverse vertebrate radiations, but single-nucleotide polymorphism (SNP)-based studies have revealed little sequence difference, with 0.1%-0.25% pairwise divergence between Lake Malawi species. These were based on aligning short reads to a single linear reference genome and ignored the contribution of larger-scale structural variants (SVs). We constructed a pangenome graph that integrates six new and two existing long-read genome assemblies of Lake Malawi haplochromine cichlids. This graph intuitively represents complex and nested variation between the genomes and reveals that the SV landscape is dominated by large insertions, many exclusive to individual assemblies. The graph incorporates a substantial amount of extra sequence across seven species, the total size of which is 33.1% longer than that of a single cichlid genome. Approximately 4.73% to 9.86% of the assembly lengths are estimated as interspecies structural variation between cichlids, suggesting substantial genomic diversity underappreciated in SNP studies. Although coding regions remain highly conserved, our analysis uncovers a significant proportion of SV sequences as transposable element (TE) insertions, especially DNA, LINE, and LTR TEs. These findings underscore that the cichlid genome is shaped both by small-nucleotide mutations and large, TE-derived sequence alterations, both of which merit study to understand their interplay in cichlid evolution.