Understanding how biodiversity arises and how organisms adapt to different environments is fundamental to evolutionary biology. Hybridization may play an essential role in generating genetic diversity and promoting adaptation. In this work, we analyzed the population structure, demographic history, and selective landscapes of East Asian domestic pigs using a whole-genome resequencing dataset of 1,092 samples from 43 breeds. Our results indicate that North and South pigs form two deeply divergent lineages that split approximately 17,797 years ago, with no extant wild boar population identified as the direct ancestor of South pigs. In contrast, Central and Southwest pigs originated from a North ancestral background (∼9,012 years ago) with introgression from South pigs before their divergence (∼6,284 years ago), which was likely driven by bidirectional migrations between ancient northern and southern human populations in China. The North ancestry under selection contributed to increased body weight, while the South ancestry enhanced environmental adaptation through pathways involved in UV-B response, stress tolerance, and extracellular matrix remodeling. This South-derived ancestry facilitated the geographic expansion of North pigs into broader ecological zones, which led to the establishment of the Central and Southwest pigs. Specifically, we demonstrate that North-South hybridization increases genetic diversity and produces a mosaic inheritance pattern. Moreover, hybridization and adaptive introgression contribute to novel phenotypic variation in domesticated animals, offering valuable insights for genetic improvement and selective breeding.
Accurate gene annotation is essential for deciphering the mapping from genomic sequences to their functional roles. However, current methods struggle to model complex gene transmission patterns, such as vertical inheritance and horizontal gene transfer. Here we introduce ANNEVO, a mixture of experts-based genomic language model that directly models distal sequence dependencies and joint evolutionary relationships from diverse genomes, enabling precise ab initio gene annotation. Through extensive benchmarking on 566 phylogenetically diverse species, we demonstrate that ANNEVO substantially outperforms existing ab initio methods and achieves performance comparable to state-of-the-art annotation pipelines. Furthermore, ANNEVO's independence from external evidence allows it to deliver more complete annotations than reference annotations for a broad range of species while correcting errors within them. These advancements will improve genome sequence interpretation and provide a framework capable of integrating evolutionary insights.
The Vertebrate Genomes Project (VGP) aims to produce complete and near-error-free reference genomes for all ~70,000 extant vertebrate species1. Organized in four phases, it progressively targets all vertebrate orders, families, genera, and eventually all species. Here we present the completion of VGP Phase I, delivering reference genomes for ~95% of vertebrate orders, along with additional lineages within those orders, totaling 816 species and 1.6 trillion base pairs of main haplotype sequence. These genomes were assembled and annotated over an 8-year period (2018-2026) of rapid advances in genome sequencing, assembly, and annotation methods2-4, alongside the growth of associated consortium initiatives and international collaborations5-9. They represent some of the highest-quality vertebrate genomes currently available, and most have become the primary reference for their respective species in public databases. Comparative analyses across a subset of 579 species when we reached a threshold of 85% of orders allowed us to reconstruct the genome of the last common ancestor of all vertebrates 500 million years ago, identify diverse modes of sex chromosome evolution, reveal clade-specific three-dimensional genome architecture, discover methylated epigenetic landscapes across vertebrates, and provide a framework for studying gene and pseudogene evolution, immune loci, cancer-associated genes, and other trait-associated loci. Approximately a quarter of this subset are listed as Vulnerable to Critically Endangered by the IUCN Red List of Threatened Species, and have enabled more advanced genomic investigations of extinction risk. VGP Phase I delivers a reference backbone for vertebrate genomics, enabling discoveries that would otherwise remain out of reach across evolution, conservation, and medicine.
Extracellular Vesicles (EVs) derived from mesenchymal stem cells (MSCs) have gained recognition as promising therapeutic and drug delivery agents in regenerative medicine. However, their clinical application is limited by donor variability, low scalability, and inconsistent therapeutic quality. To overcome these challenges, a robust and standardized production platform is urgently needed. We developed a scalable biomanufacturing strategy by generating and expanding MSCs from extended pluripotent stem cells (EPSC) using a suspension bioreactor culture system. A fixed-bed bioreactor was integrated for automated, continuous expansion of iMSCs and downstream EV harvesting. EVs were isolated through a streamlined protocol and characterized for size, morphology, surface markers, and bioactivity. Therapeutic efficacy was assessed in a bleomycin-induced pulmonary fibrosis mouse model. iMSC-derived EVs (iMSC-EVs) exhibited comparable characteristics to primary MSC-EVs, including a size distribution of 70–80 nm, cup-shaped morphology, and expression of canonical EV markers (CD63, CD81, TSG101). iMSCs were expanded for up to 20 days in 3D culture, yielding > 5 × 10⁸ cells per batch using a suspension bioreactor culture system and producing 1.2 × 10¹³ EV particles/day in a fixed-bed bioreactor. In vivo, iMSC-EVs significantly reduced Ashcroft fibrosis scores and bronchoalveolar lavage fluid protein levels in bleomycin-injured lungs, with therapeutic efficacy comparable to primary MSC-EVs. This study establishes a scalable and standardized platform for producing high-quality iMSC-EVs using bioreactor-based systems. Our approach addresses key limitations in traditional EV production and sets the stage for AI-integrated, fully automated, GMP-compliant manufacturing of therapeutic EVs suitable for clinical translation.
>Hedgehogs, small nocturnal mammals of the Erinaceinae subfamily, play a crucial role in maintaining ecological balance(Hernandez, 2008; Taucher et al., 2020). Atelerix albiventris (A. albiventris), a species native to West and Central Africa, is the smallest of the African hedgehogs. A. albiventris has undergone domestication and is utilized in biomedical research and offered in the exotic pet trade (Santana et al., 2010). In recent years, hedgehog population numbers have shown a declining trend due to human-induced disturbance (Johnson et al., 2015).
Nomascus leucogenys is a critically endangered species of small apes. Here, we sequenced and assembled the male genome of N. leucogenys, using PacBio and Hi-C datasets, with a particular focus on its Y-chromosome. The resulting high-quality haplotype-phased assemblies are at chromosome-scale, with scaffold/contig N50 values of 124.2/102.2 Mb for Haplotype 1 and 121.2/85.67 Mb for Haplotype 2. The assembled Y-chromosome spans 16.06 Mb. BUSCO assessment indicated completeness scores exceeding 95%. We predicted 18,925 protein-coding genes (23,783 mRNAs), including 58 genes on the Y-chromosome. Approximately 50% of the genome comprises repetitive elements. These comprehensive genome datasets will serve as a valuable resource for future studies on the genetics and protection of gibbons and improve our understanding on the evolution of Y-chromosome-related genes in primates.
The epigenetic landscape of cancer is regulated by many factors, but primarily it derives from the underlying genome sequence. Chromothripsis is a catastrophic localized genome shattering event that drives, and often initiates, cancer evolution. We characterized five esophageal adenocarcinoma organoids with chromothripsis using long-read sequencing and transcriptome and epigenome profiling. Complex structural variation and subclonal variants meant that haplotype-aware de novo methods were required to generate contiguous cancer genome assemblies. Chromosomes were assembled separately and scaffolded using haplotype-resolved Hi-C reads, producing accurate assemblies even with up to 900 structural rearrangements. There were widespread differences between the chromothriptic and wild-type copies of chromosomes in topologically associated domains, chromatin accessibility, histone modifications, and gene expression. Differential epigenome peaks were most enriched within 10 kb of chromothriptic structural variants. Alterations in transcriptome and higher-order chromosome organization frequently occurred near differential epigenetic marks. Overall, chromothripsis reshapes gene regulation, causing coordinated changes in epigenetic landscape, transcription, and chromosome conformation.
Long-range sequencing grants insight into additional genetic information beyond what can be accessed by both short reads and modern long-read technology. Several new sequencing technologies, such as "Hi-C" and "Linked Reads", produce long-range datasets for high-throughput and high-resolution genome analyses, which are rapidly advancing the field of genome assembly, genome scaffolding, and more comprehensive variant identification. In this review, we focused on five major long-range sequencing technologies: high-throughput chromosome conformation capture (Hi-C), 10X Genomics Linked Reads, haplotagging, transposase enzyme linked long-read sequencing (TELL-seq), and single- tube long fragment read (stLFR). We detailed the mechanisms and data products of the five platforms and their important applications, evaluated the quality of sequencing data from different platforms, and discussed the currently available bioinformatics tools. This work will benefit the selection of appropriate long-range technology for specific biological studies.
Transmissible cancers are malignant cell lineages that spread clonally between individuals. Several such cancers, termed bivalve transmissible neoplasia (BTN), induce leukemia-like disease in marine bivalves. This is the case of BTN lineages affecting the common cockle, Cerastoderma edule, which inhabits the Atlantic coasts of Europe and northwest Africa. To investigate the evolution of cockle BTN, we collected 6,854 cockles, diagnosed 390 BTN tumors, generated a reference genome and assessed genomic variation across 61 tumors. Our analyses confirmed the existence of two BTN lineages with hemocytic origins. Mitochondrial variation revealed mitochondrial capture and host co-infection events. Mutational analyses identified lineage-specific signatures, one of which likely reflects DNA alkylation. Cytogenetic and copy number analyses uncovered pervasive genomic instability, with whole-genome duplication, oncogene amplification and alkylation-repair suppression as likely drivers. Satellite DNA distributions suggested ancient clonal origins. Our study illuminates long-term cancer evolution under the sea and reveals tolerance of extreme instability in neoplastic genomes.
Numerous novel adaptations characterise the radiation of notothenioids, the dominant fish group in the freezing seas of the Southern Ocean. To improve understanding of the evolution of this iconic fish group, here we generate and analyse new genome assemblies for 24 species covering all major subgroups of the radiation, including five long-read assemblies. We present a new estimate for the onset of the radiation at 10.7 million years ago, based on a time-calibrated phylogeny derived from genome-wide sequence data. We identify a two-fold variation in genome size, driven by expansion of multiple transposable element families, and use the long-read data to reconstruct two evolutionarily important, highly repetitive gene family loci. First, we present the most complete reconstruction to date of the antifreeze glycoprotein gene family, whose emergence enabled survival in sub-zero temperatures, showing the expansion of the antifreeze gene locus from the ancestral to the derived state. Second, we trace the loss of haemoglobin genes in icefishes, the only vertebrates lacking functional haemoglobins, through complete reconstruction of the two haemoglobin gene clusters across notothenioid families. Both the haemoglobin and antifreeze genomic loci are characterised by multiple transposon expansions that may have driven the evolutionary history of these genes.
: In this position paper we demonstrate our ongoing efforts to develop and test a number of statistical tools and methedologies which allow us to study the underlying statistical properties of a genetic sequence which has undergone chromothripsis, and hence provide some novel probes into the mechanisms which cause such catastrophic genomic rearrangement. Using these tools, we study an oesophogeal cancer sample showing more than 1000 rearrangements, with 800 of these on chromosome 6. By studying this chromosome, we challenge a prevalent idea within the literature: that chromothripsis breakpoints are non-random, finding instead that despite a high degree of clustering, the clusters themselves are uniformly distributed across the chromosome. We also show that although 3-dimensional proximity is a tempting explanation for the rearrangement pattern, the statistical evidence does not favour it at the current time. In addition, we attempt to disambiguate some of the terminology surrounding chromothripsis.
Tasmanian devils have spawned two transmissible cancer lineages, named devil facial tumor 1 (DFT1) and devil facial tumor 2 (DFT2). We investigated the genetic diversity and evolution of these clones by analyzing 78 DFT1 and 41 DFT2 genomes relative to a newly assembled, chromosome-level reference. Time-resolved phylogenetic trees reveal that DFT1 first emerged in 1986 (1982 to 1989) and DFT2 in 2011 (2009 to 2012). Subclone analysis documents transmission of heterogeneous cell populations. DFT2 has faster mutation rates than DFT1 across all variant classes, including substitutions, indels, rearrangements, transposable element insertions, and copy number alterations, and we identify a hypermutated DFT1 lineage with defective DNA mismatch repair. Several loci show plausible evidence of positive selection in DFT1 or DFT2, including loss of chromosome Y and inactivation of MGA , but none are common to both cancers. This study reveals the parallel long-term evolution of two transmissible cancers inhabiting a common niche in Tasmanian devils.
EDITORIAL article Front. Plant Sci., 06 September 2023Sec. Plant Bioinformatics Volume 14 - 2023 | https://doi.org/10.3389/fpls.2023.1273769
Transmissible cancers are malignant cell clones that spread among individuals through transfer of living cancer cells. Several such cancers, collectively known as bivalve transmissible neoplasia (BTN), are known to infect and cause leukaemia in marine bivalve molluscs. This is the case of BTN clones affecting the common cockle, Cerastoderma edule , which inhabits the Atlantic coasts of Europe and north-west Africa. To investigate the origin and evolution of contagious cancers in common cockles, we collected 6,854 C. edule specimens and diagnosed 390 cases of BTN. We then generated a reference genome for the species and assessed genomic variation in the genomes of 61 BTN tumours. Analysis of tumour-specific variants confirmed the existence of two cockle BTN lineages with independent clonal origins, and gene expression patterns supported their status as haemocyte-derived marine leukaemias. Examination of mitochondrial DNA sequences revealed several mitochondrial capture events in BTN, as well as co-infection of cockles by different tumour lineages. Mutational analyses identified two lineage-specific mutational signatures, one of which resembles a signature associated with DNA alkylation. Karyotypic and copy number analyses uncovered genomes marked by pervasive instability and polyploidy. Whole-genome duplication, amplification of oncogenes CCND3 and MDM2 , and deletion of the DNA alkylation repair gene MGMT , are likely drivers of BTN evolution. Characterization of satellite DNA identified elements with vast expansions in the cockle germ line, yet absent from BTN tumours, suggesting ancient clonal origins. Our study illuminates the evolution of contagious cancers under the sea, and reveals long-term tolerance of extreme instability in neoplastic genomes.
Genome sequences are computationally assembled from millions of much shorter sequencing reads. Although this process can be impressively accurate with long reads, it is still subject to a variety of types of errors, including large structural misassembly errors in addition to localised base pair substitutions. Recent advances in long single molecule sequencing in combination with other long-range technologies such as synthetic long read clouds and Hi-C have dramatically increased the contiguity of assembly. This makes it all the more important to be able to validate the structural integrity of the chromosomal scale assemblies now being generated. Here we describe a novel assembly evaluation tool, Asset, which evaluates the consistency of a proposed genome assembly with multiple primary long-range data sets, identifying both supported regions and putative structural misassemblies. We present tests on three de novo assemblies from a human, a goat and a fish species, demonstrating that Asset can identify structural misassemblies accurately by combining regionally supported evidence from long read and other raw sequencing data. Not only can Asset be used to assess overall assembly confidence, and discover specific problematic regions for downstream genome curation, a process that leads to improvement in genome quality, but it can also provide feedback to automated assembly pipelines.
The STORR gene fusion event is considered essential for the evolution of the promorphinan/morphinan subclass of benzylisoquinoline alkaloids (BIAs) in opium poppy as the resulting bi-modular protein performs the isomerization of (S)- to (R)-reticuline essential for their biosynthesis. Here, we show that of the 12 Papaver species analysed those containing the STORR gene fusion also contain promorphinans/morphinans with one important exception. P. californicum encodes a functionally conserved STORR but does not produce promorphinans/morphinans. We also show that the gene fusion event occurred only once, between 16.8-24.1 million years ago before the separation of P. californicum from other Clade 2 Papaver species. The most abundant BIA in P. californicum is (R)-glaucine, a member of the aporphine subclass of BIAs, raising the possibility that STORR, once evolved, contributes to the biosynthesis of more than just the promorphinan/morphinan subclass of BIAs in the Papaveraceae.
Perilla is a young allotetraploid Lamiaceae species widely used in East Asia as herb and oil plant. Here, we report the high-quality, chromosome-scale genomes of the tetraploid ( Perilla frutescens ) and the AA diploid progenitor ( Perilla citriodora ). Comparative analyses suggest post Neolithic allotetraploidization within 10,000 years, and nucleotide mutation in tetraploid is 10% more than in diploid, both of which are dominated by G:C → A:T transitions. Incipient diploidization is characterized by balanced swaps of homeologous segments, and subsequent homeologous exchanges are enriched towards telomeres, with excess of replacements of AA genes by fractionated BB homeologs. Population analyses suggest that the crispa lines are close to the nascent tetraploid, and involvement of acyl-CoA: lysophosphatidylcholine acyltransferase gene for high α-linolenic acid content of seed oil is revealed by GWAS. These resources and findings provide insights into incipient diploidization and basis for breeding improvement of this medicinal plant.
Understanding SARS-CoV-2 evolution and host immunity is critical to control COVID-19 pandemics. At the core is an arms-race between SARS-CoV-2 antibody and angiotensin-converting enzyme 2 (ACE2) recognition, a function of the viral protein spike. Mutations in spike impacting antibody and/or ACE2 binding are appearing worldwide, with the effect of mutation synergy still incompletely understood. We engineered 25 spike-pseudotyped lentiviruses containing individual and combined mutations, and confirmed that E484K evades antibody neutralization elicited by infection or vaccination, a capacity augmented when complemented by K417N and N501Y mutations. In silico analysis provided an explanation for E484K immune evasion. E484 frequently engages in interactions with antibodies but not with ACE2. Importantly, we identified a novel amino acid of concern, S494, which shares a similar pattern. Using the already circulating mutation S494P, we found that it reduces antibody neutralization of convalescent and post-immunization sera, particularly when combined with E484K and N501Y. Our analysis of synergic mutations provides a landscape for hotspots for immune evasion and for targets for therapies, vaccines and diagnostics. One-Sentence Summary Amino acids in SARS-CoV-2 spike protein implicated in immune evasion are biased for binding to neutralizing antibodies but dispensable for binding the host receptor angiotensin-converting enzyme
Background Efficient and effective genome scaffolding tools are still in high demand for generating reference-quality assemblies. While long read data itself is unlikely to create a chromosome-scale assembly for most eukaryotic species, the inexpensive Hi-C sequencing technology, capable of capturing the chromosomal profile of a genome, is now widely used to complete the task. However, the existing Hi-C based scaffolding tools either require a priori chromosome number as input, or lack the ability to build highly continuous scaffolds. Results We design and develop a novel Hi-C based scaffolding tool, pin_hic, which takes advantage of contact information from Hi-C reads to construct a scaffolding graph iteratively based on N-best neighbors of contigs. Subsequent to scaffolding, it identifies potential misjoins and breaks them to keep the scaffolding accuracy. Through our tests on three long read based de novo assemblies from three different species, we demonstrate that pin_hic is more efficient than current standard state-of-art tools, and it can generate much more continuous scaffolds, while achieving a higher or comparable accuracy. Conclusions Pin_hic is an efficient Hi-C based scaffolding tool, which can be useful for building chromosome-scale assemblies. As many sequencing projects have been launched in the recent years, we believe pin_hic has potential to be applied in these projects and makes a meaningful contribution.