African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.
Respiratory Syncytial Virus (RSV) remains a significant cause of respiratory illness in infants and older adults worldwide. The COVID-19 pandemic disrupted typical RSV fall/winter seasonality, leading to unusual patterns of viral circulation and resurgence. We investigated the evolutionary dynamics of RSV in Houston, Texas, over a nine-year period (2015–2024), encompassing pre-pandemic, pandemic, and post-pandemic phases. We sequenced 606 RSV/A and 570 RSV/B nasal swab or throat/nasal swabs samples from children outpatient clinic or in hospital with acute respiratory infections between November 1st, 2015, and Feb 1st, 2024, from the Houston site of the New Vaccine Surveillance Network. We assessed genetic diversity, lineage dynamics, and selective pressures, by phylogenetic analysis, variant calling, dN/dS ratio calculations and Shannon entropy. Phylogenetic analysis revealed distinct lineage dynamics, with RSV/A showing persistence of certain pre-pandemic lineages (e.g., A.D.1) and the emergence of new ones (e.g., A.D.3) during pandemic and post-pandemic periods. In contrast, RSV/B underwent a dramatic restructuring, with disappearance of pre-pandemic lineages and appearance of a dominant B.D.E.1 lineage in pandemic and post-pandemic periods. RSV/B exhibited higher genetic diversity and accumulated non-synonymous variants at nearly twice the rate of RSV/A, particularly in the M2-2 gene. The M2-2 gene in RSV/B showed a significant increase in non-synonymous mutations during the pandemic and post-pandemic periods, correlating with increased transcriptional activity. Mutations in antigenic site in the F protein, particularly in RSV/B, were also observed, with possible implications for immune evasion and resistance to treatments. The COVID-19 pandemic significantly impacted RSV evolution, leading to reduced genetic diversity during the pandemic and the emergence of novel lineages post-pandemic. RSV/B exhibited more dynamic evolutionary changes, particularly in the M2-2 gene, suggesting potential adaptive advantages. These findings highlight the importance of continued genomic surveillance to monitor the impact of emerging interventions, such as vaccines and monoclonal antibodies, on RSV evolution and public health. Pedro A. Piedra, MD, Gilead: Honoraria|GSK: Grant/Research Support|Icosavax: Grant/Research Support|Merck: Advisor/Consultant|Merck: Grant/Research Support|Moderna: Advisor/Consultant|Novavax: Grant/Research Support|Pfizer: Advisor/Consultant|Sanofi-Pasteur: Advisor/Consultant|Sanofi-Pasteur: Grant/Research Support|Shionogi: Advisor/Consultant|Shionogi: Grant/Research Support
Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.
Background:Accurate discrimination of true structural variants (SVs) from artifacts in long-read sequencing data remains a critical bottleneck. Numerous machine learning solutions have been proposed, ranging from classical models using engineered features to advanced deep learning and foundation model interpretability methods. However, a systematic comparison of their performance, efficiency, and practical utility is lacking. Results:We conducted a comprehensive benchmark of five machine learning paradigms for SV filtering using standardized Genome in a Bottle (GIAB) data for samples HG002 and HG005. We evaluated classical Random Forest classifiers on 15 genomic features, computer vision models (ResNet/VICReg), diffusion-based anomaly detection, sparse autoencoders (SAEs) on the Evo2-7B foundation model, and multimodal ensembles. A simple Random Forest on interpretable features achieved a peak F1-score of 95.7%, effectively matching all more complex models (ResNet50: 95.9%, Diffusion: 95.8%). This study represents the first application of diffusion-based anomaly detection and sparse autoencoders to structural variant analysis; while diffusion models learned highly discriminative, disentangled representations and SAEs uncovered biologically interpretable features (including atoms that were specific for ALU deletions, chromosome X variants and insertion events), they did not significantly surpass this classification ceiling. Ensemble methods offered no performance benefit but may have future potential given the orthogonality of vision-based and linear features. Conclusions:Our findings demonstrate that for the established task of germline SV filtering, simpler, interpretable models provide an optimal balance of accuracy, speed, and transparency. This benchmark establishes a pragmatic framework for method selection and argues that increased model complexity must be justified by clear, unmet biological needs rather than marginal predictive gains.
Long-read sequencing (LRS) technologies, namely, Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), have emerged as promising solutions to overcome the limitations of short-read sequencing (SRS). Nevertheless, the still higher sequencing error rates compared with SRS, need for customized pipelines, rapidly updating software, and incipient scalability continue to present challenges for adopting ONT in standard clinical practice. Here we assess the performance of ONT (R9 and R10 chemistries) in comparison to Illumina and MGI across 17 well-characterized reference samples with 11 clinical variants representing nine different genetic diseases. To enable this, we have implemented a production-ready pipeline including SNV, indel, STR, SV, and CNV detection, alongside reporting key summary metrics to ensure high-quality data at the production sequencing level. Our results show high accuracy of ONT across SNVs (F-score 0.978-0.983) and SVs (F-score = 0.75) but still weaknesses across indels (F-score 0.659-0.758). However, we highlight that ONT accurately detected all four pathogenic indels as well as the performance improvement in exons and with the newer R10 chemistry. We further demonstrated the importance of long reads to detect clinically impactful variants such as a FMR1 pathogenic expansion, often misclassified by SRS as being in the premutation range. Our multiplatform analysis and Sanger validation uncovered a 1 bp error in the Coriell annotation for a cystic fibrosis-causing indel in GM07829. This work underscores the growing readiness of ONT for clinical applications, highlighting both its advancements and its potential for broader adoption in clinical genomics and large-scale operations.
The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, Hi-C, and karyotyping. These data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic variants. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.
Cancer is fundamentally a disease of the genome, characterized by extensive genomic, transcriptomic, and epigenomic alterations. Most current studies predominantly use short-read sequencing, gene panels, or microarrays to explore these alterations; however, these technologies can systematically miss or misrepresent certain types of alterations, especially structural variants, complex rearrangements, and alterations within repetitive regions. Long-read sequencing is rapidly emerging as a transformative technology for cancer research by providing a comprehensive view across the genome, transcriptome, and epigenome, including the ability to detect alterations that previous technologies have overlooked. In this Perspective, we explore the current applications of long-read sequencing for both germline and somatic cancer analysis. We provide an overview of the computational methodologies tailored to long-read data and highlight key discoveries and resources within cancer genomics that were previously inaccessible with prior technologies. We also address future opportunities and persistent challenges, including the experimental and computational requirements needed to scale to larger sample sizes, the hurdles in sequencing and analyzing complex cancer genomes, and opportunities for leveraging machine learning and artificial intelligence technologies for cancer informatics. We further discuss how the telomere-to-telomere genome and the emerging human pangenome could enhance the resolution of cancer genome analysis, potentially revolutionizing early detection and disease monitoring in patients. Finally, we outline strategies for transitioning long-read sequencing from research applications to routine clinical practice.
Rare diseases are collectively common, affecting approximately 1 in 20 individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in next-generation sequencing, development of new computational and functional genomics approaches to prioritize genes and variants and increased global sharing of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Furthermore, all data generated, currently representing over 7,500 individuals from over 3,000 families, are rapidly made available to researchers worldwide through the Analysis, Visualization and Informatics Lab-space (AnVIL) to catalyse global efforts to develop approaches for genetic diagnoses in rare diseases. Most of these families have undergone previous clinical genetic testing but remained unsolved, with most being exome-negative. Here we describe the collaborative research framework, datasets and discoveries comprising GREGoR that will provide foundational resources and substrates for the future of rare disease genomics.
The advent of single cell DNA sequencing revealed astonishing dynamics of genomic variability, but failed at characterizing smaller to mid size variants that on the germline level have a profound impact. In this work we discover previously uncharacterized genomic dynamics in 18 cells from three human brains utilizing single cell long-read whole genome sequencing. This provides key insights into the dynamic of the genomes of individual cells and further highlights brain specific activity of transposable elements, but requires validation in larger studies.
Respiratory syncytial virus (RSV) is the leading cause of lower respiratory tract infections in children worldwide, while human noroviruses (HuNoV) are a leading cause of epidemic and sporadic acute gastroenteritis. Generating full-length genome sequences for these viruses is crucial for understanding viral diversity and tracking emerging variants. However, obtaining high-quality sequencing data is often challenging due to viral strain variability, quality, and low titers. Here, we present a set of comprehensive oligonucleotide probe sets designed from 1,570 RSV and 1,376 HuNoV isolate sequences in GenBank. Using these probe sets and a capture enrichment sequencing workflow, 85 RSV positive nasal swab samples and 55 (49 stool and six human intestinal enteroids) HuNoV positive samples encompassing major subtypes and genotypes were characterized. Samples with Ct values 17.0-29.9 for RSV, and 20.2-34.8 for HuNoV, with some HuNoV below the detection limit were sequenced. The percentage of reads mapped to viral genomes was 85.1% for RSV and 40.8% for HuNoV post-capture, compared to 0.08% and 1.15% in pre-capture libraries. Full-length genomes were obtained for all RSV positive samples and in 47/55 HuNoV positive samples-a significant improvement over genome recovery from pre-capture libraries. RSV transcriptome (subgenomic mRNAs) sequences were also characterized from this data.
Long-read sequencing has transformed metagenomics and improved the quality of metagenome-assembled genomes (MAGs). However, current binning methods struggle with identifying unknown species and managing imbalanced species distributions. Here, we present LorBin, an unsupervised binner specially designed to reconstruct MAGs in natural microbiomes. LorBin deploys a two-stage multiscale adaptive DBSCAN and BIRCH clustering with evaluation decision models using single-copy genes to maximize MAG recovery. LorBin outperforms six competing binners in both simulated and real microbiomes, including oral, gut, and marine samples. LorBin generated 15-189% more high-quality MAGs with high serendipity and identified 2.4-17 times more novel taxa than state-of-the-art binning methods. Together, LorBin is a promising long-read metagenomic binner for accessing species-rich samples containing unknown taxa and is efficient at retrieving more complete genomes from imbalanced natural microbiomes.
DNA methylation is a critical epigenetic mechanism in numerous biological processes, including gene regulation, development, ageing and the onset of various diseases such as cancer. Studies of methylation are increasingly using single-molecule long-read sequencing technologies to simultaneously measure epigenetic states such as DNA methylation with genomic variation. These long-read data sets have spurred the continuous development of advanced computational methods to gain insights into the roles of methylation in regulating chromatin structure and gene regulation. In this Review, we discuss the computational methods for calling methylation signals, contrasting methylation between samples, analysing cell-type diversity and gaining additional genomic insights, and then further discuss the challenges and future perspectives of tool development for DNA methylation research. Long-read sequencing technologies can directly profile methylation modifications across the genome. In this Review, Fu et al. overview the long-read computational tools to identify and compare methylation signals, as well as tools that use these methylation signals to analyse cell-type diversity and gain additional genomic insights.
Genomic structural variants (SVs) are a major source of genetic diversity in humans. Here, through long-read sequencing of 945 Han Chinese genomes, we identify 111,288 SVs, including 24.56% unreported variants, many with predicted functional importance. By integrating human population-level phenotypic and multi-omics data as well as two humanized mouse models, we demonstrate the causal roles of two SVs: one SV that emerges at the common ancestor of modern humans, Neanderthals, and Denisovans in GSDMD for bone mineral density and one modern-human-specific SV in WWP2 impacting height, weight, fat, craniofacial phenotypes and immunity. Our results suggest that the GSDMD SV could serve as a rapid and cost-effective biomarker for assessing the risk of cisplatin-induced acute kidney injury. The functional conservation from human to mouse and widespread signals of positive natural selection suggest that both SVs likely influence local adaptation, phenotypic diversity, and disease susceptibility across diverse human populations.
Variant calling using long-read RNA sequencing (lrRNA-seq) can be applied to diverse tasks, such as capturing full-length isoforms and gene expression profiling. It poses challenges, however, due to higher error rates than DNA data, the complexities of transcript diversity, RNA editing events, etc. In this paper, we propose Clair3-RNA, the first deep learning-based variant caller tailored for lrRNA-seq data. Clair3-RNA leverages the strengths of the Clair series' pipelines and incorporates several techniques optimized for lrRNA-seq data, such as uneven coverage normalization, refinement of training materials, editing site discovery, and the incorporation of phasing haplotype to enhance variant-calling performance. Clair3-RNA is available for various platforms, including PacBio and ONT complementary DNA sequencing (cDNA), and ONT direct RNA sequencing (dRNA). Our results demonstrated that Clair3-RNA achieved a ~91% SNP F1-score on the ONT platform using the latest ONT SQK-RNA004 kit (dRNA004) and a ~92% SNP F1-score in PacBio Iso-Seq and MAS-Seq for variants supported by at least four reads. The performance reached a ~95% and ~96% F1-score for ONT and PacBio, respectively, with at least ten supporting reads and disregarding the zygosity. With read phased, the performance reached ~97% for ONT and ~98% for PacBio. Extensive evaluation of various GIAB samples demonstrated that Clair3-RNA consistently outperformed existing callers and is capable of distinguishing RNA high-quality editing sites from variants accurately. Clair3-RNA is open-source and available at (https://github.com/HKU-BAL/Clair3-RNA).
Genetic mutations within select cells of a tissue, termed mosaic variants (MV), are being increasingly recognized for their role in human disease. This growing interest underscores the need for specialized tools to detect and analyze MVs. However, such detection methods still lack thorough evaluation, largely due to missing benchmarking datasets that are large, reliable, and reflective of the complexity of biological samples. To address this gap, we developed MosaicSim, a tool for simulating variants in realistic sequencing data. The TweakVar workflow is at the tool's core and represents a unique simulation pipeline that layers simulated MVs onto empirical whole genome sequencing data, generating a large, realistic ground truth dataset that combines the strengths of both simulation and biological data. To demonstrate the functionality of the workflow, we simulated 1,000 mosaic single nucleotide polymorphisms using TweakVar within whole genome sequencing files of different coverages. MVs were called with Illumina's DRAGEN and compared to the ground truth. Our results show 150×-445× coverage performed comparably, with a true-positive rate between 50.4% (300×) and 54.9% (150×) and no false-positives detected. Across all samples, increasing variant allele frequency had a significant positive effect on call success. Additionally, we observed that call rates for variants in lower complexity regions improved with increasing read depth. We did not find significant effects attributable to specific mutation patterns or mean read map quality. MosaicSim fills a critical unmet need by providing representative, customizable ground truth datasets for MV benchmarking, enabling systematic evaluation and optimization of variant calling methods.
Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs. We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to ~2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions ≥50bp was derived from haplotype-resolved assemblies. We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithm's sensitivity exceeded 80% for VAFs ≥4% and 15% for VAFs of 0.5-1% with 60x coverage. The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..
The extent of genetic variation and its influence on gene expression across multiple tissue and cellular contexts is still being characterized, with germline Structural Variants (SVs) being historically understudied. DNA methylation also represents a component of normal germline variation across individuals. Here, we combine germline SVs (by short-read sequencing) with tumor DNA methylation across 1292 pediatric brain tumor patients. For thousands of methylation probes for CpG Islands (CGIs) or enhancers, rare and common SV breakpoints upstream or downstream associate with differential methylation in tumors spanning various histologic types, a significant subset involving genes with SV-associated differential expression. Cancer predisposition genes involving SV-associated differential methylation and expression include MSH2, RSPA, and PALB2. SV breakpoints falling within CGIs or histone marks H3K36me3 or H3K9me3 associate with differential CGI methylation. Genes with SVs and CGI methylation associated with patient survival include POLD4. Our results capture a class of normal phenotypic variation having disease implications.