
The two closely related nematode species Caenorhabditis nigoni and Caenorhabditis briggsae, are commonly used to study the evolution of reproductive modes in animals, with the self-fertile C. briggsae and outcrossing C. nigoni sharing a common ancestor ~3.5 million years ago. Earlier genomic analyses revealed that selfing Caenorhabditis species have smaller genomes and proposed that at least some gene loss in C. briggsae is adaptive. However, the incomplete C. nigoni reference genome has limited most comparative analyses to genic regions. Here, we leverage long-read sequencing to generate and annotate telomere-to-telomere (T2T) assemblies for the C. nigoni strain JU1422 and the C. briggsae strain AF16. This new 139 Mb C. nigoni genome resolves 57 gaps and 149 unassigned scaffolds from the previous genome assembly. A major driver of the size difference with the 107 Mb T2T C. briggsae genome is the abundance of satellite DNA, which accounts for 12.8 Mb (9.2%) in C. nigoni and only 3. 2Mb (3.0%) in C. briggsae Notably, the C. nigoni X Chromosome is 13.4 Mb larger than in the previous assembly, making it 60% larger than the C. briggsae X Chromosome compared with 18-26% difference for the autosomes. We also document a surprising degree of plasticity in the ribosomal DNA, with the C. nigoni X Chromosome harboring a second 45S rDNA array that is absent in C. briggsae The hitherto undocumented divergence in the abundance of repetitive DNA elements makes the new genomes an invaluable resource for genomic analysis.
Small nuclear RNAs (snRNAs) are essential components of the spliceosome and are encoded by large, multicopy gene families. However, their genome-wide identification and quantification have remained challenging due to high sequence similarity among family members. To address this, we utilized RAMPAGE (Rapid Amplification of cDNA Ends) data from the ENCODE project to comprehensively profile nascent transcription of spliceosomal snRNAs across 115 human biosamples. We identified 74 expressed snRNA variants, characterized by canonical promoter features including bidirectional transcription flanking a positioned nucleosome, active histone modifications, and evolutionary conservation- features largely absent from unexpressed variants. These transcriptional events were corroborated by total RNA-seq and Bru-seq data, yet the majority of these variants showed extremely low levels in small RNA-seq, indicating post-transcriptional bottlenecks for snRNA processing and maturation. Our findings reveal new layers of regulation in snRNA variant expression and suggest that selective post-transcriptional processing plays a critical role in shaping the functional snRNA repertoire and its contribution to splicing regulation.
Inversions are a unique type of balanced structural variants (SVs) with important consequences in multiple organisms. However, despite considerable effort, these and other complex SVs remain poorly characterized due to the presence of large repeats. New techniques are finally allowing us to identify the full spectrum of human inversions, but the number of individuals analyzed is still quite limited. Here, we take advantage of Oxford Nanopore Technologies (ONT) long reads to characterize an exhaustive catalogue of 612 candidate inversions between 197 bp and 4.4 Mb of length and flanked by <190-kb long inverted repeats (IRs). For that, we have developed a bioinformatic package to identify inversion alleles reliably from long-read data. Next, using a combination of different DNA extraction, library preparation, and ONT sequencing protocols, we show that ultra-long reads (50-100 kb) and adaptive sampling are an efficient method to detect most human inversions. Lastly, by analyzing ONT data from 54 diverse individuals, 87-99% of the inversions can be genotyped in each sample, depending mainly on read and IR length and genome coverage. Both orientations have been observed for 155 of the analyzed regions (frequency 0.01-0.49), which multiplies by three the number of polymorphic IR-mediated inversions studied in detail so far. Moreover, we have found more than 300 additional independent SVs in the studied regions and resolved several complex rearrangements. Therefore, our work provides an accurate benchmark of those inversions that typically escape most analyses, and it demonstrates the potential of nanopore sequencing to characterize missing human genomic variation.
Hybridization between species is an important force in evolution, commonly modeled by the network multispecies coalescent. Reconstructing evolutionary histories under this model is computationally challenging, even for level-1 networks where hybridization events are isolated. Divide-and-conquer is a promising path forward, but current methods with statistical guarantees rely on an estimated tree of blobs (TOB) for the network, which compresses each nontree-like part into a single vertex. TOB reconstruction is itself challenging, with the only available method TINNiK having time complexity O(n 5 + n 4 k) for k genes and n species. Here, we present a new framework for scalable TOB reconstruction with statistical guarantees. Our approach operates by (1) seeking a refinement of the TOB and then (2) contracting edges in it. For step (1), we show that any optimal solution to Weighted Quartet Consensus is a TOB refinement almost surely, as the number of genes goes to infinity, motivating the use of methods, such as ASTRAL or TREE-QMC. For step (2), we show that applying the same hypothesis tests as TINNiK to just O(n) four-taxon subsets around each edge is sufficient for statistically consistent TOB reconstruction when the underlying network is level-1. Leveraging TREE-QMC for the first step gives our method time complexity O(n 3 k) and its name: TOB-QMC. On simulated data, TOB-QMC typically matches or exceeds TINNiK in accuracy while being more scalable. TOB-QMC also enables fast exploration of nontreelike evolution, as demonstrated through reanalysis of three phylogenomic data sets. Lastly, our study clarifies the theoretical utility of quartet-based species tree methods in the context of hybridization, which is critical given the recent result that ASTRAL can be misleading.
Accurate cell type annotation is essential for revealing the dynamic, cell type-specific accessibility of regulatory elements from single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. However, unlike the more mature single-cell RNA-sequencing (scRNA-seq) cell type annotation workflows, scATAC-seq cell type annotation remains challenging due to extreme sparsity, high dimensionality, the scarcity of labelled scATAC references, and pronounced batch effects across datasets. To enable annotation without relying on extensive scATAC labels, we introduce CARA, a cross-omics Bayesian framework that transfers cell type knowledge from scRNA-seq to scATAC-seq. CARA simultaneously enables cell type annotation, batch correction, and detection of cell types absent from the reference by combining pretraining on scRNA-seq data and semisupervised learning on target scATAC-seq data, along with distribution alignment, dynamic class weighting, and novel cell type detection. Across diverse benchmark datasets, CARA consistently outperforms baseline methods, remaining robust to batch effects. CARA's cross-modal alignment and robust annotation strategy preserve fine-grained lineage structure, enabling reconstruction of the hematopoietic differentiation trajectory. Through multidimensional uncertainty and latent-space clustering, CARA identifies novel, rare, or aberrant populations outside the reference cell type space, providing candidates for further biological validation and perturbation. Using an omics-specific generative framework and distribution alignment, CARA delivers accurate knowledge transfer and detects novel cell types in single-cell DNA methylation data, demonstrating seamless extensibility to new modalities. Ultimately, CARA offers a powerful and flexible solution for cross-modal cell type annotation in complex single-cell settings, facilitating the discovery of novel cell types and mechanistic insight into cell type-specific regulation across diverse analyses.
Topologically associating domains (TADs) are fundamental units of 3D genome architecture that shape gene regulation. Comparative analyses of TADs across biological conditions have revealed their involvement in development and disease. However, accurately identifying differential TADs from low sequencing depth and pseudo-bulk chromatin contact maps remains challenging. Here, we present HiDT, a graph neural network-based algorithm with an attention-based, edge-enhanced layer to capture structural differences between TADs. HiDT integrates a depth-specific normalization module and is trained across a wide range of sequencing depths, enabling robust detection of differential TADs under low sequencing depth conditions. Comprehensive benchmarking demonstrates that HiDT consistently outperforms existing methods at both TAD and subTAD levels, maintaining accuracy even in datasets with only a few million contacts. We further apply it to multiple low sequencing depth and pseudo-bulk datasets that are challenging for existing methods, revealing TAD reorganization linked to oncogene dysregulation during tumor progression, capturing differential TADs associated with underlying transcriptional heterogeneity in single-cell Hi-C data, and identifying haplotype-specific TADs associated with allele-specific structural variations. Overall, HiDT provides a robust tool for differential TAD analysis and facilitates insights into chromatin structure-function relationships.
The widespread transcriptomic diversity driven by alternative splicing (AS) contributes to all hallmarks of cancer and represents a critical source of neoantigens for personalized immunotherapy. However, unlike other major malignancies, the full repertoire of AS in nasopharyngeal carcinoma (NPC) remains underexplored. Here, we employ long-read sequencing (LR-seq) to generate a high-resolution, isoform-level transcriptomic atlas from a cohort of 14 NPC tumor samples and four immortalized nasopharyngeal epithelial cell lines. We identify a substantial number of full-length novel transcripts (22,687; ∼44.38%), which reveal diverse splicing patterns and previously unannotated splicing events. By integrating short-read RNA-seq data to quantify isoform expression, we discover a subset of novel transcripts that are differentially expressed between tumor samples and immortalized nasopharyngeal epithelial cell lines. Furthermore, LR-seq enables precise identification of chimeric readthrough fusion transcripts, such as CLDN15-FIS1 and FOXRED2-TXN2 Finally, we develop a computational framework, tumor-specific splicing neoantigen detection (TS-SNAD), to predict neoantigens originating from novel exon-exon junctions (neojunctions) in tumor-specific novel transcripts. Using this framework, we identify neojunction-derived neoantigens and experimentally validate the immunogenicity of selected HLA-B*40:01-restricted neoantigens. These neojunction-derived peptides constitute a new class of noncanonical neoantigens with significant potential for developing personalized cancer vaccines for NPC.
Single-cell RNA sequencing (scRNA-seq) pipelines rely on the assumption that sequencing reads possess correct structural architecture, a premise we show is incomplete. Standard quantification tools treat errors exclusively as base mismatches, failing to identify structural aberrations arising from off-target priming or nonspecific amplification. We demonstrate that these pervasive artifacts, reads lacking essential anchor motifs like poly(T) tracts, linkers, or template-switching oligos, generate large numbers of spurious barcodes, artificially inflate cell counts, and substantially affect biological interpretation. To resolve this, we developed Pattern-Filter, a universal preprocessing tool that systematically validates read integrity before alignment. It functions by detecting platform-specific anchor sequences and applying strict base-composition filtering to ensure barcodes and UMIs contain only canonical nucleotides. When applied across diverse platforms, including 10x Genomics, Drop-seq, BD Rhapsody, and SPLiT-seq, Pattern-Filter systematically removes 2%-18% of total reads yet reduces spurious barcode diversity by up to 80%. This asymmetric reduction confirms that a small fraction of invalid reads drives the majority of technical noise, compromising cluster stability. Consequently, this targeted removal enhances data reproducibility and recovers biologically relevant cell types, such as dopaminergic neurons in mouse striatum, which were previously obscured by artifact-induced noise. These findings establish structural validation as an essential prerequisite for analysis, positioning Pattern-Filter as a useful standard for ensuring molecular fidelity and reliable biological discovery in single-cell transcriptomics.
Adolescent idiopathic scoliosis (AIS) is a common pediatric musculoskeletal disorder characterized by lateral spinal curvature, often leading to chronic pain and deformity. While a significant genetic component to AIS is recognized, the functional impact of most associated genetic variants, particularly those in noncoding regions, remains largely unknown. Using massively parallel reporter assays, we examined the regulatory activity of 1,664 variants in linkage disequilibrium with 26 AIS lead variants identified through genome-wide association studies (GWAS). Candidate regulatory sequences containing the reference or alternative alleles were tested in two human chondrocyte cell lines, TC28a2 and SW1353, as chondrocytes are a major cell type implicated in AIS pathogenesis. We identified 92 variants with significant allele-specific regulatory activity, 79 of which are predicted to disrupt transcription factor binding sites, often correlating with their observed regulatory effect. Notably, we validate rs9496392, a single-nucleotide variant near the ADGRG6 locus, which shows consistent differential regulatory activity in both cell lines. ADGRG6 is a key regulator of cartilage homeostasis, and its cartilage-specific knockout in mice results in a scoliosis-like phenotype. The AIS risk allele of rs9496392 (T) is predicted to strongly disrupt several TFBSs, including SP1. This study provides a foundational catalog of functional AIS-associated regulatory variants active in chondrocytes, offering crucial insights into the perturbed gene regulatory networks in AIS. These findings lay the groundwork for identifying biomarkers and potential therapeutic targets for this complex childhood disease.
Spatial omic technologies have revolutionized tissue analysis by enabling multimodal molecular coprofiling within their native tissue context. Integrating multislice spatial multiomic data offers unprecedented opportunities to reconstruct three-dimensional (3D) tissue landscapes from multimodal molecular perspectives. However, current spatial omic integration methods remain narrowly focused on either vertical (cross-omic) or horizontal (cross-slice) integration, leaving a critical gap for a unified framework that simultaneously addresses both dimensions. Here we present SpatialMOSI, a unified framework for mosaic integration that concurrently resolves cross-modality and cross-section variations. At its core, SpatialMOSI employs a hierarchical graph contrastive learning (HiGCL) strategy that coordinates three integrative objectives: cross-omic alignment and fusion, cross-slice batch correction, and spatial microenvironment preservation. This approach operates on modality-specific latent representations while maintaining feature fidelity through decoding reconstruction. We demonstrate SpatialMOSI's versatility across multiple biological systems, accurately identifying spatially conserved domains, imputing missing omic layers, revealing B cell dynamics in germinal centers, delineating tumor-immune interactions, and reconstructing embryonic developmental trajectories. SpatialMOSI provides a critical computational foundation for constructing integrative 3D molecular atlases from complex multimodal spatial data sets.
Advances in single-cell sequencing technologies greatly enhance our understanding of molecular and cellular features. However, effectively leveraging these data to uncover key biological factors remains a major challenge in integrative analyses across multiomic types and comparative studies across species, particularly livestock species such as pigs and cattle. To address this, we develop AlignCell, a deep learning model designed to learn robust biological features by integrating multisource omic data across platforms, omic types, and species, thereby facilitating the discovery of key factors, such as conserved and species-specific genes in cross-species comparative studies. Across various applications and benchmarking compared with existing tools, AlignCell performs well. Notably, using AlignCell to integrate female gonad data across four species (human, mouse, pig, and cattle), including the bovine single-cell data generated in this study, AlignCell reveals the unexplored species-conserved gene CCT2 in primordial germ cells (PGCs). Additionally, it identifies unexplored species-specific genes PRICKLE4 and CTSV in pig and cattle PGCs, providing important insights for reproductive and developmental research.
Transfer RNAs (tRNAs) are central to protein synthesis and are increasingly recognized as dynamic regulators of gene expression whose abundance and chemical modifications are subject to precise biological control. Here, we systematically investigate how two distinct dietary interventions, low-protein and high-fat diets, reshape the tRNA landscape across multiple mouse tissues, using RNA mass spectrometry and Ordered Two-Template Relay sequencing (OTTR-seq) to comprehensively profile cytosolic and mitochondrial tRNAs at single-nucleotide resolution. We reveal pronounced tissue-specific biases in tRNA isodecoder expression, including the unexpected presence of full-length cytosolic tRNAs in mature sperm with a distinct isotype composition. In somatic tissues such as liver and heart, dietary conditions alter both tRNA abundance and key modifications known to regulate decoding efficiency, whereas in reproductive tissues diet primarily affects the abundance of select tRNAs with comparatively limited changes in modification profiles. We further demonstrate that mitochondrial tRNAs are subject to diet-responsive changes in both abundance and modification status, and that even subtle differences in dietary fat composition are sufficient to alter tRNA modification signatures. Together, these findings establish the tRNA epitranscriptome as a sensitive and tissue-specific sensor of nutritional state, and provide a resource for understanding how dietary cues interface with translational regulation in somatic and reproductive tissues.
Partial order alignment (POA) has emerged as a fundamental component in long-read error correction, assembly and pangenomics. However, conventional POA algorithms are limited by high time and memory requirements, making them inefficient for large-scale data sets. Here, we present minipoa, a fast and memory-efficient POA tool that incorporates seed-chain-align heuristics, adaptive or static banding strategies, and single-instruction multiple-data optimizations. Minipoa achieves up to a fivefold speedup over abPOA, reduces memory usage by up to 16-fold, and improves correction accuracy, while maintaining strong performance on both Pacific Biosciences and Oxford Nanopore Technologies simulated data sets, and can be readily integrated into existing long-read error correction and assembly workflows. In multiple sequence alignment data sets, minipoa demonstrates highly competitive computational efficiency and alignment accuracy, achieving total column scores up to 2.5-fold higher compared with MAFFT in low-similarity scenarios. Moreover, minipoa enables multiple sequence alignment of megabase-long genomes and million-sequence data sets, demonstrated by 342 Mycobacterium tuberculosis sequences and 1 million SARS-CoV-2 sequences, respectively. Collectively, minipoa is well positioned to become a cornerstone in the era of large-scale pangenomics.
Analyzing omics data in the context of pathway knowledge is critical for understanding the molecular mechanisms underlying pathological changes. However, current pathway analysis methods do not model the detailed mechanistic nature of biological interactions, limiting the understanding of pathway behavior to a relatively shallow level. To address this issue, we present a knowledge-driven machine learning framework that embeds features into pathway graphs and models reactions analytically, producing interpretable feature hierarchies and subnetworks in which functional associations are estimated to model biological interactions. The approach is agnostic to feature selection, enabling the use of full omics data sets without discarding weak signals. Applications to breast cancer microRNA-gene regulation data and COVID-19 metabolomic data highlight immune and metabolic pathways relevant to disease progression. This framework bridges predictive modeling with mechanistic interpretation and offers a foundation for integrative pathway analysis.
Multiple myeloma (MM) is a type of hematological cancer that arises from uncontrolled proliferation of plasma cells. In addition to frequent genetic mutations, malignant plasma cells are characterized by alterations to the epigenome. Myeloma cells display a genome-wide loss of DNA methylation and a corresponding increase in "active" chromatin modifications. The epigenetic remodeling that occurs in cancer genomes is associated with loss of silencing at transposable elements, which can impact genome regulation. Through paired epigenome and transcriptome profiling of patient-derived MM samples, we have found that loss of DNA methylation in MM genomes results in the formation of partially methylated domains that is variable across patients. This loss of DNA methylation coincides with the expression of hundreds of transcripts driven by LINE-1 (L1) retrotransposons that are epigenetically silenced in normal cells. MM samples can be stratified based on L1 transcriptional activity with distinct gene expression signatures. The high L1 samples are characterized by a more proliferative, less differentiated state as well as inhibition of interferon and genome defense pathways. Several L1 promoters generate chimeric transcripts with adjacent oncogenes. We further find that KRAB-containing zinc finger proteins (KZFPs) that are responsible for the epigenetic silencing of L1s have abnormally low abundance in MM samples with high L1 transcriptional activity. These results indicate that cell proliferation in MM is associated with a loss of KZFP expression and transcriptional activation of L1 elements.
Single-cell multiomic technology makes it possible to profile both the transcriptome and epigenome within individual cells. The detection of rare cell states from such data is crucial for exploring novel disease biomarkers with clinical potential. However, existing methods typically focus on the gene and peak counts while ignoring the underlying genomic sequence at accessible sites. Here, we propose ComicGTN, an innovative computational framework that integrates single-cell multiomic data with DNA sequence information via enhanced graph transformer networks to accurately identify rare cell clusters. ComicGTN consistently outperforms nine state-of-the-art methods in identifying rare cells across multiple complex scenarios. In mouse breast cancer data, ComicGTN detects several functionally distinct immune subpopulations. In cerebral cortex of epilepsy patients, it uncovers oligodendrocytes undergoing particular differentiation states. For polycystic kidney disease patient-derived organoids, ComicGTN elaborates pathological pathways in a captured rare glomerular subgroup. Overall, ComicGTN accelerates the localization of disease-associated rare cell populations, facilitating the derivation of clinical insights in development and disease progression.
Meiosis relies on programmed DNA double-strand breaks (DSBs) to initiate recombination between homologous chromosomes. Only a fraction of these breaks mature into crossovers (COs), creating the chiasmata essential for physically linking homologs and ensuring their accurate segregation in meiosis I. Errors in CO number or placement underlie a large fraction of human infertility and aneuploidies, yet how DSBs are designated to become COs and how the mechanisms ensuring COs form on all chromosomes (CO assurance) are not well understood, particularly in systems with a large number of chromosomes. Here, we investigate CO formation in the pantry moth, Plodia interpunctella, a novel invertebrate model system with n = 31 chromosomes. Using a combination of sequencing approaches, Oligopaints FISH, and statistical Weinstein tetrad analysis, we find robust CO assurance and interference in Plodia, with most chromosomes harboring only a single distal CO and few bivalents lacking COs. We employ CUT&RUN, ATAC-seq, and RNA-seq to profile the epigenomic landscape in Plodia testes, revealing that the distal chromosome arms in which COs form are enriched for heterochromatin and devoid of accessible chromatin, suggesting a bias for COs to form in repressed regions in the genome. These studies pave the way for future work diving deeper into the molecular regulation of CO formation in Plodia and other systems, in which the large number of chromosomes may be the key to revealing novel insights about CO regulation.
Understanding the consequences of individual DNA methylation variation is crucial for advancing our knowledge of human biology and disease, yet the collective impact of individual traits on DNA methylation and their downstream effects on gene expression across human tissues remains poorly understood. Here, we quantify the contributions of sex, age, genetic ancestry, and BMI on autosomal DNA methylation variation across nine human tissues and 424 individuals from the Genotype-Tissue Expression project. We show that genetic ancestry and age have a greater impact on DNA methylation compared with sex, with aging effects being more widespread but less pronounced. On average, <10% of the gene expression variation in sex, age, and ancestry is mediated by DNA methylation differences, with ancestry showing the largest proportion of mediation. We further show that ancestry-associated DNA methylation differences accumulate at CpG sites with extreme methylation states and are largely under genetic control. The female autosomal genome exhibits consistent hypermethylation across tissues at Polycomb-repressed regions. Ultimately, we show that age-related Polycomb target hypermethylation is observed across multiple tissues but not in the gonads. Our multi-individual, multitissue approach defines the key drivers of human DNA methylation variation in healthy conditions, establishing a baseline for the interpretation of DNA methylation changes in disease contexts.
Medaka (Oryzias latipes) is a small freshwater teleost widely used as a vertebrate model organism. Existing medaka reference genomes, however, contain many gaps and unresolved repetitive regions, hindering precise genome annotation and comparative analyses. Here we present one complete and two near-complete genome assemblies for three inbred medaka strains derived from geographically distant populations. These assemblies provide a comprehensive view of highly repetitive sequences and chromosome-scale genome architecture in medaka. The fully resolved centromeres reveal an intriguing sequence organization characterized by short, distinct SF1+3 satellite arrays flanked by larger homogenized repeats. These short arrays are putatively hypomethylated and conserved across all acrocentric chromosomes, suggesting a functional role in centromere stability. The reconstructed 121 copies of the giant mobile element Teratorn retain complete genes of both a transposon and a herpesvirus, highlighting its unique persistence and impact on host genomes. Moreover, our assemblies reveal extensive structural divergence of medaka Y Chromosomes, yet identify a small (∼24 kb) conserved region encompassing Dmy that may suffice for male determination. Collectively, these (near-)complete medaka genomes provide a powerful resource for exploring the biology of uncharacterized repetitive regions and the molecular basis of phenotypic diversity in vertebrates.
Polygenic Risk Scores (PRSs) estimate the likelihood of individuals to develop complex diseases based on their genetic variations. While their use in clinical practice and direct-to-consumer genetic testing is growing, the privacy implications of public PRS sharing are often underestimated. In this work, we demonstrate that PRSs can be exploited to recover genotypes and to de-anonymize individuals. We describe how to reconstruct a portion of an individual's genome from a single PRS value by using dynamic programming and population-based likelihood estimation, which we experimentally demonstrate on PRS panels of up to 50 variants. We highlight the risks of combining multiple, even larger-panel PRSs to improve genotype-recovery accuracy, which can enable the reidentification of individuals or their relatives in genomic databases or the prediction of additional health risks, not originally associated with the disclosed PRSs. We then develop an analytical framework to assess the privacy risk of releasing individual PRS values and provide a potential solution for sharing PRS models without decreasing their utility.