
Identification of unique and shared genomic regions between organisms has substantial translational potential for the development of marker-based diagnostic assays and sequence homology-driven taxonomic classification. An automated pipeline capable of performing genome comparisons at both the intra- and inter-species levels with minimal computational requirements can significantly advance genome-driven translational research. Species-specific genomic regions are particularly valuable for sequence-based species identification and for developing DNA amplification- or hybridization-based diagnostic assays. Here, we present SURE-Pipe, an automated and flexible pipeline for genome comparison and extraction of unique and shared genomic regions (https://github.com/BPaul-bioinfoLAB/SURE-Pipe). Benchmarking of this pipeline using simulated datasets demonstrated high accuracy for shared and unique region identification. Using the pairwise genome comparison module, six genome pairs from diverse microorganisms were analysed, and identified the unique and shared regions. In addition, the multigenome comparison module was applied to 96 genomes representing 24 Bacillus species and identified species-specific genomic regions. These regions were highly conserved among four strains of a species (>98% sequence identity) and exhibit little to no similarity with other species. Species-specific primers designed for all 24 Bacillus species showed no off-target amplification in in-silico polymerase chain reaction analysis, indicating their specificity. Overall, SURE-Pipe provides a robust and multipurpose framework for comparative genomics, and the outcomes can be used for species identification and the development of genome-based diagnostic approaches.
Bacterial genomic surveillance requires balancing computational efficiency with genetic resolution for effective cluster investigation. cgMLST distance calculations treat all allelic differences as equivalent units, obscuring nucleotide-level variation. Furthermore, single nucleotide polymorphism-based pipelines provide finer resolution at substantially higher computational cost, which limits their routine deployment in surveillance laboratories. We present cgDist, an algorithm that calculates nucleotide-level distances directly from cgMLST allelic profiles, providing finer resolution than allele-count distances by leveraging within-allele nucleotide variation. The cache architecture stores alignment statistics, enabling distance calculation modes without computation and supporting both dataset-specific and schema-complete cache generation. This design enables incremental surveillance analysis, with performance benefits as laboratories accumulate alignment data. cgDist functions as a precision 'zoom lens' for the investigation of clusters identified through initial cgMLST screening. Rather than restructuring population relationships, this targeted approach concentrates enhanced resolution where it is most informative. The algorithm ensures that cgDist distances are greater than or equal to corresponding cgMLST distances, preserving epidemiological interpretability while adding genetic discrimination. By increasing resolution within identified clusters, cgDist may also support outbreak investigation, a potential application that remains to be evaluated on outbreak-derived data.
Small interfering RNAs (siRNAs) are a clinically validated therapeutic modality with eight FDA-approved drugs, yet designing effective siRNAs remains computationally challenging due to complex dependencies on sequence composition, thermodynamic properties, target-site accessibility, and off-target interactions. Over two decades, computational approaches have evolved from empirical heuristics to deep learning systems integrating physical priors with learned representations. We review the complete landscape of machine learning methods for siRNA design, spanning classical scoring rules, pretrained RNA foundation models, transformer-based efficacy predictors, graph neural networks encoding siRNA/messenger RNA interaction topology, off-target prediction frameworks, and chemical modification-aware architectures. Across over 40 studies, we identify convergent findings: hybrid models integrating thermodynamic features with learned representations are among the strongest performers, although this evidence rests largely on single-model ablations and does not establish that foundation-model embeddings specifically are required; graph neural networks with leakage-aware data splitting address pervasive benchmark inflation; and off-target prediction has matured through empirical RNA-seq frameworks and structure-based features. We distinguish throughout between chemically unmodified siRNAs, which dominate public benchmarks, and the fully modified siRNAs used therapeutically, whose efficacy data remain scarce and whose prediction is correspondingly harder. We provide a taxonomy of methods, head-to-head performance comparisons, benchmark dataset descriptions, code availability, biology-informed interpretability analysis with formal saliency validation protocols, and concrete recommendations for advancing siRNA design. Critical gaps in uncertainty quantification, active learning, and prospective experimental validation are identified as priorities for clinical translation.
RNA-RNA interactions (RRI) can now be probed on a transcriptome-wide scale using proximity ligation methods such as SPLASH, PARIS, LIGR-seq, RIC-seq, and others. While individual pipelines for computationally processing the corresponding raw duplex reads exist, there is currently no method for comparing and exploring RRI networks across experimental protocols and cellular conditions. It thus remains difficult to discover biologically relevant interactions, to assess reproducibility, and to evaluate protocol-specific differences. Here, we present InteRRact - a web server for interactively exploring human RRI datasets derived from published duplex probing experiments. One key feature is that all available datasets have been processed uniformly. InteRRact visualizes intermolecular RRIs as gene-level networks and intramolecular interactions as linear, locus-specific genomic tracks. All RRIs are annotated with multiple quantitative metrics and evaluated statistically, allowing the user to readily filter interactions by strength of evidence and confidence. Moreover, any two datasets can be compared to identify shared and dataset-specific interactions across cell types, conditions, and experimental protocols. InteRRact thereby enables the discovery of biologically relevant interactions. InteRRact is available at https://e-rna.org/interract.
A fundamental understanding of genome organization relies on accurately annotating topologically associating domains (TADs) and their boundaries. This is crucial for understanding how cis-regulatory elements regulate gene expression. To go beyond calling TADs and boundaries from Hi-C data, several machine learning-based methods have been proposed to go the step further and predict TAD boundaries from genomic sequences. As the growing evidence of TADs and their boundaries, TADs have been proved exhibiting diverse properties, such as differences in replication timing and epigenetic patterns. However, existing methods do not take this heterogeneity into account. To address this, we propose a method called TADBpred for TAD boundary prediction in a large genomic context in humans. TADBpred focuses on TAD boundaries in active and inactive chromatin across cell-lines and tissues, which are GC-rich and AT-rich, respectively. By integrating genomic elements and sequence composition, we designed two models for GC-rich and AT-rich boundaries, respectively. When testing the performance on respective independent held-out datasets, we obtain AUC scores of 0.91 and 0.80. Our results indicate that TADBpred excels in TAD boundary prediction. Additionally, feature importance analysis highlights the essential features for different classes of TAD boundaries, thereby enhancing our understanding of these TAD boundaries.
Metagenomic sequencing is transforming diverse areas of health and biological sciences, including pathogen surveillance, clinical diagnostics, and microbiome research. However, the inherent complexity of metagenomic data limits most computational tools to species-level classification and abundance estimation, overlooking within-species genetic diversity that drives key phenotypes. We present metaWEPP, a novel computational pipeline that achieves near-haplotype resolution in metagenomic analysis for species with adequate representation in reference genome biobanks and having sufficient sequencing depth and genome coverage. Specifically, metaWEPP assigns sequencing reads to species using standard taxonomic classifiers, phylogenetically places them onto species-specific mutation-annotated trees of publicly available sequences, and selects the haplotypes that best explain the sample. It also reports unaccounted alleles indicative of novel variants and provides an interactive dashboard for read-level visualization. Applied to diverse metagenomic and mixed-genome samples from prior studies, metaWEPP produced concordant species-level results, while revealing finer lineage- and haplotype-level insights not captured by existing tools. On various clinical samples, metaWEPP identified infecting pathogens and additionally provided credible lineage- and haplotype-level information that can support clinical decision-making. On wastewater samples, metaWEPP uncovered previously undetected haplotype clusters of epidemiological relevance. These findings demonstrate metaWEPP's ability to advance various clinical, epidemiological, and research applications with deeper, actionable insights.
Topologically associating domains (TADs) are generally considered as a homogeneous basic units of genome folding, which is critical for transcriptional regulation. However, recent studies indicate that both the TAD domain structures and the boundaries between them are not as homogeneous as originally recognized. Here, we address the heterogeneity of the TAD boundaries in the human genome at a large scale, which varies between active and inactive chromatin and across cell lines and tissues. To address this, based on the well-annotated TAD boundaries extracted from multiple cell lines and tissues, we examine their nucleotide content, resulting in two main clusters, one GC-rich and one AT-rich, which are mainly distributed in active and inactive chromatin, respectively. Also, they contain different types of repetitive sequences and have different epigenetic patterns, with more CTCF binding motifs in the GC-rich cluster. Hence, our observations of the TAD boundary content provide novel insights into TAD genomic architecture. In addition, we find that cell- or tissue-specific boundaries are less evolutionarily conserved than other boundaries. We highlight the importance of TAD boundary diversity in different functional contexts and discuss the importance of the different types of repetitive sequences and epigenetic patterns in the two main types of boundaries.
Messenger RNA (mRNA) decapping mediated by DCP2 is a key mechanism controlling RNA stability and gene expression in eukaryotes, including plants. Despite its central role in regulating plant development and stress responses, the repertoire of mRNA 5' caps targeted by DCP2 remains undefined. Here, we combined in vitro decapping with 5'-end enriched and full-length transcriptome sequencing of DCP2-deficient mutants to comprehensively characterize the mRNA capping landscape in Arabidopsis thaliana. We mapped over 13 000 high-confidence capped transcripts at nucleotide resolution, revealing distinct 5' cap signatures in wild-type and dcp2 seedlings. Most caps localized near annotated transcription start sites, validating the accuracy of our approach. Loss of DCP2 led to substantial expansion of detectable capped 5'-ends, including 275 caps from previously unannotated loci, and an increased prevalence of multi-capped genes, highlighting the role of DCP2-mediated decapping in removing unwanted transcripts. Integration with degradome resources revealed extensive convergence between DCP2-sensitive capped RNAs and substrates of co-translational, cytosolic XRN4-dependent decay, as well as nonsense-mediated decay pathways. Our results suggest that DCP2 shapes the steady-state 5' cap landscape in vivo by limiting the accumulation of unstable and cryptic capped transcripts. Also, this study provides a valuable resource for transcript annotation and isoform-aware analysis of RNA turnover.
Understanding how gene perturbations reshape cellular states across diverse contexts is fundamental to functional genomics and therapeutic discovery, yet experimental profiling across all perturbations and cell types remains infeasible. Computational approaches promise scalable inference but often rely on discrete mappings or per-cell-type models that fail to capture continuous and shared biological dynamics. We introduce CFM-GP, a conditional flow-matching framework that learns a continuous vector field transforming control expression profiles into perturbed states, explicitly conditioned on cell type. This unified design models both common regulatory programs and type-specific responses within a single architecture, removing the need to train separate models. Across five single-cell perturbation datasets, CFM-GP consistently outperformed existing methods in predictive accuracy, distributional alignment, and cross-species generalization. The inferred flow trajectories recovered canonical signaling pathways and context-dependent transcriptional cascades, demonstrating mechanistic interpretability. By coupling principled generative dynamics with biological conditioning, CFM-GP offers a scalable foundation for modeling cellular perturbation responses, enabling data-driven exploration of gene function and intervention strategies across heterogeneous cellular systems.
Many tumours show deficiencies in DNA damage response (DDR), not only driving tumorigenesis but also exposing vulnerabilities with therapeutic potential. Assessing which patients might benefit from DDR-targeting therapy requires knowledge of tumour DDR deficiency (DDRd) status, with mutational signatures reportedly better predictors than loss-of-function mutations. Existing DDRd models offer effective prediction for pathways with well-characterized processes and mutational signatures. Nevertheless, development of models for additional DDRd and clinically relevant mechanisms could be hampered by the fact that most mutational signatures have unknown etiology. Using supervised non-negative matrix factorization (SNMF), we integrate mutational signature learning with multiclass DDR-deficiency prediction to enable etiology-guided learning of signatures from cell lines with confirmed gene knockouts. Applied to DDR gene knockout human-induced pluripotent stem cell lines, SNMF identified etiology-aligned representations of deficiency in homologous recombination, mismatch repair, and base excision repair. Even guided by pathway-level labels, SNMF captured gene-specific base excision repair submechanisms, showing the integration offered added granularity. Learned cell line signatures showed high similarity to tumour-derived COSMIC signatures, revealed associations with mutations in DDR genes, and enabled high recall of tumours with DDR deficiencies. We envision that SNMF-like methods could leverage knockout screens to learn etiology-guided signatures for improved DDRd annotation and treatment optimization. SNMF is available at: https://github.com/joanagoncalveslab/SNMF.
Visualizing relationships within massive biological datasets remains a significant challenge, particularly as sequence length and volume increase. We introduce CAPASYDIS (Cartesian Projections of Asymmetric Distances), a scalable approach designed to map the explored regions of a given sequence space. Unlike traditional dimensionality reduction methods, CAPASYDIS calculates asymmetric distances which account for both the position and type of sequence differences. It projects sequences into a fixed, low-dimensional coordinate system, termed a "seqverse," where each sequence occupies a permanent and unique location. This design allows for the instant mapping of new sequences without the need to recalculate the global coordinate system, enabling the exploration of the diversity space on a standardized map. We applied this method to a large rRNA sequence dataset spanning the three domains of life. Our results demonstrate that the sequences of Bacteria, Archaea, and Eukaryota occupy spatially distinct regions characterized by fundamentally different shapes and patterns of variation. Furthermore, the resulting seqverses retain high amount of taxonomic information, when analyzed from broad domain levels to single-base differences. Overall, CAPASYDIS provides a reproducible, scalable framework for defining the boundaries and topography of biological sequence universes.
Time-resolved transcriptomic profiling has been used to study phage-host interactions for more than a decade. However, the resulting datasets are not readily accessible for custom re-analysis, and resources are lacking that provide standardized processing, storage, and analysis of transcriptomes from phage infections. Here, we present the PhageExpressionAtlas, the first bioinformatics resource for storing time-resolved dual RNA-sequencing data from phage infections. This data was processed uniformly using a custom analysis pipeline and is presented for interactive exploration through visualization. The PhageExpressionAtlas currently hosts 42 datasets from 23 studies. Using the PhageExpressionAtlas, we replicate key findings from original publications and extend hypothesis testing across multiple phage-host systems. By systematically querying and analyzing the underlying database, we evaluate approaches to phage gene classification and find that uncharacterized phage genes dominate all infection phases, with their distribution depending strongly on the classification strategy. Moreover, we provide a comprehensive view of the expression dynamics of anti-phage defenses as well as host- and phage-encoded anti-defense systems in the infection context, indicating unique and conserved patterns of transcriptional regulation underlying bacterial anti-phage immunity and phage counter-strategies. Together, the PhageExpressionAtlas is a unifying resource that democratizes transcriptomics-driven analyses of phage-host interactions and supports integrative cross-study assessment.
Bacterial pathogens spread in clinical and environmental settings, and mobile genetic elements (MGEs), such as plasmids and phages, mediate the transfer of virulence factor genes (VFGs) and antimicrobial resistance genes (ARGs) among bacterial communities. Metagenomic analysis of environmental and wastewater samples using highly accurate long-read sequencing technologies, such as Pacific Biosciences (PacBio) HiFi sequencing, provides valuable insights into monitoring the regional spread of VFGs and ARGs, including dissemination mediated by MGEs. No visualization tool is currently available for the comprehensive display of numerous resulting circular metagenome-assembled genomes (cMAGs) with functional gene annotations. Here, we developed visualization of circular metagenome-assembled genome (VicMAG), a visualization tool for highly complex cMAGs derived from long-read metagenome assemblies annotated using updated databases of VFGs, ARGs, and MGEs. Using 353 cMAGs from PacBio HiFi sequencing of a wastewater sample, we demonstrated the utility of VicMAG for metagenome visualization. VicMAG provides comprehensive, size-aware visualization of cMAGs representing bacterial chromosomes and plasmids, annotated with VFGs, ARGs, and phages. By simultaneously visualizing all cMAGs in a framework, VicMAG facilitates a holistic understanding of the distribution and genomic context of VFGs and ARGs across complex microbial communities. This tool supports integrated surveillance of bacteria associated with virulence and antimicrobial resistance across clinical, environmental, and One Health contexts.
Allele-specific methylation (ASM) refers to the differential DNA methylation between two alleles at a specific locus. This phenomenon can be driven by genomic imprinting (where gene expression depends on the parent of origin) or by genetic variants (that may affect DNA-protein binding), both of which play crucial roles in gene regulation and contribute to normal biological variation and disease. ASM can be tissue- or cell-type specific, which adds an extra layer of complexity to its analysis. Detection methods typically rely on phasing using genetic variants, a process that is computationally intensive and can fail in regions with low heterozygosity. To overcome these limitations, we developed asms (Allele-Specific Methylation Scanner), a tool that identifies potential ASM loci from methylation data without requiring prior phasing by looking at heterogeneous methylation across reads. Although we tested it using Oxford Nanopore Technologies (ONT) data, the tool can be used with any platform using MM/ML tags to store methylation in BAM files. asms can examine efficiently thousands of loci (either from a list or through an automatic genomic scan) by segregating reads based on their methylation patterns. When information on variants is available, asms can verify whether distinct alleles correspond to different base modification profiles. asms is available at https://github.com/ecmra/asms.
Population aging is increasing the burden of age-related disease, highlighting the need for accurate molecular tools to estimate aging. DNA methylation, a heritable epigenetic mark linked to development and age-related disease, can capture biological aging and predict chronological age; however, most existing epigenetic clocks were developed primarily in non-East Asian populations and may show reduced calibration across ancestries. We analyzed methylation profiles from East Asian cohorts from Taiwan, Japan, and China to develop an East Asian-specific epigenetic clock, termed the EAS Clock. After quality control and probe filtering, a parsimonious model based on 38 CpG sites was constructed using stepwise multivariate regression with forward selection and Bayesian Information Criterion. The EAS Clock showed significant correlations between predicted and chronological age in both training and independent testing datasets, with residuals centered near zero. Functional enrichment analyses indicated that genes associated with the selected CpGs were involved in neurobiological, musculoskeletal, immune-inflammatory, and lipid-metabolic pathways. These findings support the EAS Clock as a compact methylation-based estimator of chronological age tailored to East Asian populations.
Mitochondrial dysfunction and fragmentation are observed in various circumstances, such as neurodegeneration and aging. Studies have shown that altered mitochondrial function activates the integrated stress response (ISR), with ATF4 serving as a major mediator of adaptation to stress. Presently, little is known about the role of ATF4 in neurons under mitochondrial stress. Using primary cortical neurons, we demonstrate that inhibiting ATF4 under OPA1-mediated mitochondrial stress accelerates the impairment of neuronal differentiation, as evidenced by smaller dendrites and lower dendritic spine density. To better understand the role of ATF4 in this context, we investigated the global binding sites of ATF4 using chromatin immunoprecipitation sequencing (ChIP-seq) and examined the chromatin accessibility changes that occur following the loss of ATF4 in neurons under conditions of mitochondrial stress. We found that ATF4 binds to a wide range of targets and alters the chromatin accessibility of genes involved in metabolism, neuronal fate, and neuron maturation. The downstream targets of ATF4 identified in this study can reveal novel and direct targets of ATF4 in neuronal survival and maturation. These adaptations are the hallmarks of stress response in mitochondrial dysfunction-mediated neurodegeneration.
SLNCR is a long non-coding RNA (lncRNA) that promotes melanoma formation. Previously, we have shown that SLNCR regulates gene expression via its interactions with different transcription factors. Here we show that SLNCR is associated with several genomic sites that contain RNA•DNA triplex target sites (TTSs). The primary sequence of SLNCR contains four triplex-forming regions (TFRs), which are homologous to the TFRs of MEG3, HOTAIR, and PARTICLE lncRNAs. Full-length SLNCR overexpression promotes proliferation, migration, and invasion of melanoma cells by inducing characteristic gene expression changes of genes that are enriched in pathways involved in cell cycle regulation and cell migration. The transcriptomic signature of full-length SLNCR is reversed by deletions of TFR1, 2, 3, or 4. Full-length SLNCR is structurally flexible and permissive of different interactions with the DNA and/or proteins. Deletion of any individual TFR causes SLNCR to gain structural definition, potentially preventing some of its native interactions. Using RNA sequencing and the Triplex Domain Finder tool, we demonstrate that SLNCR TFRs form RNA•DNA:DNA triplexes with the TTSs in promoters of key genes (nodes) predicted to regulate melanoma migration and invasion networks. We further validate the effect on invasion using A375 cells in functional assays. Finally, inhibitory oligos that block SLNCR•DNA:DNA triplex formation reversed the effect of SLNCR-mediated migration ex vivo, demonstrating that triplex formation plays a role in driving melanoma progression.
Scalable proxies for 3D genome contacts-such as single-cell co-accessibility and deep learning predictions-have emerged as powerful alternatives to chromatin capture-based methods, but predictions systematically overestimate long-range interactions. Here we show how to correct this bias using distance-based penalty functions informed by Gaussian mixture modeling and polymer-physics scaling. Using Hi-C datasets from maize, rice, and soybean, we derive tissue-specific and global consensus penalties parameterized by multiregime power-law exponents. Applying these corrections to single-cell ATAC sequencing co-accessibility scores improves their distance profiles in concordance with Hi-C and reduces long-range false positives by an average of 73% with tissue-specific penalties and 66% with the global consensus. We provide open-source code and fitted parameters to support adoption in maize, rice, and soybean.
Differences in enhancer activity between species can help drive phenotypic diversity, yet enhancers often have conserved functions despite rapid sequence evolution, posing a challenge for quantifying their functional differences between species. Previous machine learning models have focused on the binary task of predicting differences in the presence of enhancers between species but have yet to demonstrate an ability to predict continuous differences in enhancer activity. Here, we trained convolutional neural networks on a regression task to predict chromatin accessibility-a proxy for enhancer activity-in the liver across five mammals, and we developed a novel framework to evaluate cross-species performance. We demonstrated that training on multiple species improves model generalization to both species used in training and held-out species. However, the models consistently achieved poor performance in predicting quantitative differences in accessibility between species at orthologous regions. Our study highlights the challenges in using regression models to predict chromatin accessibility changes between species. All data and code are available at http://daphne.compbio.cs.cmu.edu/files/azstephe/liver_regression_resource/ and https://figshare.com/projects/liverRegression/274293.
Shared epitopes pose safety and efficacy issues for T-cell immunotherapy. To characterize the extent of this problem, we performed a computational analysis establishing a complete atlas of shared identical sequences across the human and murine proteomes. Unlike bacterial or viral antigens, self-antigens, including tumor-associated antigens (TAAs), frequently contain sequences of sufficient length to generate identical epitopes in other self-proteins. Epitopes from these shared sequences can theoretically reduce target specificity, confound immunomonitoring studies, and contribute to pre-existing immune tolerance toward TAAs. Notably, a subset of TAAs identified in this atlas is free of this drawback, providing a new criterion for antigen prioritization in cancer immunotherapy. To facilitate the detection of shared sequences, a web server has been made available at https://epitopemapper.ircan.org/ and the open-source code at https://github.com/IRCAN/EpitopeMapper.