
Biological function is mediated by the hierarchical organization of cell types and states within tissue ecosystems. Identifying interpretable composite marker sets that both define and distinguish hierarchical cell identities is essential for decoding biological complexity yet remains a major challenge. Here, we present RECOMBINE, an algorithm that identifies recurrent composite marker sets to define hierarchical cell identities. Validation using both simulated and biological data sets demonstrates that RECOMBINE is robust to hyperparameter variation and data sparsity, and achieves higher accuracy in identifying discriminant markers compared with existing approaches. As a partition-free framework, RECOMBINE is particularly powerful for data sets characterized by continuous cell-state transitions, in which defining discrete boundaries is inappropriate. This capability is demonstrated by its application to zebrafish development, revealing gradual transcriptional transitions across embryonic stages, and to the mouse cerebellum, in which it uncovers transcriptional variation shaped by spatial gradients. When applied to single-cell data and validated with spatial transcriptomic data from the mouse visual cortex, RECOMBINE identifies key cell-type markers and generates a robust gene panel for targeted spatial profiling. It also uncovers markers of CD8+ T cell states, including GZMK+HAVCR2- effector memory cells associated with anti-PD-1 therapy response. Finally, using data from the Tabula Sapiens project, RECOMBINE identifies composite marker sets across a broad range of human tissues. Together, these results highlight RECOMBINE as a robust, data-driven framework for optimized marker selection, enabling the discovery and validation of hierarchical cell identities across diverse tissue contexts.
Advances in high-throughput sequencing have lowered the cost and complexity of genome sequencing, making it possible for the first time to assemble large pangenomic data sets for many species. These data sets, comprising thousands of individuals, already span from hundreds of gigabytes to petabytes, far exceeding the memory capacity of most machines, and are expected to continue growing in scale over time. Already, many traditional bioinformatics tools fail on inputs at this scale because they cannot construct their necessary data structures within memory limits. There is a growing need for methods that can construct these structures directly from compressed representations. Prefix-free parsing (PFP) addresses this challenge. PFP serves as a preprocessing step that compresses sufficiently repetitive text, yet still permits building important data structures for the original data set from its compressed output. This survey offers an overview of PFP, covering its core principles, the primary data structures it enables, current applications, and future research directions.
Gene transcription is activated through the interaction between cis-regulatory elements (CREs) and transcription factors (TFs). CREs serve as templates to provide binding sites, whereas TFs provide functions to directly initiate transcription. Current research mainly focuses on deciphering cis-regulatory code, while neglecting TF regulatory code. However, CREs alone are not sufficient to determine the binding of TFs, which makes the interpretation of cis-regulatory code ambiguous. In this study, we systematically analyze 13 TF binding profiles associated with transcription initiation and encode them as TF sequence to explore the TF regulatory code. Furthermore, we propose a deep learning model named DeepTF to predict gene expression from TF sequence. Results show that TF binding exhibits conserved positional preferences and combinatorial patterns in promoters and DeepTF is able to predict gene expression with high accuracy (AUROC = 0.97). Meanwhile, cross-cell-line validation (AUROC > 0.90) further confirms the model's transferability. Model interpretation reveals DeepTF successfully captures the TF regulatory grammar associated with gene expression. Compared with the cis-regulatory code, our proposed TF regulatory code is better suited for investigating the relationship between TF binding positions or combinatorial patterns and gene expression. Collectively, DeepTF simplifies gene expression prediction and provides clear biological insights of transcriptional regulation.
Cytosine methylation, a crucial epigenetic modification, plays a vital role in genomic regulation. Leveraging the advancements in long-read sequencing, we investigate the methylation patterns of polymorphic transposable element (TE) insertions of human lymphoblastoid cell lines (LCLs). We validate the high concordance between long-read methylation calls and the conventional whole-genome bisulfite sequencing (WGBS) method. We then aim to establish general rules of TE methylation with our data by addressing three key questions: (1) what is the methylation profile of each insertion; (2) do newly inserted TEs adopt the methylation pattern of their genomic context; and (3) do new TE insertions affect the methylation of their flanking regions. Although most non-TE insertions exhibit DNA methylation patterns consistent with their genomic context, TE insertions are generally highly methylated, exhibiting distinct, class-specific patterns with some variation within TE bodies. A small percentage of Alu insertions are hypomethylated, particularly those inserted within hypomethylated CpG islands. We also reveal that majority of TEs exhibited minimal impact on nearby regions, although numerous exceptions exist in which the methylation status of TEs spread into nearby regions. In conclusion, although TE insertions primarily exhibit methylation patterns restricted within their boundaries, some TEs are able to affect the methylation level of their genomic neighborhoods. At last, our findings are limited to human LCLs, and more comprehensive analysis would be needed to test the rules we found here on a broader spectrum of cell types and developmental stages.
Precise transcription factor (TF) binding to DNA governs gene regulation, yet nucleotide sequence alone often fails to fully capture binding specificity. Although static DNA shape is a recognized determinant of indirect readout, the role of intrinsic conformational flexibility remains underexplored across TF families. Here, we demonstrate that integrating sequence-derived DNA flexibility descriptors into predictive models improves both prediction and mechanistic interpretability of TF-DNA affinity. Across large-scale in vitro data sets encompassing HT-SELEX and protein-binding microarrays for mammalian and Drosophila TFs, flexibility-augmented models consistently outperform sequence-only baselines and complement DNA shape models. Cross-platform analyses further indicate that flexibility features capture structural information that is robust to platform-specific biases. Using a position-resolved interpretation framework, we uncover family-specific "flexibility footprints," including recurrent hotspots in core motifs and flanks that align with DNA structural deformations from TF-DNA complex structures. Extending to ENCODE ChIP-seq and DNase-seq data, flexibility augmentation improves classification of functional TF binding sites across diverse TFs and cellular contexts. Collectively, these results underscore the insufficiency of sequence-only models and highlight the utility of the flexibility descriptors as an interpretable component of the TF recognition code.
Spatial transcriptomics enable fine-scale characterization of spatial heterogeneity and cellular niches within tissues, and have substantially advanced our understanding of tissue architecture and functional organization. However, existing spatial transcriptomic integration methods often struggle to effectively capture the rich morphological information provided by the histology and thus further limit their capacity for comprehensive cross-modality learning. In this paper, we present SYMOL, a unified synergistic self-supervised multimodal framework that integrates spatial coordinates, gene expression, and histological images covering both multichannel immunohistochemistry (IHC) and hematoxylin and eosin (H&E) stains for effective spatial transcriptomic integration and representation learning. Specifically, SYMOL extracts distinct visual characteristics via several pretrained large vision models and synergistically aggregates cross-modal features into unified morphology-aware embeddings. Comprehensive benchmarking on multiple publicly available spatial transcriptomic data sets with multichannel IHC images and H&E images shows that SYMOL consistently surpasses state-of-the-art methods in various downstream tasks, including cellular niche identification, multislice integration, cross-data set label transfer, and gene expression enhancement. In addition, SYMOL accurately delineates tumor microenvironment in lung tissues with histopathological imaging and enables fine-scale mapping of cellular niches in the mouse brain, thereby demonstrating both clinical relevance and robustness in complex neuroanatomical settings.
Vitamins are essential metabolic cofactors, yet their roles in epitranscriptomic regulation, particularly N 6-methyladenosine (m6A) modification, remain unclear. Here, we investigate the effects of various vitamins (VB2, VB6, and VB12) on the mRNA m6A epitranscriptome in multiple mouse tissues (brain, liver, and testis). Clustering analyses reveal closer similarity between the brain and testis m6A profiles, whereas the liver exhibits a unique pattern, reflecting tissue-specific regulatory dynamics. More than 90% of m6A sites are independent of mRNA abundance, highlighting the post-transcriptional role of m6A modification. Additionally, alternative splicing (AS) variations reveal complex interactions between vitamins, m6A modification, and AS, with tissue- and vitamin-specific effects on biological pathways. We identify 22 comethylation modules, associated with pathways such as neurodegenerative diseases and immune regulation. Key vitamin-responsive genes are found as central regulators of m6A dynamics, aligning with known roles of vitamins in metabolism, neural plasticity, and gene expression. Together, our findings provide the first comprehensive atlas of vitamin-driven, tissue-specific m6A modifications, offering new insights into the post-transcriptional regulatory mechanisms underlying vitamin-mediated cellular functions and their implications for nutritional or pharmacological modulation of the epitranscriptome in health and disease.
R-loops, RNA:DNA hybrids that often form cotranscriptionally, are emerging as key regulators of genome function, yet their roles in shaping chromatin architecture and developmental potential remain incompletely defined. Here, we use inducible Rnaseh1 expression in mouse embryonic stem cells (mESCs) to achieve short-term, global R-loop depletion and to systematically interrogate their impact on chromatin structure and lineage specification. We find that R-loop loss has a minimal effect on steady-state gene expression or self-renewal. Instead, it leads to a striking reduction in H2A.Z occupancy at both active and bivalent promoters, accompanied by increased nucleosome density, revealing a previously unrecognized role for R-loops in maintaining promoter architecture. During gastruloid differentiation, R-loop-depleted mESCs exhibit accelerated ectodermal differentiation, along with dysregulation of lineage-specific transcription factors and impaired cell-cell signaling. Consistent with these alterations, R-loop-depleted cells show widespread perturbations in gene regulatory networks across several early cell types. These findings uncover a critical role for R-loops in shaping the H2A.Z chromatin landscape and preserving balanced lineage trajectories during early development, offering new insights into the epigenomic regulation of stem cell fate.
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
Cell lines are invaluable tools for biomedical and evolutionary studies, but their genomic stability over time is often assumed rather than systematically assessed. In this study, we investigate the dynamics of genomic instability and structural rearrangements across multiple batches of a Nomascus siki cell line using a combination of single-cell template strand sequencing (Strand-seq), whole-genome sequencing (WGS), and fluorescence in situ hybridization (FISH). We identify 22 shared inversions in all the Strand-sequenced batches, confirming a common clonal origin. However, we detect additional large-scale rearrangements in all the batches, including trisomy of Chromosome 14 and the formation of isochromosomes of the same chromosome (iso-q and iso-p), leading to the rise of distinct subclonal populations. These rearrangements show evidence of clonal expansion, suggesting a proliferative advantage under in vitro conditions. From an evolutionary perspective, the gibbon genome is known for its exceptional level of chromosomal reshuffling, and this inherent plasticity may have contributed to the cell line's sensitivity to culture-induced structural changes. Despite extensive structural variation, the cell line remains stable at the nucleotide level, with ∼99% of SNPs shared across all batches. Our results illustrate how cell culture can recapitulate aspects of karyotypic evolution and underscore the need for regular genomic surveillance, particularly in long-term cultures. Furthermore, this study demonstrates the power of combining Strand-seq and cytogenetic approaches to detect both balanced and unbalanced rearrangements, especially those present in subclonal populations that would be missed by standard WGS.
Evolutionary divergence in body size is common in animal adaptive radiations and is often associated with differences in key ecological traits, including habitat use and prey consumption. Here we characterize a notable case of body size-associated adaptive radiation in a group of predatory open water cichlid fish species from the Lake Malawi catchment. Using whole-genome sequences, we show that body size differences have evolved multiple times in the focal genus, Rhamphochromis, and that the group possesses well-defined signals of ancient interspecific hybridization. We identify genetic variants strongly associated with body size and show that these variants are connected to genes enriched for functions in vertebrate skeletal and nervous system development. We focus our analyses on two species of Rhamphochromis endemic to Lake Kingiri, a small (600 m diameter) crater lake geographically isolated from the main body of Lake Malawi but within the catchment. We show that these two ecomorphologically divergent sympatric species-one small-bodied, the other larger-bodied-share a unique common ancestor and diverged from one another ∼2000 years ago. We demonstrate strong directional selection focused on the larger-bodied Kingiri species, specifically on genetic variants connected to genes with anatomical development and nervous system function. Collectively, these results are supportive of body size-associated speciation taking place rapidly in the Lake Malawi cichlid fish superradiation. We conclude that body size-associated genetic variants have been important targets of selection during large-scale cichlid fish diversification, including in a crater lake sympatric speciation context.
Accurate protein function prediction is fundamental to advancing drug discovery and precision medicine and understanding complex biological systems. Although Gene Ontology (GO) provides a standardized framework for protein annotation, a critical challenge persists: the imbalance between low-specificity GO terms and high-specificity GO terms. This imbalance creates blind spots in our understanding of protein function landscapes, particularly in clinically relevant pathways. Here, we present ProGO-PSL, a novel large graph architecture designed to resolve this imbalance. ProGO-PSL simultaneously leverages explicit domain identifiers from InterPro and implicit evolutionary contexts from multiple sequence alignments, fusing these complementary data sources within a powerful imbalance learning framework. Our model consistently outperforms state-of-the-art methods by 5%-15% across all specificity levels and on both a benchmark data set and an independent test set, demonstrating robust generalization. Furthermore, ProGO-PSL generates interpretable representations that clarify relationships between low- and high-specificity GO terms, enabling a more complete functional characterization of the proteome. This work accelerates the identification of therapeutic targets in previously uncharacterized biological pathways.
Homology-directed repair (HDR) enables precise genome editing; however, its application in mammalian cells is limited by low efficiency owing to competition from error-prone repair pathways and intrinsically restricted HDR activity. Existing HDR-enhancement strategies, including small-molecule treatments and marker-based selection, are constrained by cytotoxicity, genomic scarring, and inconsistent performance. Here, we present essential-gene-supported scarless HDR (ESS-HDR), a robust, drug- and marker-free platform that selectively enriches HDR-proficient cells. By leveraging essential-gene coediting, ESS-HDR enables precise and scarless genome modification with enhanced efficiency. CRISPR-Cas9 induces double-strand breaks at both the target locus and an essential gene, accompanied by two donor templates: one introducing the desired edit and the other restoring essential-gene function. Only cells that undergo accurate HDR at the essential locus survive, providing endogenous selection without exogenous markers. Single-cell clone analysis confirms that enrichment of HDR-proficient cells enhances editing at the target locus. Using ssODN donors carrying a 1 nucleotide substitution or a 10 nucleotide insertion, ESS-HDR increases HDR efficiencies by sevenfold to 16-fold in HEK293 cells and 41-fold in primary epidermal keratinocytes compared with conventional single-site HDR. With plasmid donors targeting TUBA1B, LMNB1, or ACTB, ESS-HDR improves knock-in efficiencies by sixfold to 34-fold across HEK293, U2OS, and HeLa cells. ESS-HDR also outperforms chemical enhancers including RS-1, SCR7, nocodazole, and AZD7648. Together, these findings establish ESS-HDR as a broadly applicable strategy for efficient, scarless genome editing without external selection markers.
Intronic polyadenylation (IPA) is a key mechanism driving transcriptome diversity, yet its detection and functional characterization remain challenging owing to complex splicing patterns and the complexity of intronic regions. Here, we introduce IPAseek, a dynamic programming based computational framework that leverages the pruned exact linear time (PELT) algorithm and changepoints over a range of penalties (CROPS) to enable de novo identification of IPA events from bulk RNA-seq data. IPAseek robustly detects both composite and skipped IPA isoforms. Applying IPAseek to bulk RNA-seq of hematopoietic cell types reveals lineage and stage-specific IPA signatures, with lymphoid cells exhibiting higher IPA site usage compared with myeloid cells. Temporal profiling during megakaryocyte differentiation uncovers dynamic, gene-specific IPA regulation linked to functional pathways including peroxisomal metabolism and autophagy, which are known to play a crucial role in megakaryocytic differentiation, impacting the development and maturation of megakaryocytes. Further, integrative analysis demonstrates that IPA site usage is associated with lower DNA methylation within introns, supporting a regulatory axis connecting epigenetic state and IPA. This finding aligns with emerging evidence that DNA methylation modulates alternative polyadenylation via CTCF-mediated chromatin looping. Thus, IPAseek provides a platform to characterize IPA across physiological systems and disease contexts using widely available bulk RNA-seq data. These IPA events can be further integrated with other regulatory data sets to elucidate their interplay and functional significance.
Ligand-receptor interactions mediate intercellular communication, inducing transcriptional changes that regulate physiological and pathological processes. Ligand-induced transcriptomic signatures can be used to infer ligand activity; however, the absence of a comprehensive set of ligand-response signatures has limited their practical application in predicting ligand-receptor interactions. To bridge this gap, we develop Lignature, a curated database encompassing intracellular transcriptomic signatures for 362 human ligands, significantly expanding the repertoire of ligands with available intracellular response signatures such as CytoSig and ImmuneDictionary. Lignature compiles signatures from published transcriptomic data sets, generating both gene- and pathway-based signatures for each ligand. We apply Lignature to prioritize ligand-associated transcriptional activity in controlled in vitro experiments and real-world single-cell sequencing data sets. Across these settings, Lignature consistently improves the prioritization of experimentally supported ligands compared with existing approaches. We additionally develop a regression-based framework to model combinatorial regulation by multiple ligands. These results establish Lignature as a robust platform for ligand signaling inference, providing a powerful tool to explore ligand-receptor interactions across diverse experimental and physiological contexts.
Tandem repeat (TR) analysis is crucial for understanding genome structure and variation. However, string decomposition, a key challenge in TRs analysis, remains computationally demanding. In this study, we introduce Wavefront-based String Decomposer (WSD), a novel algorithm that enhances efficiency and accuracy in TRs decomposition. By integrating wavefront techniques, WSD significantly reduces computational and memory costs. Additionally, two adaptive strategies minimize parameter sensitivity and further improve efficiency. Through extensive experiments, we demonstrate that WSD outperforms current state-of-the-art (SOTA) methods, achieving an average speedup of ∼2.33× and reducing memory usage by two orders of magnitude when analyzing human TRs.
The development of 3C-based techniques for analyzing three-dimensional chromatin structure dynamics has driven significant interest in computational methods for 3D chromatin reconstruction. In particular, models based on Hi-C and its single-cell variants, such as scHi-C, have gained widespread popularity. Current approaches for reconstructing the chromatin structure from scHi-C data typically operate by processing one scHi-C map at a time, generating a corresponding 3D chromatin structure as output. Here, we introduce an alternative approach to the whole-genome 3D chromatin structure reconstruction that builds upon existing methods while incorporating the broader context of dynamic cellular processes, such as the cell cycle or cell maturation. Our approach integrates scHi-C contact data with single-cell trajectory information and is based on applying simultaneous modeling of a number of cells ordered along the progression of a given cellular process. The approach is able to successfully recreate known nuclear structures while simultaneously achieving smooth, continuous changes in chromatin structure throughout the cell cycle trajectory. Although both Hi-C-based chromatin reconstruction and cellular trajectory inference are well-developed fields, little effort has been made to bridge the gap between them. To address this, we present ChromMovie, a comprehensive molecular dynamics framework for modeling 3D chromatin structure changes in the context of cellular trajectories. To our knowledge, no existing method effectively leverages both the variability of single-cell Hi-C data and explicit information from estimated cellular trajectories, such as cell cycle progression, to improve chromatin structure reconstruction.
ACTB is a cytoskeletal protein involved in intracellular trafficking. In recent years, it has become evident that, in addition to its established roles in these compartments, ACTB also participates in the regulation of transcription. However, the molecular mechanisms underlying this function remain poorly understood. The methyltransferase SETD3 has previously been shown to methylate ACTB at H73, thereby regulating ACTB polymerization and smooth muscle contraction. Here, we show that the genomic distribution of ACTB is SETD3-dependent and that this regulation modulates the transcription of genes involved in cell adhesion and mRNA translation in colorectal cancer cells. Proteomic analyses reveal that ACTB and SETD3 interact with multiple large protein complexes, including complexes associated with transcriptional regulation. Specifically, we demonstrate that SETD3-mediated ACTB methylation is required for the colocalization of SMARCA4, a subunit of the SWI/SNF BAF complex, at specific genomic loci. Genomic analyses further show that this colocalization enables the coordinated occupancy of SMARCA4 and H73-methylated ACTB at genes involved in cell adhesion and mRNA translation. Finally, phenotypic assays confirm these regulatory effects. Together, these findings uncover a new mechanistic layer of selective transcriptional regulation mediated by an ACTB-SETD3-SMARCA4 axis in colorectal cancer cells.
Small supernumerary marker chromosomes (sSMCs) remain a diagnostic challenge despite sequencing advances. As the field shifts toward cytogenomics, there is a need to establish methodologies to resolve these complex genetic variants at base pair resolution, as well as to identify their chromosomal origin and formation mechanism. Here, we apply long-read genome sequencing (lrGS) in combination with the telomere-to-telomere (T2T-CHM13) assembly to characterize the structure and genomic content of 10 clinically detected sSMCs. We use sequencing data to reconstruct the derivative chromosomes, identify breakpoint junctions (BPJs), and infer formation mechanisms. We resolve the BPJs of nine of the 10 sSMCs at base pair resolution. The analysis reveals six simple intrachromosomal rearrangements (one continuous and five discontinuous) with one to three BPJs, one complex three-way translocation with two BPJs, and two highly complex intrachromosomal rearrangements with five and nine BPJs, respectively. Breakpoint analysis reveals distinct mechanistic signatures: Simple sSMCs show features consistent with microhomology-mediated end joining (MMEJ) or microhomology-mediated break-induced replication (MMBIR), whereas complex sSMCs demonstrate evidence of translocation, chromoanasynthesis, and breakage-fusion-bridge (BFB) cycles. Haplotype analysis supports trisomy rescue in four cases, including all three complex sSMCs. In summary, our study demonstrates that lrGS combined with T2T-CHM13 enables detailed structural and mechanistic characterization of sSMCs, providing experimental support for disruption of trisomy rescue as a key formation mechanism. This work illustrates the feasibility of resolving highly challenging chromosomal abnormalities using long-read sequencing technologies.
Cellular identity is dictated by precise transcriptomic programs; however, the mechanisms that dynamically shape the transcriptome by producing diverse RNA isoforms remain incompletely understood. Here, through longitudinal transcriptomic profiling of cardiomyogenesis, we delineate pervasive RNA isoform switching across the developmental trajectory of cardiac differentiation. We show that the changes in isoform accumulation are largely decoupled from shifts in overall gene expression levels and arise predominantly from alternative transcription start and termination rather than splicing. This regulatory layer preferentially targets genes essential for cardiac function, specifically those encoding contractile machinery and ion channel complexes. We demonstrate that genes undergoing isoform switching without changing overall expression constitute a functionally distinct cohort, establishing isoform switching as an independent layer of gene regulation. Furthermore, we characterize stage-specific expression dynamics for numerous RNA-binding proteins and link their expression to specific isoform switching events. Our findings support a coordinated mechanism for the targeted regulation of RNA diversity, positioning isoform switching as a primary driver of transcriptomic maturation during lineage commitment.