Information bottleneck (IB) leverages information theory to guide the learning process of deep multi-view clustering (MVC). It optimizes the trade-off between multi-view compression and preservation by minimizing and maximizing mutual information (MI). Although existing deep MVC based on IB has witnessed great achievements, they usually resort to variational inference to estimate the MI lower bound, which typically introduces estimation errors and results in an unstable lower bound of MI. In this study, we propose a novel Cauchy-Schwarz Mutual Information Maximin (CS-MIM) method, which directly estimates MI with closed-form expressions without requiring variational inference, possessing explicit multi-view information modeling capabilities. Specifically, we first present a non-parametric MI estimation method with Cauchy-Schwarz (CS) divergence, which leverages multi-kernel Gram matrices to capture distributional similarities and avoids the approximation errors introduced by the neural estimators of variational inference. Then, based on the new estimation method, a MI maximin mechanism is devised to parameterize the IB principle with analytical gradients, which facilitates effective compression of multi-view data while preserving the relevant features. Finally, we design a cross-view adaptive attention (CAA) mechanism constrained by the MI based on CS divergence, which further captures the complementarity across views under the guidance of the fused multi-view representation. Extensive evaluation results on 12 public available datasets demonstrate that the CS-MIM remarkably outperforms existing SOTA approaches.
Clustering is a common technique for statistical data analysis and is essential for developing precision medicine. Numerous computational methods have been proposed for integrating multi-omics data to identify cancer subtypes. However, most existing clustering models based on network fusion fail to preserve the consistency of the distribution of the data before and after fusion. Motivated by this observation, we would like to measure and minimize the distribution difference between networks, which may not be in the same space, to improve the performance of data fusion. We were therefore motivated to develop a flexible clustering model, based on network fusion, that minimizes the distribution difference between the data before and after fusion by co-regularization; the model can be applied to both single- and multi-omics data. We propose a new network fusion model for single- and multi-omics data clustering for identifying cancer or cell subtypes based on co-regularized network fusion (SMCC). SMCC integrates low-rank subspace representation and entropy to fuse networks. In addition, it measures and minimizes the distribution difference between the similarity networks and the fusion network by co-regularization. The model can both reduce the noise interference in the source data and make the statistical characteristics of the fusion result closer to those of the source data. We evaluated the clustering performance of SMCC across 16 real single- and multi-omics dataset. The experimental results demonstrated that SMCC is superior to 17 state-of-the-art clustering methods. Moreover, it is effective for identifying cancer or cell subtypes, thereby promoting the development of precision medicine.
Accurate cell type annotation across datasets is a key challenge in single-cell analysis. snRNA-seq enables profiling of frozen or difficult-to-dissociate tissues, complementing scRNA-seq by capturing fragile or rare cell types. However, cross-annotation between these two datasets remains largely unexplored, as existing methods treat them independently. We introduce ScNucAdapt, a method designed for cross-annotation between paired and unpaired scRNA-seq and snRNA-seq datasets. To address distributional and cell composition differences, ScNucAdapt employs partial domain adaptation. Experiments across both unpaired and paired scRNA-seq and snRNA-seq show that ScNucAdapt achieves robust and accurate cell type annotation, outperforming existing approaches. Therefore, ScNucAdapt provides a practical framework for the cross-domain cell type annotation between scRNA-seq and snRNA-seq data.
Molecular property prediction is essential in drug discovery for early-stage compound evaluation. Recently, contrastive learning has demonstrated significant potential under limited labeled data by constructing augmented views. However, current augmentation strategies often disrupt molecular semantics and ignore chemical priors, limiting representation quality. Moreover, molecular data is inherently multimodal, including graphs, fingerprints, and sequences, yet how to effectively integrate their complementary information remains challenging. Therefore, we propose MPMFMol, a unified framework that integrates multitask self-supervised pretraining with multimodal fine-tuning for molecular property prediction. During pretraining, we construct heterogeneous augmented views based on molecular fragments to preserve original molecular semantics, enabling the graph encoder to capture fragment-level information. Meanwhile, fingerprint features are integrated into a multitask learning objective, reducing reliance on negative sampling and enhancing the encoder's representation capability. During fine-tuning, we further incorporate functional group and SMILES sequence information and design a stage-aware modality fusion strategy. Specifically, pretrained graph features are injected into the initial representation of functional groups to guide feature extraction and then fused with SMILES features to enable deep cross-modal interaction and enhance downstream predictive performance. Experimental results on six classification and three regression data sets demonstrate that MPMFMol outperforms state-of-the-art baselines.
Gram-negative bacterial secreted effectors are translocated through specialized secretion systems to manipulate host cellular processes, and their accurate identification is crucial for understanding bacterial pathogenesis. Recent deep learning methods have significantly advanced this field, yet current approaches primarily rely on global sequence representations, overlooking the biological significance of terminal regions where secretion signals reside. Moreover, severe class imbalance among different secreted effector types remains a critical challenge for multi-class prediction. Here, we propose TermSE, a terminal signal-aware framework for multi-class secreted effector identification. TermSE explicitly captures N-terminal and C-terminal sequence features through convolutional neural networks applied to protein language model embeddings, and integrates them with global sequence representations for multi-view sequence characterization. To address class imbalance, TermSE employs a cosine-normalized classifier combined with weighted sampling to mitigate feature magnitude bias and ensure sufficient learning from minority classes. Extensive experiments demonstrate that TermSE outperforms existing methods in both cross-validation and independent test settings, with robust generalization across varying sequence identity levels. Furthermore, interpretability analysis confirms that TermSE learns to focus on biologically meaningful terminal patterns specific to each secreted effector type. These results highlight the potential of TermSE as an effective and interpretable tool for secreted effector discovery.
Neoadjuvant immune checkpoint blockade (NICB) therapy has shown significant efficacy in oral squamous cell carcinoma (OSCC); however, the mechanisms by which intracellular microbiota influence immune function within the tumor microenvironment remain unclear. In this study, we employ invasion-adhesion-directed expression sequencing (INVADEseq) technology to simultaneously capture host single-cell RNA sequencing data and bacterial signatures, revealing that intracellular bacteria inhibit PDCD1 expression, enhance infection responses, promote antigen presentation and macrophage activation, and reduce T cell-related immune gene expression. The key genera, Fusobacterium, Streptococcus, and Capnocytophaga, are associated with increased risks and adverse outcomes in immunotherapy across multiple tumor types. We identify GZMK+ CD8+ T cells and ZBTB16+ TAMs as markers of complete response to NICB. However, intracellular bacteria weaken the communication between cDC1 and these immune cells, potentially reducing therapeutic efficacy. Our study provides a foundation for further investigation into the mechanisms by which intracellular bacteria mediate effective responses to NICB in OSCC.
The rapid advancement of high-throughput sequencing technologies has resulted in an explosive growth of biolog ical sequence data, making sequence clustering a fundamental task in large-scale bioinformatics analyses. Unlike traditional clustering problems, biological sequence clustering faces unique challenges arising from the absence of direct similarity measures, strict biological constraints, and demanding requirements for scalability and accuracy. Over the past decades, numerous methods have been proposed, differing in how they model sequence similarity, construct clusters, and prioritize optimization objectives. In this review, we present a methodological overview of biological sequence clustering algorithms. We first summarize major strategies for similarity modeling, which can be divided into three stages: sequence encoding, feature generation, and similarity measurement. We then review the principal clustering paradigms, including greedy incremental, hierarchical, graph-based, model-based, partitional, and deep learning approaches, highlighting their methodological features and practical trade-offs. In addition, we discuss clustering objectives from three perspectives: scalability and resource efficiency, biological interpretability, and robustness of clustering quality. By organizing existing methods along these dimensions, we elucidate the fundamental trade-offs underlying biological sequence clustering and clarify the application con texts in which different approaches are most suitable. Finally, we discuss current limitations and open challenges, offering guidance for future method development.
2'-O-methylation (2OM) of ribose is a widespread RNA modification that significantly impacts RNA stability, structure, and function. Accurately predicting 2OM sites is crucial for understanding RNA's biological functions and related pathologies. Traditional detection methods pose challenges such as resource intensiveness, potential RNA sample damage, and high costs. However, recent advancements in machine learning, particularly deep learning techniques, offer rapid and cost-effective prediction solutions. In this study, we introduce DeepR2OM, a novel method integrating feature selection and deep learning for 2OM sites prediction. DeepR2OM encodes sequences using eight RNA descriptors, employs feature selection algorithms to reduce dimensions, and then utilizes a deep learning network for training. After evaluating various deep learning architectures, we selected Convolutional Neural Network (CNN), Multi-Head Self-Attention mechanism, and Deep Neural Network (DNN) as our final prediction models. Experimental results demonstrate DeepR2OM's effectiveness, achieving 87.1% accuracy (ACC), 85.5% recall rate (Recall), 87.9% precision (PRE), and a Matthews correlation coefficient (MCC) of 75.7% on an independent test set. This tool serves as a valuable resource for exploring the functional and bioinformatic aspects of 2OM sites.
Liquid-liquid phase separation (LLPS) is a key mechanism driving the assembly of membrane-less organelles and is increasingly recognized for its involvement in essential cellular functions and various diseases. However, existing computational approaches largely rely on sequence-level descriptors and often fail to explicitly incorporate structural topology information, limiting their ability to capture the complex determinants of LLPS behavior. Accurate identification of LLPS-capable proteins remains challenging due to their sequence diversity and complex structural determinants. Here, we present MuFGPS (Multi-level Feature Graph-based Predictor for Phase-Separating proteins), a predictive framework integrating sequence-derived physicochemical features, Define Secondary Structure of Proteins-annotated secondary structures, and graph-based structural embeddings from AlphaFold residue contact maps via a multi-head Graph Attention Network. Class imbalance is addressed using Synthetic Minority Oversampling Technique (SMOTE), and classification is performed through a stacking ensemble of Random Forest, XGBoost, and LightGBM. Benchmarks against six representative methods demonstrate that MuFGPS achieves superior performance across all metrics, with notable gains in F1-score and matthews correlation coefficient (MCC). Ablation analyses confirm the synergistic contributions of structural features and ensemble learning to accuracy and robustness. MuFGPS offers a scalable and high-accuracy framework for proteome-wide LLPS protein prediction.
DNA N4-methylcytosine (4mC), a key epigenetic modification regulating DNA repair and replication, requires efficient computational detection methods due to experimental limitations. Although machine learning predictors have been proposed, their performance could be enhanced through systematic optimization of feature encoding schemes. Here, we propose EnDeep4mC, a dual-adaptive framework integrating species-specific modeling with ensemble deep learning architectures to systematically optimize feature encoding schemes. Evaluated across six species, EnDeep4mC demonstrates commendable prediction performance and significantly outperforms current state-of-the-art predictors. Cross-species validation confirms its robust transferability from animal to microbe groups. Evolutionary analysis further uncovers the functional differentiation of 4mC sequences in biological evolution: Prokaryotic 4mC relies on stable patterns, whereas eukaryotes achieve regulatory plasticity through dynamic sequence combinations, which provides experimental evidence for species-adaptive encoding strategies.
Partial order alignment (POA) has emerged as a fundamental component in long-read error correction, assembly and pangenomics. However, conventional POA algorithms are limited by high time and memory requirements, making them inefficient for large-scale datasets. Here, we present minipoa, a fast and memory-efficient POA tool that incorporates seed-chain-align heuristics, adaptive or static banding strategies, and single-instruction multiple-data optimizations. Minipoa achieves up to a 5-fold speedup over abPOA, reduces memory usage by up to 16-fold, and improves correction accuracy, while maintaining strong performance on both PacBio and ONT simulated datasets, and can be readily integrated into existing long-read error correction and assembly workflows. In multiple sequence alignment datasets, minipoa demonstrates superior computational efficiency and alignment accuracy compared with all other tested tools, achieving Total Column scores up to 2.5-fold higher than MAFFT in low-similarity scenarios. Moreover, minipoa enables multiple sequence alignment of megabase-long genomes and million-sequence datasets, demonstrated by 342 Mycobacterium tuberculosis sequences and one million SARS-CoV-2 sequences respectively. Collectively, minipoa is well positioned to become a cornerstone in the era of large-scale pangenomics.
Motivation The expression of circular RNAs (circRNAs) has been shown to be strongly correlated with drug sensitivity in human cells. However, experimental validation using wet-lab techniques is costly and inefficient, leaving a substantial portion of circRNA-drug sensitivity associations undiscovered. Therefore, improving the prediction efficiency of circRNA and sensitivity associations remains critical.Methods Here, we describe a method that integrates collaborative feature learning and graph structure learning to predict associations between circRNAs and drug sensitivity (CFGSCDSA). Specifically, collaborative learning integrated heterogeneous features from diverse data sources, thereby addressing the issue of data sparsity. Furthermore, graph structure learning with a confidence-guided pseudo-labeling strategy was employed to mitigate the detrimental effect of excessive negative samples. Results: Experimental evaluation revealed that CFGSCDSA attained superior performance compared to all competing models. Moreover, case studies provided further evidence of its capability to accurately predict both novel associations and new drug-related links.
Motivation Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines.Results We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency.Availability The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998
Polypharmacy has become essential in managing complex and chronic conditions, yet it introduces significant clinical risks due to potential drug-drug interactions (DDIs). Existing computational models struggle to provide accurate and interpretable predictions, largely due to fragmented integration of molecular structures and biomedical knowledge. These limitations arise from static graph designs, inadequate substructure modeling, and insufficient incorporation of semantic context, highlighting the need for a unified, explainable solution. In this study, we propose MKGFlow-DDI, a multi-view knowledge-guided framework that jointly leverages drug-drug interaction networks and biomedical knowledge graphs to dynamically construct drug-flow subgraphs. The model incorporates a dual-channel encoder designed to capture atom-level information and substructure-level features, integrating them with the global semantic embeddings of the composite network to derive novel feature representations. These representations are fused to initialize node features within each subgraph, which are iteratively optimized through similarity-based edge refinement to reduce noise and enhance biological relevance. To further improve generalization and stability, a contrastive learning module is introduced to align representations of perturbed subgraphs by maximizing consistency across positive and negative sample pairs. Experimental results on DrugBank and TWOSIDES demonstrate that MKGFlow-DDI outperforms state-of-the-art baselines, especially in scenarios involving previously unseen drugs. Additionally, the model produces interpretable semantic pathways that align with known pharmacological mechanisms, enabling clinically meaningful insights. Overall, MKGFlow-DDI establishes a robust and biologically grounded approach to DDI prediction, offering a promising direction for computational pharmacovigilance and personalized therapy optimization.
Accurate prediction of drug-target affinity (DTA) is essential for accelerating drug discovery. Although pretrained protein language models have achieved significant progress, existing methods predominantly focus on bottom-up sequence patterns and lack explicit constraints from high-level biological functions. We propose GoMA-DTA, a framework integrating gene ontology (GO) functional annotations with protein semantic features. GoMA-DTA introduces a channelwise gating mechanism that uses functional semantics as anchors to dynamically recalibrate ESM-2embeddings, achieving adaptive semantic filtering. For drugs, the model integrates Molformer-based semantic and TransConv-derived structural features. These dual-modality drug representations interact with calibrated protein features through a parallel synergistic architecture of cross-attention and Mamba modules, ensuring precise cross-modal alignment and efficient long-range dependency modeling. Evaluations on PDBBind, BindingDB, and ChEMBL benchmarks demonstrate that GoMA-DTA significantly outperforms state-of-the-art models across various evaluation scenarios. Its superior screening power is further validated on CASF-2016. Moreover, virtual screening of 200 000compounds against the SARS-CoV-2Spike protein, supported by experimental evidence (ZINC2111387), underscores its practical utility as a robust and biologically reliable tool. The datasets and codes are publicly available at https://github.com/xa-123955/GoMA-DTA.
Drug-drug interactions (DDIs) can lead to severe adverse reactions, and accurate prediction of DDI events is crucial for ensuring the safety of combination therapies and supporting drug development. Although deep learning-based approaches have achieved promising progress, existing models remain limited in modeling local-global dependencies, integrating multimodal information, and capturing cross-level molecular relationships. To address these challenges, we propose Multi-modal Hierarchical Attention Fusion and Relation-aware Architecture for DDI Event Prediction (MHAFR-DDI), a multimodal hierarchical attention fusion and relation-aware framework that enables unified modeling from intra-molecular representation to inter-molecular interaction. MHAFR-DDI adopts a two-stage pretraining-finetuning paradigm. In the pretraining stage, the model learns complementary representations from molecular sequences, 2D topological structures, and 3D spatial conformations, with modality-specific encoding mechanisms designed to capture both local structural characteristics and global semantic dependencies. Localized chemical primitives within each modality are first stabilized and then integrated into higher-level representations to ensure intra-modality stability and representational completeness. Subsequently, by introducing attention-guided data augmentation and multi-level contrastive learning, the model establishes alignment constraints across different modalities and their augmented views, thereby achieving cross-modal semantic consistency and effectively alleviating data sparsity. During the finetuning stage, the pretrained molecular representations are hierarchically fused and propagated over the drug-drug interaction graph, enabling interaction-aware information sharing among drugs and improving prediction reliability for rare drugs and long-tail interaction types. Experiments on benchmarks with 65 and 86 DDI types show that MHAFR-DDI outperforms state-of-the-art methods under the standard split, achieving macro-F1 gains of 9.5% and 6.7%, while remaining robust in weakly supervised long-tail and cold-start settings.
As viral sequencing datasets grow to hundreds of thousands or even millions of genomes, widely used MSA pipelines require substantial computational resources, limiting their practical scalability. Reference-based MSA helps by using a reference sequence as a fixed coordinate system: each sequence is mapped or pairwise-aligned to the reference, and insertions relative to the reference are usually removed during merging to keep coordinates consistent. Even so, at the million-sequence scale, widely used tools such as MAFFT can still be limited by runtime and memory use. We therefore updated HAlign4 by adding a keep-length mode to preserve reference coordinates, replacing suffix-array–based homologous segment search with minimizer seeding and chaining, switching the segment alignment kernel from the wavefront alignment to the ksw2 pairwise aligner, and allowing a user-provided or consensus-derived center sequence. Experiments show that the updated HAlign4 improves alignment quality on low-similarity datasets and can align one million SARS-CoV-2 genomes using only 6.1 GB RAM and 0.21 h. At comparable alignment quality, it is 7.1× faster and uses 54.8× less memory than MAFFT.
Abstract Accurate prediction of effector proteins secreted by Gram-negative bacteria is important for elucidating bacterial pathogenic mechanisms and developing precise anti-infective strategies. Although existing methods have benefited from the strong sequence feature extraction capacity of pretrained protein language models, reliance on linear sequence information alone often fails to fully capture the three-dimensional conformational signals required for virulence functions. Meanwhile, conventional structure-based methods are limited by the scarcity of experimentally resolved protein structures. To address these challenges, We propose GeoEPred, a multimodal deep learning framework designed for the synergistic modeling of protein sequence and structure to identify Gram-negative bacterial effector proteins. Specifically, the model integrates sequence-contextual embeddings from a pretrained protein language model with three-dimensional structural representations predicted by ESMFold. A feature projection network refines fine-grained sequence signals associated with effector functions, while geometric vector perceptrons characterize inter-residue orientations, distances, and local spatial topology to capture potential structural conformational motifs. To further enable effective cross-modal fusion, we design a cross-modal alignment and feature-tokenized self-attention module. This module enhances consistency between the sequence-semantic and structural-geometric spaces through contrastive learning and models associations between linear functional motifs and spatial conformational patterns at a fine-grained token level. Extensive evaluations on multiple benchmark datasets show that GeoEPred achieves better predictive performance than existing leading models in T3SE, T4SE, and T6SE prediction tasks, while maintaining stable performance in remote homolog recognition scenarios. Moreover, the modular and extensible architecture of GeoEPred demonstrates strong generalization ability and substantial application potential for genome-scale effector protein discovery. Author summary Secreted effector proteins are central virulence factors used by many Gram-negative bacterial pathogens to execute infection strategies. Their functions are governed not only by secretion signals and short linear motifs in the amino acid sequence, but also by three-dimensional folds, local domains, and surface geometric patterns. However, current predictors mainly exploit sequence-contextual features, limiting their ability to model the correspondence between linear sequence signals and spatial conformational motifs, and thereby constraining accuracy and interpretability. Here, we present GeoEPred, a multimodal deep learning framework for secreted effector protein identification. GeoEPred couples sequence-semantic embeddings from a pretrained protein language model with structural representations learned by geometric vector perceptrons. A cross-modal alignment and interaction module uses contrastive learning to improve functional consistency between sequence and structure modalities, while feature-token attention captures fine-grained links between key linear and conformational motifs. Across benchmark datasets covering multiple effector types, GeoEPred outperforms existing state-of-the-art methods and provides interpretable evidence from sequence fragments, structural regions, and cross-modal associations, supporting functional annotation, pathogenic mechanism analysis, and experimental validation.