Motivation Protein language models are critical for modeling antibody-antigen interactions, yet sequence-based affinity prediction remains a key challenge, particularly when structural data are scarce. Existing methods often struggle to fully exploit sequence information, limiting their applicability across diverse antibody formats such as single-domain antibodies (sdAbs).Results We propose dual-level protein representation for affinity prediction (DLP-Affinity), a dual-level deep learning framework for accurate sequence-based affinity prediction. It leverages two complementary modules: residue-to-residue to capture local interface contacts, and global stochastic projection embedding to represent global protein properties. Utilizing a fine-tuned protein language model, our approach achieves state-of-the-art performance on the general AB-Bind dataset (reducing mean absolute error by up to 20.9%) and delivers highly competitive results on the sdAb-DB dataset. This provides a robust tool for sequence-based antibody affinity prediction.Availability and implementation The source code and datasets for DLP-Affinity are freely available at https://github.com/Zy-Wang-bit/DLP_Affinity and archived on Zenodo at https://doi.org/10.5281/zenodo.18437656
Peptides are attractive candidates for drug development because of their low toxicity and relatively small binding interfaces, making accurate protein-peptide binding prediction crucial. In this study, we propose a flexible transformer-based framework with mutual attention that integrates protein pocket structural information and can be instantiated with different pocket-structure encoders. Within this unified framework, we systematically compare three encoders: an attention-based SE(3)-Transformer, a geometric graph neural network ProtGVP, and the large-scale structure-based pretrained model ESM-IF1. Using rigorous data partitioning with strict separation of training and test sets, we show that incorporating pocket structural information consistently improves binding prediction over sequence-only models, with GVP-GNN providing particularly effective pocket representations and structure-based variants exhibiting superior robustness on previously unseen data.
Accurate prediction of drug-target affinity (DTA) is crucial for drug discovery. Although deep learning-based methods have achieved promising results, most existing approaches rely on latent-space interaction modeling, which often ignores explicit geometric constraints and fails to capture higher-order interactions between protein residues and drug atoms. To address these limitations, we propose CHIMNet, a pose-aware multimodal hierarchical network with hypergraph interaction and confidence gates for DTA prediction. CHIMNet models interactions at multiple levels of abstraction: representation-level interactions are captured via bilinear attention over multimodal target and drug features, while interaction-level dependencies are explicitly modeled using pose-aware drug-target hypergraphs constructed from predicted binding conformations. To mitigate the uncertainty introduced by structural prediction, a confidence-gated fusion mechanism is further designed to adaptively regulate cross-level feature fusion. Experiments on the Davis and KIBA benchmarks demonstrate that CHIMNet consistently outperforms state-of-the-art methods, validating the effectiveness of hierarchical interaction modeling and uncertainty-aware feature fusion for robust DTA prediction.
Accurate drug-drug interaction (DDI) prediction is crucial for optimizing the efficacy of combination therapies and minimizing adverse effects. Most existing methods rely on single features and struggle to integrate structural and sequential drug information. Additionally, prediction bias caused by class imbalance remains a significant challenge. To address these issues, this study proposes a multi-source example-driven learning framework for DDI (MEDL-DDI) that jointly models structural and sequential drug representations to achieve robust multimodal fusion and mitigate class imbalance. MEDL-DDI enriches SMILES with chemical knowledge, extracts global semantic features via a Transformer, and identifies key substructures through a graph information bottleneck. Moreover, an example-driven mechanism guided by example centers enhances the model's ability to recognize minority classes. Experimental results on three benchmark datasets validate that MEDL-DDI outperforms state-of-the-art methods. The case study on cardiovascular drug interactions further highlights MEDL-DDI's practical value and applicability.
The prediction of binding free energy changes ($\Delta \Delta G$) caused by mutations in protein complexes is crucial for understanding disease mechanisms and designing antibodies. Approximately 60% of pathogenic missense mutations lead to functional abnormalities by disrupting molecular interactions. However, although existing $\Delta \Delta G$ predictors exhibit strong performance in benchmarks, they suffer from inadequate generalization, a misalignment between evaluation metrics and practical needs, and poor adaptability to complex mutation scenarios. This study systematically assessed eight mainstream predictors, covering both physical energy function-based and machine learning-based methods, and constructed an independent evaluation set. This study employed multi-dimensional metrics, including regression accuracy and classification capability, while also analyzing the performance variations of predictors across different mutation types, stability categories, and microenvironments of protein mutation sites. The results indicate that >60% of predictors (5 out of 8) predictors exhibit a systematic bias toward overestimating mutational instability. In the three-class classification task, predictors demonstrate a limited ability to identify stabilizing mutations ($\Delta \Delta G< -0.5$ kcal/mol), with recall rates <0.1 for this class, and overall predictive efficacy depends on the protein local structure. In summary, this study reveals the limitations of current $\Delta \Delta G$ predictors in terms of generalization and adaptability to complex scenarios, thus providing a reference for the optimization and practical application of $\Delta \Delta G$ prediction methods. It suggests that future breakthroughs can be achieved by constructing balanced and standardized datasets alongside developing local-global fusion algorithms.
Liquid-liquid phase separation plays a critical role in cellular processes, including protein aggregation and RNA metabolism, by forming membraneless subcellular structures. Accurate identification of phase-separated proteins is essential for understanding and controlling these processes. Traditional identification methods are effective but often costly and time-consuming. The recent machine learning methods have reduced these costs, but most models are restricted to classifying scaffold and client proteins with limited experimental conditions. To address this limitation, we developed a Mamba-based encoder using contrastive learning that incorporates separation probability, protein type, and experimental conditions. Our model achieved 95.2% accuracy in predicting phase-separated proteins and an ROCAUC score of 0.87 in classifying scaffold and client proteins. Further validation in the DgHBP-2 drug delivery system demonstrated its potential for condition modulation in drug development. This study provides an effective framework for the accurate identification and control of phase separation, facilitating advancements in biomedical research and therapeutic applications.
The success of drug discovery relies on predicting the binding affinity of protein-ligand. Applying deep learning to this field can expedite the process and reduce resource consumption. Recently, researchers have employed graph neural networks for predicting protein-ligand binding affinitiy, showcasing remarkable performance. However, this is largely attributed to the natural representation of biomolecule by graph neural networks, rather than a rational modeling of interactions within protein-ligand complex. In this regard, we have developed an Equivariant Interaction-aware Graph Network (EIGN), capable of learning 3D geometric structural information of complex while perceiving interactions related to protein-ligand binding affinity between nodes. Specifically, we designed distance-inspired edge-gated attention layer for inter-node interactions within the complex, uniformly learning interactions within and between molecules. To precisely simulate interactions between nodes, we considered local structural information around nodes when interactions occur. Leveraging equivariant convolutional layer to harness the advantages of learning geometric structure and drawing insights from existing work, we developed EIGN. Demonstrated on two benchmark sets, EIGN presents exceptional performance and generalization, highlighting the importance of accurate interaction modeling in drug discovery.
Protein-protein interactions (PPIs) refer to the phenomenon of protein binding through various types of bonds to execute biological functions. These interactions are critical for understanding biological mechanisms and drug research. Among these, the protein binding interface is a critical region involved in protein-protein interactions, particularly the hotspot residues on it that play a key role in protein interactions. Current deep learning methods trained on large-scale data can characterize proteins to a certain extent, but they often struggle to adequately capture information about protein binding interfaces. To address this limitation, we propose the PPI-Graphomer module, which integrates pretrained features from large-scale language models and inverse folding models. This approach enhances the characterization of protein binding interfaces by defining edge relationships and interface masks on the basis of molecular interaction information. Our model outperforms existing methods across multiple benchmark datasets and demonstrates strong generalization capabilities.
Proteins with specific functions and characteristics play a crucial role in biomedicine and nanotechnology. De novo protein design enables the customization of sequences to produce proteins with desired structures that do not exist in the nature. In recent years, with the rapid development of artificial intelligence (AI), deep learning-based generative models have increasingly become powerful tools, enabling the design of functional proteins with atomic-level precision. This article provides an overview of the evolution of de novo protein design, with focus on the latest algorithmic models, and then analyzes existing challenges such as low design success rates, insufficient accuracy, and dependence on experimental validation. Furthermore, this article discusses the future trends in protein design, aiming to provide insights for researchers and practitioners in this field.
Protein-protein interactions (PPI) play a crucial role in numerous key biological processes, and the structure of protein complexes provides valuable clues for in-depth exploration of molecular-level biological processes. Protein-protein docking technology is widely used to simulate the spatial structure of proteins. However, there are still challenges in selecting candidate decoys that closely resemble the native structure from protein-protein docking simulations. In this study, we introduce a docking evaluation method based on three-dimensional point cloud neural networks named SurfPro-NN, which represents protein structures as point clouds and learns interaction information from protein interfaces by applying a point cloud neural network. With the continuous advancement of deep learning in the field of biology, a series of knowledge-rich pre-trained models have emerged. We incorporate protein surface representation models and language models into our approach, greatly enhancing feature representation capabilities and achieving superior performance in protein docking model scoring tasks. Through comprehensive testing on public datasets, we find that our method outperforms state-of-the-art deep learning approaches in protein-protein docking model scoring. Not only does it significantly improve performance, but it also greatly accelerates training speed. This study demonstrates the potential of our approach in addressing protein interaction assessment problems, providing strong support for future research and applications in the field of biology.
Introduction: Protein engineering, which aims to improve the properties and functions of proteins, holds great research significance and application value. However, current models that predict the effects of amino acid substitutions often perform poorly when evaluated for precision. Recent research has shown that ProteinMPNN, a large-scale pre-training sequence design model based on protein structure, performs exceptionally well. It is capable of designing mutants with structures similar to the original protein. When applied to the field of protein engineering, the diverse designs for mutation positions generated by this model can be viewed as a more precise mutation range.Methods: We collected three biological experimental datasets and compared the design results of ProteinMPNN for wild-type proteins with the experimental datasets to verify the ability of ProteinMPNN in improving protein fitness.Results: The validation on biological experimental datasets shows that ProteinMPNN has the ability to design mutation types with higher fitness in single and multi-point mutations. We have verified the high accuracy of ProteinMPNN in protein engineering tasks from both positive and negative perspectives.Discussion: Our research indicates that using large-scale pre trained models to design protein mutants provides a new approach for protein engineering, providing strong support for guiding biological experiments and applications in biotechnology.
Background Natural proteins occupy a small portion of the protein sequence space, whereas artificial proteins can explore a wider range of possibilities within the sequence space. However, specific requirements may not be met when generating sequences blindly. Research indicates that small proteins have notable advantages, including high stability, accurate resolution prediction, and facile specificity modification. Results This study involves the construction of a neural network model named TopoProGenerator(TPGen) using a transformer decoder. The model is trained with sequences consisting of a maximum of 65 amino acids. The training process of TopoProGenerator incorporates reinforcement learning and adversarial learning, for fine-tuning. Additionally, it encompasses a stability predictive model trained with a dataset comprising over 200,000 sequences. The results demonstrate that TopoProGenerator is capable of designing stable small protein sequences with specified topology structures. Conclusion TPGen has the ability to generate protein sequences that fold into the specified topology, and the pretraining and fine-tuning methods proposed in this study can serve as a framework for designing various types of proteins.
With advanced computational methods, it is now feasible to modify or design proteins for specific functions, a process with significant implications for disease treatment and other medical applications. Protein structures and functions are intrinsically linked to their backbones, making the design of these backbones a pivotal aspect of protein engineering. In this study, we focus on the task of unconditionally generating protein backbones. By means of codebook quantization and compression dictionaries, we convert protein backbone structures into a distinctive coded language and propose a GPT-based protein backbone generation model, PB-GPT. To validate the generalization performance of the model, we trained and evaluated the model on both public datasets and small protein datasets. The results demonstrate that our model has the capability to unconditionally generate elaborate, highly realistic protein backbones with structural patterns resembling those of natural proteins, thus showcasing the significant potential of large language models in protein structure design.
Protein-protein interactions (PPIs) play essential roles in many vital movements and the determination of protein complex structure is helpful to discover the mechanism of PPI. Protein-protein docking is being developed to model the structure of the protein. However, there is still a challenge to selecting the near-native decoys generated by protein-protein docking. Here, we propose a docking evaluation method using 3D point cloud neural network named PointDE. PointDE transforms protein structure to the point cloud. Using the state-of-the-art point cloud network architecture and a novel grouping mechanism, PointDE can capture the geometries of the point cloud and learn the interaction information from the protein interface. On public datasets, PointDE surpasses the state-of-the-art method using deep learning. To further explore the ability of our method in different types of protein structures, we developed a new dataset generated by high-quality antibody-antigen complexes. The result in this antibody-antigen dataset shows the strong performance of PointDE, which will be helpful for the understanding of PPI mechanisms.
The principal goal of drug design is to find ligand molecules that exhibit affinity to a given target protein. In recent years, deep generative methods have shown their promise in de novo drug design. However, most of these methods design molecules based on target-specific ligand datasets instead of targets’ features and fail to design drugs against novel target proteins that barely have active ligand datasets. A fast and relatively accurate evaluation method is needed to evaluate algorithms capable of generating large numbers of molecules. In this work, we treat target-specific de novo drug design as a sequence-to-sequence generation task and propose a Transformer architecture that compensates for the lack of training data with a BERT pretraining approach to generate protein sequence-conditioned Target Ligand Molecules SMILES. First, we pre-train two self-attention blocks of Transformer on the large-scale amino acid sequence dataset and molecular SMILES dataset, respectively, to capture the feature representation of the target. Then we fine-tune the Transformer’s encoder-decoder mutual attention block on the protein-ligand complex dataset to learn conditional generation using autoregressive supervised learning. The individual results do not demonstrate the effect of the generative algorithm, so we propose to evaluate the model by calculating the affinity distribution of the molecules. We also evaluate our method by designing ligands against three well-studied proteins. Furthermore, our model proposes molecules with binding affinities exceeding certain FDA-approved drugs in docking experiments.
Artificial intelligence, such as deep generative methods, represents a promising solution to de novo design of molecules with the desired properties. However, generating new molecules with biological activities toward two specific targets remains an extremely difficult challenge. In this work, we conceive a novel computational framework, herein called dual-target ligand generative network (DLGN), for the de novo generation of bioactive molecules toward two given objectives. Via adversarial training and reinforcement learning, DLGN treats a sequence-based simplified molecular input line entry system (SMILES) generator as a stochastic policy for exploring chemical spaces. Two discriminators are then used to encourage the generation of molecules that belong to the intersection of two bioactive-compound distributions. In a case study, we employ our methods to design a library of dual-target ligands targeting dopamine receptor D2 and 5-hydroxytryptamine receptor 1A as new antipsychotics. Experimental results demonstrate that the proposed model can generate novel compounds with high similarity to both bioactive datasets in several structure-based metrics. Our model exhibits a performance comparable to that of various state-of-the-art multi-objective molecule generation models. We envision that this framework will become a generally applicable approach for designing dual-target drugs in silico.
Enhancer-promoter interactions (EPIs) play an important role in transcriptional regulation. Recently, machine learning-based methods have been widely used in the genome-scale identification of EPIs due to their promising predictive performance. In this paper, we propose a novel method, termed EPI-DLMH, for predicting EPIs with the use of DNA sequences only. EPI-DLMH consists of three major steps. First, a two-layer convolutional neural network is used to learn local features, and an bidirectional gated recurrent unit network is used to capture long-range dependencies on the sequences of promoters and enhancers. Second, an attention mechanism is used for focusing on relatively important features. Finally, a matching heuristic mechanism is introduced for the exploration of the interaction between enhancers and promoters. We use benchmark datasets in evaluating and comparing the proposed method with existing methods. Comparative results show that our model is superior to currently existing models in multiple cell lines. Specifically, we found that the matching heuristic mechanism introduced into the proposed model mainly contributes to the improvement of performance in terms of overall accuracy. Additionally, compared with existing models, our model is more efficient with regard to computational speed.
Enhancer-promoter interactions (EPIs) in the human genome are of great significance to transcriptional regulation, which tightly controls gene expression. Identification of EPIs can help us better decipher gene regulation and understand disease mechanisms. However, experimental methods to identify EPIs are constrained by funds, time, and manpower, while computational methods using DNA sequences and genomic features are viable alternatives. Deep learning methods have shown promising prospects in classification and efforts that have been utilized to identify EPIs. In this survey, we specifically focus on sequence-based deep learning methods and conduct a comprehensive review of the literature. First, we briefly introduce existing sequence-based frameworks on EPIs prediction and their technique details. After that, we elaborate on the dataset, pre-processing means, and evaluation strategies. Finally, we concluded with the challenges these methods are confronted with and suggest several future opportunities. We hope this review will provide a useful reference for further studies on enhancer-promoter interactions.
Accurate prioritization of potential disease genes is a fundamental challenge in biomedical research. Various algorithms have been developed to solve such problems. Inductive Matrix Completion (IMC) is one of the most reliable models for its well-established framework and its superior performance in predicting gene-disease associations. However, the IMC method does not hierarchically extract deep features, which might limit the quality of recovery. In this case, the architecture of deep learning, which obtains high-level representations and handles noises and outliers presented in large-scale biological datasets, is introduced into the side information of genes in our Deep Collaborative Filtering (DCF) model. Further, for lack of negative examples, we also exploit Positive-Unlabeled (PU) learning formulation to low-rank matrix completion. Our approach achieves substantially improved performance over other state-of-the-art methods on diseases from the Online Mendelian Inheritance in Man (OMIM) database. Our approach is 10 percent more efficient than standard IMC in detecting a true association, and significantly outperforms other alternatives in terms of the precision-recall metric at the top-k predictions. Moreover, we also validate the disease with no previously known gene associations and newly reported OMIM associations. The experimental results show that DCF is still satisfactory for ranking novel disease phenotypes as well as mining unexplored relationships. The source code and the data are available at https://github.com/xzenglab/DCF.
Anticancer peptides (ACPs) eliminate pathogenic bacteria and kill tumor cells, showing no hemolysis and no damages to normal human cells. This unique ability explores the possibility of ACPs as therapeutic delivery and its potential applications in clinical therapy. Identifying ACPs is one of the most fundamental and central problems in new antitumor drug research. During the past decades, a number of machine learning-based prediction tools have been developed to solve this important task. However, the predictions produced by various tools are difficult to quantify and compare. Therefore, in this article, a comprehensive review of existing machine learning methods for ACPs prediction and fair comparison of the predictors is provided. To evaluate current prediction tools, a comparative study was conducted and analyzed the existing ACPs predictor from the 10 public works of literature. The comparative results obtained suggest that the Support Vector Machine-based model with features combination provided significant improvement in the overall performance when compared to the other machine learning method-based prediction models.