Accurately predicting virus-host protein interactions(VH-PPI) and their binding sites is essential for understanding viral pathogenic mechanisms and developing drugs and vaccines. Existing sequence- or structure-based approaches still have limitations in extracting dynamic information and identifying binding sites, which constrains their practical applications. In this study, we propose a multimodal data-driven graph convolutional neural network model named DMVHP-IBS, which integrates dynamic and static protein information. Through the fusion of protein dynamic and structural attributes into graph data, alongside the encoding and feature extraction from protein sequences, the model assembles multimodal data for the prediction of viral-host protein interactions and binding sites. To further elucidate the binding mechanisms between viral-host proteins, we introduce a biologically inspired binding site prediction method called Gradient-Enhanced Interaction Contribution Analysis (GEICA) which can highlight key binding residues. The results show that DMVHP-IBS outperforms state-of-the-art methods across various viral datasets, demonstrating its generalizability and robustness. In the binding site prediction, we successfully identified key binding sites between VH-PPI, and between proteins and drugs by leveraging the self-attention mechanisms of GEICA and ProtBERT. DMVHP-IBS is useful in the design of drugs, targeted therapeutics, and antibodies.
Protein S-acylation is the addition of fatty acids to the cysteine residues in a protein, catalyzed by protein S-acyltransferases (PATs). Despite extensive research on protein S-acylation in animals, our understanding of this process in plants remains limited. In this study, we sought to characterize the S-acylproteome of membrane proteins in Arabidopsis and identify potential substrates for two important plant immunity-related PATs (PAT5 and PAT9). To achieve this, S-acylated membrane proteins were first enriched via our optimized acyl-biotinyl exchange strategies at both the protein-level and peptide-level. The enriched samples were then analyzed by label-free quantitative liquid chromatography-mass spectrometry. The results from the two enrichment methods demonstrated that they were complementary in identifying S-acylated proteins and S-acylation sites. Using these methods, over 2500 S-acylation sites in more than 2000 putative S-acylated proteins were identified. Proteins involved in vesicle trafficking, plant phosphorylation, immune responses, and signal transduction pathways were significantly enriched. Additionally, certain amino acid patterns surrounding the S-acylation sites were identified. Comparisons of the S-acylproteomes between the wild type and the PAT5 and PAT9 mutants revealed over 100 potential substrates for both S-acyltransferases. The high quality of our data was supported by the significant overlap with the previously reported data and successful experimental verification of selected candidate proteins. Overall, our study revealed a well-represented S-acylproteome for Arabidopsis (especially its membrane proteins) and identified potential substrates for PAT5 and PAT9. These findings will facilitate the functional characterization of S-acylated proteins in plants.
Protein Representation Learning (PRL) holds signif-icant value in various fields. However, existing methods primarily focus on static amino acid sequences or structures of proteins, paying less attention to their dynamic behaviors, which limits their ability to capture the intrinsic properties of proteins. In this paper, a novel multimodal protein representation learning method named IMPDI is proposed, which integrates amino acid sequences, structural information, and multiple dynamic information to construct a unified protein representation learning framework. To validate the effectiveness of the IMPDI model, we applied it to three typical protein-related downstream tasks: Protein-Protein Interaction (PPI) Prediction, Protein Secondary Structure Prediction, and Protein Thermostability Prediction. The results indicate that IMPDI outperforms baseline methods in various metrics, demonstrating strong generalizability and application potential. This study not only offers a new perspective on protein representation methods, but also provides robust technical support for related downstream tasks.
Drug-Target Interaction (DTI) prediction plays a pivotal role in accelerating drug discovery and development by identifying novel interactions between drugs and targets. Most previous studies on Drug-Protein Pair (DPP) networks have primarily focused on learning their topological structures. However, two key challenges remain: the integration of topological and semantic information is often insufficient, and the representation diversity may be diminished during graph convolution operations, affecting the expressiveness of learned features. To address the above challenges, we propose a novel paradigm named Multi-view Based Heterogeneous Graph Contrastive Learning for Drug-Target Interaction Prediction (HGCML-DTI). Specifically, we initially establish a drug-protein heterogeneous graph, followed by employing a weighted Graph Convolutional Network (GCN) to derive vector representations for both drug and protein nodes. Subsequently, we individually construct the topology and semantic graphs for DPP and integrate them to form a unified public graph. A multi-channel graph neural network is employed to learn DPP representations. To preserve representation diversity and enhance discriminative ability, a multi-view contrastive learning strategy is introduced. Then, a Multilayer Perceptron (MLP) neural network is used to recognize DTI. To prove the effectiveness of this work, extensive experiments are conducted on six real-world datasets, and comparisons are made with seven competitive baselines. The results demonstrate that the proposed HGCML-DTI significantly outperforms state-of-the-art methods. This work highlights the importance of combining multi-view learning and contrastive strategies to advance the field of DTI prediction. Source codes are available at https://github.com/7A13/HGCML-DTI.
In eukaryotic organisms, protein kinases regulate diverse protein activities and signaling pathways through phosphorylation of specific protein substrates. Isolating and characterizing kinase substrates is vital for defining downstream signaling pathways. The Kinase Client (KiC) assay is an in vitro synthetic peptide LC-MS/MS phosphorylation assay that has enabled identification of protein substrates (i.e., clients) for various protein kinases. For example, previous use of a 2,100-member (2k) peptide library identified substrates for the extracellular ATP receptor-like kinase, P2K1. Many P2K1 clients were confirmed by additional in vitro and in planta studies, including Integrin-Linked Kinase 4 (ILK4), for which we provide the evidence herein. In addition, we developed a new KiC peptide library containing 8,000 (8k) peptides based on phosphorylation sites primarily from Arabidopsis thaliana datasets. The 8k peptides are enriched for sites with conservation in other angiosperm plants, with the paired goals of representing functionally conserved sites and usefulness for screening kinases from diverse plants. Screening the 8k library with the active P2K1 kinase domain identified 177 phosphopeptides, including calcineurin B-like protein (CBL9) and G protein alpha subunit 1 (GPA1), which functions in cellular calcium signaling. We confirmed that P2K1 directly phosphorylates CBL9 and GPA1 through in vitro kinase assays. This expanded 8k KiC assay will be a useful tool for identifying novel substrates across diverse plant protein kinases, ultimately facilitating the exploration of previously undiscovered signaling pathways.
Existing breast cancer prognosis prediction models often rely on single-type data such as clinical or mRNA expression data, neglecting the prognostic value of gene mutations linked to tumor aggressiveness. Multi-omics integration remains difficult due to the sparsity of mutation data and the high dimensionality of expression data. To address these limitations, we propose ICMProg, an interpretable framework for breast cancer prognosis prediction by integrating clinical and multi-omics data. ICMProg uses a Frequency-aware Sparse Variational Autoencoder (FS-VAE) to extract informative features from sparse mutation data, and a Dynamic Pathway-Guided Variational Autoencoder (DPG-VAE) to capture biologically meaningful features from high-dimensional expression data. For model interpretability, we develop an interpretable method based on Grad-SHAP, systematically identifying critical biomarkers in clinical, mutational, gene, and pathway data. Experimental results on two breast cancer datasets show that ICM-Prog outperforms state-of-the-art methods, and it can significantly separate high-risk and low-risk patient groups. These results suggest that ICM-Prog can not only improve prognostic accuracy, but also support the discovery of potential therapeutic targets.
Graph contrastive learning, which aims to learn supervised signals from unlabeled graph data, has gained popularity as an effective method for learning node representations. However, most existing methods leverage random edge dropping to obtain the augmented view, which results in many isolated nodes and leads to limited performance. Moreover, how to reasonably and accurately identify important topology-feature-level positive samples with graph homophily is still an interesting and challenging problem. To address these issues, we propose a novel graph contrastive learning method with adaptive graph augmentation and topology-feature-level homophily, named GCL-GATH. Specifically, GCL-GATH assigns different weights to edges during the graph augmentation process, aiming to preserve the global topological structure as much as possible. Moreover, it simultaneously utilizes both structural and feature information to select positive samples from neighboring nodes. Extensive experimental results fully demonstrate that the proposed GCL-GATH outperforms the state-of-the-art methods. The source codes of this work are available at https://github.com/ZZY-GraphMiningLab/GCL-GATH .
Heterogeneous Graph (HG) is a data structure composed of various types of nodes and rich relational information, which can accurately show complex application scenarios in the real world. Although heterogeneous graph neural networks (HGNNs) have been widely applied to model HGs, there are still some issues that need to be addressed. On the one hand, most of HGNNs ignore the fine-grained information when modeling HGs, such as attribute and topology overcoupling due to the accumulation of multi-source heterogeneous information in message passing. On the other hand, HGNNs are designed from a single view (based on metapath or relation awareness), which undoubtedly leads to information loss and makes it difficult to fully extract potential interactions in HGs. To tackle the aforementioned limitations, a Multi-view fusion based Heterogeneous Graph Neural Network (MHGNN) is proposed, which is modeled from node view, network schema view, and semantics view to mine the information from different granularity in HGs. MHGNN extracts the fine-gained information of nodes, heterogeneous interaction of neighboring nodes, and mutual influence between different semantics from three views respectively. Then, the model integrates these information as the final node representation. To prove the effectiveness of this work, extensive experiments are conducted on four real-world datasets, and comparisons are made with seven competitive baselines. The results demonstrate that the proposed MHGNN significantly outperforms state-of-the-art methods. Source codes are available at https://github.com/ZZY-GraphMiningLab/MHGNN.
Although growing evidence shows that microRNA (miRNA) regulates plant growth and development, miRNA regulatory networks in plants are not well understood. Current experimental studies cannot characterize miRNA regulatory networks on a large scale. This information gap provides a good opportunity to employ computational methods for global analysis and to generate useful models and hypotheses. To address this opportunity, we collected miRNA-target interactions (MTIs) and used MTIs from Arabidopsis thaliana and Medicago truncatula to predict homologous MTIs in soybeans, resulting in 80,235 soybean MTIs in total. A multi-level iterative bi-clustering method was developed to identify 483 soybean miRNA-target regulatory modules (MTRMs). Furthermore, we collected soybean miRNA expression data and corresponding gene expression data in response to abiotic stresses. By clustering these data, 37 MTRMs related to abiotic stresses were identified including stress-specific MTRMs and shared MTRMs. These MTRMs have gene ontology (GO) enrichment in resistance response, iron transport, positive growth regulation, etc. Our study predicts soybean miRNA-target regulatory modules with high confidence under different stresses, constructs miRNA-GO regulatory networks for MTRMs under different stresses and provides miRNA targeting hypotheses for experimental study. The method can be applied to other biological processes and other plants to elucidate miRNA co-regulation mechanisms.
Accurately predicting the binding affinities between Human Leukocyte Antigen (HLA) molecules and peptides is a crucial step in understanding the adaptive immune response. This knowledge can have important implications for the development of effective vaccines and the design of targeted immunotherapies. Existing sequence-based methods are insufficient to capture the structure information. Besides, the current methods lack model interpretability, which hinder revealing the key binding amino acids between the two molecules. To address these limitations, we proposed an interpretable graph convolutional neural network (GCNN) based prediction method named GIHP. Considering the size differences between HLA and short peptides, GIHP represent HLA structure as amino acid-level graph while represent peptide SMILE string as atom-level graph. For interpretation, we design a novel visual explanation method, gradient weighted activation mapping (Grad-WAM), for identifying key binding residues. GIHP achieved better prediction accuracy than state-of-the-art methods across various datasets. According to current research findings, key HLA-peptide binding residues mutations directly impact immunotherapy efficacy. Therefore, we verified those highlighted key residues to see whether they can significantly distinguish immunotherapy patient groups. We have verified that the identified functional residues can successfully separate patient survival groups across breast, bladder, and pan-cancer datasets. Results demonstrate that GIHP improves the accuracy and interpretation capabilities of HLA-peptide prediction, and the findings of this study can be used to guide personalized cancer immunotherapy treatment. Codes and datasets are publicly accessible at: https://github.com/sdustSu/GIHP.
Plants are remarkable in their ability to adapt to changing environments, with receptor-like kinases (RLKs) playing a pivotal role in perceiving and transmitting environmental cues into cellular responses. Despite extensive research on RLKs from the plant kingdom, the function and activity of many kinases, i.e., their substrates or “clients”, remain uncharted. To validate a novel client prediction workflow and learn more about an important RLK, this study focuses on P2K1 (DORN1), which acts as a receptor for extracellular ATP (eATP), playing a crucial role in plant stress resistance and immunity. We designed a Kinase-Client (KiC) assay library of 225 synthetic peptides, incorporating previously identified P2K phosphorylated peptides and novel predictions from a deep-learning phosphorylation site prediction model (MUsite) and a trained hidden Markov model (HMM) based tool, HMMER. Screening the library against purified P2K1 cytosolic domain (CD), we identified 46 putative substrates, including 34 novel clients, 27 of which may be novel peptides, not previously identified experimentally. Gene Ontology (GO) analysis among phosphopeptide candidates revealed proteins associated with important biological processes in metabolism, structure development, and response to stress, as well as molecular functions of kinase activity, catalytic activity, and transferase activity. We offer selection criteria for efficient further in vivo experiments to confirm these discoveries. This approach not only expands our knowledge of P2K1’s substrates and functions but also highlights effective prediction algorithms for identifying additional potential substrates. Overall, the results support use of the KiC assay as a valuable tool in unraveling the complexities of plant phosphorylation and provide a foundation for predicting the phosphorylation landscape of plant species based on peptide library results.
Protein-Protein Interactions (PPIs) involves in various biological processes, which are of significant importance in cancer diagnosis and drug development. Computational based PPI prediction methods are more preferred due to their low cost and high accuracy. However, existing protein structure based methods are insufficient in the extraction of protein structural information. Furthermore, most methods are less interpretable, which hinder their practical application in the biomedical field. In this paper, we propose MGPPI, which is a Multiscale graph convolutional neural network model for PPI prediction. By incorporating multiscale module into the Graph Neural Network (GNN) and constructing multi convolutional layers, MGPPI can effectively capture both local and global protein structure information. For model interpretability, we introduce a novel visual explanation method named Gradient Weighted interaction Activation Mapping (Grad-WAM), which can highlight key binding residue sites. We evaluate the performance of MGPPI by comparing with state-of-the-arts methods on various datasets. Results shows that MGPPI outperforms other methods significantly and exhibits strong generalization capabilities on the multi-species dataset. As a practical case study, we predicted the binding affinity between the spike (S) protein of SARS-COV-2 and the human ACE2 receptor protein, and successfully identified key binding sites with known binding functions. Key binding sites mutation in PPIs can affect cancer patient survival statues. Therefore, we further verified Grad-WAM highlighted residue sites in separating patients survival groups in several different cancer type datasets. According to our results, some of the highlighted residues can be used as biomarkers in predicting patients survival probability. All these results together demonstrate the high accuracy and practical application value of MGPPI. Our method not only addresses the limitations of existing approaches but also can assists researchers in identifying crucial drug targets and help guide personalized cancer treatment.
Community detection methods based on attribute network representation learning are receiving increasing attention. However, few existing works are focused exclusively on unsupervised network representation learning for the task of community detection. They mainly capture information about the topology or attributes of the network, but do not fully utilize clustering-oriented information. In this paper, we present a community detection algorithm based on unsupervised attributed network embedding (CDBNE) to resolve the above issues. To be specific, we propose a framework that learns the representation based on network structure and attribute information and the clustering-oriented representation simultaneously. The framework includes the graph attention auto-encoder module, the modularity maximization module, and the self-training clustering module. Firstly, CDBNE encodes the topology structure and the node attribute with the graph attention mechanism. Secondly, it captures the mesoscopic community structure with modularity maximization. Finally, the self-training clustering module optimizes the representation learning process in a self-supervised manner to obtain high-quality node representation. The performance of CDBNE is verified with experiments on community detection tasks. According to the results on three datasets, CDBNE outperforms the state-of-the-art methods. The implementation of CDBNE is available at https://github.com/xidizxc/CDBNE.
Plant tissues are distinguished by their gene expression patterns, which can help identify tissue-specific highly expressed genes and their differential functional modules. For this purpose, large-scale soybean transcriptome samples were collected and processed starting from raw sequencing reads in a uniform analysis pipeline. To address the gene expression heterogeneity in different tissues, we utilized an adversarial deconfounding autoencoder (AD-AE) model to map gene expressions into a latent space and adapted a standard unsupervised autoencoder (AE) model to help effectively extract meaningful biological signals from the noisy data. As a result, four groups of 1,743, 914, 2,107, and 1,451 genes were found highly expressed specifically in leaf, root, seed and nodule tissues, respectively. To obtain key transcription factors (TFs), hub genes and their functional modules in each tissue, we constructed tissue-specific gene regulatory networks (GRNs), and differential correlation networks by using corrected and compressed gene expression data. We validated our results from the literature and gene enrichment analysis, which confirmed many identified tissue-specific genes. Our study represents the largest gene expression analysis in soybean tissues to date. It provides valuable targets for tissue-specific research and helps uncover broader biological patterns. Code is publicly available with open source at https://github.com/LingtaoSu/SoyMeta.
Nodule organogenesis in legumes is regulated temporally and spatially through gene networks. Genome-wide transcriptome, proteomic, and metabolomic analyses have been used previously to define the functional role of various plant genes in the nodulation process. However, while significant progress has been made, most of these studies have suffered from tissue dilution since only a few cells/root regions respond to rhizobial infection, with much of the root non-responsive. To partially overcome this issue, we adopted translating ribosome affinity purification (TRAP) to specifically monitor the response of the root cortex to rhizobial inoculation using a cortex-specific promoter. While previous studies have largely focused on the plant response within the root epidermis (e.g., root hairs) or within developing nodules, much less is known about the early responses within the root cortex, such as in relation to the development of the nodule primordium or growth of the infection thread. We focused on identifying genes specifically regulated during early nodule organogenesis using roots inoculated with Bradyrhizobium japonicum. A number of novel nodulation gene candidates were discovered, as well as soybean orthologs of nodulation genes previously reported in other legumes. The differential cortex expression of several genes was confirmed using a promoter-GUS analysis, and RNAi was used to investigate gene function. Notably, a number of differentially regulated genes involved in phytohormone signaling, including auxin, cytokinin, and gibberellic acid (GA), were also discovered, providing deep insight into phytohormone signaling during early nodule development.
Although growing evidence shows that microRNA (miRNA) regulates plant growth and development, miRNA regulatory networks in plants are not well understood. Current experimental studies cannot characterize miRNA regulatory networks on a large scale. This information gap provides an excellent opportunity to employ computational methods for global analysis and generate valuable models and hypotheses. To address this opportunity, we collected miRNA–target interactions (MTIs) and used MTIs from Arabidopsis thaliana and Medicago truncatula to predict homologous MTIs in soybeans, resulting in 80,235 soybean MTIs in total. A multi-level iterative bi-clustering method was developed to identify 483 soybean miRNA–target regulatory modules (MTRMs). Furthermore, we collected soybean miRNA expression data and corresponding gene expression data in response to abiotic stresses. By clustering these data, 37 MTRMs related to abiotic stresses were identified, including stress-specific MTRMs and shared MTRMs. These MTRMs have gene ontology (GO) enrichment in resistance response, iron transport, positive growth regulation, etc. Our study predicts soybean MTRMs and miRNA-GO networks under different stresses, and provides miRNA targeting hypotheses for experimental analyses. The method can be applied to other biological processes and other plants to elucidate miRNA co-regulation mechanisms.
Community detection methods based on attribute network representation learning is receiving increasing attention. However, existing methods still face the following challenges: How to fully utilize the topological structure and node attribute, and how to design an end-to-end network representation learning method for the task of community detection. In this paper, we present a Community Detection algorithm based on attributed Network Embedding (CDBNE) to resolve the above issues. Firstly, CDBNE encodes the topology structure and the node attribute with the graph attention mechanism. Secondly, it captures the mesoscopic community structure with the modularity maximization. Finally, the self-training clustering module optimizes the representation learning process in a self-supervised manner to obtain high-quality node representation. The performance of CDBNE is verified with experiments on community detection tasks. According to the results on three datasets, CDBNE outperforms the state-of-the-art methods. The implementation of CDBNE is available at https://github.com/xidizxc/CDBNE.
22 We present DeepMAPS (Deep learning-based Multi-omics Analysis Platform for Single-23 cell data) for biological network inference from single-cell multi-omics (scMulti-omics). 24 DeepMAPS includes both cells and genes in a heterogeneous graph to simultaneously 25 infer cell-cell, cell-gene, and gene-gene relations. The multi-head attention mechanism in 26 a graph transformer considers the heterogeneous relation among cells and genes within 27 both local and global context, making DeepMAPS robust to data noise and scale. We 28 benchmarked DeepMAPS on 18 scMulti-omics datasets for cell clustering and biological 29 network inference, and the results showed that our method outperformed various existing 30 tools. We further applied DeepMAPS on lung tumor leukocyte CITE-seq data and matched 31 diffuse small lymphocytic lymphoma scRNA-seq and scATAC-seq data. In both cases, 32 DeepMAPS showed competitive performance in cell clustering and predicted biologically 33 meaningful cell-cell communication pathways based on the inferred gene networks. Note 34 that we deployed a webserver using DeepMAPS implementation equipped with multiple 35 functions and visualizations to improve the feasibility and reproducibility of scMulti-omics 36 data analysis. Overall, DeepMAPS represents a heterogeneous graph transformer for 37 single-cell study and may benefit the use of scMulti-omics data in various biological 38 systems. 39
SARS-CoV-2, responsible for the current COVID-19 pandemic that claimed over 5.0 million lives, belongs to a class of enveloped viruses that undergo quick evolutionary adjustments under selection pressure. Numerous variants have emerged in SARS-CoV-2, posing a serious challenge to the global vaccination effort and COVID-19 management. The evolutionary dynamics of this virus are only beginning to be explored. In this work, we have analysed 1.79 million spike glycoprotein sequences of SARS-CoV-2 and found that the virus is fine-tuning the spike with numerous amino acid insertions and deletions (indels). Indels seem to have a selective advantage as the proportions of sequences with indels steadily increased over time, currently at over 89%, with similar trends across countries/variants. There were as many as 420 unique indel positions and 447 unique combinations of indels. Despite their high frequency, indels resulted in only minimal alteration of N-glycosylation sites, including both gain and loss. As indels and point mutations are positively correlated and sequences with indels have significantly more point mutations, they have implications in the evolutionary dynamics of the SARS-CoV-2 spike glycoprotein.
Detecting gene sets that serve as biomarkers for differentiating patient survival groups may help diagnose diseases robustly and develop multi-gene targeted therapies. However, due to the exponential growth of search space imposed by gene combinations, the performance of existing methods is still far from satisfactory. In this study, we developed a new method called BISG (BIclustering based Survival-related Gene sets detection) based on a rectified factor network (RFN) model, which allows efficiently biclustering gene subsets. By correlating genes in each significant bicluster with patient survival outcomes using a log-rank test and multi-sampling strategy, multiple survival-related gene sets can be detected. We applied BISG on three different cancer types, and the resulting gene sets were tested as biomarkers for survival analyses. Secondly, we systematically analyzed 12 different cancer datasets. Our analysis shows that the genes in all the survival-related gene sets are mainly from five gene families: microRNA protein coding host genes, zinc fingers C2H2-type, solute carriers, CD (cluster of differentiation) molecules, and ankyrin repeat domain containing genes. Moreover, we found that they are mainly enriched in heme metabolism, apoptosis, hypoxia and inflammatory response-related pathways. We compared BISG with two other methods, GSAS and IPSOV. Results show that BISG can better differentiate patient survival groups in different datasets. The identified biomarkers suggested by our study provide useful hypotheses for further investigation. BISG is publicly available with open source at https://github.com/LingtaoSu/BISG.