
Homology-based scaffolding orders contigs based on conserved collinearity of homologous sequences across related species. Existing methods often rely on costly whole-genome alignments or show limited robustness when integrating multiple references. Here, we introduce an anchor-based scaffolding framework that adapts synteny anchors to efficiently infer contig order and orientation relative to one or more reference genomes. Our approach leverages precomputed, sufficiently unique anchors and their respective high-confidence homology matches in a greedy approach, combining single-reference to multi-reference scaffolds using a maximum matching. Across simulated and real datasets, anchor-based scaffolding achieves accuracy comparable to state-of-the-art methods. Notably, the approach shows particular strengths in multi-reference settings. These results demonstrate that synteny-anchorbased scaffolding provides an additional tool for homology-based scaffolding with robust accuracy and uperior performance in multi-reference scenarios.
Copy number variation (CNV), as a major type of DNA structural variations (SVs), plays a key role in causing human diseases and contributing to genetic diversity. Accurate identification of CNVs is significant for disease mechanism analysis, personalized diagnosis and treatment, and drug development. Although next-generation sequencing (NGS) technology has greatly promoted the development of CNV detection methods, the existing methods generally have problems such as high false positives and inaccurate boundaries. Therefore, a new method is proposed for detecting CNVs in a single sample of NGS data, called CNV-ECOD. The method first employs the empirical-cumulative-distribution-based outlier detection (ECOD) algorithm to identify abnormal signals of read depth (RD) for preliminary detection of CNVs. To correct false positives and refine CNV boundaries further, it integrates paired-end mapping (PEM) and split read (SR) strategies. The integration of the RD-PEM-SR hierarchical progressive framework and the anomaly scoring mechanism based on ECOD can effectively improve the accuracy of CNV detection. Comparing our approach to four peer methods, simulation results demonstrate that it achieves the best balance between precision and sensitivity. Also, the proposed method has the best F1-scores and the highest overlap density scores (ODSs) in real-sample experiments. Therefore, CNV-ECOD is expected to develop into an efficient and robust CNV detection tool.
Accurate protein structure prediction is critical for rational enzyme engineering, which requires high-fidelity models. This study benchmarks three distinct structure prediction paradigms against the experimental crystal structure of IsPETase, serving as a diagnostic case study. The evaluated approaches include classical homology modeling (SWISS-MODEL), MSA-conditioned diffusion (AlphaFold 3), and generative language modeling (ESM-3). Predicted models were evaluated using stereochemical validation, molecular docking with a PET dimer analogue, and molecular dynamics simulations. While all approaches reproduced the overall fold and preserved the catalytic triad geometry, notable differences were observed in atomic clashes and hydrogen bonding patterns. ESM-3 showed elevated steric clashes and reduced hydrogen bond counts. Molecular dynamics indicated that the experimental structure maintained the highest stability, with SWISS-MODEL closely following, while ESM-3 displayed greater fluctuations, particularly in loop regions. Crucially, blind docking simulations revealed that the ESM-3 active site was sterically occluded, rendering it inaccessible to the PET dimer. This inaccessibility persisted even after targeted energy minimization. These findings suggest that while generative language models represent a powerful capability for rapid scaffold exploration, they do not yet achieve the thermodynamic precision of established homology and evolutionary approaches required for functional active site engineering.
Boolean networks provide a qualitative framework for modelling regulatory systems when kinetic parameters are unavailable, with cellular phenotypes represented as attractors of the induced dynamics [J. D. Schwab et al., Concepts in Boolean network modeling, Comput. Struct. Biotechnol. J. 18:571-582, 2020]. A central challenge is phenotypic reachability: determining whether asynchronous dynamics can connect invariant regions of the state space, a problem that becomes computationally intractable in large networks [L. Cifuentes-Fontanals, M. Noual, and E. Remy, Revisiting trap spaces in Boolean networks, Theor. Comput. Sci. 915:1-20, 2022; K. Perrot and C. Paulevé, Complexity of asynchronous reachability in Boolean networks, Theor. Comput. Sci. 1000:114650, 2024.]. We develop a structural theory of reachability in which trap spaces are identified with labelled order ideals of SCC-posets. The SCC-poset determines the order of commitment events, while admissible evaluations encode branching within regulatory modules, so that multistability appears as an intrinsic feature of the theory. Within this framework, we establish necessary and sufficient conditions for reachability, introduce the commitment depth, and show that deciding non-trivial branching is computationally intractable. We further demonstrate that effective interaction structure is jointly determined by topology and Boolean logic. We validate the framework on a Boolean model of CD4[Formula: see text] T-cell differentiation, where refinement chains recover the observed ordering of cytokine response, lineage commitment, and phenotypic branching. In the absence of multistability the structure collapses to a distributive lattice, a non-generic limiting regime.
Copy number variation (CNV) is one of the most imperative forms of structural variations that can span over the coding and non-coding regulatory regions of an individual's genome. Copy number variations (CNV) can significantly impact the genotype and phenotype traits by altering the gene dosage, consequently affecting the gene expression landscape concerning various cellular functions and are the cause behind complex diseases in an individual. Exceptionally fast advancement in Next Generation Sequencing (NGS) technology has led to massive growth of DNA-seq data, which contains both Whole Genome Sequence (WGS) and targeted Exome Sequence data of various species including H.sapiens, and precise detection of the DNA region affected by CNV enables the copy number profiling of a genome, thereby understanding our genome. This work has proposed a methodology named ReinVar, which can accurately determine and analyze the underlying copy number profile of the whole genome by adopting a model-free reinforcement learning paradigm. The methodology involves a novel approach to model the problem of identifying CNV as a Markov decision process (MDP), followed by determination of CNV under Reinforcement Learning framework. ReinVar also adopted a Map-Reduce programming paradigm to provide a big data solution to address the issue of exponential growth of NGS read sequence data. ReinVar has shown strong performance in detecting CNV gains and losses across diverse ethnic groups, with a high number of shared variant calls. ReinVar's ability to accurately identify both CNV gains and losses, coupled with consistent detection across ethnic groups and strong ROC characteristics, underscores ReinVar's effectiveness as a robust and sensitive CNV detection method.
Some post-GWAS analysis software can run to completion without reporting an error while producing results that are not biologically valid. We call this failure false locality: a result appears to be local to a gene or protein because it was produced from a regional window, but the window or its metadata points to the wrong genomic address. We identify three mechanisms. First, a genome-build address mismatch occurs when GRCh38 protein-QTL coordinates are used with GRCh37 outcome files; in our 91-sentinel audit, 66 coordinates moved by more than 100[Formula: see text]kb and 54 moved by more than 200[Formula: see text]kb after remapping. Second, non-specific regional noise occurs when significant P values persist after the analysis window is deliberately moved to a zero-overlap variant set. Third, location-label blindness occurs when software returns the same output after the declared coordinate label is changed while the SNP table is unchanged; in an official SMR/HEIDI CXCL10 test, the top SNP, SMR P value, and HEIDI P value remained identical across correct, wrong, and shifted labels. We propose a simple Change Test: a result should not be treated as local evidence unless the numerical output changes, weakens, or disappears when the analysis window or coordinate label is intentionally moved to a biologically wrong location. This standard turns software execution from a passive success signal into an explicit spatial-specificity check.
Protein-protein interactions (PPIs) play crucial roles in diverse cellular functions and biological processes, and structural knowledge of the protein complexes is valuable for the elucidation of those functions and designing new drugs. Due to the limitations of experimental methods, computational modeling approaches capable of producing reliable protein complex models using molecular docking tools are of considerable practical interest. The success of protein docking largely depends on an accurate scoring function that can pick out good protein docking models. In this work, we present a neural network-based scoring function for scoring protein-protein docking models, NNDock2, the updated version of our previous scoring function, NNDock1. To improve NNDock1, we augmented the training decoys by adding a large number of more distant decoys. In addition, instead of interface root mean square deviation (iRMSD) in NNDock1, we used the fraction of native contact ([Formula: see text] as a target function, which shows better correlation with true model quality. We also applied regularization during training to avoid overfitting. We tested NNDock2 on the protein-protein docking benchmark version 5.0 (BM5), DOCKGROUND dataset, and the CAPRI score set and compared the performance of NNDock2 with other state-of-the-art scoring functions. NNDock2 performed comparably to other state-of-the-art scoring functions, despite the simplicity of the method and low computational costs. We envision that NNDock2 could be used as an independent scoring function or as an element or feature of composite or deep learning-based scoring functions for protein complex model quality estimation.
This study addresses the problem of protein function annotation and proposes a multi-source biological information-fusion framework called Functional co-Occurrence Probability Estimation (FOPE) for estimating functional co-occurrence probabilities. The framework integrates Protein-Protein Interaction (PPI) network topology and protein domain information, quantifying the functional synergy between protein pairs through bidirectional functional participation modeling. Experiments on four model organisms (A. thaliana, C. elegans, D. melanogaster, and S. cerevisiae) show that FOPE delivers effective predictions for the three main categories of Gene Ontology (GO). Compared to existing representative methods, its macro-Fmax values improved by 25.3%, 19.3%, and 20.9% on average for Biological Process (BP), Cellular Component (CC), and Molecular Function (MF), respectively. Ablation studies further reveal the functional-specific contributions of different information sources: domain information plays a dominant role in MF prediction, while PPI network features are more critical for BP prediction. The effective integration of both is key to achieving comprehensive prediction performance. Robustness tests demonstrate that FOPE maintains strong stability in BP and CC predictions even under significant noise in the PPI network (adding or removing 30% of interactions), verifying the error-tolerance advantages of multi-source information fusion. The FOPE framework proposed in this study provides a feasible information fusion approach for protein function prediction. Preliminary experimental results demonstrate the applicability of the method across different functional categories and multiple model organisms, offering a potential computational pathway for exploring functional synergy relationships among proteins from a system-level perspective.
In this paper, we develop a connectivity-exact bond-resolved framework for quantifying structural accessibility and persistence in drug molecules. Representing a molecule as a heavy-atom graph [Formula: see text] each bond is classified uniquely as either a connectivity-critical Entry-point bond ([Formula: see text]) or a connectivity-preserving Fortress bond ([Formula: see text]) yielding the exact identity [Formula: see text] This exhaustive partition separates fragmentation capacity from invariant scaffold structure without adjustable parameters. To refine the invariance coordinate we introduce Total Structural Entrenchment (TSE) a persistence-weighted functional defined over cycle-supported bonds and modulated by local steric wall contributions [Formula: see text]. The resulting two-parameter embedding [Formula: see text] distinguishes superficial cyclic extent from deeply embedded structural reinforcement and resolves degeneracies inherent in raw bond counts. Metabolic progression is formalized as recursive bridge depletion generating a directed migration across the architectural plane toward a bridge-depleted refractory core. Within this framework scaffold persistence is interpreted as a connectivity-driven contraction governed strictly by graph topology. The resulting invariance-variation embedding establishes a mathematically controlled bond-level representation of structural accessibility and cyclic entrenchment.
Human Genome Project (HGP), Genome Wide Association Studies (GWAS) and The Cancer Genome Atlas (TCGA) are some of the remarkable research endeavors that generated massive amounts of information about Single Nucleotide Polymorphisms (SNPs) and other genetic variations, providing valuable insights for understanding the association of SNPs with diseases. It enables early diagnosis, prevention, and treatment planning for diseases. In this study, a novel approach is proposed for the identification of SNPs. This approach consists of two techniques: technique I introduces a modified matching strategy for chosen matching algorithms and technique II combines the Divide-and-Conquer technique with technique 1. Performance evaluation of the proposed techniques is performed using performance metrics such as Precision, Recall, F-measure, Execution Time, and Resource Utilization (including CPU Utilization and RAM Usage). The proposed techniques overcome most of the research gaps and shortcomings of the existing techniques.
Deep learning reports over 90% DNA-binding protein (DBP) prediction performance on common benchmarks, but these results are usually obtained on balanced test sets and may not translate to proteome-wide scans with extreme class imbalance. Here, we use KANBind as a diagnostic probe to stress-test sequence-based DBP prediction under strict homology control and realistic prevalence. Evaluated on the homology-controlled HBTD benchmark with prevalence-calibrated reporting, KANBind achieves a calibrated precision of 0.0558 at a realistic bacterial prevalence ([Formula: see text]), implying an expected false discovery rate (FDR) of 94.42%. In a proteome-scale scan, this corresponds to approximately 95 false positives per 100 predicted DBPs. Interpretability analysis indicates that predictions are driven mainly by coarse physicochemical cues such as electrostatics, which may be necessary for DNA binding but are insufficient to determine DBP function. Together, these results suggest that apparent benchmark gains can be dominated by homology leakage and evaluation on balanced sets rather than by generalizable functional rules, motivating stress-test benchmarks with strict homology control and realistic negative backgrounds.
The dynamic p53 response is a known determinant of cell fate. However, its temporal control, specifically the mechanisms regulating the delay time ([Formula: see text]) of the graded p53 pulse following single-stranded breaks (SSBs), remains poorly understood. To systematically dissect this timing mechanism, we developed and analyzed a mechanistic ordinary differential equation (ODE) model of the p53-Mdm2-ATR network. We first established that increasing damage intensity reliably shortens the delay time, accelerating the cellular decision-making process. Our analysis revealed a critical finding: the delay time is most acutely sensitive to the p53-dependent Mdm2 production rate ([Formula: see text]), highlighting the dominant role of the negative feedback loop in setting the pace. We further classified the model parameters into functional roles: accelerators (e.g. Ataxia-Telangiectasia and RAD3-related (ATR) production rate dependent on damage ([Formula: see text]), p53 activation rate dependent on ATR ([Formula: see text]), p53-dependent Mdm2 production rate ([Formula: see text]), p53-dependent Wip1 production rate ([Formula: see text], ATR degradation rate ([Formula: see text]) and Mdm2-dependent p53 degradation rate ([Formula: see text], which shorten the delay time ([Formula: see text]), and brakes (e.g. ATR-dependent Mdm2 degradation rate ([Formula: see text]), self-degradation rate of Mdm2 ([Formula: see text]) and self-degradation rate of Wip1 ([Formula: see text]), which prolong it. Sensitivity analysis showed that as parameter values increase, [Formula: see text] becomes less sensitive to [Formula: see text]. The sensitivity to [Formula: see text] exhibited an initial increase followed by a decrease, whereas the opposite trend was observed for [Formula: see text]. The remaining parameters ([Formula: see text], [Formula: see text], [Formula: see text], [Formula: see text], [Formula: see text], [Formula: see text]) all showed a monotonic decrease in sensitivity. This work provides a quantitative blueprint for therapeutic interventions, suggesting that targeting the p53-Mdm2 feedback strength is the most effective strategy to sensitize cancer cells and shorten the critical delay to cell death.
The performance of classifiers can decline due to irrelevant or non-informative features. Therefore, extracting meaningful high-level features from raw data through effective feature selection and construction is vital. This challenge is particularly significant in bioinformatics, especially for predicting protein-peptide interactions. To address this, we propose IntPPPred, a computational method designed to enhance residue-level prediction of such interactions. IntPPPred leverages evolutionary computation and ensemble learning to build high-level features from selected informative ones, reducing computational complexity and feature space. It first identifies unique and effective features before applying multiple feature constructions using a gravitational search algorithm. This enhances the prediction capability of a stacking-based ensemble classifier, particularly on imbalanced datasets. Experimental results show that IntPPPred improves Matthews correlation coefficient (MCC), F-measure, and precision by at least 1.2%, 19.6%, and 3.9%, respectively, over existing methods. When evaluated on a second dataset, further gains of 5.2%, 24%, 16.7%, and 1.1% were observed in precision, sensitivity, F-measure, and MCC. Additionally, performance consistency between cross-validation and independent test sets confirms the method's robustness. Overall, IntPPPred serves as an effective computational tool that supports experimental research and enhances machine learning performance in protein-peptide interaction prediction.
Given the rapid advancements in high-throughput technology, multi-omics data have become essential for identifying cancer subtypes and providing accurate medical treatments for patients. However, integrating multi-omics data and collecting patient information pose complex and challenging tasks. Although numerous integration techniques have emerged in recent years to address the challenges posed by heterogeneity and noise in omics data, most of these algorithms are based on unsupervised methods due to the lack of labeled data. This indicates there is still potential for enhancing the extraction of valuable information from omics data. This study introduces a novel framework, namely Deep Subspace Fusion based on Integrated Self-supervision (DSFIS), for the recognition of cancer subtypes. DSFIS is built on the autoencoder with a self-representation layer and guides the autoencoder to generate the most representative sample subspace structure by integrating self-supervision. This framework can not only create a comprehensive representation of the differences and similarities among patients but also more fully uncover the potential information from omics data. The DSFIS was compared to eight cutting-edge approaches for integrating multi-omics data. The experimental findings demonstrated that DSFIS effectively identified cancer subtypes according to the omics data. It achieved significant results superior to other algorithms in survival prognosis analysis and clinical correlation analysis, demonstrating that DSFIS has great potential in identifying cancer subtypes through multi-omics data.
Predicting metastasis in early stages of cancer plays a crucial role in effectively controlling cancer progression and thereby improving patient survival outcomes. Although several computational methods have been developed to predict cancer metastasis, most focus on lymph node metastasis. Distant metastasis is more difficult to detect or predict than lymph node metastasis. In this study, we developed a multilayer perceptron (MLP) model to predict distant cancer metastases. We constructed a weighted gene interaction network and computed sample-specific differential gene correlations for individual cancer samples. The MLP model was trained on sample-specific differential gene correlations and tested on independent datasets of differential gene correlations from samples that were not used in training the model. The MLP model is capable of predicting whether or not distant metastasis will occur and potential distant metastatic sites. In independent testing, it predicted distant metastasis with a high performance (AUC of 0.95) and predicted metastatic sites with an average AUC of 0.97. In comparison of our model with other state-of-the-art methods using the same data set, our model showed better performance than the others. The prediction model developed in this study may help clinicians determine site-specific testing and treatment options.
Introduction Sorafenib remains the only approved treatment for advanced hepatocellular carcinoma (HCC), yet its clinical use is hindered by toxicity and the emergence of drug resistance. Sorafenib's anticancer effects are largely attributed to its inhibition of multiple kinases, including c-Raf, a key player in the Ras-Raf-MEK-ERK signalling cascade that promotes cell growth and survival. Given the critical role of c-Raf in tumor progression, targeting this kinase offers a promising strategy for improving therapeutic outcomes. Developing new analogues with stronger c-Raf inhibition, better pharmacokinetics, and reduced side effects could help address the current limitations of sorafenib. Objectives This study aimed to design novel sorafenib analogues with enhanced binding affinity and favourable pharmacokinetic profiles, specifically targeting the c-Raf kinase to increase therapeutic efficacy against HCC. By using a fragment replacement approach combined with computational methods, the goal was to identify candidates capable of forming stronger, more stable interactions with c-Raf, potentially overcoming resistance linked to sorafenib treatment. Methods A total of 84 sorafenib analogues (A1-A84) were generated by modifying key functional groups, including the 2-picolinamide and substituted phenyl moieties known to influence kinase binding and anticancer activity. These analogues were evaluated through chemoinformatics and pharmacokinetic screening to assess their drug-likeness and safety. Molecular docking was performed to estimate their binding affinity toward c-Raf. Six top-performing analogues (A2, A6, A9, A20, A22, A63) were selected for further analysis. To evaluate their dynamic behaviour, 100 ns all-atom molecular dynamics simulations were conducted, followed by MM-PBSA (Molecular Mechanics Poisson-Boltzmann Surface Area) calculations to determine binding free energies. Principal component analysis (PCA) was carried out to explore key motion patterns within the protein-ligand complexes. Results Molecular docking showed that the selected analogues exhibited stronger binding affinities (-11.6 to -10.9 kcal/mol) compared to sorafenib (-9.3 kcal/mol) and regorafenib (-9.5 kcal/mol). Molecular dynamics simulations substantiated the docking results. MM-PBSA results revealed that at 100 ns, the binding free energy for the c-Raf-sorafenib complex was 86.751 kJ/mol, while the c-Raf complexes with A2, A6, A9, A20, A22, and A63 demonstrated significantly lower free energies of -129.114, -135.637, -136.242, -127.178, -94.25, and -123.176 kJ/mol, respectively, indicating stronger and more stable binding. PCA further confirmed the stability and favourable dynamic profiles of these analogues trajectory with c-Raf. Discussion The improved binding affinities and lower free energies of the top analogues indicate that specific structural changes to sorafenib can enhance its effectiveness against c-Raf. Molecular dynamics and MM-PBSA results suggest the stability and strength of these interactions, particularly for A2, A6, and A9. Conclusion This study identified six promising sorafenib analogues with improved binding affinity, favourable pharmacokinetic characteristics, and stable interactions with c-Raf. By focusing on c-Raf inhibition, the combined use of computational modelling, molecular simulations and mmPBSA analysis provided valuable insights for drug design. Among the candidates, A2, A6, and A9 emerged as promising drug candidates for further development, supporting the potential of targeting c-Raf to enhance therapeutic strategies against hepatocellular carcinoma
Graves’ disease (GD) is a common autoimmune disorder. However, the circulating proteins that causally drive its pathogenesis remain largely unconfirmed by genetic evidence. Identifying such proteins is critical for developing novel therapeutics. Methods: We performed a two-sample Mendelian randomization (MR) analysis using genetic instruments for 4907 plasma proteins and summary statistics from a GD genome-wide association study (GWAS) of European ancestry to identify causal proteins. We subsequently conducted protein–protein interaction (PPI) network analysis and used AlphaFold3 to model the structural impact of a key variant, rs41271951, in the top candidate protein, cathepsin S (CTSS). Results: MR analysis identified 23 plasma proteins with putative causal effects on GD risk. Among these, CD5L showed the strongest evidence for colocalization (posterior [Formula: see text]), suggesting a shared causal variant. Network analysis revealed that these proteins converge on a novel complement–ECM–coagulation axis in GD pathogenesis. CTSS emerged as a central hub in this network. AlphaFold3 modeling suggested that the CTSS variant rs41271951 (p.Val7Ala), located within the signal peptide, induces subtle structural perturbations. The primary and most plausible consequence is a reduction in CTSS secretion and circulating levels, as supported by the pQTL data. Conclusion: This multi-omics analysis proposes a novel complement–ECM–coagulation axis in GD. By structurally and functionally linking reduced CTSS abundance and secretion to genetic variation, we identify CTSS as a potential candidate for therapeutic repurposing in GD.
The emergence of SARS-CoV-2 has highlighted the need for computational methods to identify neutralizing antibodies. Existing sequence-based tools for predicting antigen-antibody interactions struggle to effectively identify antibodies capable of neutralizing different variants due to high sequence similarity among SARS-CoV-2 strains and the similarity in the framework regions (FWRs) of antibodies. To address this challenge, particularly the issue of high sequence similarity among homologous antigens that impedes accurate prediction of antigen-antibody interactions, we developed a deep learning framework named PLMABFW. It differentiates homologous antigens using encoding techniques and network architecture design. It employs pre-trained protein language models ESM-2 for antigens and AntiBERTy for antibodies to encode sequences and capture additional features. The framework also incorporates both antigen features and their transposed versions to enhance antigen information capture. To validate the performance of PLMABFW, we collected a SARS-CoV-2 neutralization dataset. PLMABFW outperformed existing neutralizing antibody prediction tools (AbAgIntPre, DeepAAI) and docking tools (HDOCK, LSTM-PHV) in predicting neutralizing antibodies for homologous antigens. Furthermore, it effectively learned the interactions between the antibody's CDR-H3 region and antigens via a partial masking strategy. The model code is available on GitHub for customization and adaptation to diverse research needs.
This study investigates the thermodynamic behavior of molecular self-assembly along biochemical pathways leading to the formation of higher-order complexes. We specifically examine how thermodynamic parameters evolve - such as the dissociation constant [Formula: see text], the entropic contribution [Formula: see text], and the stability parameter of the interaction matrix [Formula: see text] - as molecular complexity increases from monomers to dimers, trimers, and tetramers. A central hypothesis is that stepwise thermodynamic modeling allows prediction of assembly pathways, identification of dead ends points, and the co-directional changes of thermodynamic variables during complex formation which reflect a preference for main biocomplex formation direction. We also introduce a practical rule to classify dead-end intermediates: a pathway step is considered a dead-end if the minimum [Formula: see text] occurs at a non-final intermediate or if [Formula: see text] falls below zero, indicating an entropic barrier. This criterion provides a reproducible way to flag non-viable assembly routes. We apply this analysis to several biologically relevant molecular systems, including the complex of LGP2 bound to an 8-base pair double-stranded RNA molecule, the dimer of VP35 protein interacting with double-stranded RNA and hexamer formations.
Multi-drug therapy has become more common in recent years, especially among older people who have many illnesses. However, patients are put at risk when unanticipated drug-drug interactions (DDIs) result in negative reactions or serious toxicity. Predicting possible DDI by computational model improves the drug design process and minimizes unexpected drug interactions and research expenses. In this paper, the proposed model is constructed by a complete graph convolutional neural (GCN) network on publicly DDI data from the DrugBank, including 65 classes for DDI prediction. In this data, the number of samples is 37,264 for two drugs with three optimal features, including the chemical, target, and enzyme. This multi-classification model consists of three phases, including drug preprocessing, three layers of GCN, and a fully connected network. The findings confirmed the performance of this proposed model, achieving an accuracy of 95.12%, which is the best result compared with previous works on the same data. Although the data was imbalanced, this paper primary contribution was to enhance both the computational time and the classification evaluation metrics rather than other state-of-the-art models. Explainable artificial intelligence (XAI) is applied SHapley Additive exPlainations (SHAP) to the proposed model to avoid misclassification and produce easily comprehensible results. This proposed model will help to explore the possible drug hazards and support intelligent pharmaceutical management.