O-linked glycosylation is a widespread protein post-translational modification with site selection strongly shaped by local sequence context. Threonine is a key acceptor residue, yet computational methods for predicting human O-linked threonine glycosites remain limited and often lack interpretability. Here, we present DeepO-GlyThr, an interpretable deep learning framework for O-linked threonine glycosite prediction in human proteins. DeepO-GlyThr integrates sequence features and employs CNN, BiGRU, and attention modules to learn discriminative patterns. On the independent test set, DeepO-GlyThr achieved a sensitivity of 0.885, an accuracy of 0.900, an F1-score of 0.898, and an MCC of 0.800, outperforming existing methods under the default classification threshold. Interpretability analyses showed that predictions were mainly driven by the local physicochemical environment and positional context around the central threonine. Volume, hydrophobicity, polarity, and positional features contributed most strongly. Attention, Integrated Gradients, and mutagenesis analyses further revealed context-dependent sequence patterns in local flanking regions. In summary, DeepO-GlyThr provides an accurate and interpretable framework for predicting O-linked threonine glycosites. Additionally, a user-friendly web server has been developed to facilitate its use, available at http://i-health.info/DeepO-GlyThr.
Traditional Chinese medicine (TCM) serves as a treasure trove of ancient knowledge, holding a crucial position in the medical field. However, the exploration of TCM's extensive information has been hindered by challenges related to data standardization, completeness, and accuracy, primarily due to the decentralized distribution of TCM resources. To address these issues, we developed a platform for TCM knowledge discovery (TCMKD, https://cbcb.cdutcm.edu.cn/TCMKD/). Seven types of data, including syndromes, formulas, Chinese patent drugs (CPDs), Chinese medicinal materials (CMMs), ingredients, targets, and diseases, were manually proofread and consolidated within TCMKD. To strengthen the integration of TCM with modern medicine, TCMKD employs analytical methods such as TCM data mining, enrichment analysis, and network localization and separation. These tools help elucidate the molecular-level commonalities between TCM and contemporary scientific insights. In addition to its analytical capabilities, a quick question and answer (Q&A) system is also embedded within TCMKD to query the database efficiently, thereby improving the interactivity of the platform. The platform also provides a TCM text annotation tool, offering a simple and efficient method for TCM text mining. Overall, TCMKD not only has the potential to become a pivotal repository for TCM, delving into the pharmacological foundations of TCM treatments, but its flexible embedded tools and algorithms can also be applied to the study of other traditional medical systems, extending beyond just TCM.
Polycystic ovary syndrome (PCOS) is a prevalent endocrine and metabolic disorder characterised by heterogeneous clinical and molecular phenotypes. Flaxseed, widely used in traditional Chinese medicine and as a nutritional supplement, has shown promising therapeutic potential for PCOS. In this study, we integrated transcriptomic data with machine learning-based analytical approaches and network pharmacology to investigate the molecular mechanisms underlying PCOS and to identify the potential targets and pathways modulated by flaxseed. Differentially expressed genes (DEGs) and PCOS-related targets were systematically identified from GEO, GeneCards and DisGeNet databases. Bioactive compounds in flaxseed were predicted using TCMSP, SwissTargetPrediction and INPUT2.0. Functional and pathway enrichment analyses were conducted to explore mechanistic insights. Core targets were prioritised using Centiscape network topology parameters and LASSO regression, followed by molecular docking validation using AutoDock. Our results revealed that flaxseed's therapeutic action may primarily involve modulation of immune regulation, insulin signalling, apoptosis and inflammation pathways. Key active compounds, notably β-sitosterol and stigmasterol, exhibited strong binding affinities with critical targets, such as IL1B, GSK3B and HMGCR, suggesting potential anti-inflammatory and antioxidant effects. The findings provide a theoretical foundation for future experimental studies and support the development of flaxseed-based therapeutic strategies for PCOS through precision medicine frameworks.
Certain RNAs exhibit both protein-coding and regulatory non-coding functions, termed bifunctional RNAs or coding and non-coding RNAs. Long non-coding RNAs (lncRNAs), which play crucial roles in gene regulation and cellular processes, represent a major subset of bifunctional RNAs. Accurate identification of bifunctional lncRNAs is critical for advancing RNA biology and uncovering opportunities for biomarker discovery and therapeutic development. Here, we present cncFinder, a graph-attention-network-based model for predicting bifunctional lncRNAs. It transforms lncRNA sequences into k-mer graphs, encodes node features with Word2Vec, and employs graph attention network to capture higher-order sequence dependencies. On the testing dataset, cncFinder achieved superior performance, significantly outperforming state-of-the-art models. Its robustness and broad applicability were further confirmed through validation on cross-species datasets from mouse and fruit fly. Interpretability analysis revealed that cncFinder captured biologically meaningful motifs, including canonical start codons and Kozak-like elements. In a case study of LINC00961, cncFinder precisely detected an experimentally validated translation initiation motif, highlighting its biological relevance. To support broad accessibility, we developed a user-friendly web server. In summary, cncFinder advances predictive accuracy and interpretability, providing a powerful tool for systematic discovery of bifunctional lncRNAs and enabling new insights into RNA multifunctionality.
INTRODUCTION:The blood-brain barrier (BBB) serves as a critical structural barrier and impedes the entry of most neurotherapeutic drugs into the brain. This poses substantial challenges for central nervous system (CNS) drug development, as there is a lack of efficient drug delivery technologies to overcome this obstacle. BBB penetrating peptides (BBBPs) hold promise in overcoming the BBB and facilitating the delivery of drug molecules to the brain. Therefore, precise identification of BBBPs has become a crucial step in CNS drug development. However, most computational methods are designed based on conventional models that inadequately capture the intricate interaction between BBBPs and the BBB. Moreover, the performance of these methods was further hampered by unbalanced datasets. OBJECTIVES:This study addresses the problem of unbalanced datasets in BBBP prediction and proposes a powerful predictor for efficiently and accurately identifying BBBPs, as well as generating analogous BBBPs. METHODS:A transformer-based deep learning model, DeepB3P, was proposed for predicting BBBP. The feedback generative adversarial network (FBGAN) model was employed to effectively generate analogous BBBPs, addressing data imbalance. RESULTS:The FBGAN model possesses the ability to generate novel BBBP-like peptides, effectively mitigating the data imbalance in BBBP prediction. Extensive experiments on benchmarking datasets demonstrated that DeepB3P outperforms other BBBP prediction models by approximately 9.09%, 4.55% and 9.41% in terms of specificity, accuracy, and Matthew's correlation coefficient, respectively. For accelerating the progress in BBBP identification and CNS drug design, the proposed DeepB3P was implemented as a webserver, which is accessible at http://cbcb.cdutcm.edu.cn/deepb3p/. CONCLUSION:The interpretable analyses provided by DeepB3P offer valuable insights and enhance downstream analyses for BBBP identification. Moreover, the BBBP-like peptides generated by FBGAN hold potential as candidates for CNS drug development.
Lung cancer is one of the most common malignant tumors around the world, which has the highest mortality rate among all cancers. Traditional Chinese medicine (TCM) has attracted increased attention in the field of lung cancer treatment. However, the abundance of ingredients in Chinese medicines presents a challenge in identifying promising ingredient candidates and exploring their mechanisms for lung cancer treatment. In this work, two network-based algorithms were combined to calculate the network relationships between ingredient targets and lung cancer targets in the human interactome. Based on the enrichment analysis of the constructed disease module, key targets of lung cancer were identified. In addition, molecular docking and enrichment analysis of the overlapping targets between lung cancer and ingredients were performed to investigate the potential mechanisms of ingredient candidates against lung cancer. Ten potential ingredients against lung cancer were identified and they may have similar effect on the development of lung cancer. The results obtained from this study offered valuable insights and provided potential avenues for the development of novel drugs aimed at treating lung cancer.
Heat shock proteins (HSPs) are crucial cellular stress proteins that react to environmental cues, ensuring the preservation of cellular functions. They also play pivotal roles in orchestrating the immune response and participating in processes associated with cancer. Consequently, the classification of HSPs holds immense significance in enhancing our understanding of their biological functions and in various diseases. However, the use of computational methods for identifying and classifying HSPs still faces challenges related to accuracy and interpretability. In this study, we introduced MulCNN-HSP, a novel deep learning model based on multi-scale convolutional neural networks, for identifying and classifying of HSPs. Comparative results showed that MulCNN-HSP outperforms or matches existing models in the identification and classification of HSPs. Furthermore, MulCNN-HSP can extract and analyze essential features for the prediction task, enhancing its interpretability. To facilitate its accessibility, we have made MulCNN-HSP available at http://cbcb.cdutcm.edu.cn/HSP/. We hope that MulCNN-HSP will contribute to advancing the study of HSPs and their roles in various biological processes and diseases.
O-linked glycosylation is one of the most complex post-translational modifications (PTM) of human proteins modulating various cellular metabolic and signaling pathways. Unlike N-glycosylation, the O-glycosylation has non-specific sequence features and unstable glycan core structure, which makes identification of O-glycosites more challenging either by experimental or computational methods. Biochemical experiments to identify O-glycosites in batches are technically and economically demanding. Therefore, development of computation-based methods is greatly warranted. This study constructed a prediction model based on feature fusion for O-glycosites linked to the threonine residues in Homo sapiens. In the training model, we collected and sorted out high-quality human protein data with O-linked threonine glycosites. Seven feature coding methods were fused to represent the sample sequence. By comparison of different algorithms, random forest was selected as the final classifier to construct the classification model. Through 5-fold cross-validation, the proposed model, namely O-GlyThr, performed satisfactorily on both training set (AUC: 0.9308) and independent validation dataset (AUC: 0.9323). Compared with previously published predictors, O-GlyThr achieved the highest ACC of 0.8475 on the independent test dataset. These results demonstrated the high competency of our predictor in identifying O-glycosites on threonine residues. Furthermore, a user-friendly webserver named O-GlyThr (http://cbcb.cdutcm.edu.cn/O-GlyThr/) was developed to assist glycobiologists in the research associated with glycosylation structure and function.
A new network platform named ACU&MOX-DATA (http://cbcb.cdutcm.edu.cn/ACU&MOX-DATA/) was recently set up and introduced. The purpose of this website is to integrate and analyze multi-omics data of acupuncture and moxibustion and serve as a disease, clinical, and laboratory experiment public database. Some existing functions and the future plan of the website were also discussed. This platform can help researchers to integrate analyses and visual interpretation of key scientific issues in acupuncture & moxibustion under the guidance of traditional Chinese medicine theory through heterogeneous and multi-level data fusion analysis.
Background: Yi-Jing decoction (YJD), a traditional Chinese medicine prescription, has been reported to be effective in the treatment of polycystic ovary syndrome (PCOS). However, the underlying mechanisms of YJD in treating PCOS are still unclear. Objective: In the present work, the effective ingredients of YJD and their treatment mechanisms on PCOS were systematically analyzed. Methods: The effective ingredients of YJD and targets of PCOS were selected from public databases. The network pharmacology method was used to analyze the ingredients, potential targets, and pathways of YJD for the treatment of PCOS. Results: One hundred and three active ingredients were identified from YJD, of which 82 were hit by 65 targets associated with PCOS. By constructing the disease-common targetcompound network, five ingredients (quercetin, arachidonate, beta-sitosterol, betacarotene, and cholesterol) were selected out as the key ingredients of YJD, which can interact with the 10 hub genes (VEGFA, AKT1, TP53, ALB, TNF, PIK3CA, IGF1, INS, IL1B, PTEN) against PCOS. These genes are mainly involved in prostate cancer, steroid hormone biosynthesis, and EGFR tyrosine kinase inhibitor resistance pathways. In addition, the results of molecular docking showed that the ingredients of YJD have a good binding affinity with the hub genes. Conclusion: These results demonstrate that the treatment of PCOS by YJD is through regulating the levels of androgen and insulin and improving the inflammatory microenvironment.
Billions of people worldwide have experienced irreversible kidney injuries, which is mainly attributed to the complexity of drug-induced nephrotoxicity. Consequently, there is an urgent need for uncovering the mechanisms of nephrotoxicity caused by compounds. In the present study, a network-based methodology was applied to explore the mechanisms of nephrotoxicity induced by specific compounds. Initially, a total of 42 nephrotoxic compounds and 60 kinds of syndromes associated with nephrotoxicity were collected from public resources. Afterward, network localization and separation algorithms were used to map the targets of compounds and diseases into the human interactome. By doing so, 199 statistically significant nephrotoxic networks displaying the interaction between compound targets and disease genes were obtained, which played pivotal roles in compounds-induced nephrotoxicity. Subsequently, enrichment analysis pinpointed core Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways that highlight commonalities in nephrotoxicity induced by nephrotoxic compounds. It was found that nephrotoxic compounds primarily induce nephrotoxicity by mediating the advanced glycosylation end products-receptor for advanced glycosylation end products signaling pathway in diabetic complications, human cytomegalovirus infection, lipid and atherosclerosis, Kaposi sarcoma-associated herpesvirus infection, apoptosis, and the phosphatidylinositol 3-kinase-Akt pathways. These results provide valuable insights for preventing drug-induced nephrotoxicity. Furthermore, the approaches we used are also helpful in conducting research on other kinds of toxicities.
The global pandemic of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection has generated tremendous concern and poses a serious threat to international public health. Phosphorylation is a common post-translational modification affecting many essential cellular processes and is inextricably linked to SARS-CoV-2 infection. Hence, accurate identification of phosphorylation sites will be helpful to understand the mechanisms of SARS-CoV-2 infection and mitigate the ongoing COVID-19 pandemic. In the present study, an attention-based bidirectional gated recurrent unit network, called IPs-GRUAtt, was proposed to identify phosphorylation sites in SARS-CoV-2-infected host cells. Comparative results demonstrated that IPsGRUAtt surpassed both state-of-the-art machine-learning methods and existing models for identifying phosphorylation sites. Moreover, the attention mechanism made IPsGRUAtt able to extract the key features from protein sequences. These results demonstrated that the IPs-GRUAtt is a powerful tool for identifying phosphorylation sites. For facilitating its academic use, a freely available online web server for IPs-GRUAtt is provided at http://cbcb.cdutcm. edu.cn/phosphory/.
The application of network pharmacology has greatly promoted the scientific interpretation of disease treatment mechanism of traditional Chinese medicine (TCM). However, the data required by network pharmacology analysis were scattered in different resources. In the present work, by integrating and reorganizing the data from multiple resources, we developed the intelligent network pharmacology platform unique for traditional Chinese medicine, called INPUT (http://cbcb.cdutcm.edu.cn/INPUT/), for automatically performing network pharmacology analysis. Besides the curated data collected from multiple resources, a series of bioinformatics tools for network pharmacology analysis were also embedded in INPUT, which makes it become the first automatic platform able to explore the disease treatment mechanisms of TCM. With the built-in tools, researchers can also analyze their own in-house data and obtain the results of pivotal ingredients, GO and KEGG pathway, protein-protein interactions, etc. In addition, as a proof-of-principle, INPUT was applied to decipher the antidepressant mechanism of a commonly used prescription. In summary, INPUT is a powerful platform for network pharmacology analysis and will facilitate the researches on drug discovery.
The dynamic RNA modifications were orchestrated by a series of enzymes, namely “writer”, “reader” and “eraser”, which can install, recognize and remove the modifications, respectively. However, only a very small number of experimentally validated RNA modification enzymes have been identified and reported. Therefore, there is an urgent need to develop a database to deposit RNA modification enzymes. In the present work, we developed the RNAME database (https://chenweilab.cn/rname/) to provide a comprehensive resource for RNA modification enzymes. The current version of RNAME deposits more than 21,000 manually curated RNA modification enzymes, which are from 456 species and covers the 7 common kinds of RNA modifications (i.e., adenosine to inosine, N1-methyladenosine, N6-methyladenosine, 5-methylcytidine, N7-methylguanosine, mRNA cap modification, and pseudouridine). The 3D structures, domains, subcellular locations, and biological functions of these enzymes were also integrated in RNAME. It is anticipated that RNAME will facilitate the researches on RNA modifications.
Background:Wujiayizhi granule (WJYZG) is a kind of traditional Chinese medicine, which is used for treating Alzheimer's disease (AD). Although the clinical effect of WJYZG for AD is obvious, its underlying mechanism is still obscure.Objective:Explore the mechanism of WJYZG in the treatment of AD by using bioinformatics methods.Methods:Traditional Chinese Medicine Systems Pharmacology Database and Analysis Platform (TCMSP), Traditional Chinese Medicine Integrated Database (TCMID) and Encyclopedia Database of Chinese Medicine (ETCM) were used to search the ingredients and targets of WJYZG. DisGeNET, Drugbank, Online Mendelian Inheritance in Man (OMIM), and Terapeutic Target Database (TTD) were used to retrieve the targets of AD. The Cytoscape3.6.1 software was used to construct the interaction network of herbs-ingredients-targets. The Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the treatment mechanism of WJYZG on AD. Molecular docking was used to validate the interactions between the ingredients and targets.Results:One hundred and thirty-three ingredients were identified from WJYZG. According to the herbingredient- targets network, quercetin, kaempferol, luteolin, anhydroicaritin, and 8-prenyl-flavone were screened out as the key ingredients, which can interact with the core targets encompassing INS, IL6, TNF, IL1B, CASP3, PTGS2, VEGFA, and PPARG. The enrichment analysis indicates that the treatment of AD by WJYZG was through inhibiting inflammation and neurocyte apoptosis, regulating the calcium ion signaling pathway and adjusting INS levels.Conclusion:The underlying mechanisms of WJYZG in the treatment of AD were theoretically illustrated. We hope these results will enlighten the researches on AD.
Because of their great therapeutic and economic value, medicinal plants have attracted increasing scientific attention. With the rapid development of high-throughput sequencing technology, the genomes of many medicinal plants have been sequenced. Storing and analyzing the increasing volume of genomic data has become an urgent task. To solve this challenge, we have proposed the Traditional Chinese Medicine Plant Genome database (TCMPG, http://cbcb.cdutcm.edu.cn/TCMPG/), an integrative database for storing the scattered genomes of medicinal plants. TCMPG currently includes 160 medicinal plants, 195 corresponding genomes, and 255 herbal medicines. Detailed information on plant species, genomes, and herbal medicines is also integrated into TCMPG. Popular genomic analysis tools are embedded in TCMPG to facilitate the systematic analysis of medicinal plants. These include BLAST for identifying orthologs from different plants, SSR Finder for identifying simple sequence repeats, JBrowse for browsing genomes, Synteny Viewer for displaying syntenic blocks between two genomes, and HmmSearch for identifying protein domains. TCMPG will be continuously updated by integrating new data and tools for comparative and functional genomic analysis.
Long non-coding RNA (lncRNA) plays important roles in a series of biological processes. The transcription of lncRNA is regulated by its promoter. Hence, accurate identification of lncRNA promoter will be helpful to understand its regulatory mechanisms. Since experimental techniques remain time consuming for gnome-wide promoter identification, developing computational tools to identify promoters are necessary. However, only few computational methods have been proposed for lncRNA promoter prediction and their performances still have room to be improved. In the present work, a convolutional neural network based model, called DeepLncPro, was proposed to identify lncRNA promoters in human and mouse. Comparative results demonstrated that DeepLncPro was superior to both state-of-the-art machine learning methods and existing models for identifying lncRNA promoters. Furthermore, DeepLncPro has the ability to extract and analyze transcription factor binding motifs from lncRNAs, which made it become an interpretable model. These results indicate that the DeepLncPro can server as a powerful tool for identifying lncRNA promoters. An open-source tool for DeepLncPro was provided at https://github.com/zhangtian-yang/DeepLncPro.
Establishing an RNA-associated interaction repository facilitates the system-level understanding of RNA functions. However, as these interactions are distributed throughout various resources, an essential prerequisite for effectively applying these data requires that they are deposited together and annotated with confidence scores. Hence, we have updated the RNA-associated interaction database RNAInter (RNA Interactome Database) to version 4.0, which is freely accessible at http://www.rnainter.org or http://www.rna-society.org/rnainter/. Compared with previous versions, the current RNAInter not only contains an enlarged data set, but also an updated confidence scoring system. The merits of this 4.0 version can be summarized in the following points: (i) a redefined confidence scoring system as achieved by integrating the trust of experimental evidence, the trust of the scientific community and the types of tissues/cells, (ii) a redesigned fully functional database that enables for a more rapid retrieval and browsing of interactions via an upgraded user-friendly interface and (iii) an update of entries to >47 million by manually mining the literature and integrating six database resources with evidence from experimental and computational sources. Overall, RNAInter will provide a more comprehensive and readily accessible RNA interactome platform to investigate the regulatory landscape of cellular RNAs.
The ability of a compound to permeate across the blood-brain barrier (BBB) is a significant factor for central nervous system drug development. Thus, for speeding up the drug discovery process, it is crucial to perform high-throughput screenings to predict the BBB permeability of the candidate compounds. Although experimental methods are capable of determining BBB permeability, they are still cost-ineffective and time-consuming. To complement the shortcomings of existing methods, we present a deep learning-based multi-model framework model, called Deep-B3, to predict the BBB permeability of candidate compounds. In Deep-B3, the samples are encoded in three kinds of features, namely molecular descriptors and fingerprints, molecular graph and simplified molecular input line entry system (SMILES) text notation. The pre-trained models were built to extract latent features from the molecular graph and SMILES. These features depicted the compounds in terms of tabular data, image and text, respectively. The validation results yielded from the independent dataset demonstrated that the performance of Deep-B3 is superior to that of the state-of-the-art models. Hence, Deep-B3 holds the potential to become a useful tool for drug development. A freely available online web-server for Deep-B3 was established at http://cbcb.cdutcm.edu.cn/deepb3/, and the source code and dataset of Deep-B3 are available at https://github.com/GreatChenLab/Deep-B3.
Background Agarwood, generated from the Aquilaria sinensis , has high economic and medicinal value. Although its genome has been sequenced, the ploidy of A. sinensis paleopolyploid remains unclear. Moreover, the expression changes of genes associated with agarwood formation were not analyzed either. Results In the present work, we reanalyzed the genome of A. sinensis and found that it experienced a recent tetraploidization event ~ 63–71 million years ago (Mya). The results also demonstrated that the A. sinensis genome had suffered extensive gene deletion or relocation after the tetraploidization event, and exhibited accelerated evolutionary rates. At the same time, an alignment of homologous genes related to different events of polyploidization and speciation were generated as well, which provides an important comparative genomics resource for Thymelaeaceae and related families. Interestingly, the expression changes of genes related to sesquiterpene synthesis in wounded stems of A. sinensis were also observed. Further analysis demonstrated that polyploidization promotes the functional differentiation of the key genes in the sesquiterpene synthesis pathway. Conclusions By reanalyzing its genome, we found that the tetraploidization event shaped the A. sinensis genome and contributed to the ability of sesquiterpenes synthesis. We hope that these results will facilitate our understanding of the evolution of A. sinensis and the function of genes involved in agarwood formation.