The current reliance on wet-lab experiments for evaluating the efficacy and adverse effects of natural products remains a major obstacle in new drug discovery. We propose KAN-PROSPECT, a novel deep learning framework that integrates transfer learning with Kolmogorov-Arnold Networks (KAN). This model can simultaneously predict the efficacy and adverse effects of natural products based solely on molecular SMILES representations, thereby addressing the limited generalizability of existing models in predicting these aspects for natural products. Beyond methodological advances, KAN-PROSPECT also contributes to more sustainable and resource-efficient drug discovery. Leveraging a cross-modal transfer learning strategy pretrained on approximately 3,800 drugs and fine-tuned on 400 natural products, KAN-PROSPECT consistently outperforms baseline models in dual-label prediction tasks. Notably, it demonstrates exceptional robustness in addressing data scarcity, excelling particularly in few-shot and zero-shot scenarios. Through the transfer learning strategy, the model partially alleviates the data scarcity issue commonly encountered in natural product prediction tasks. In addition, the incorporation of KAN layers enhances the ability to model complex nonlinear relationships between molecular structures and associated pharmacological or adverse reaction profiles, contributing to improved predictive performance. Furthermore, the framework demonstrates strong generalization ability, enabling high-accuracy predictions for the efficacy and adverse effects of entirely new natural products. KAN-PROSPECT was further applied to comprehensively predict natural products from the MEC and NPASS databases, with Icaritin from Epimedium used as a representative case study. Overall, KAN-PROSPECT is the first framework to unify transfer learning with the KAN architecture for dual-label prediction of natural products, showing great potential for large-scale bioactivity and toxicity prediction, new drug development, and drug repositioning of natural products.
Chronic infection with Helicobacter pylori (H. pylori) is a major environmental risk factor for gastric carcinogenesis. Malignancy is largely driven by variations within the virulence factor CagA, with East Asian lineages exhibiting higher oncogenic potential than Western ones. However, how these variants modulate cellular crosstalk remains poorly understood. We integrated molecular dynamics (MD) simulations with single-cell transcriptomics across progressive disease stages, including chronic atrophic gastritis, intestinal metaplasia, and gastric cancer. Local niche remodeling was evaluated via cell–cell communication profiling among epithelial, stromal, and immune circuits, while simulations of MARK2 kinase bound to distinct CagA lineages determined binding affinities. Single-cell analysis revealed that H. pylori toxicity progressively dampens epithelial–stromal crosstalk, marked by severe epithelial polarity aberrations that disrupt neuroendocrine-like secretory and synaptic pathways during malignant transformation. Mechanistically, MD simulations and MM/GBSA calculations demonstrated that East Asian CagA lineages exhibit higher binding affinity toward host MARK2 than Western lineages. Specific East Asian amino acid substitutions dramatically tighten the protein interface, driving stronger signaling perturbations. This study bridges atomistic structural virulence with microenvironmental shifting, establishing geographic CagA toxicity divergence as a critical determinant for pathogen-driven gastric cancer risk.
The prediction of drug response would significantly improve the treatment of lung cancer. Tumor heterogeneity and complex signal transduction pathways lead to varied treatment effects among patients, but traditional computational approaches struggle to model the nonlinear, high-dimensional relationship between genes and drug responses. In order to develop a Generative Adversarial Network (GAN)-based model that can predict drug-induced gene expression profiles from lung cancer cell lines, we developed GRIP-Lung (Generative Model of Response to Drug-Induced Perturbation in Lung Cancer). By making use of biologically informed embeddings of cell line identity as well as drug treatment conditions, this model is able to gain a fairly good understanding of cell types and their transcriptional perturbations induced by different drugs. The GRIP-Lung model displayed reasonably good prediction ability in terms of predictive accuracy and showed high concordance between the predicted and experimental expression profiles. We not only predicted transcriptional changes induced by drug therapy but also used single-sample Gene Set Enrichment Analysis (ssGSEA) to classify post-treatment response states based on characteristic molecular biomarkers, offering a means for selecting effective drugs to target specific heterogeneity within lung tumors. The proposed GRIP-Lung framework faithfully reproduces drug-induced transcriptional perturbations in lung cell line models. By integrating biologically informed embeddings and adversarial learning, the model advances drug response prediction. This makes it a flexible computational tool for drug repositioning.
Metastasis, the spread of cancer cells from the primary tumor to distant organs, is the leading cause of mortality in cancer patients. This process often exhibits a preference for specific organs, a phenomenon known as tumor organotropism. This study focuses on the organotropism of breast cancer and analyzes its genomic alterations following metastasis to four organs (bone, brain, liver, and lung). The research aims to explore the intrinsic characteristics of primary breast cancer and the interactions between tumor cells and the tumor microenvironment (TME) within these target organs. Building upon this foundation, we developed a deep learning model to identify organ-specific metastatic genes, providing insights into the molecular mechanisms of metastasis. To investigate the mechanisms of organ-specific metastasis in breast cancer, we employed an integrative approach combining single-cell RNA sequencing, bulk RNA sequencing, ChIP-seq data, and deep learning techniques. Single-cell analysis provided detailed insights into cellular heterogeneity and microenvironment interactions at metastatic sites. Bulk RNA sequencing enabled the identification of gene expression patterns associated with metastatic propensity. A deep neural network (DNN) model was developed to analyze these complex datasets and identify key predictors of organ-specific metastasis. Our integrative analysis revealed distinct gene expression profiles and cellular compositions in metastatic lesions across different organs. We have identified that, regardless of the target organ, breast cancer metastasis critically depends on specific biological signaling pathways, including the MAPK signaling pathway, metabolic pathways, the PI3K-Akt signaling pathway, and the positive regulation of cell adhesion. Single-cell sequencing highlighted unique interactions between tumor cells and the microenvironment, which varied significantly depending on the metastatic site. Fibroblasts play a critical role in facilitating the colonization of breast cancer cells in metastatic organs. The deep learning models effectively identified key molecular signatures and pathways associated with organ-specific metastasis, providing insights into the metastatic process. The study underscores the importance of the tumor microenvironment in influencing breast cancer metastasis to distant organs. We also established a comprehensive framework for understanding the mechanisms driving organotropism metastasis in breast cancer. Additionally, we identified key genes and signaling pathways associated with organ-specific metastasis, providing insights that may inform future studies on risk assessment and potential therapeutic targets for metastatic breast cancer.
Background: The oligometastatic disease has been proposed as an intermediate state between primary tumor and systemically metastatic disease, which has great potential curable with locoregional therapies. However, since no biomarker for the identification of patients with true oligometastatic disease is clinically available, the diagnosis of oligometastatic disease remains controversial. Objective: We aim to identify potential biomarkers of colorectal cancer patients with true oligometastatic states, who will benefit most from local therapy. Methods: This study retrospectively analyzed the transcriptome profiles and clinical parameters of 307 metastatic colorectal cancer patients. A novel network propagation method and network-based strategy were combined to identify oligometastatic biomarkers to predict the prognoses of metastatic colorectal cancer patients. Results: We defined two metastatic risk groups according to twelve oligometastatic biomarkers, which exhibit distinct prognoses, clinicopathological features, immunological characteristics, and biological mechanisms. The metastatic risk assessment model exhibited a more powerful capacity for survival prediction compared to traditional clinicopathological features. The low-MRS group was most consistent with an oligometastatic state, while the high-MRS might be a potential polymetastatic state, which leads to the divergence of their prognostic outcomes and response to treatments. We also identified 22 significant immune check genes between the high-MRS and low- MRS groups. The difference in molecular mechanism between the two metastatic risk groups was associated with focal adhesion, nucleocytoplasmic transport, Hippo, PI3K-Akt, TGF-β, and EMCreceptor interaction signaling pathways. Conclusion: Our study provided a molecular definition of the oligometastatic state in colorectal cancer, which contributes to precise treatment decision-making for advanced patients.
The phenotypes of drug action, including therapeutic actions and adverse drug reactions (ADRs), are important indicators for evaluating the druggability of new drugs and repositioning the approved drugs. Here, we provide a user-friendly database, DAPredict (http://bio-bigdata.hrbmu.edu.cn/DAPredict), in which our novel original drug action phenotypes prediction algorithm (Yang,J., Zhang,D., Liu,L. et al. (2021) Computational drug repositioning based on the relationships between substructure-indication. Brief. Bioinformatics, 22, bbaa348) was embedded. Our algorithm integrates characteristics of chemical genomics and pharmacogenomics, breaking through the limitations that traditional drug development process based on phenotype cannot analyze the mechanism of drug action. Predicting phenotypes of drug action based on the local active structures of drugs and proteins can achieve more innovative drug discovery across drug categories and simultaneously evaluate drug efficacy and safety, rather than traditional one-by-one evaluation. DAPredict contains 305 981 predicted relationships between 1748 approved drugs and 454 ADRs, 83 117 predicted relationships between 1478 approved drugs and 178 Anatomical Therapeutic Chemicals (ATC). More importantly, DAPredict provides an online prediction tool, which researchers can use to predict the action phenotypic spectrum of more than 110 000 000 compounds (including about 168 000 natural products) and corresponding proteins to analyze their potential effect mechanisms. DAPredict can also help researchers obtain the phenotype-corresponding active structures for structural optimization of new drug candidates, making it easier to evaluate the druggability of new drug candidates and develop more innovative drugs across drug categories.Database URL: http://bio-bigdata.hrbmu.edu.cn/DAPredict/
At the beginning of the "Disease X" outbreak, drug discovery and development are often challenged by insufficient and unbalanced data. To address this problem and maximize the information value of limited data, we propose a drug screening model, LGCNN, based on convolutional neural network (CNN), which enables rapid drug screening by integrating features of drug molecular structures and drug-target interactions at both local and global (LG) levels. Experimental results show that LGCNN exhibits better performance compared to other state-of-the-art classification methods under limited data. In addition, LGCNN was applied to anti-SARS-CoV-2 drug screening to realize therapeutic drug mining against COVID-19. LGCNN transcends the limitations of traditional models for predicting interactions between single drug targets and shows new advantages in predicting multi-target drug-target interactions. Notably, the cross-coronavirus generalizability of the model is also implied by the analysis of targets, drugs, and mechanisms in the prediction results. In conclusion, LGCNN provides new ideas and methods for rapid drug screening in emergency situations where data are scarce.
This study investigates the impact of Hashimoto's thyroiditis (HT), an autoimmune disorder, on the papillary thyroid cancer (PTC) microenvironment using a dataset of 140,456 cells from 11 patients. By comparing PTC cases with and without HT, we identify HT-specific cell populations (HASCs) and their role in creating a TSH-suppressive environment via mTE3, nTE0, and nTE2 thyroid cells. These cells facilitate intricate immune-stromal communication through the MIF-(CD74+CXCR4) axis, emphasizing immune regulation in the TSH context. In the realm of personalized medicine, our HASC-focused analysis within the TCGA-THCA dataset validates the utility of HASC profiling for guiding tailored therapies. Moreover, we introduce a novel, objective method to determine K-means clustering coefficients in copy number variation inference from bulk RNA-seq data, mitigating the arbitrariness in conventional coefficient selection. Collectively, our research presents a detailed single-cell atlas illustrating HT-PTC interactions, deepening our understanding of HT's modulatory effects on PTC microenvironments. It contributes to our understanding of autoimmunity-carcinogenesis dynamics and charts a course for discovering new therapeutic targets in PTC, advancing cancer genomics and immunotherapy research.
BACKGROUND:Non-small cell lung cancer (NSCLC) is a prevalent and heterogeneous disease with significant genomic variations between the early and advanced stages. The identification of key genes and pathways driving NSCLC tumor progression is critical for improving the diagnosis and treatment outcomes of this disease.METHODS:In this study, we conducted single-cell transcriptome analysis on 93,406 cells from 22 NSCLC patients to characterize malignant NSCLC cancer cells. Utilizing cNMF, we classified these cells into distinct modules, thus identifying the diverse molecular profiles within NSCLC. Through pseudotime analysis, we delineated temporal gene expression changes during NSCLC evolution, thus demonstrating genes associated with disease progression. Using the XGBoost model, we assessed the significance of these genes in the pseudotime trajectory. Our findings were validated by using transcriptome sequencing data from The Cancer Genome Atlas (TCGA), supplemented via LASSO regression to refine the selection of characteristic genes. Subsequently, we established a risk score model based on these genes, thus providing a potential tool for cancer risk assessment and personalized treatment strategies.RESULTS:We used cNMF to classify malignant NSCLC cells into three functional modules, including the metabolic reprogramming module, cell cycle module, and cell stemness module, which can be used for the functional classification of malignant tumor cells in NSCLC. These findings also indicate that metabolism, the cell cycle, and tumor stemness play important driving roles in the malignant evolution of NSCLC. We integrated cNMF and XGBoost to select marker genes that are indicative of both early and advanced NSCLC stages. The expression of genes such as CHCHD2, GAPDH, and CD24 was strongly correlated with the malignant evolution of NSCLC at the single-cell data level. These genes have been validated via histological data. The risk score model that we established (represented by eight genes) was ultimately validated with GEO data.CONCLUSION:In summary, our study contributes to the identification of temporal heterogeneous biomarkers in NSCLC, thus offering insights into disease progression mechanisms and potential therapeutic targets. The developed workflow demonstrates promise for future applications in clinical practice.
Combination therapy is a promising strategy for cancers, increasing therapeutic options and reducing drug resistance. Yet, systematic identification of efficacious drug combinations is limited by the combinatorial explosion caused by a large number of possible drug pairs and diseases. At present, machine learning techniques have been widely applied to predict drug combinations, but most studies rely on the response of drug combinations to specific cell lines and are not entirely satisfactory in terms of mechanism interpretability and model scalability. Here, we proposed a novel network propagation-based machine learning framework to predict synergistic drug combinations. Based on the topological information of a comprehensive drug-drug association network, we innovatively introduced an affinity score between drug pairs as one of the features to train machine learning models. We applied network-based strategy to evaluate their therapeutic potential to different cancer types. Finally, we identified 17 specific-, 21 general- and 40 broad-spectrum antitumor drug combinations, in which 69% drug combinations were validated by vitro cellular experiments, 83% drug combinations were validated by literature reports and 100% drug combinations were validated by biological function analyses. By quantifying the network relationships between drug targets and cancer-related driver genes in the human protein-protein interactome, we show the existence of four distinct patterns of drug-drug-disease relationships. We also revealed that 32 biological pathways were correlated with the synergistic mechanism of broad-spectrum antitumor drug combinations. Overall, our model offers a powerful scalable screening framework for cancer treatments.
Background: Long non-coding RNAs (lncRNAs) play an important role in the immune regulation of gastric cancer (GC). However, the clinical application value of immune-related lncRNAs has not been fully developed. It is of great significance to overcome the challenges of prognostic prediction and classification of gastric cancer patients based on the current study.Methods: In this study, the R package ImmLnc was used to obtain immune-related lncRNAs of The Cancer Genome Atlas Stomach Adenocarcinoma (TCGA-STAD) project, and univariate Cox regression analysis was performed to find prognostic immune-related lncRNAs. A total of 117 combinations based on 10 algorithms were integrated to determine the immune-related lncRNA prognostic model (ILPM). According to the ILPM, the least absolute shrinkage and selection operator (LASSO) regression was employed to find the major lncRNAs and develop the risk model. ssGSEA, CIBERSORT algorithm, the R package maftools, pRRophetic, and clusterProfiler were employed for measuring the proportion of immune cells among risk groups, genomic mutation difference, drug sensitivity analysis, and pathway enrichment score.Results: A total of 321 immune-related lncRNAs were found, and there were 26 prognostic immune-related lncRNAs. According to the ILPM, 18 of 26 lncRNAs were selected and the risk score (RS) developed by the 18-lncRNA signature had good strength in the TCGA training set and Gene Expression Omnibus (GEO) validation datasets. Patients were divided into high- and low-risk groups according to the median RS, and the low-risk group had a better prognosis, tumor immune microenvironment, and tumor signature enrichment score and a higher metabolism, frequency of genomic mutations, proportion of immune cell infiltration, and antitumor drug resistance. Furthermore, 86 differentially expressed genes (DEGs) between high- and low-risk groups were mainly enriched in immune-related pathways.Conclusion: The ILPM developed based on 26 prognostic immune-related lncRNAs can help in predicting the prognosis of patients suffering from gastric cancer. Precision medicine can be effectively carried out by dividing patients into high- and low-risk groups according to the RS.
Introduction: Target therapy for cancer cell mutation has brought attention to several challenges in clinical applications, including limited therapeutic targets, less patient benefits, and susceptibility to acquired due to their clear biological mechanisms and high specificity in targeting cancers with specific mutations. However, the identification of truly lethal synthetic lethal therapeutic targets for cancer cells remains uncommon, primarily due to compensatory mechanisms.Methods: In our pursuit of core therapeutic targets (CTTs) that exhibit extensive synthetic lethality in cancer and the corresponding potential drugs, we have developed a machine-learning model that utilizes multiple levels and dimensions of cancer characterization. This is achieved through the consideration of the transcriptional and post-transcriptional regulation of cancer-specific genes and the construction of a model that integrates statistics and machine learning. The model incorporates statistics such as Wilcoxon and Pearson, as well as random forest. Through WGCNA and network analysis, we identify hub genes in the SL network that serve as CTTs. Additionally, we establish regulatory networks for non-coding RNA (ncRNA) and drug-target interactions.Results: Our model has uncovered 7277 potential SL interactions, while WGCNA has identified 13 gene modules. Through network analysis, we have identified 30 CTTs with the highest degree in these modules. Based on these CTTs, we have constructed networks for ncRNA regulation and drug targets. Furthermore, by applying the same process to lung cancer and renal cell carcinoma, we have identified corresponding CTTs and potential therapeutic drugs. We have also analyzed common therapeutic targets among all three cancers.Discussion: The results of our study have broad applicability across various dimensions and histological data, as our model identifies potential therapeutic targets by learning multidimensional complex features from known synthetic lethal gene pairs. The incorporation of statistical screening and network analysis further enhances the confidence in these potential targets. Our approach provides novel theoretical insights and methodological support for the identification of CTTs and drugs in diverse types of cancer.
Identifying drug phenotypic effects, including therapeutic effects and adverse drug reactions (ADRs), is an inseparable part for evaluating the potentiality of new drug candidates (NDCs). However, current computational methods for predicting phenotypic effects of NDCs are mainly based on the overall structure of an NDC or a related target. These approaches often lead to inconsistencies between the structures and functions and limit the prediction space of NDCs. In this study, first, we constructed quantitative associations of substructure-domain, domain-ADR, and domain-ATC (Anatomical Therapeutic Chemical Classification System code) through L1LOG and L1SVM machine learning models. These associations represent relationships between phenotypes (ADRs and ATCs) and local structures of drugs and proteins. Then, based on these established associations, substructure-phenotype relationships were constructed which were utilized to quantify drug-phenotype relationships. Thus, this approach could achieve high-throughput and effective evaluations of the druggability of NDCs by referring to the established substructure-phenotype relationships and structural information of NDCs without additional prior knowledge. Using this computational pipeline, 83,205 drug-ATC relationships (including 1,479 drugs and 178 ATCs) and 306,421 drug-ADR relationships (including 1,752 drugs and 454 ADRs) were predicted in total. The prediction results were validated at four levels: five-fold cross validation, public databases, literature, and molecular docking. Furthermore, three case studies demonstrated the feasibility of our method. 79 ATCs and 269 ADRs were predicted to be related to Maraviroc, an approved drug, including the existing antiviral effect in clinical use. Additionally, we also found risk substructures of severe ADRs, for example, SUB215 (>= 1, saturated or only aromatic carbon ring size 7) can result in shock. And we analyzed the mechanism of action (MOA) of interested drugs based on the established drug-substructure-domain-protein associations. In a word, this approach through establishing drug-substructure-phenotype relationships can achieve quantitative prediction of phenotypes for a given NDC or drug without any prior knowledge except its structure information. Using that way, we can directly obtain the relationships between substructure and phenotype of a compound, which is more convenient to analyze the phenotypic mechanism of drugs and accelerate the process of rational drug design.
At present, most patients with oral squamous cell carcinoma (OSCC) are in the middle or advanced stages at the time of diagnosis. Advanced OSCC patients have a poor prognosis after traditional therapy, and the complex heterogeneity of OSCC has been proven to be one of the main reasons. Single-cell sequencing technology provides a powerful tool for dissecting the heterogeneity of cancer. However, most of the current studies at the single-cell level are static, while the development of cancer is a dynamic process. Thus, understanding the development of cancer from a dynamic perspective and formulating corresponding therapeutic measures for achieving precise treatment are highly necessary, and this is also one of the main study directions in the field of oncology. In this study, we combined the static and dynamic analysis methods based on single-cell RNA-Seq data to comprehensively dissect the complex heterogeneity and evolutionary process of OSCC. Subsequently, for clinical practice, we revealed the association between cancer heterogeneity and the prognosis of patients. More importantly, we pioneered the concept of pseudo-time score of patients, and we quantified the levels of heterogeneity based on the dynamic development process to evaluate the relationship between the score and the survival status at the same stage, finding that it is closely related to the prognostic status. The pseudo-time score of patients could not only reflect the tumor status of patients but also be used as an indicator of the effects of drugs on the patients so that the medication strategy can be adjusted on time. Finally, we identified candidate drugs and proposed precision medication strategies to control the condition of OSCC in two respects: treatment and blocking.
Background Breast cancer (BC) is a complex disease with high heterogeneity, which often leads to great differences in treatment results. Current common molecular typing method is PAM50, which shows positive results for precision medicine; however, room for improvement still remains because of the different prognoses of subtypes. Therefore, in this article, we used lncRNAs, which are more tissue-specific and developmental stage-specific than other RNAs, as typing markers and combined single-cell expression profiles to retype BC, to provide a new method for BC classification and explore new precise therapeutic strategies based on this method. Methods Based on lncRNA expression profiles of 317 single cells from 11 BC patients, SC3 was used to retype BC, and differential expression analysis and enrichment analysis were performed to identify biological characteristics of new subtypes. The results were validated for survival analysis using data from TCGA. Then, the downstream regulatory genes of lncRNA markers of each subtype were searched by expression correlation analysis, and these genes were used as targets to screen therapeutic drugs, thus proposing new precision treatment strategies according to the different subtype compositions of patients. Results Seven lncRNA subtypes and their specific biological characteristics are obtained. Then, 57 targets and 210 drugs of 7 subtypes were acquired. New precision medicine strategies were proposed according to the different compositions of patient subtypes. Conclusions For patients with different subtype compositions, we propose a strategy to select different drugs for different patients, which means using drugs targeting multi subtype or combinations of drugs targeting a single subtype to simultaneously kill different cancer cells by personalized treatment, thus reducing the possibility of drug resistance and even recurrence.
At present, computational methods for drug repositioning are mainly based on the whole structures of drugs, which limits the discovery of new functions due to the similarities between local structures of drugs. In this article, we, for the first time, integrated the features of chemical-genomics (substructure-domain) and pharmaco-genomics (domain-indication) based on the assumption that drug-target interactions are mediated by the substructures of drugs and the domains of proteins to identify the relationships between substructure-indication and establish a drug-substructure-indication network for predicting all therapeutic effects of tested drugs through only information on the substructures of drugs. In total, 83 205 drug-indication relationships with different correlation scores were obtained. We used three different verification methods to indicate the accuracy of the method and the reliability of the scoring system. We predicted all indications of olaparib using our method, including the known antitumor effect and unknown antiviral effect verified by literature, and we also discovered the inhibitory mechanism of olaparib toward DNA repair through its specific sub494 (o = C-C: C), as it participates in the low synthesis of the poly subfunction of the apoptosis pathway (hsa04210) by inhibiting the Inositol 1,4,5-trisphosphate receptor(s) (ITPRs) and hydrolyzing poly (ADP ribose) polymerases. ElectroCardioGrams of four drugs (quinidine, amiodarone, milrinone and fosinopril) demonstrated the effect of anti-arrhythmia. Unlike previous studies focusing on the overall structures of drugs, our research has great potential in the search for more therapeutic effects of drugs and in predicting all potential effects and mechanisms of a drug from the local structural similarity.
Background The analysis of cancer diversity based on a logical framework of hallmarks has greatly improved our understanding of the occurrence, development and metastasis of various cancers. Methods We designed Cancer Hallmark Genes (CHG) database which focuses on integrating hallmark genes in a systematic, standard way and annotates the potential roles of the hallmark genes in cancer processes. Following the conceptual criteria description of hallmark function the keywords for each hallmark were manually selected from the literature. Candidate hallmark genes collected were derived from 301 pathways of KEGG database by Lucene and manually corrected. Results Based on the variation data, we finally identified the hallmark genes of various types of cancer and constructed CHG. And we also analyzed the relationships among hallmarks and potential characteristics and relationships of hallmark genes based on the topological structures of their networks. We manually confirm the hallmark gene identified by CHG based on literature and database. We also predicted the prognosis of breast cancer, glioblastoma multiforme and kidney papillary cell carcinoma patients based on CHG data. Conclusions In summary, CHG, which was constructed based on a hallmark feature set, provides a new perspective for analyzing the diversity and development of cancers.
Colorectal cancer (CRC) is one of the leading causes of cancer-related death worldwide. Due to the lack of early diagnosis methods and warning signals of CRC and its strong heterogeneity, the determination of accurate treatments for CRC and the identification of specific early warning signals are still urgent problems for researchers. In this study, the expression profiles of cancer tissues and the expression profiles of tumor-adjacent tissues in 28 CRC patients were combined into a human protein-protein interaction (PPI) network to construct a specific network for each patient. A network propagation method was used to obtain a mutant giant cluster (GC) containing more than 90% of the mutation information of one patient. Next, mutation selection rules were applied to the GC to mine the mutation sequence of driver genes in each CRC patient. The mutation sequences from patients with the same type CRC were integrated to obtain the mutation sequences of driver genes of different types of CRC, which provide a reference for the diagnosis of clinical CRC disease progression. Finally, dynamic network analysis was used to mine dynamic network biomarkers (DNBs) in CRC patients. These DNBs were verified by clinical staging data to identify the critical transition point between the pre-disease state and the disease state in tumor progression. Twelve known drug targets were found in the DNBs, and 6 of them have been used as targets for anticancer drugs for clinical treatment. This study provides important information for the prognosis, diagnosis and treatment of CRC, especially for pre-emptive treatments. It is of great significance for reducing the incidence and mortality of CRC.
Chemotherapy agents can cause serious adverse effects by attacking both cancer tissues and normal tissues. Therefore, we proposed a synthetic lethality (SL) concept-based computational method to identify specific anticancer drug targets. First, a 3-step screening strategy (network-based, frequency-based and function-based screening) was proposed to identify the SL gene pairs by mining 697 cancer genes and the human signaling network, which had 6306 proteins and 62937 protein-protein interactions. The network-based screening was composed of a stability score constructed using a network information centrality measure (the average shortest path length) and the distance-based screening between the cancer gene and the non-cancer gene. Then, the non-cancer genes were extracted and annotated using drug-target interaction and drug description information to obtain potential anticancer drug targets. Finally, the human SL data in SynLethDB, the existing drug sensitivity data and text-mining were utilized for target validation. We successfully identified 2555 SL gene pairs and 57 potential anticancer drug targets. Among them, CDK1, CDK2, PLK1 and WEE1 were verified by all three aspects and could be preferentially used in specific targeted therapy in the future.
Hepatocellular carcinoma (HCC) is the most frequent type of liver cancer with poor survival rate and high mortality. Despite efforts on the mechanism of HCC, new molecular markers are needed for exact diagnosis, evaluation and treatment. Here, we combined transcriptome of HCC with networks and pathways to identify reliable molecular markers. Through integrating 249 differentially expressed genes with syncretic protein interaction networks, we constructed a HCC-specific network, from which we further extracted 480 pivotal genes. Based on the cross-talk between the enriched pathways of the pivotal genes, we finally identified a HCC signature of 45 genes, which could accurately distinguish HCC patients with normal individuals and reveal the prognosis of HCC patients. Among these 45 genes, 15 showed dysregulated expression patterns and a part have been reported to be associated with HCC and/or other cancers. These findings suggested that our identified 45 gene signature could be potential and valuable molecular markers for diagnosis and evaluation of HCC.