Tumorigenesis arises from the dysfunction of cancer genes, leading to uncontrolled cell proliferation through various mechanisms. Establishing a complete cancer gene catalogue will make precision oncology possible. Although existing methods based on graph neural networks (GNN) are effective in identifying cancer genes, they fall short in effectively integrating data from multiple views and interpreting predictive outcomes. To address these shortcomings, an interpretable representation learning framework IMVRL-GCN is proposed to capture both shared and specific representations from multiview data, offering significant insights into the identification of cancer genes. Experimental results demonstrate that IMVRL-GCN outperforms state-of-the-art cancer gene identification methods and several baselines. Furthermore, IMVRL-GCN is employed to identify a total of 74 high-confidence novel cancer genes, and multiview data analysis highlights the pivotal roles of shared, mutation-specific, and structure-specific representations in discriminating distinctive cancer genes. Exploration of the mechanisms behind their discriminative capabilities suggests that shared representations are strongly associated with gene functions, while mutation-specific and structure-specific representations are linked to mutagenic propensity and functional synergy, respectively. Finally, our in-depth analyses of these candidates suggest potential insights for individualized treatments: afatinib could counteract many mutation-driven risks, and targeting interactions with cancer gene SRC is a reasonable strategy to mitigate interaction-induced risks for NR3C1, RXRA, HNF4A, and SP1.
Abstract Long non-coding RNAs (lncRNAs) act as versatile regulators of many biological processes and play vital roles in various diseases. lncRNASNP is dedicated to providing a comprehensive repository of single nucleotide polymorphisms (SNPs) and somatic mutations in lncRNAs and their impacts on lncRNA structure and function. Since the last release in 2018, there has been a huge increase in the number of variants and lncRNAs. Thus, we updated the lncRNASNP to version 3 by expanding the species to eight eukaryotic species (human, chimpanzee, pig, mouse, rat, chicken, zebrafish, and fruitfly), updating the data and adding several new features. SNPs in lncRNASNP have increased from 11 181 387 to 67 513 785. The human mutations have increased from 1 174 768 to 2 387 685, including 1 031 639 TCGA mutations and 1 356 046 CosmicNCVs. Compared with the last release, updated and new features in lncRNASNP v3 include (i) SNPs in lncRNAs and their impacts on lncRNAs for eight species, (ii) SNP effects on miRNA−lncRNA interactions for eight species, (iii) lncRNA expression profiles for six species, (iv) disease & GWAS-associated lncRNAs and variants, (v) experimental & predicted lncRNAs and drug target associations and (vi) SNP effects on lncRNA expression (eQTL) across tumor & normal tissues. The lncRNASNP v3 is freely available at http://gong_lab.hzau.edu.cn/lncRNASNP3/.
MicroRNAs (miRNAs) are crucial regulators in various diseases. The identification of associations between miRNAs and diseases could greatly facilitate the investigation of disease mechanisms and drug development. Limited by time and cost efficiency, conventional experimental techniques are inadequate for this purpose. With the extensive advance and application of deep learning, developing efficient and accurate computational models for predicting miRNA‒disease associations has a vital role and is feasible. In this study, we proposed a meta-path-aggregated multilevel graph embedding model for miRNA‒ disease association prediction. The model first calculated the multiple similarities among miRNAs and diseases, respectively. Then, the node features were extracted from similarity matrices for miRNAs and diseases. Furthermore, we integrated four types of meta-paths from the miRNA‒lncRNA‒disease heterogeneous graph and learned node embeddings by hierarchical graph attention modules. Finally, the model predicted the miRNA‒ disease associations using two-layer graph convolution networks (GCNs). Compared with six state-of-the-art models, the experimental results demonstrated that our model achieved higher prediction performance with an AUC of 0.9892 and an AUPR of 0.9898 for the 5-fold cross-validation on the HMDDv3.2 dataset. With the case study, the model’s performance was further validated, and the top 20 predicted associations could be experimentally confirmed. All in all, it implies the predictive power of our model 1 and the potential value in understanding disease pathology.
Alternative polyadenylation (APA) is an important post-transcription regulatory mechanism widely occurring in eukaryotes and has been associated with special traits/diseases by several studies. However, the dynamic roles and patterns of APA in cell differentiation remain largely unknown. Here, we systematically characterized the APA profiles during the differentiation of induced pluripotent stem cells (iPSCs) to cardiomyocytes by the previously published RNA-seq data across 16 time points. We identified 950 differential APA events and found five dynamic APA patterns with fuzzy c-means clustering analysis. Among them, 3′UTR progressive lengthening is the main APA pattern over time, the genes of which are enriched in cell cycle and mRNA metabolic process pathways. By constructing the linear mixed-effects model, we also indicated that TIA1 plays an important role in regulating APA events with this pattern, including genes essential to cardiac function. Additionally, APA and polyA machinery activity with another pattern can immediately respond to developmental signal-mediated stimuli at the early differentiation stage and result in a sharp shortening of the 3′UTR. Finally, a miRNA-APA network is constructed and several hub miRNAs potentially regulating cardiomyocyte differentiation are detected. Our results show the complex APA mechanisms during the differentiation of iPSCs into cardiomyocytes and provide further insights for the understanding of APA regulation and cell differentiation.
Genome-wide association study (GWAS) has identified thousands of single nucleotide polymorphisms (SNPs) associated with complex diseases and traits. However, deciphering the functions of these SNPs still faces challenges. Recent studies have shown that SNPs could alter chromatin accessibility and result in differences in tumor susceptibility between individuals. Therefore, systematically analyzing the effects of SNPs on chromatin accessibility could help decipher the functions of SNPs, especially those in non-coding regions. Using data from The Cancer Genome Atlas (TCGA), chromatin accessibility quantitative trait locus (caQTL) analysis was conducted to estimate the associations between genetic variants and chromatin accessibility. We analyzed caQTLs in 23 human cancer types and identified 9,478 caQTLs in breast carcinoma (BRCA). In BRCA, these caQTLs tend to alter the binding affinity of transcription factors, and open chromatin regions regulated by these caQTLs are enriched in regulatory elements. By integrating with eQTL data, we identified 141 caQTLs showing a strong signal for colocalization with eQTLs. We also identified 173 caQTLs in genome-wide association studies (GWAS) loci and inferred several possible target genes of these caQTLs. By performing survival analysis, we found that ~10% caQTLs potentially influence the prognosis of patients. To facilitate access to relevant data, we developed a user-friendly data portal, BCaQTL ( http://gong_lab.hzau.edu.cn/caqtl_database ), for data searching and downloading. Our work may facilitate fine-map regulatory mechanisms underlying risk loci of cancer and discover the biomarkers or therapeutic targets for cancer prognosis. The BCaQTL database will be an important resource for genetic and epigenetic studies.
Enhancer RNAs (eRNAs) are a class of non-coding RNAs transcribed from enhancers. As the markers of active enhancers, eRNAs play important roles in gene regulation and are associated with various complex traits and characteristics. With increasing attention to eRNAs, numerous eRNAs have been identified in different human tissues. However, the expression landscape, regulatory network and potential functions of eRNAs in animals have not been fully elucidated. Here, we systematically characterized 185 177 eRNAs from 5085 samples across 10 species by mapping the RNA sequencing data to the regions of known enhancers. To explore their potential functions based on evolutionary conservation, we investigated the sequence similarity of eRNAs among multiple species. In addition, we identified the possible associations between eRNAs and transcription factors (TFs) or nearby genes to decipher their possible regulators and target genes, as well as characterized trait-related eRNAs to explore their potential functions in biological processes. Based on these findings, we further developed Animal-eRNAdb (http://gong_lab.hzau.edu.cn/Animal-eRNAdb/), a user-friendly database for data searching, browsing and downloading. With the comprehensive characterization of eRNAs in various tissues of different species, Animal-eRNAdb may greatly facilitate the exploration of functions and mechanisms of eRNAs.