Alzheimer’s disease (AD) is a severe neurodegenerative disorder whose pathological progression is closely associated with the reduced binding affinity of apolipoprotein E ε4 (ApoE4) for amyloid-β (Aβ), which impairs Aβ clearance. Existing computational molecular design approaches are largely limited to single-target optimization and therefore lack the capacity for synergistic dual-target modulation. Herein, we proposed EvoPlay-MuZero, a hybrid computational framework incorporating a dual-peptide bridging (DPB) strategy. The framework adopted latent-state planning in MuZero reinforcement learning to enhance exploration and sequence-generation efficiency in high-dimensional sequence spaces. For the first time, it enabled the automated design of bispecific peptides targeting ApoE4 and Aβ, thereby forming a synergistic molecular bridge via a flexible linker. A full-process pipeline for structural and energetic evaluation was established by integrating AlphaFold3 and PDBePISA. Benchmark experiments on the 1SSC, 2CNZ, and 3R7G datasets demonstrated that EvoPlay-MuZero substantially outperformed the vanilla EvoPlay in convergence speed and the yield of valid generated sequences. Specifically, the optimal DPB molecule (15 × 25-3A) achieved a calculated interfacial solvation energy score (ΔiG) of −29.5 kcal/mol, demonstrating substantially enhanced interface-stabilization properties compared to the native baseline control. This study provides a novel molecular intervention strategy for adjuvant therapy in AD and highlights the considerable potential of reinforcement learning in multi-target drug design.
Pleiotropic genetic loci have been increasingly reported in cancer, and identifying genetic variants with pleiotropic associations can reveal shared biological pathways influencing multiple cancers. Using summary statistics from genome-wide association studies for 37 cancer types (N = 433 836), we identified extensive genome-wide and local genetic correlations among cancers. Through pairwise pleiotropic analysis, we identified 75 243 significant pleiotropic single nucleotide polymorphisms (SNPs) across 372 cancer pairs, among which 3472 were lead SNPs with potential regulatory functions. Using FUMA and MAGMA, we identified 2527 pleiotropic risk loci and 4272 candidate pleiotropic genes. Notably, genes such as TERT (5p15.33), POU5F1B (8q24.21), and FANCA (16q24.3) exhibited widespread pleiotropy across multiple cancer types. Pathway enrichment analysis highlighted the critical roles of pigment synthesis, metabolism, and apoptosis in skin-related cancers, while cross-cancer enrichment analysis emphasized pathways related to apoptosis, chromatin structure, and intermediate filaments. We also identified 33 novel functional genes harboring previously unreported cancer risk variants. Drug-gene interaction analysis revealed several repositionable FDA-approved drugs. Importantly, drug sensitivity assays demonstrated that bosutinib and cobimetinib exhibited promising therapeutic potential in breast cancer cell lines. Finally, we developed the PleioCancer database (https://gonglab.hzau.edu.cn/PleioCancer/), providing a comprehensive resource for cancer pleiotropy research. These findings have important implications for carcinogenesis cancer, prevention and treatment.
Increasing evidence shows that epistasis, defined as interactive effects between genetic loci, may contribute to the missing heritability of cancer. However, systematic genome-wide epistasis identification in cancer remains challenging. Here, by leveraging genotype and clinical data from 380,983 samples in the UK Biobank, we identified 202,032 candidate epistatic single nucleotide polymorphism (epiSNP) pairs associated with cancer risk across 16 cancer types. Notably, multivariable Cox regression identified 123 epiSNP pairs with significant interaction effects on overall survival, suggesting that interaction-level genetic signals can provide prognostic information beyond individual SNP effects. Through functional analysis of the 202,032 candidate epiSNP pairs, we identified 7152 pairs supported by gene co-expression data and 12,326 pairs with protein–protein interaction (PPI) evidence. By mapping epiSNP pairs to corresponding gene pairs and then linking these gene pairs to drug–target databases, we identified 1040 epistatic gene pairs with FDA-approved drug–target records. Additionally, through KM survival analysis of the candidate epiSNP pairs, we detected 7068 pairs significantly associated with patient overall survival. Finally, we constructed an open-access database, EpiSNPdb, to facilitate cancer epistasis research.
The explosive expansion of fish multi-omics data is reshaping basic and applied research in fisheries and aquaculture. Here, we introduce iFish, a rigorously curated and comprehensive resource that integrates genome, transcriptome, epigenome, and proteome datasets from 88 fish species. From 884 whole-genome sequencing datasets, we identified 137.04 million single-nucleotide polymorphisms and 45.52 million insertions/deletions. Leveraging 9797 bulk RNA-seq, 293 single-cell RNA-seq, and 840 microRNA (miRNA)-seq datasets, we quantified 1.87 million messenger RNAs, 287 873 long noncoding RNAs, 197 554 circular RNAs, and 6068 mature miRNAs across different tissues of multiple species. By integrating 1563 epigenetics sequencing datasets, we provided genome-wide maps of histone modifications, chromatin accessibility, DNA methylation, and three-dimensional chromatin interactions. Furthermore, 50 proteomics projects were included for protein expression profiling. In addition, to facilitate gene analyses, we also curated 2.7 million gene annotation entries, identified homologous genes, annotated 192 107 transcription factors (TFs), constructed TF regulatory networks and gene co-expression networks, and performed cross-species conservation analyses of noncoding RNAs. Finally, we established the iFish database (https://gonglab.hzau.edu.cn/iFish/), a user-friendly platform that supports interactive browsing, visualization, and download. With comprehensive data and useful tools, iFish will advance genetic, transcriptional-regulatory, and epigenetic investigations, and accelerate genetic improvement and practical applications in aquaculture.
Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.
BACKGROUND Response to neoadjuvant chemoradiotherapy (nCRT) varies substantially among patients with rectal cancer (RC). Identifying clinically accessible pretreatment predictors is essential for optimizing risk stratification and individualized treatment planning. AIM To investigate whether pretreatment sarcopenia and tumor-related imaging parameters assessed on magnetic resonance imaging can predict treatment response, tumor downstaging, and overall survival (OS) in patients with RC undergoing nCRT. METHODS This retrospective study included 137 patients with RC who underwent standard nCRT. Clinical characteristics, tumor-related imaging parameters derived from pretreatment magnetic resonance imaging, and pretreatment sarcopenia were analyzed as potential prognostic factors. Sarcopenia was assessed using the skeletal muscle index measured on computed tomography at the time of diagnosis. Univariate and multivariate logistic regression analyses were performed to identify predictors of treatment response and tumor downstaging. OS was analyzed using the Kaplan-Meier method and Cox proportional hazards regression. RESULTS Pretreatment sarcopenia was non-significantly associated with response to nCRT or OS. Conversely, tumor length [hazard ratio (HR) = 1.029, 95%CI: 1.011-1.056; P = 0.004] and mesorectal fascia (MRF) involvement (HR = 1.853, 95%CI: 0.873-3.970; P = 0.025) independently predicted treatment response. Clinical T stage (HR = 4.928, 95%CI: 2.170-12.340; P < 0.001) and MRF status (HR = 4.456, 95%CI: 1.881-11.532; P = 0.001) were significantly associated with tumor downstaging. Nodal status (HR = 2.655, 95%CI: 1.281-5.503; P = 0.009) and MRF status (HR = 2.149, 95%CI: 1.005-4.596; P = 0.049) were associated with OS in univariate analysis; however, they did not remain independent predictors in multivariate models. CONCLUSION Pretreatment sarcopenia is not an independent predictor of treatment response or survival in patients with RC undergoing nCRT. Conversely, tumor-related parameters - including tumor length, MRF involvement, and clinical staging - have greater prognostic value and may assist in pretreatment risk stratification and in the individualized management of RC.
The identification of cancer prognostic biomarkers is crucial for predicting disease progression, optimizing personalized therapies, and improving patient survival. Molecular biomarkers are increasingly being identified for cancer prognosis estimation. However, existing studies and databases often focus on single-type molecular biomarkers, deficient in comprehensive multi-omics data integration, which constrains the comprehensive exploration of biomarkers and underlying mechanisms. To fill this gap, we conducted a systematic prognostic analysis using over 10,000 samples across 33 cancer types from The Cancer Genome Atlas (TCGA). Our study integrated nine types of molecular biomarker-related data: single-nucleotide polymorphism (SNP), copy number variation (CNV), alternative splicing (AS), alternative polyadenylation (APA), coding gene expression, DNA methylation, lncRNA expression, miRNA expression, and protein expression. Using log-rank tests, univariate Cox regression (uni-Cox), and multivariate Cox regression (multi-Cox), we evaluated potential biomarkers associated with four clinical outcome endpoints: overall survival (OS), disease-specific survival (DSS), disease-free interval (DFI), and progression-free interval (PFI). As a result, we identified 4,498,523 molecular biomarkers significantly associated with cancer prognosis. Finally, we developed SurvDB, an interactive online database for data retrieval, visualization, and download, providing a comprehensive resource for biomarker discovery and precision oncology research.
The pathogenesis of Parkinson's disease (PD) was recently hypothesized to change along with the disease course. Given the fact that transcriptional changes in blood can provide insightful clues for PD pathogenesis, we performed case-control and longitudinal whole blood transcriptome analyses to identify the signature genes underlying the hypothesized dynamic pathogenesis of PD. In the case-control study, we compared the gene expression patterns in healthy control (N = 189), prodromal (N = 58) and de novo idiopathic PD subjects (N = 390). The results showed that the prodromal subjects were at the tipping-point stage, which is characterized by the abnormal expression patterns of 414 genes associated with oxygen transport and reactive oxygen species metabolic process. We next performed a longitudinal transcriptome analysis on 255 PD patients from the baseline to the third year, and identified 203 genes related to immune and inflammatory responses during disease progression. These findings not just offer deeper insights into the dynamic pathogenesis of PD, but also help to find potential drugs to prevent the early neurodegeneration process.
Single nucleotide polymorphisms (SNPs) within microRNAs (miRNAs) and their target binding sites can influence miRNA biogenesis and target regulation, thereby participating in a variety of diseases and biological processes. Current miRNA-related SNP databases are often species-limited or based on outdated data. Therefore, we updated our miRNASNP database to version 4 by updating data, expanding the species from Homo sapiens to 17 species, and introducing several new features. In miRNASNP-v4, 82 580 SNPs in miRNAs and 24 836 179 SNPs in 3'UTRs of genes across 17 species were identified and their potential effects on miRNA secondary structure and target binding were characterized. In addition, compared to the last release, miRNASNP-v4 includes the following improvements: (i) gene enrichment analysis for gained or lost miRNA target genes; (ii) identification of miRNA-related SNPs associated with drug response and immune infiltration in human cancers; (iii) inclusion of experimentally supported immune-related miRNAs and (iv) online prediction tools for 17 animal species. With the extensive data and user-friendly web interface, miRNASNP-v4 will serve as an invaluable resource for functional studies of SNPs and miRNAs in multiple species. The database is freely accessible at http://gong_lab.hzau.edu.cn/miRNASNP/.
The prevalence of rheumatoid arthritis (RA) subtypes, including seropositive and seronegative, is influenced by lifestyle factors and exhibits high heterogeneity, resulting in reduced drug efficacy. This study aims to identify cytokines mediating the effects of different lifestyles on RA subtypes and to discover new drugs for personalized treatment. Mendelian randomization revealed that three cytokines (MIP1b, SCGFb, and TRAIL) partially mediated the effects of different lifestyles on RA overall or its subtypes. The pretrained model, i.e., DrugBAN, predicted the probability of 723,000 small molecule drugs binding to these three targets. In molecules with high binding rates, we calculated the structural similarity between known drugs for RA and other drugs to screen for new drugs, followed by molecular docking and molecular dynamics simulations for validation. The results indicate that these targets had promising binding affinity with known drugs and other drugs with high similarity. Our findings may guide therapeutic approaches for heterogeneous RA patients with specific lifestyle habits.
Alternative promoter (AP) events, as a major pre-transcriptional mechanism, can initiate different transcription start sites to generate distinct mRNA isoforms and regulate their expression. At present, hundreds of thousands of APs have been identified across human tissues, and a considerable number of APs have been demonstrated to be associated with complex traits and diseases. Recent researches have also proven important effects of APs on animals. However, the landscape of APs in animals has not been fully recognized. In this study, 102,349 AP profiles from 23,077 samples across 12 species were systematically characterized. We further identified tissue-specific APs and investigated trait-related promoters among various species. In addition, we analyzed the associations between APs and enhancer RNAs (eRNA)/transcription factors (TF) as a means of identifying potential regulatory factors. Integrating these findings, we finally developed Animal-APdb, a database for the searching, browsing, and downloading of information related to Animal APs. Animal-APdb is expected to serve as a valuable resource for exploring the functions and mechanisms of APs in animals.
Predicting individual patient responses to anticancer drugs is a central challenge in precision oncology, hindered by the scarcity of clinical pharmacogenomic data and substantial biological dissimilarity between preclinical models and patient tumors. Patient-derived xenograft (PDX) models offer significantly enhanced tumor biological fidelity compared to in vitro cancer cell line models, yet computational methods to translate PDX-based drug response predictions (DRP) into clinical settings remain limited. We developed TRANSPIRE-DRP, a deep learning framework that bridges the translational gap between PDX models and patient tumors through unsupervised domain adaptation. The framework employs a two-stage architecture: first, an autoencoder-based pretraining phase learns domain-invariant genomic representations from large-scale unlabeled data; second, an adversarial adaptation phase aligns these representations while preserving drug response signals from PDX models. We evaluated TRANSPIRE-DRP across three therapeutic agents—Cetuximab, Paclitaxel, and Gemcitabine—in real-life clinical prediction scenarios. TRANSPIRE-DRP consistently outperformed both cell line-based state-of-the-art models and PDX-based baselines, demonstrating superior translational capacity. Notably, the learned representations preserved tumor-specific molecular features and spontaneously recapitulated established drug-cancer type associations without requiring explicit histological annotations. Interpretability analyses revealed biologically coherent pathway enrichments consistent with known drug mechanisms of action, including EGFR-Wnt signaling crosstalk for Cetuximab, mitotic arrest mechanism for Paclitaxel, and NF-κB-mediated immunomodulation for Gemcitabine. TRANSPIRE-DRP establishes a scalable, interpretable, and clinically relevant framework for translating preclinical PDX data into personalized therapeutic predictions, providing a robust computational foundation for advancing precision oncology beyond the inherent limitations of traditional in vitro systems.
Tumorigenesis arises from the dysfunction of cancer genes, leading to uncontrolled cell proliferation through various mechanisms. Establishing a complete cancer gene catalogue will make precision oncology possible. Although existing methods based on graph neural networks (GNN) are effective in identifying cancer genes, they fall short in effectively integrating data from multiple views and interpreting predictive outcomes. To address these shortcomings, an interpretable representation learning framework IMVRL-GCN is proposed to capture both shared and specific representations from multiview data, offering significant insights into the identification of cancer genes. Experimental results demonstrate that IMVRL-GCN outperforms state-of-the-art cancer gene identification methods and several baselines. Furthermore, IMVRL-GCN is employed to identify a total of 74 high-confidence novel cancer genes, and multiview data analysis highlights the pivotal roles of shared, mutation-specific, and structure-specific representations in discriminating distinctive cancer genes. Exploration of the mechanisms behind their discriminative capabilities suggests that shared representations are strongly associated with gene functions, while mutation-specific and structure-specific representations are linked to mutagenic propensity and functional synergy, respectively. Finally, our in-depth analyses of these candidates suggest potential insights for individualized treatments: afatinib could counteract many mutation-driven risks, and targeting interactions with cancer gene SRC is a reasonable strategy to mitigate interaction-induced risks for NR3C1, RXRA, HNF4A, and SP1.
BACKGROUND:The colon cancer prognosis is influenced by multiple factors, including clinical, pathological, and non-biological factors. However, only a few studies have focused on computed tomography (CT) imaging features. Therefore, this study aims to predict the prognosis of patients with colon cancer by combining CT imaging features with clinical and pathological characteristics, and establishes a nomogram to provide critical guidance for the individualized treatment. AIM:To establish and validate a nomogram to predict the overall survival (OS) of patients with colon cancer. METHODS:A retrospective analysis was conducted on the survival data of 249 patients with colon cancer confirmed by surgical pathology between January 2017 and December 2021. The patients were randomly divided into training and testing groups at a 1:1 ratio. Univariate and multivariate logistic regression analyses were performed to identify the independent risk factors associated with OS, and a nomogram model was constructed for the training group. Survival curves were calculated using the Kaplan-Meier method. The concordance index (C-index) and calibration curve were used to evaluate the nomogram model in the training and testing groups. RESULTS:Multivariate logistic regression analysis revealed that lymph node metastasis on CT, perineural invasion, and tumor classification were independent prognostic factors. A nomogram incorporating these variables was constructed, and the C-index of the training and testing groups was 0.804 and 0.692, respectively. The calibration curves demonstrated good consistency between the actual values and predicted probabilities of OS. CONCLUSION:A nomogram combining CT imaging characteristics and clinicopathological factors exhibited good discrimination and reliability. It can aid clinicians in risk stratification and postoperative monitoring and provide important guidance for the individualized treatment of patients with colon cancer.
Wound infection is a serious complication in burn injury, which is a common form of trauma and an important public health issue. We investigated samples from burn and nonburn wounds for microbial characteristics and temporal trends of antibiotic resistance. Wound samples were collected from 369 burned patients and 927 non-burned individuals admitted from 2007 to 2017. Higher frequency of Acinetobacter baumannii, Klebsiella pneumonia, and Pseudomonas aeruginosa was observed in samples from burned individuals when compared to those from non-burned. The prevalence of different groups of bacteria varied when the samples were stratified according to age and sex. The antimicrobial resistance profiles showed a significant difference between burned and non-burned patients. The different temporal trends of antimicrobial resistance rates were also found, which may be critical for the proper selection of antibiotics in burn treatment. The present study suggested that frequent pathogens and antibacterial resistance evolution could differ between burn wounds and other wounds. Therefore, periodic surveillance of antibiotic resistance patterns in the burn unit might help physicians properly select antibiotics for treatment.
Single-nucleotide polymorphisms (SNPs) as the most important type of genetic variation are widely used in describing population characteristics and play vital roles in animal genetics and breeding. Large amounts of population genetic variation resources and tools have been developed in human, which provided solid support for human genetic studies. However, compared with human, the development of animal genetic variation databases was relatively slow, which limits the genetic researches in these animals. To fill this gap, we systematically identified ∼ 499 million high-quality SNPs from 4784 samples of 20 types of animals. On that basis, we annotated the functions of SNPs, constructed high-density reference panels and calculated genome-wide linkage disequilibrium (LD) matrixes. We further developed Animal-SNPAtlas, a user-friendly database (http://gong_lab.hzau.edu.cn/Animal_SNPAtlas/) which includes high-quality SNP datasets and several support tools for multiple animals. In Animal-SNPAtlas, users can search the functional annotation of SNPs, perform online genotype imputation, explore and visualize LD information, browse variant information using the genome browser and download SNP datasets for each species. With the massive SNP datasets and useful tools, Animal-SNPAtlas will be an important fundamental resource for the animal genomics, genetics and breeding community.
Learning representations from data is a fundamental step for machine learning. High-quality and robust drug representations can broaden the understanding of pharmacology, and improve the modeling of multiple drug-related prediction tasks, which further facilitates drug development. Although there are a number of models developed for drug representation learning from various data sources, few researches extract drug representations from gene expression profiles. Since gene expression profiles of drug-treated cells are widely used in clinical diagnosis and therapy, it is believed that leveraging them to eliminate cell specificity can promote drug representation learning. In this paper, we propose a three-stage deep learning method for drug representation learning, named DRLM, which integrates gene expression profiles of drug-related cells and the therapeutic use information of drugs. Firstly, we construct a stacked autoencoder to learn low-dimensional compact drug representations. Secondly, we utilize an iterative clustering module to reduce the negative effects of cell specificity and noise in gene expression profiles on the low-dimensional drug representations. Thirdly, a therapeutic use discriminator is designed to incorporate therapeutic use information into the drug representations. The visualization analysis of drug representations demonstrates DRLM can reduce cell specificity and integrate therapeutic use information effectively. Extensive experiments on three types of prediction tasks are conducted based on different drug representations, and they show that the drug representations learned by DRLM outperform other representations in terms of most metrics. The ablation analysis also demonstrates DRLM's effectiveness of merging the gene expression profiles with the therapeutic use information.
MicroRNAs (miRNAs) are crucial regulators in various diseases. The identification of associations between miRNAs and diseases could greatly facilitate the investigation of disease mechanisms and drug development. Limited by time and cost efficiency, conventional experimental techniques are inadequate for this purpose. With the extensive advance and application of deep learning, developing efficient and accurate computational models for predicting miRNA‒disease associations has a vital role and is feasible. In this study, we proposed a meta-path-aggregated multilevel graph embedding model for miRNA‒ disease association prediction. The model first calculated the multiple similarities among miRNAs and diseases, respectively. Then, the node features were extracted from similarity matrices for miRNAs and diseases. Furthermore, we integrated four types of meta-paths from the miRNA‒lncRNA‒disease heterogeneous graph and learned node embeddings by hierarchical graph attention modules. Finally, the model predicted the miRNA‒ disease associations using two-layer graph convolution networks (GCNs). Compared with six state-of-the-art models, the experimental results demonstrated that our model achieved higher prediction performance with an AUC of 0.9892 and an AUPR of 0.9898 for the 5-fold cross-validation on the HMDDv3.2 dataset. With the case study, the model’s performance was further validated, and the top 20 predicted associations could be experimentally confirmed. All in all, it implies the predictive power of our model 1 and the potential value in understanding disease pathology.
Alternative polyadenylation (APA) is an important post-transcription regulatory mechanism widely occurring in eukaryotes and has been associated with special traits/diseases by several studies. However, the dynamic roles and patterns of APA in cell differentiation remain largely unknown. Here, we systematically characterized the APA profiles during the differentiation of induced pluripotent stem cells (iPSCs) to cardiomyocytes by the previously published RNA-seq data across 16 time points. We identified 950 differential APA events and found five dynamic APA patterns with fuzzy c-means clustering analysis. Among them, 3′UTR progressive lengthening is the main APA pattern over time, the genes of which are enriched in cell cycle and mRNA metabolic process pathways. By constructing the linear mixed-effects model, we also indicated that TIA1 plays an important role in regulating APA events with this pattern, including genes essential to cardiac function. Additionally, APA and polyA machinery activity with another pattern can immediately respond to developmental signal-mediated stimuli at the early differentiation stage and result in a sharp shortening of the 3′UTR. Finally, a miRNA-APA network is constructed and several hub miRNAs potentially regulating cardiomyocyte differentiation are detected. Our results show the complex APA mechanisms during the differentiation of iPSCs into cardiomyocytes and provide further insights for the understanding of APA regulation and cell differentiation.
Multi-nucleotide variants (MNVs) are defined as clusters of two or more nearby variants existing on the same haplotype in an individual. Recent studies have identified millions of MNVs in human populations, but their functions remain largely unknown. Numerous studies have demonstrated that single-nucleotide variants could serve as quantitative trait loci (QTLs) by affecting molecular phenotypes. Therefore, we propose that MNVs can also affect molecular phenotypes by influencing regulatory elements. Using the genotype data from The Cancer Genome Atlas (TCGA), we first identified 223 759 unique MNVs in 33 cancer types. Then, to decipher the functions of these MNVs, we investigated the associations between MNVs and six molecular phenotypes, including coding gene expression, miRNA expression, lncRNA expression, alternative splicing, DNA methylation and alternative polyadenylation. As a result, we identified 1 397 821 cis- MNVQTLs and 402 381 trans- MNVQTLs. We further performed survival analysis and identified 46 173 MNVQTLs associated with patient overall survival. We also linked the MNVQTLs to genome-wide association studies (GWAS) data and identified 119 762 MNVQTLs that overlap with existing GWAS loci. Finally, we developed Pancan-MNVQTLdb (http:// gong lab. hzau.edu.cn/mnvQTLdb/) for data retrieval and download. Pancan-MNVQTLdb will help decipher the functions of MNVs in different cancer types and be an important resource for genetic and cancer research.