The accumulation of DNA damage on the cancer genome, often summarized as mutational profiles, is influenced by a complex interplay of environmental and cellular mutagenic conditions. Despite cancer genomics having some of the largest biological datasets in existence collected by pan-cancer consortiums, a significant portion of the diverse spectrum of cancer etiologies is still underrepresented and understudied. We leveraged the power of generative artificial intelligence (Gen AI) to develop MutaGen, a conditional variational autoencoder and generative adversarial network hybrid model (CVAE-GAN), to generate realistic mutational profiles conditioned on various clinical and genetic features. Trained on the PCAWG cohort, MutaGen can generate pan-cancer mutational profiles across four mutation types: single base substitutions (SBS96), small insertions and deletions (ID83), structural variations (SV32), and copy number alterations (CN48). MutaGen's cancer-conditioned generations show high similarities to the real PCAWG samples, with median cosine similarities of 0.98, 0.97, 0.92, and 0.96 for SBS96, ID83, SV32, and CN48 respectively. When conditioned on homologous recombination deficiency (HRD) status, MutaGen can generate zero-shot (i.e. cases not seen in training set) HRD lung cancers with high similarity to a real BRCA2-null lung cancer patient in both SBS96 and ID83 (0.93 and 0.78 cosine similarities respectively). The generation is significantly more similar than non-HRD lung cancer or non-lung HRD cancer in the PCAWG training set. Inspection of the generated profiles revealed the generations harbouring both lung cancer-associated and HRD-associated mutational patterns, showing that MutaGen can generate novel and biologically plausible profiles. By providing a versatile tool for generating synthetic mutational data, MutaGen can facilitate the exploration of understudied cancer subtypes and enable the training and benchmarking of predictive models. Andy Wu, Hannan Wong, Aakansha Narain, Vedant Sandhu, Jason Pitt. Conditional generation of mutational profiles in cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2491.
Supplementary methods. Supplementary Figure 1. Phenotyping of B-cells in non-malignant tissues. A, Quantitation of marker positivity across ten tonsil and two reactive lymph node samples (rLN). Analysis is spatially resolved between the GC and extra-GC zones. B, Spatial map of cellular coordinates based on cell segmentation of images in Figure 1B. Marker-positivity is indicated, and a total proportion of positive and negative cells is depicted as a pie chart. These maps were used to derive sub-population phenotypes depicted in Figure 1C. Scale bar is 100μm. C, Proliferation analysis (i.e., Ki67-positivity) among sub-populations in five tonsil samples. Median with interquartile range, whiskers denote 10th and 90th percentile. Supplementary figure 2. Example pseudo-colored mfIHC images for MYC, BCL2, BCL6 cases in DLBCL. Images of a range of mean fluorescent intensities are shown with equal scaling for reference. Supplementary figure 3. Global distribution of MYC, BCL2 and BCL6 sub-populations within DLBCL cohorts. Heat-maps displaying the percentage extent of individual markers and each sub-population within the DLBCL NUH, CMMC, SGH and MDA cohorts. Hierarchical k-means clustering of patients according to sub-population extent is applied. Positivity shading for single markers ranges between 0-100% positivity, whereas shading for sub-populations reflects 0-50% positivity and remains fully saturated until 100%. IPI Risk Group - International Prognostic Index Risk Group, FISH - fluorescence in situ hybridization. Supplementary figure 4. Intra-tumor heterogeneity of sub-populations. A, Correlation of sub-population extent quantification between two biopsies of the same patient for which at least two tissue microarray (TMA) biopsies are available. Correlation is shown separately for lymph node and extranodal biopsies. Spearman rho is indicated for each correlation. Axes are in exponential and equivalent in all panels. B, Sub-population percentage extent quantification across multiple TMA cores (columns) of the same patient (rows). Pie charts are ordered according to decreasing cell numbers evaluated per core. All patients from the NUH cohort with at least five cores are evaluated. A heterogenous cluster is highlighted by the red box. Supplementary figure 5. Spatial heterogeneity of sub-population interactions. A, Conceptual schematic of pair correlation function (PCF) plots depicting a clustered distribution (left, green) and a random distribution (right, grey). Representative counterpart spatial maps are above each plot. B, PCF analysis for sub-populations to investigate spatial clustering (top). Mean results for two independent cohorts (shading is cohort standard deviation). An example tissue microarray core is shown as physical distance reference for spatial analyses (bottom left). Absolute number of neighboring cells expected within a given radius (data from 3500 randomly selected cells across all images, mean with standard deviation) (bottom right). C, Actual spatial map of sub-populations of an example DLBCL case (top). Extent of all sub-populations within the sample is shown on the left. Simulated, hypothetical random distribution of cells for the same case (middle). PCF analysis for the shown sample and its matched simulated random distribution (bottom). Scale bars in B and C are 100µm. D, Mean deviations from expected neighbor abundance (Δ%) summarizing cell-cell interactions between sub-populations for the sample shown in (C). E, Sub-population interaction matrices from spatially distinct biopsies (cores in tissue microarray) for example DLBCL patients. Biopsies of stable, spatially homogenous, sub-population interaction profiles are grouped (top), whereas biopsies of a differing, heterogenous, interaction profile are grouped separately (bottom). Supplementary figure 6. Global deviations from expected spatial neighbor abundance (Δ%). Hierarchical clustering (minimum variance method) of measured Δ% for all cases in the SGH and MDA cohorts. Extents of sub-populations are indicated for reference (top). For the MDA cohort, multiple biopsies (n = 1-3) from the same patient were included in the analysis to determine spatial interaction similarity across spatially distinct regions (bottom). Supplementary figure 7. Correlation of predicted MYC, BCL2 and BCL6 sub-population percentage extent based on single oncogene positivity and observed percentage extent in DLBCL cohorts. Spearman rho, axes are equivalent in all panels. Supplementary figure 8. Variance of M+2+6- percentage extent in the context of positivity calling across a 15% cut-off. A, M+2+6- scoring variance across multiple pathological imaging fields. All whole-tissue DLBCL sections from University of Palermo (UP), and samples from the NUH TMA with at least four fields scored per patient and a mean M+2+6- score above 5% are shown. Mean with SD. Ordinates between 50-100% are compressed for clarity. Dashed line denotes M+2+6- 15% positivity. B, Stability of M+2+6- case positivity calling across scoring increasing number of imaging fields. All cases from panel A with at least five fields scored in this study are shown. Only one case is called M+2+6- Low (<15%) at the first image scored, and subsequently called M+2+6- High (≥15%) after two or more fields scored. Supplementary figure 9. Mapping of mRNA expression data into percentage extent data. A, Cumulative histogram of MYC, BCL2 and BCL6 protein percentage extent positivity in DLBCL cohorts (data transformed from Figure 4A) (top). B, Aggregated single oncogene cumulative distribution of MYC, BCL2 and BCL6 protein percentage extent positivity across all measured protein cohorts and its smoothed empirical cumulative distribution function (eCDF).C, Distribution of inferred single oncogene percentage extent in GEP cohorts. (see Supplementary table 6 for all values). Supplementary figure 10. Analysis of the GOYA clinical trial. A, Correlation of MYC mRNA with quantitative IHC score. Linear regression (left) and Wilcoxon rank sum test (right). B, Analysis as in (A) for BCL2. C, Kaplan-Meier curves for PFS and OS for patients stratified across the 15% M+2+6- metric (GEP-derived). Multivariate Cox proportional hazards model is available in Supplementary table 9. PFS - progression free survival, OS - overall survival. Supplementary figure 11. Proliferative advantage of cyclin D2 (CCND2) overexpressing B-cells. Representative FACS plots documenting to the expansion over time of the cyclin D2 positive GC B-cell population in cyclin D2 overexpressing GC B-cells (CCND2-Lyt2) and non-cyclin D2 overexpressing GC B-cells (Empty vector Lyt2). All GC B-cells co-overexpress BCL2, BCL6, MYC and GFP. Supplementary table 3. Non-parametric correlation of sub-population percentage extent with clinicopathological features. Supplementary table 4. Pooled univariate analysis for MYC, BCL2 and BCL6 single oncogene and sub-populations percentage extents as a continuous variable at 5% increments as predictors for overall survival (OS) in mfIHC cohorts of DLBCL (Cox proportional hazards model). Supplementary table 5. Univariate analysis of clinicopathological features as a predictor of overall survival (OS) after first-line R-CHOP treatment in the NUH, SGH and MDA cohorts of DLBCL (Cox proportional hazards model). Supplementary table 7. Pooled univariate analysis for sub-population metrics as a continuous variable at 5% increments as predictors for overall survival (OS) in GEP DLBCL cohorts (Cox proportional hazards model). Supplementary table 8. Multivariate analysis of continuous M+2+6- metric at 5% increments as a predictor of overall survival (OS) in cohorts with gene-expression data (Cox proportional hazards model). Supplementary table 9. Univariate and multivariate analysis of continuous M+2+6- metric as a continuous variable at 5% increments as predictor of progression-free survival (PFS) and overall survival (OS) in the GOYA trial cohort (Cox proportional hazards model). Supplementary table 10. Multivariate analysis of M+2+6- metric dichotomized at 15% as a predictor of overall survival (OS) in cohorts with gene-expression data (Cox proportional hazards model). Supplementary table 15. Clinicopathologic characteristics of DLBCL patients evaluated by multiplexed fluorescent immunohistochemistry (mfIHC) in this study. Supplementary table 16. Manual multiplexed fluorescent immunohistochemistry (mfIHC) staining protocol performed on the NUH and CMMC cohort TMA. Supplementary table 17. Automated multiplexed fluorescent immunohistochemistry (mfIHC) staining protocol performed on the SGH, MDA and BCA cohort TMA.
Next-generation sequencing (NGS) is increasingly utilized in oncological practice; however, only a minority of patients benefit from targeted therapy. Developing drug response prediction (DRP) models is important for the "untargetable" majority. Prior DRP models typically use whole-transcriptome and whole-exome sequencing data, which are clinically unavailable. We aim to develop a DRP model toward the repurposing of chemotherapy, requiring only information from clinical-grade NGS (cNGS) panels of restricted gene sets. Data sparsity and limited patient drug response information make this challenging. We firstly show that existing DRPs perform equally with whole-exome versus cNGS (∼300 genes) data. Drug IDentifier (DruID) is then described, a DRP model for restricted gene sets using transfer learning, variant annotations, domain-invariant representation learning, and multi-task learning. DruID outperformed state-of-the-art DRP methods on pan-cancer data and showed robust response classification on two real-world clinical datasets, representing a step toward a clinically applicable DRP tool.
The metazoan lifespan is determined in part by a complex signaling network that regulates energy metabolism and stress responses. Key signaling hubs in this network include insulin/IGF-1, AMPK, mTOR, and sirtuins. The Hippo/Mammalian Ste20-like Kinase1 (MST1) pathway has been reported to maintain lifespan in Caenorhabditis elegans, but its role has not been studied in higher metazoans. In this study, we report that overexpression of Hpo, the MST1 homolog in Drosophila melanogaster, decreased lifespan with concomitant changes in lipid metabolism and aging-associated gene expression, while RNAi Hpo depletion increased lifespan. These effects were mediated primarily by Hpo-induced transcriptional activation of the RNA-binding protein maternal expression at 31B (Me31b)/RCK, resulting in stabilization of mRNA-encoding a lipolytic hormone, Akh. In mouse adipocytes, Hpo/Mst1 mediated adipocyte differentiation, phosphorylation of RNA-binding proteins such as Rck, decapping MRNA 2 (Dcp2), enhancer Of MRNA decapping 3 (Edc3), nucleolin (NCL), and glucagon mRNA stability by interacting with Rck. Decreased lifespan in Hpo-overexpressing Drosophila lines required expression of Me31b, but not DCP2, which was potentially mediated by recovering expression of lipid metabolic genes and formation of lipid droplets. Taken together, our findings suggest that Hpo/Mst1 plays a conserved role in longevity by regulating adipogenesis and fatty acid metabolism.
Supplementary table 1. Per-patient mfIHC MYC, BCL2 and BCL6 single oncogene and subpopulation scores for normal tonsil tissue and reactive lymph node tissue. Supplementary table 2. Per-patient mfIHC MYC, BCL2 and BCL6 single oncogene and subpopulation scores for DLBCL tissue (NUH, CMMC, SGH, MDA, BCA and UP). Supplementary table 6. Inferred percentage extents of MYC, BCL2, BCL6 and sub-population metrics in GEP cohorts. Supplementary table 11. Correlation of M+2+6- metric with gene expression in GEP cohorts. Supplementary table 12. Differential gene expression analysis of primary germinal center (GC) B-cells with M+2+ and M+2+6+ overexpression. Supplementary table 13. Differentially expressed genes between M+2+6- and all other malignant cells in scRNA-seq samples of DLBCL. Dichotomized non-parametric comparison, Wilcoxon rank sum test. Supplementary table 14. Analysis of positive enrichment of Wikipathways terms between M+2+6- and all other malignant cells in scRNA-seq samples of DLBCL by gprofiler2.
Cancer remains a global challenge due to its growing clinical and economic burden. Its uniquely personal manifestation, which makes treatment difficult, has fuelled the quest for personalized treatment strategies. Thus, genomic profiling is increasingly becoming part of clinical diagnostic panels. Effective use of such panels requires accurate drug response prediction (DRP) models, which are challenging to build due to limited labelled patient data. Previous methods to address this problem have used various forms of transfer learning. However, they do not explicitly model the variable length sequential structure of the list of mutations in such diagnostic panels. Further, they do not utilize auxiliary information (like patient survival) for model training. We address these limitations through a novel transformer-based method, which surpasses the performance of state-of-the-art DRP models on benchmark data. Code for our method is available at https://github.com/CDAL-SOC/PREDICT-AI. We also present the design of a treatment recommendation system (TRS), which is currently deployed at the National University Hospital, Singapore and is being evaluated in a clinical trial. We discuss why the recommended drugs and their predicted scores alone, obtained from DRP models, are insufficient for treatment planning. Treatment planning for complex cancer cases, in the face of limited clinical validation, requires assessment of many other factors, including several indirect sources of evidence on drug efficacy. We discuss key lessons learnt on model validation and use of indirect supporting evidence to build clinicians' trust and aid their decision making.
Knudson’s “two-hit” paradigm posits that carcinogenesis requires inactivation of both copies of an autosomal tumor suppressor gene. Here, we report that the glycolytic metabolite methylglyoxal (MGO) transiently bypasses Knudson’s paradigm by inactivating the breast cancer suppressor protein BRCA2 to elicit a cancer-associated, mutational single-base substitution (SBS) signature in nonmalignant mammary cells or patient-derived organoids. Germline monoallelic BRCA2 mutations predispose to these changes. An analogous SBS signature, again without biallelic BRCA2 inactivation, accompanies MGO accumulation and DNA damage in Kras-driven, Brca2-mutant murine pancreatic cancers and human breast cancers. MGO triggers BRCA2 proteolysis, temporarily disabling BRCA2’s tumor suppressive functions in DNA repair and replication, causing functional haploinsufficiency. Intermittent MGO exposure incites episodic SBS mutations without permanent BRCA2 inactivation. Thus, a metabolic mechanism wherein MGO-induced BRCA2 haploinsufficiency transiently bypasses Knudson’s two-hit requirement could link glycolysis activation by oncogenes, metabolic disorders, or dietary challenges to mutational signatures implicated in cancer evolution.
Single-base substitution (SBS) mutational signatures have become standard practice in cancer genomics. In lieu of de novo signature extraction, reference signature assignment allows users to estimate the activities of pre-established SBS signatures within individual malignancies. Several tools have been developed for this purpose, each with differing methodologies. However, due to a lack of standardization, there may be inter-tool variability in signature assignment. We deeply characterized three assignment strategies and five SBS signature assignment tools. We observed that assignment strategy choice can significantly influence results and interpretations. Despite varying recommendations by tools, Refit performed best by reducing overfitting and maximizing reconstruction of the original mutational spectra. Even after uniform application of Refit, tools varied remarkably in signature assignments both qualitatively (Jaccard index = 0.38-0.83) and quantitatively (Kendall tau-b = 0.18-0.76). This phenomenon was exacerbated for 'flat' signatures such as the homologous recombination deficiency signature SBS3. An ensemble approach (EnsembleFit), which leverages output from all five tools, increased SBS3 assignment accuracy in BRCA1/2-deficient breast carcinomas. After generating synthetic mutational profiles for thousands of pan-cancer tumors, EnsembleFit reduced signature activity assignment error 15.9-24.7% on average using Catalogue of Somatic Mutations In Cancer and non-standard reference signature sets. We have also released the EnsembleFit web portal (https://www.ensemblefit.pittlabgenomics.com) for users to generate or download ensemble-based SBS signature assignments using any strategy and combination of tools. Overall, we show that signature assignment heterogeneity across tools and strategies is non-negligible and propose a viable, ensemble solution.
Despite large collections of cancer genomics data being openly available, the inability to quickly interrogate this information remains a barrier for researchers and oncologists. Here we present Melvin, an Amazon Alexa skill to explore cancer genomics data through simple conversations. Melvin is a multi-modal Amazon Alexa skill that allows users to quickly explore cancer genomics data from TCGA through simple conversations.
<p>Supplementary Figure S1 PDF file - 1274K, Schematic presentation of the workflow of experiments and analytical procedures</p>
Supplementary Figure S3 PDF file - 2243K, RNAi analysis and measurement of cell viability of amplified and overexpressed chromosome 13 candidate genes
<p>Supplementary Figure S9 PDF file - 1478K, Expression levels of the LNX2 module</p>
Supplementary Table S4 PDF file - 358K, List of filtered genes used in our experimental setup
Evaluation of M+2+6− cells in scRNA-seq datasets of DLBCL. A, Uniform Manifold Approximation and Projection (UMAP) of malignant B cells from the Roider and Steen cohorts. B, Proportion of subpopulations across samples and annotation of M+2+6− cells in UMAP. C, Correlation of genes enriched in the M+2+6− subpopulation as evaluated by scRNA-seq with hits from the bulk GEP cohorts (Fig. 5E). CCND2 is highlighted and is among the concordant hits (see also Supplementary Table S13). D, WikiPathways terms enrichment among genes positively associated with M+2+6− cells. Both axes in C and D are scaled exponentially for clarity (see also Supplementary Table S14).
Supplementary Figure S2 PDF file - 1800K, Corroboration of gene expression levels by qRT-PCR
Transcriptomic analysis of M+2+6− high cases and potential role of CCND2. A, Correlation of observed M+2+6− percentage extent in the BCA cohort with the cell-of-origin (COO) DLBCL90-COO signature. Bonferroni corrected Kruskal–Wallis test for ABC vs. others. B, Correlation of M+2+6− metric in GEP cohorts with COO signatures. Mean M+2+6− metric value per group per cohort is shown. Bonferroni corrected paired-samples t test. C, Correlation of the M+2+6− percentage extent and metric evaluated by mfIHC and mRNA inference, respectively, with genetic subtypes (LymphGen classification). D, Sankey plot of M+2+6− dichotomized grouping matched with molecular features. E, Volcano plot of pooled direct correlation of gene mRNA expression and M+2+6− metric across seven GEP cohorts. Genes highly correlated with M+2+6− metric across datasets at absolute Spearman rho ≥0.2 and FDR≤0.001 are shown (see also Supplementary Table S11). The abscissa is scaled exponentially. F, Differential gene expression between primary GC B cells overexpressing M+2+ and M+2+6+ (see also Supplementary Table S12). Analysis is generated from 4 biological replicates from each condition, from cells of independent donors. G, Genes highly enriched in M+2+6− cells: correlation of results from E and F. H,CCDN2 gene expression in GEP cohorts in patients dichotomized by M+2+6− 15% metric (left) and in primary B cells (right). Paired t test (left); mean with standard deviation and FDR (FDR as per Supplementary Table S12) for t test (right). I, Single-cell RNA-seq of GC primary B cells transduced either with BCL2 and MYC (MYC-transduced) or BCL2 and BCL6 (BCL6-transduced). Untransduced GC primary B cells are also included. Expression of CCND2 is indicated in color. J, Proliferation analysis of M+2+6+ primary GC B cells overexpressing cyclin D2 (CCND2) compared with M+2+6+ primary GC B cells transduced with an empty vector (EV). Analysis performed with 3 biological replicates for each condition, using cells from 3 independent patients; mean with standard deviation; t test. UMAP, Uniform Manifold Approximation and Projection.
Abstract Cancers often overexpress multiple clinically relevant oncogenes, but it is not known if combinations of oncogenes in cellular subpopulations within a cancer influence clinical outcomes. Using quantitative multispectral imaging of the prognostically relevant oncogenes MYC, BCL2, and BCL6 in diffuse large B-cell lymphoma (DLBCL), we show that the percentage of cells with a unique combination MYC+BCL2+BCL6− (M+2+6−) consistently predicts survival across four independent cohorts (n = 449), an effect not observed with other combinations including M+2+6+. We show that the M+2+6− percentage can be mathematically derived from quantitative measurements of the individual oncogenes and correlates with survival in IHC (n = 316) and gene expression (n = 2,521) datasets. Comparative bulk/single-cell transcriptomic analyses of DLBCL samples and MYC/BCL2/BCL6-transformed primary B cells identify molecular features, including cyclin D2 and PI3K/AKT as candidate regulators of M+2+6− unfavorable biology. Similar analyses evaluating oncogenic combinations at single-cell resolution in other cancers may facilitate an understanding of cancer evolution and therapy resistance. Significance: Using single-cell–resolved multiplexed imaging, we show that selected subpopulations of cells expressing specific combinations of oncogenes influence clinical outcomes in lymphoma. We describe a probabilistic metric for the estimation of cellular oncogenic coexpression from IHC or bulk transcriptomes, with possible implications for prognostication and therapeutic target discovery in cancer. This article is highlighted in the In This Issue feature, p. 1027
Supplementary Figure S8 PDF file - 391K, Reduction of CTNNB1 after siRNA against LNX2