To identify cancer-associated gene regulatory changes, we generated single-cell chromatin accessibility landscapes across eight tumor types as part of The Cancer Genome Atlas. Tumor chromatin accessibility is strongly influenced by copy number alterations that can be used to identify subclones, yet underlying cis-regulatory landscapes retain cancer type–specific features. Using organ-matched healthy tissues, we identified the “nearest healthy” cell types in diverse cancers, demonstrating that the chromatin signature of basal-like–subtype breast cancer is most similar to secretory-type luminal epithelial cells. Neural network models trained to learn regulatory programs in cancer revealed enrichment of model-prioritized somatic noncoding mutations near cancer-associated genes, suggesting that dispersed, nonrecurrent, noncoding mutations in cancer are functional. Overall, these data and interpretable gene regulatory models for cancer and healthy tissue provide a framework for understanding cancer-specific gene regulation.
Benign prostatic hyperplasia (BPH) may decrease patient quality of life and often leads to acute urinary retention and surgical intervention. While effective treatments are available, many BPH patients do not respond or develop resistance to treatment. To understand molecular determinants of clinical symptom persistence after initiating BPH treatment, we investigated gene expression profiles before and after treatments in the prostate transitional zone of 108 participants in the Medical Therapy of Prostatic Symptoms (MTOPS) Trial. Unsupervised clustering revealed molecular subgroups characterized by expression changes in a large set of genes associated with resistance to finasteride, a 5α-reductase inhibitor. Pathway analyses within this gene cluster found finasteride administration induced changes in fatty acid metabolism, amino acid metabolism, immune response, steroid hormone metabolism, and kinase activity within the transitional zone. We found that patients without this transcriptional response were highly likely to develop clinical progression, which is expected in 13.2% of finasteride-treated patients. Importantly, a patient's transcriptional response to finasteride was associated with their pre-treatment kinase expression. Further, we identified novel expression signatures of finasteride resistance among the transcriptionally responded patients. These patients showed different gene expression profiles at baseline and increased prostate transitional zone volume compared to the patients who responded to the treatment. Our work suggests molecular mechanisms of clinical resistance to finasteride treatment that could be potentially helpful for personalized BPH treatment as well as new drug development to increase patient drug response.
Abstract Background: Fresh frozen (FF) and formalin-fixed paraffin-embedded (FFPE) samples are primary resources for archival tissues in cancer studies. Despite the advantages of easy storage and cost-effectiveness, FFPE samples have the disadvantage of inevitable chemical-induced RNA degradation. Although, in mRNA-seq data, the 3' bias due to the characteristics of the mRNA-seq platform allows the measure of RNA degradation levels, for total RNA-seq and FFPE samples, there is still no clear measure for RNA degradation. Methods: Our study, utilizing The Cancer Genome Atlas (TCGA) pilot data comprised of 26 paired samples of fresh frozen mRNA-seq [FFM], fresh frozen total RNA-seq [FFT], and FFPE total RNA-seq [PET], investigated RNA degradation patterns in FFPE samples. Upon investigating the read coverage of transcripts in FFPE data, we observed increased noise parameters in selected samples compared to others, suggesting a method for assessment of RNA degradation in FFPE. To measure these noise patterns, we developed a method called windowCV (wCV) utilizing coefficient of variance (CV) along the transcript length. Also, based on reports that membrane-surrounded RNA such as lncRNA and mitochondrial RNA (mtRNA) might be less affected by formalin fixation, we developed the simple measure dividing expression value of other proteins with certain lncRNA or mtRNA inferring the degree of RNA degradation. Results: We first compared the shape of transcript read coverage from 3 different RNA-seq data and observed 3’bias pattern in FFM which was not observed in FFT and PET. For FFT and PET, we applied wCV and executed hierarchical clustering. We observed that a subset of samples and genes were clustered by high wCV value suggesting notable RNA degradation. To support the hypothesis of RNA degradation, we utilized Trimmed Mean of M values (TMM) expression ratio of paired FFT and PET to determine which samples and genes were associated with underestimation of expression in PET. Interestingly, we observed that PET/FFT TMM ratio showed negative correlation with wCV. Additionally, when we investigated hierarchical clustering result of PET/FFT TMM ratio, we found that lncRNA and mtRNA were less affected by FFPE. We then calculated expression ratio between membrane-surrounded RNA and a set of protein coding genes largely affected by FFPE. As a result, we found that this ratio positively correlated with paired PET/FFT TMM ratio. We conclude that this simple measurement can provide information of FFPE induced underestimated expression with limited transcriptome data as well as RNA degradation. Conclusion: Our method can provide gene-level CV while avoiding false positive variance signal originating from transcript model and simple way to detect the degree of RNA degradation across FFPE samples. Citation Format: Wonyoung Choi, Miyeon Yeon, Hyo Young Choi, David Neil Hayes. Deciphering RNA degradation: Insights from a comparative analysis of paired fresh frozen/FFPE total RNA-seq [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 2323.
Background: RNA-seq is now the most widely used technique for gene expression profiling, which generates nucleotide level genome coverage as well as summary gene expression values. In general, low-expressed genes are excluded in the data processing due to their low signal-to-noise ratio. Nonetheless, it has been shown that low-expressed genes can provide crucial information such as presence of rare cells in bulk tissue samples. To optimize signal to noise in low-expressed genes, we applied a novel approach in which we transform low-expressed genes to a robust dichotomized state of being either “on” or “off”. Methods: To determine the status of genes, we use the shape of base-level read coverage which is expected to be homogeneous across the samples if they are “on”, whereas appear to be random noise if they are “off”. We model base-resolution RNA-seq data as vectors in high dimensional data space and measure their level of shape similarity (LSS) using angles between samples with lower angles indicating higher similarity which is more likely to be “on” status. Applying this approach to 3 human cancer samples (head and neck squamous cell carcinoma (HNSC), lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC)) from the Cancer Genome Atlas (TCGA), we identified lists of genes, the OFFONOME, that were either always or sometimes “off”. Differential OFFONOME genes were queried to address supervised and unsupervised analyses. Results: Using our technique, we characterized the OFFONOME of HNSC (5851 genes), a set which would typically have been filtering for removal because of low expression. In the HNSC OFFONOME, we observed five gene clusters, each of which was strongly associated with specific gene ontology. One of the clusters identified a rare population of normal myocytes infiltrating otherwise invasive tumors. For the OFFONOME of LUAD (5435 genes) and LUSC (5292 genes), we found clusters enriched with cilia and keratinization-related genes. In the result of integrated analysis, we observed that squamous cell carcinoma (SCC) tumor types shared “on” status for keratinization-related genes known for the cause of SCC. By comparison, LUAD had “on” status genes related to microtubule-based movement related with cilia structure. Strikingly, clustering the OFFONOME with the LSS distinguished 3 tumor types with almost perfect separation, outperforming several competing gene expression measures. Conclusion: In this study, we applied a new notion of gene expression. This approach enables a robust characterization of “on” and “off” status which can be especially effective for the genes expressed at low level. The OFFONOME from 3 cancer types revealed not only the tissue-specific genes but also the genes shared in similar tumor types. Collectively, OFFONOME can provide new insights into genes expressed at vanishingly low levels, such as from minor cell populations within bulk tumor analysis. Citation Format: Wonyoung Choi, Hyo Young Choi, Daivd Neil Hayes. OFFONOME: a new notion of genes’ on/off. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 4297.
Abstract Background: RNA-seq is now the most widely used technique for gene expression profiling, which generates nucleotide level genome coverage as well as summary gene expression values. In general, low-expressed genes are excluded in the data processing due to their low signal-to-noise ratio. Nonetheless, it has been shown that low-expressed genes can provide crucial information such as presence of rare cells in bulk tissue samples. To optimize signal to noise in low-expressed genes, we applied a novel approach in which we transform low-expressed genes to a robust dichotomized state of being either “on” or “off”. Methods: To determine the status of genes, we use the shape of base-level read coverage which is expected to be homogeneous across the samples if they are “on”, whereas appear to be random noise if they are “off”. We model base-resolution RNA-seq data as vectors in high dimensional data space and measure their level of shape similarity (LSS) using angles between samples with lower angles indicating higher similarity which is more likely to be “on” status. Applying this approach to 3 human cancer samples (head and neck squamous cell carcinoma (HNSC), lung adenocarcinoma (LUAD) and lung squamous cell carcinoma (LUSC)) from the Cancer Genome Atlas (TCGA), we identified lists of genes, the OFFONOME, that were either always or sometimes “off”. Differential OFFONOME genes were queried to address supervised and unsupervised analyses. Results: Using our technique, we characterized the OFFONOME of HNSC (5851 genes), a set which would typically have been filtering for removal because of low expression. In the HNSC OFFONOME, we observed five gene clusters, each of which was strongly associated with specific gene ontology. One of the clusters identified a rare population of normal myocytes infiltrating otherwise invasive tumors. For the OFFONOME of LUAD (5435 genes) and LUSC (5292 genes), we found clusters enriched with cilia and keratinization-related genes. In the result of integrated analysis, we observed that squamous cell carcinoma (SCC) tumor types shared “on” status for keratinization-related genes known for the cause of SCC. By comparison, LUAD had “on” status genes related to microtubule-based movement related with cilia structure. Strikingly, clustering the OFFONOME with the LSS distinguished 3 tumor types with almost perfect separation, outperforming several competing gene expression measures. Conclusion: In this study, we applied a new notion of gene expression. This approach enables a robust characterization of “on” and “off” status which can be especially effective for the genes expressed at low level. The OFFONOME from 3 cancer types revealed not only the tissue-specific genes but also the genes shared in similar tumor types. Collectively, OFFONOME can provide new insights into genes expressed at vanishingly low levels, such as from minor cell populations within bulk tumor analysis. Citation Format: Wonyoung Choi, Hyo Young Choi, Daivd Neil Hayes. OFFONOME: a new notion of genes’ on/off. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 4297.
Abstract Background: Alternative transcription initiation (ATI) has been frequently observed in cancer suggesting that it contributes to the malignant transformation of the cells. However, ATI remains largely unexplored mainly due to the lack of tools for detecting ATI. We propose a computational method for integrating bulk ATAC-seq and bulk RNA-seq to identify ATI and understand their differential usage between tissues. We hypothesize that differential ATAC-seq intensity can be used as a guide for looking for differential promoter usage, which might enable the identification of novel ATI events in a transcript-agnostic way. Methods: We recently published methods for RNA-seq base-level analysis for identifying structural variations in transcripts. Building on this, we developed a supervised sparse non-negative matrix factorization approach that integrates base-level RNA-seq and ATAC-seq, which aims to identify ATIs as well as characterize the latent structure of the underlying isoforms. This method equips many unique features that can be useful for the identification of novel ATIs. Because the method scans an entire collection of DNA accessible regions provided by ATAC-seq, it enables a comprehensive screening of novel ATI candidates that are not limited to the known promoters. Additionally, the uncompressed view of base-level RNA-seq allows us to infer the structure of individual isoforms independently of known gene annotation. The predicted isoforms can be further used to deconvolute base-level RNA-seq of individual cases into the isoforms and thus infer expression levels of each isoform. Results: We applied this method to a sub-cohort (N=350) of TCGA pan-cancer samples in which both ATAC-seq and RNA-seq are available. Empiric comparison to existing methods confirmed known true positive ATIs including important cancer genes such as CDKN2A and ALK. In particular, the method successfully identified the ATIs including some challenging cases such as ATIs located at internal introns or at constitutive exons in other transcripts. The deconvolution analysis applied to the extended cohort to all available ~10,000 TCGA pan-caner samples across 32 tissue types revealed that ATIs were predominantly differentiated across tissue types. By applying the method to a set of transcription regulator genes, we identified ~1% of the genes had ATIs including the novel cancer-specific genes with known ATIs that had not been previously reported as relevant to cancer. Additionally, we provide examples in which the isoform’s function relative to cancer appears mechanistically interesting. Conclusion: We propose a multi-omics integration method that is independent of known gene annotation, enabling a robust identification of ATIs. Our results strongly demonstrate our ability to pick up known as well as novel ATIs that are otherwise difficult to identify by existing methods. Citation Format: Hyo Young Choi, Won-Young Choi, David N. Hayes. Novel framework for systematically detecting alternative transcript initiation by integrating ATAC-seq and RNA-seq [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 2076.