5084 Background: The androgen receptor pathway inhibitor (ARPI) enzalutamide is one of the principal treatments for metastatic hormone-naïve and castration-resistant prostate cancer (CRPC). Most patients respond to enzalutamide. However, tumors from a subset of patients exhibit extreme non-response and are primary refractory to treatment. We sought to understand the gene expression program of enzalutamide extreme non-response (ENR) and identify alternate therapeutic approaches for tumors driven by this program. Methods: We analyzed gene expression by RNA-sequencing in pre-treatment metastatic biopsies from men with CRPC treated on a prospective enzalutamide clinical trial (NCT02099864). We focused on those with ENR (progression within 3 months) vs. long-term response (progression after 24 months) and identified a gene program linked to enzalutamide ENR. We validated the utility of this program in additional patient cohorts using a multivariable analysis and in preclinical models. Results: Unsupervised clustering correctly classified ENR patients whose tumors harbored proliferative, epithelial-to-mesenchymal transition, and stemness genes sets. Using a supervised approach, we developed a gene signature to measure the ENR program. High expression of this program in CRPC patient validation cohorts was independently associated with poor tumor control with AR targeting in multivariable analysis. Conversely, high expression of the program was independently associated with benefit with docetaxel chemotherapy, suggesting the ENR program is predictive and not merely prognostic. In support of our findings, high expression of the ENR program was strongly linked to docetaxel sensitivity in a large panel of CRPC models. Finally, we identified putative regulators of the ENR program—several of which can be targeted pharmacologically with agents that are FDA-approved or in clinical trials. Conclusions: The enza ENR program we identified is independently predictive of ENR to AR targeting. However, patients whose tumors harbor this program may be good candidates for docetaxel chemotherapy or clinical trials testing agents that block putative regulators of this program.
Human pluripotent stem cell-derived tissue engineering offers great promise in designer cell-based personalized therapeutics. To harness such potential, a broader approach requires a deeper understanding of tissue-level interactions. We previously developed a manufacturing system for the ectoderm-derived skin epithelium for cell replacement therapy. However, it remains challenging to manufacture the endoderm-derived esophageal epithelium, despite both possessing similar stratified structure. Here we employ single cell and spatial technologies to generate a spatiotemporal multi-omics cell atlas for human esophageal development. We illuminate the cellular diversity, dynamics and signal communications for the developing esophageal epithelium and stroma. Using the machine-learning based Manatee, we prioritize the combinations of candidate human developmental signals for in vitro derivation of esophageal basal cells. Functional validation of the Manatee predictions leads to a clinically-compatible system for manufacturing human esophageal mucosa. Our approach creates a versatile platform to accelerate human tissue manufacturing for future cell replacement therapies to treat human genetic defects and wounds.
Understanding transcriptional heterogeneity in cancer cells and its implication for treatment response is critical to identify how resistance occurs and may be targeted. Such heterogeneity can be captured by in vitro studies through clonal barcoding methods. We present TraCSED (Transformer-based modeling of Clonal Selection and Expression Dynamics), a dynamic deep learning approach for modeling clonal selection. Using single-cell gene expression and the fitness of barcoded clones, TraCSED identifies interpretable gene programs and the time points at which they are associated with clonal selection. When applied to cells treated with either giredestrant, a selective estrogen receptor (ER) antagonist and degrader, or palbociclib, a CDK4/6 inhibitor, pathways dynamically associated with resistance are revealed. For example, ER activity is associated with positive selection around day four under palbociclib treatment and this adaptive response can be suppressed by combining the drugs. Yet, in the combination treatment, one clone still emerged. Clustering based on partial least squares regression found that high baseline expression of both SNHG25 and SNCG genes was the primary marker of positive selection to co-treatment and thus potentially associated with innate resistance — an aspect that traditional differential analysis methods missed. In conclusion, TraCSED enables associating features with phenotypes in a time-dependent manner from scRNA-seq data.
As genes tend to be co-regulated as gene modules, feature selection in machine learning (ML) on gene expression data can be challenged by the complexity of gene regulation. Here, we present a protocol for reconciling differences in classifier features identified using different ML approaches. We describe steps for loading the PathwaySpace R package, preparing input for analysis, and creating density plots of gene sets. We then detail procedures for testing whether apparently distinct feature sets are related in pathway space. For complete details on the use and execution of this protocol, please refer to Ellrott et al.1.
The androgen receptor inhibitor enzalutamide is one of the principal treatments for metastatic prostate cancer. Most patients respond. However, a subset is primary refractory. Seeking to understand enzalutamide extreme non-response (ENR), we analyzed RNA-sequencing in biopsies from men treated prospectively on an enzalutamide clinical trial. We focused on those with ENR (progression within 3 months) vs. long-term response (progression after 24 months). We identified an ENR program linked to proliferation, epithelial-to-mesenchymal transition, and stemness. High expression of this program in additional datasets was independently linked to poor tumor control with AR targeting but favorable tumor control with docetaxel, another standard treatment. CDK2 was implicated in the ENR program. CDK2 suppression reduced the ENR program and viability of ENR program-high prostate cancer models. The ENR gene program is predictive of non-response to AR targeting. Patients whose tumors harbor this program may be good candidates for docetaxel or CDK2 inhibitor clinical trials.
Abstract The NCI's The Cancer Genome Atlas (TCGA) project profiled over 10,000 tumor samples over the course of 10 years. As different tissue-specific working groups reviewed all of the available data, these patient samples were separated into distinct molecular subtypes, and these clusters were reported in various marker papers. While these assignments provided invaluable information about the common patterns of molecular characteristics in different types of cancer there was no consistent methodology for assigning new samples to these defined molecular subtypes.The NCI's Tumor Molecular Pathology group was formulated to create machine learning-based models that could be applied to non-TCGA samples and determine their TCGA mapped subtypes. Five modeling systems, JADBio, SKGrid by the Oregon Health and Science University, CloudForest by the Institute of Systems Biology, AKLIMATE by University of California Santa Cruz and subSCOPE by BC Cancer’s Genome Sciences Centre, were trained to recognize TCGA subtypes using multi-omic measurements from gene expression, DNA methylation, miRNA expression, copy number, and somatic mutation calls. While the TCGA samples were profiled using multi-omic technologies, single platform and/or compact feature set models also were assessed for their ability to assign these classifications. Each machine learning system created predictive models for 106 subtypes from 26 cancer types using as few features as possible, with a maximum of 100 features allowed for scored models. A set of 411,706 models was developed, composed of results of each of the learning methods across the various omic platforms. Top models, both multi-omic and single platform, were selected for each cancer type. On average, models were able to achieve an overall weighted F1 score of 0.895 with 42 features. While the top models for each cancer type had an overall weighted F1 mean performance of 0.936 with a mean of 29 features, in 20 of the 26 cancer types models using only gene expression provided the best performance. Analysis of features selected by the models showed some known onco-drivers were selected by many models, but many times different models would utilize features of different genes with similar levels of performance. Network-level analysis revealed that many genes of these selected features operated within the same pathways.Transferability of these models to external datasets was tested, taking TCGA breast cancer trained models and applying them to AURORA and METABRIC datasets. Interestingly, despite the data platform difference between TCGA (RNAseq) and METABRIC (microarray), model performance saw only minimal degradation of F1 values in transfer. This set of models and the training dataset will provide new opportunities for researchers and translational scientists to connect new tumors to the subtypes seen in the TCGA cohorts. Citation Format: Kyle Ellrott, Chris K. Wong, Christina Yau, Mauro A. Castro, Jordan Lee, Brian Karlberg, Jasleen K. Grewal, Vincenzo Lagani, Bahar Tercan, Verena Friedl, Toshinori Hinoue, Vladislav Uzunangelov, Lindsay Westlake, Xavier Loinaz, Ina Felau, Peggy Wang, Anab Kemal, Samantha J. Caesar-Johnson, Ilya Shmulevich, Alexander J. Lazar, Ioannis Tsamardinos, Katherine A. Hoadley, The Cancer Genome Atlas Analysis Network, Gordon A. Robertson, Theo A. Knijnenburg, Christopher C. Benz, Joshua M. Stuart, Jean C. Zenklusen, Andrew D. Cherniack, Peter W. Laird. Leveraging compact feature sets for TCGA-based molecular subtype classification on new samples [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 6548.
Inferring gene regulatory networks from single-cell RNA-sequencing trajectories has been an active area of research yet methods are still needed to identify regulators governing cell transitions. We developed DREAMIT (Dynamic Regulation of Expression Across Modules in Inferred Trajectories) to annotate transcription-factor activity along single-cell trajectory branches, using ensembles of relations to target genes. Using a benchmark representing several different tissues, as well as external validation with ATAC-Seq and Perturb-Seq data on hematopoietic cells, the method was found to have higher tissue-specific sensitivity and specificity over competing approaches.
We present the Manatee variational autoencoder model to predict transcription factor (TF) perturbation-induced transcriptomes. We demonstrate that the Manatee in silico perturbation analysis recapitulates target transcriptomic phenotypes in diverse cellular lineage transitions. We further propose the Manatee in silico screening analysis for prioritizing TF combinations targeting desired transcriptomic phenotypes.
The cellular components of tumors and their microenvironment play pivotal roles in tumor progression, patient survival, and the response to cancer treatments. Unveiling a comprehensive cellular profile within bulk tumors via single-cell RNA sequencing (scRNA-seq) data is crucial, as it unveils intrinsic tumor cellular traits that elude identification through conventional cancer subtyping methods. Our contribution, scBeacon, is a tool that derives cell-type signatures by integrating and clustering multiple scRNA-seq datasets to extract signatures for deconvolving unrelated tumor datasets on bulk samples. Through the employment of scBeacon on the The Cancer Genome Atlas (TCGA) cohort, we find cellular and molecular attributes within specific tumor categories, many with patient outcome relevance. We developed a tumor cell-type map to visually depict the relationships among TCGA samples based on the cell-type inferences.
Tissue and molecular subtype distribution among the samples in the entire cohort (A-B), and in the pan-cancer cluster (C-D). The pie charts represent the number of samples from each tissue of origin in the entire cohort (A) and the integrated pan-cancer cluster (C). Black and white matrices illustrate the presence of molecular features of each platform (x-axis) across samples (y-axis), in the entire cohort (B) or in the integrated pan-cancer cluster (D). Data available for this sample for a given platform is marked black, otherwise the entry is white.
Table S7 contains microbe screening results.
Table S1 contains cohort description, Master Patient Table and MutSigCV results.
<p>Supplementary Figure Legend contains the description of supplementary Figure S1-S4 and supplementary Datasets S1-S2.</p>
Supplementary Methods from Voltage-Gated Na<sup>+</sup> Channel <i>SCN5A</i> Is a Key Regulator of a Gene Transcriptional Network That Controls Colon Cancer Invasion
Table S2 contains BAP1 analysis results, as well as detailed lists of YY1 and IRF8 target genes.
Supplementary Figure 4 from Voltage-Gated Na+ Channel SCN5A Is a Key Regulator of a Gene Transcriptional Network That Controls Colon Cancer Invasion
Supplementary Figure 2 from Voltage-Gated Na<sup>+</sup> Channel <i>SCN5A</i> Is a Key Regulator of a Gene Transcriptional Network That Controls Colon Cancer Invasion
Table S6 contains results from the analysis of DNA methylation in SETD2 mutated and BAP1 inactivated samples.