Abstract Trastuzumab deruxtecan (TDXD) is approved for HER2 low or positive metastatic breast cancer. HER2-status is assessed through HER IHC and ERBB2 FISH assays. HER2-status has been shown to be correlated to RNA expression. This study aimed to assess if the addition of DNA to an RNA-based model can improve HER2-status prediction performance. Furthermore, to investigate the relationship between model score and outcomes for TDXD patients according to their overall survival (OS). Training and test sets of breast cancer samples were collected for HER2 status prediction (n = 1275, n = 397 respectively). HER2 status was reported positive: 11%, 12%, low: 61%, 60%, negative: 28%, 29% for each set respectively. HER2 testing and NGS were performed on samples with the same collection date. A TDXD discovery cohort of 284 patients included patients who received TDXD after RNA/DNA collection and had at least a 30-day follow-up. Metastatic disease was reported for 98% of the TDXD cohort. Receptor status was HR+/HER2- 39%, HR-/HER2- 17%, HR+/HER2+ 16%, HR-/HER2+ 9%, unknown 19%, and HER2-status was negative 11%, low 42%, and positive 27%, unknown 20%. OS was measured from medication start date. We trained a 12-gene linear model based on RNA expression and DNA copy number to predict HER2-positivity. The genes were selected by stepwise-selection on RNA and DNA features when predicting HER2-positivity using a random forest model. A baseline model consisting only of ERBB2 RNA expression was compared. Setting a HER2-positivity score threshold and allowing for an indeterminate group of 15% of the validation cohort, the RNA-DNA model achieved 91% PPA and 91% PPV in validation, whereas the RNA-only model achieved 83% PPA and 85% PPA. This indicates that the addition of DNA data improved the RNA-only HER2 predictor. The Concordance index of the RNA-DNA model score and OS for TDXD was 0.64. Characterizing this relationship, we partitioned the TDXD cohort to three equal sized sub-cohorts (33 1/3%) ordered by model score: low, intermediate, and high. Median OS in months for groups was: low 13.6, intermediate 18.5, and high 20.4. High-group patients had significantly better OS on TDXD compared to the low-group (HR = 0.39, log-rank p-value < 0.002). For patients with reported status, HER2-positivity rate in each group respectively was 6%, 16%, 85% (n=78), and HR-positivity in each group was 70%, 68%, 67% (n=156). Together, these findings indicate that HER2 status prediction improved by adding DNA to an RNA base predictor, and the predicted score was positively correlated with OS for TDXD. Citation Format: Kaveri Nadhamuni, Talal Ahmed, Mark Carty, Ben Terdich, Whitney L. Hensing, Timothy Taxter, Calvin Chao, Raphael Pelossof. Identification of poor responders to trastuzumab-deruxtecan with a multi-modal HER2-status predictor [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 2502.
Abstract Background: The prognostic and predictive value of the PAM50 intrinsic subtypes, namely Luminal A, Luminal B, Her2, and Basal-like subtypes, is well-studied in primary as well as metastatic breast cancer settings. Prosigna has emerged as a rapid PAM50 subtype predictor based on the NanoString nCounter assay. However, assay reproducibility across various RNASeq or qRT-PCR platforms can be challenging, especially when applying the predictor on metastatic breast cancer tumors. Here, we used SpinAdapt to create an intrinsic subtype predictor that works on RNA sequencing data, and validates on multiple tumor-sites. We evaluate real-world outcomes for our intrinsic subtype predictions across various immunohistochemical (IHC) labels and metastatic sites. Methods: We trained the subtype predictor on a cohort of 2,497 breast cancer patients using the PAM50 genes, profiled using Nanostring RNA nCounter assay (GSE148426). Approximately 5,423 de-identified records of breast cancer patients sequenced using whole-exome capture RNA-seq were included in our reference dataset. The reference dataset contained samples collected from various sites including breast (n=2440), liver (n=936), lymph node (n=577), lung (n=540), and bone (n=304). For subtype classification, we first batch-corrected the external dataset to Tempus RNA-seq reference dataset using SpinAdapt, then trained a Support-Vector Classifier (SVC) on the corrected data. A 10-fold CV experiment was performed on the corrected dataset to analytically validate the intrinsic subtype predictions. We retrospectively analyzed 7,021 de-identified breast cancer patients with known hormone receptor (HR) or HER2 status and a matched RNASeq sample. The concordance between HR/HER2 status and PAM50 prediction was analyzed, and these patients were excluded from training. Real-world overall survival (rwOS) was evaluated from the time of first diagnosis. The outcomes across intrinsic subtypes were further assessed according to tumor collection site and HR/HER2 IHC status. Results: The 10-fold CV experiment on the Tempus-adapted GSE148426 dataset achieved F-1 scores of 0.97, 0.86, 0.94, and 0.87 on Basal, HER2-like, Luminal A, and Luminal B PAM50 subtypes, respectively. On the Tempus evaluation dataset, 85.3% of HR+/HER2- patients (n=4,366), 65.4% of HR-/HER2+ patients (n=240), and 75.4% of HR-/HER2- patients (n=1,930) were classified as Luminal, HER2-like, and Basal, respectively. Evaluating outcomes on Tempus patients for each PAM50 group, the rwOS for the basal group was significantly shorter than patients not predicted to be basal (n=5,845, p< 2e-90). The rwOS for the predicted PAM50 basal patients remained significantly shorter than non-basal patients even when stratified by site of metastasis: breast (n=2,405; p< 1e-27), lymph node (n=577, p< 1e-7), liver (n=936; p< 1e-25), lung (n=540, p< 1e-13), and bone (n=304, p< 1e-3). Interestingly, within both HR+/HER2- (n=3,653) and HR-/HER2- (n=1,664) IHC cohorts with available outcomes data, the predicted PAM50 basal subtype could further stratify each of these populations with basal-subtype showing significantly worse prognosis than the non-basal subtype (p< 1e-27 and p< 1e-7, respectively). Conclusions: We retrospectively analyzed Tempus multimodal RWD to validate an in-house breast intrinsic subtype predictor that is agnostic to the site of metastasis. The prognostic value of the basal subtype was significant for breast cancer patients across various sites of metastasis including lymph node, liver, lung, and bones and IHC groups. For patients in each of the HR+/HER2- and triple negative IHC groups, the intrinsic molecular subtypes provided an additional level of prognostic detail with statistical significance. These data emphasize the importance of combining molecular subtypes with IHC-based diagnostics to fully characterize clinically relevant subpopulations and risk. Citation Format: Talal Ahmed, Mark Carty, Kaveri Nadhamuni, Raphael Pelossof. Breast cancer intrinsic subtypes predict outcomes in primary and metastatic samples [abstract]. In: Proceedings of the 2023 San Antonio Breast Cancer Symposium; 2023 Dec 5-9; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2024;84(9 Suppl):Abstract nr PO5-24-03.
Reproducibility of results obtained using ribonucleic acid (RNA) data across labs remains a major hurdle in cancer research. Often, molecular predictors trained on one dataset cannot be applied to another due to differences in RNA library preparation and quantification, which inhibits the validation of predictors across labs. While current RNA correction algorithms reduce these differences, they require simultaneous access to patient-level data from all datasets, which necessitates the sharing of training data for predictors when sharing predictors. Here, we describe SpinAdapt, an unsupervised RNA correction algorithm that enables the transfer of molecular models without requiring access to patient-level data. It computes data corrections only via aggregate statistics of each dataset, thereby maintaining patient data privacy. Despite an inherent trade-off between privacy and performance, SpinAdapt outperforms current correction methods, like Seurat and ComBat, on publicly available cancer studies, including TCGA and ICGC. Furthermore, SpinAdapt can correct new samples, thereby enabling unbiased evaluation on validation cohorts. We expect this novel correction paradigm to enhance research reproducibility and to preserve patient privacy.
Abstract The reproducibility of results obtained using RNA data across labs is a major hurdle in cancer research. Difference in library preparation methods and gene expression quantification platforms prevent the application of trained models to new data across labs. SpinAdapt is a novel unsupervised domain adaptation algorithm that enables the transfer of existing molecular models across labs and technological platforms, without requiring re-training or calibration of existing models for future prospective data. Furthermore, SpinAdapt uses summary statistics (independent latent space representations) to calculate data corrections, rather than requiring full data access. This allows for transfer of molecular models across sequencing platforms and between labs without loss of data ownership or compromise of data privacy. To evaluate SpinAdapt, we performed two sets of experiments: A) We transferred molecular tumor subtype classifiers across four pairs of publicly available cancer datasets (bladder, breast, colorectal, pancreatic), covering 4,076 samples across 18 different tumor subtypes and three technological platforms (RNASeq, Affymetrix U133plus2, and HE1ST). For each pair of datasets we trained a subtype classifier on one dataset (target) according to well-accepted subtyping annotations (Zea Tan et al. 2019; Prat et al. 2012; Guinney et al. 2015; Bailey et al. 2016), and then evaluated the classifier accuracy on the other dataset (source). For each tumor subtype, we quantified the classification performance using mean AUC score across random subsets of the source dataset, where each subset was corrected using SpinAdapt. We aggregated performance across all subtypes and report the average mean AUC score for each cancer type: bladder 0.95, breast 0.98, colorectal 0.98, pancreatic 0.96; demonstrating high accuracy on all diagnostic tasks. B) To demonstrate the transferability of prognostic models, we trained five Cox survival models on five target cancer datasets respectively (breast, lung, colorectal, liver, pancreatic) ranging from 186 to 2,919 RNASeq samples. We used SpinAdapt to adapt five source cancer datasets to the target datasets, ranging from 226 to 1,038 samples across different platforms (RNASeq, Affymetrix U133Plus2 and HG-U133A, Illumina HumanHT-12v4). For every cancer type, we trained a Cox model on the target dataset, and measured its performance by predicting survival risk on the corresponding adapted source dataset. We show high survival prediction accuracy for all datasets (Log-rank P-values and c-index): lung [1e-6, 01.661], breast [5e-5, 0.626], liver [1e-4, 0.708], pancreatic [2e-4, 0.629], colorectal [9e-4, 0.661]. SpinAdapt transferred diagnostic and prognostic models over 14 cancer datasets covering 7,146 samples across six different cancer types and various platforms (RNASeq, microarray), while maintaining model accuracy and statistical significance. Citation Format: Talal Ahmed, Stephane Wenric, Mark Carty, Rafael Pelossof. Transferring diagnostic and prognostic molecular models across technological platforms [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 242.
Reproducibility of results obtained using RNA data across labs remains a major hurdle in cancer research. Often, molecular predictors trained on one dataset cannot be applied to another due to differences in RNA library preparation and quantification. While current RNA correction algorithms may overcome these differences, they require access to all patient-level data, which necessitates the sharing of training data for predictors when sharing predictors. Here, we describe SpinAdapt, an unsupervised RNA correction algorithm that enables the transfer of molecular models without requiring access to patient-level data. It computes data corrections only via aggregate statistics of each dataset, thereby maintaining patient data privacy. Furthermore, SpinAdapt can correct new samples, thereby enabling evaluation of validation cohorts. Despite an inherent tradeoff between privacy and performance, SpinAdapt outperforms current correction methods that require patient-level data access. We expect this novel correction paradigm to enhance research reproducibility and patient privacy. Finally, SpinAdapt lays a mathematical framework that can be extended to other -omics modalities.
e13507 Background: Recent advances in transcriptomics have resulted in the emergence of several publicly available breast cancer RNA-Seq datasets, such as TCGA, SCAN-B, and METABRIC. However, molecular predictors cannot be applied across datasets without the correction of batch differences. In this study, we demonstrate a homogenization algorithm that allows the transfer of molecular subtype predictors from one RNA-Seq cohort to another. The algorithm only uses cohort-level RNA-Seq summary statistics, and therefore, does not require joint normalization of both datasets nor the transfer of patient information. Using this approach, we transferred a breast cancer subtype (Luminal A, Luminal B, HER2+, Basal) predictor trained on SCAN-B data to accurately predict subtypes from TCGA. Methods: First, we randomly split the TCGA cohort (n = 481 Luminal A, n = 189 Luminal B, n = 73 Her2+, n = 168 Basal) into two sets: TCGA-train and held-out TCGA-test (n = 455 and n = 456, respectively). Second, the SCAN-B cohort (n = 837) was homogenized with the TCGA-train set. Third, a molecular subtype predictor, based on a logistic regression model, was trained on homogenized SCAN-B RNA-Seq samples and used to predict the subtypes of TCGA-test RNA-Seq samples. For baseline comparison, a similar predictor trained on the non-homogenized SCAN-B cohort was tested on the TCGA-test set. The experimental framework was iterated 250 times. Reported P-values reflect a paired one-sided t-test. Results: To quantify model performance, we measured the average F1 score for each tumor subtype prediction from the held-out TCGA test set with and without cohort homogenization. The average F1 scores with vs. without homogenization were: Luminal A, 0.88 vs. 0.85 ( P< 1e-69); Luminal B, 0.74 vs. 0.51 ( P< 1e-183); Her2+, 0.73 vs. 0.53 ( P< 1e-99); Basal, 0.98 vs. 0.97 ( P< 1e-53). Overall, homogenization significantly outperformed no homogenization. Conclusions: We developed a novel homogenization algorithm that accurately transfers subtype predictors across diverse, independent breast cancer cohorts.
Here we present HiC-DC, a principled method to estimate the statistical significance (P values) of chromatin interactions from Hi-C experiments. HiC-DC uses hurdle negative binomial regression account for systematic sources of variation in Hi-C read counts-for example, distance-dependent random polymer ligation and GC content and mappability bias-and model zero inflation and overdispersion. Applied to high-resolution Hi-C data in a lymphoblastoid cell line, HiC-DC detects significant interactions at the sub-topologically associating domain level, identifying potential structural and regulatory interactions supported by CTCF binding sites, DNase accessibility, and/or active histone marks. CTCF-associated interactions are most strongly enriched in the middle genomic distance range (∼700 kb-1.5 Mb), while interactions involving actively marked DNase accessible elements are enriched both at short (<500 kb) and longer (>1.5 Mb) genomic distances. There is a striking enrichment of longer-range interactions connecting replication-dependent histone genes on chromosome 6, potentially representing the chromatin architecture at the histone locus body.
The three dimensional (3D) architecture of chromosomes is not random but instead tightly organized due to chromatin folding and chromatin interactions between genomically distant loci. By bringing genomically distant functional elements such as enhancers and promoters into close proximity, these interactions play a key role in regulating gene expression. Some of these interactions are dynamic, that is, they differ between cell types, conditions and can be induced by specific stimuli or differentiation events. Other interactions are more structural and stable, that is they are constitutionally present across several cell types. Genome contact interactions can occur via recruitment and physical interaction between chromatin-binding proteins and correlate with epigenetic marks such as histone modifications. Absence of a contact can occur due to presence of insulators, that is, chromatin-bound complexes that physically separate genomic loci. Understanding which contacts occur or do not occur in a given cell type is important since it can help explain how genes are regulated and which functional elements are involved in such regulation. The analysis of genome contact interactions has been greatly facilitated by the relatively recent development of chromosome conformation capture (3C). In an even more recent development, 3C was combined with next generation sequencing and led to Hi-C, a technique that in theory queries all possible pairwise interactions both within the same chromosome (intra) and between chromosomes (inter). Hi-C has now been used to study genome contact interactions in several human and mouse cell types as well as in animal models such as Drosophila and yeast. While it is fair to say that Hi-C has revolutionized the study of chromatin interactions, the computational analysis of Hi-C data is extremely challenging due to the presence of biases, artifacts, random polymer ligation and the huge number of potential pairwise interactions. In this chapter, we outline a strategy for analysis of genome contact experiments based on Hi-C using R and Bioconductor.
Aging hematopoietic stem cells (HSCs) exhibit numerous functional alterations including reduced capacity for self-renewal, myeloid-biased differentiation, and reduced production of mature lymphocytes and red blood cells. Interventions such as calorie restriction (CR) and rapamycin (Rapa) treatment have been shown to increase lifespan and to delay the onset of age-related diseases, and some studies have demonstrated that they may improve HSC function through poorly defined mechanisms. We and others have shown that microRNAs (miRNAs) are potent cell-intrinsic regulators of HSC self-renewal and lineage specification and also contribute to age-related disorders such as acute myeloid leukemia (AML) and the myelodysplastic syndromes (MDS). We hypothesized that miRNAs may underlie the recovery of HSC function observed in anti-aging mouse models, and thus we characterized miRNA expression profiles from HSCs (Lin-c-Kit+Sca-1+CD34-CD150+) from young mice (12-16 weeks old), old mice (20-22 months), and old mice that had been treated with anti-aging interventions. Evaluation of HSCs from CR and Rapa treated old mice revealed numerous changes consistent with inhibition/reversal of age-related HSC changes including a 5-fold reduction in HSC frequency (p=0.04), 2-fold increase in erythroid progenitors (pro-erythroblasts, p=0.04), 2.5 fold increase in common lymphoid progenitors (CLP; Lin-c-Kit+Sca-1+CD127+FLK2+, p=0.05), as well as 3.5-fold increase in peripheral blood B cells (p=0.002), 2.2 fold decrease in platelets (p=0.01), and increased red blood cells (p=0.04). These changes were associated with statistically significant increases in the percentage of HSCs in S/M/G2 (p=0.045), and undergoing apoptosis (p=0.05). Using a TaqMan-based qPCR expression profiling method evaluating 750 miRNAs, we found that old HSCs exhibited altered expression of 91 miRNAs compared to young (FDR <0.1, P <0.05). Moreover, HSCs from both CR and Rapa treated old mice exhibited expression of 60 miRNAs at levels similar to young, normal HSCs. miR-125b, a miRNA we and others previously showed to positively regulate HSC self-renewal, was reduced 2.2-fold in old mice, and its expression was restored in CR and Rapa treated HSCs. Lentivirally mediated expression of miR-125b in old HSCs increased their long-term reconstitution capacity 8.1-fold compared to control old HSCs based on donor chimerism levels at 16 weeks post-transplantation, resulting in chimerism levels similar to mice transplanted with young HSCs expressing miR-125b. The improved HSC engraftment capacity of old HSCs transduced with miR-125b was accompanied by statistically significant increases in the frequencies of lymphoid biased HSCs (Lin-c-Kit+Sca-1+CD34-CD150neg-low), megakaryocyte-erythroid progenitors (MEPs), CLPs, and peripheral blood B- and T-cells, compared to old HSCs transduced with control lentivirus (p<0.05 for all indicated cell types). While enforced expression of high levels of miR-125b in mouse HSPCs has been reported to induce myeloid leukemias, there was no evidence of a hematologic malignancy in mice transplanted with miR-125b transduced old HSCs up to 6 months post-transplantation. Overall, these results demonstrate that functional HSC aging phenotypes can be that inhibited/reversed by anti-aging interventions, that miR-125b regulates HSC aging, and that anti-aging interventions may exert their positive effects on HSC function by regulating miR-125b expression.
The posttranscriptional control of gene expression by microRNAs (miRNAs) is highly redundant, and compensatory effects limit the consequences of the inactivation of individual miRNAs. This implies that only a few miRNAs can function as effective tumor suppressors. It is also the basis of our strategy to define functionally relevant miRNA target genes that are not under redundant control by other miRNAs. We identified a functionally interconnected group of miRNAs that exhibited a reduced abundance in leukemia cells from patients with T cell acute lymphoblastic leukemia (T-ALL). To pinpoint relevant target genes, we applied a machine learning approach to eliminate genes that were subject to redundant miRNA-mediated control and to identify those genes that were exclusively targeted by tumor-suppressive miRNAs. This strategy revealed the convergence of a small group of tumor suppressor miRNAs on the Myb oncogene, as well as their effects on HBP1, which encodes a transcription factor. The expression of both genes was increased in T-ALL patient samples, and each gene promoted the progression of T-ALL in mice. Hence, our systematic analysis of tumor suppressor miRNA action identified a widespread mechanism of oncogene activation in T-ALL.
Bias created by a range of measurement disparities