Fig S2. Clinical/immunohistochemical feature distributions. Histograms of clinical/immunohistochemical features: (A) age, in years, (B) ATF6 status (histology score), (C) H3K14Ace status (histology score), (D) DUSP1 status (histology score), (E) CBX2 status (histology score), and (F) counts of BRCA Mutation status, where an N/A value indicates that the patient was not tested.
Fig S20. Random forest feature importance results for only primary tumor samples. Main text Figure 7 with random forest models trained only on primary tumor samples for the same binary outcome variables of low vs. high OS and PFS (n = 69).
Table S4. Definitions of cellular phenotypes identified with unsupervised clustering.
Table S10. Correlation coefficients between other cell type percentages and tumor cell percentages. Full list of Pearson correlation coefficients between tumor cell percentage and cell type percentage, two-sided p-values and false discovery rate adjusted p-values are reported, with significance determined by the latter.
Table S15. Ranking of feature importance in the random forest model main text results. Ranking of median Gini importance values for all features across 500 evaluations predicting OS and PFS.
Supplementary Figure S7 shows the analysis from the multispectral immunohistochemistry of ex vivo tumors treated with single agents and in combination.
Supplementary Figure S4 shows the in vivo model of HGS2 olaparib-resistant cells at day 0 and day 14 of treatment. It also shows mouse weight over the course of the treatment.
Fig S15. Cell type region size distributions. (A) The proportion of cells of each type found within a region of equal or larger size, displayed with a logarithmic x-axis. (B) The log-log complementary cumulative distribution function of cell type region sizes aggregated across all connected regions in all samples included in the final analysis.
Fig S12. Cell type counts across all 83 samples. Counts of each cell type across all identified cells.
Fig S7. Random forest predictive performance results with features derived from missing cell types all treated as NA. Main text Figure 6 repeated with missing cell types handled differently in data preprocessing.
Fig S19. Random forest predictive performance results for only primary tumor samples. Main text Figure 6 repeated with random forest models trained only on primary tumor samples for the same binary outcome variables of low vs. high OS and PFS (n = 69).
Cancer biomarker discovery is limited by small cohort sizes and the development of incompatible RNA assay platforms. Here, we systematically evaluate how to integrate data from bulk and single-cell (sc) RNA sequencing (RNA-seq), NanoString, and microarray for predictive modeling in cancer. We use high-grade serous carcinoma (HGSC) as a model system, as it is highly heterogeneous in both biology and assay data. We show that fold-change of gene expression from matched pre- and post-chemotherapy samples reduces inter-patient and inter-assay variability, but platform-specific biases persist, particularly in scRNA-seq and microarray. To optimize joint-modeling between RNA-seq and NanoString, we generate a new data set of tumor samples profiled with both assays, identifying detection limits and optimal harmonization strategies. Our approaches enable integration of cohorts for separate and combined RNA-seq and NanoString predictive models of disease recurrence (test-set AUROCs > 0.8), validated in external microarray cohorts. We leverage RNA-seq network-based analyses to provide mechanistic context for model genes, finding that GBP4 expression is a key predictor of recurrence and marks immune remodeling towards cytotoxicity. We provide an interactive web portal for data exploration. These findings establish a generalizable cross-assay harmonization of transcriptomic data and enable improved predictive modeling in heterogeneous cancers.
Hallmark pathways enriched in ID8-R cells treated with combinatory EZM8266/Olaparib compared to Olaparib alone.
Supplementary Figure S1 shows the EHMT-inhibition mediated loss of H3K9Me2 in the different HGSC models. It also shows the transcriptional changes via volcano plot and pathway analysis due to the treatment of single treatments and combination treatments.
Supplementary Figure S2 shows the transcriptional changes of transposable elements via volcano plot due to the treatment of single treatments and combination treatments. It also shows the chromatin profiling of EHMT2 bound to LTR8B elements in the genome after single and combination treatment.
Fig S22. Random forest feature importance results split by feature type - PFS. Feature importance results for progression-free survival show in Figure 7B, split by feature type and split by subtype for spatial network features.
Hallmark pathways enriched in ID8-R cells treated with combinatory EZM8266/Olaparib compared to DMSO control.
Table S8. Univariate Cox regression results with features derived from missing cell types all treated as NA. Univariate Cox regression results for OS and PFS for all features if all missing data values replaced with NA.
Table S14. Ranking of feature importance in the random forest model for only primary tumor samples. Ranking of median Gini importance values for all features across 500 evaluations predicting OS and PFS, with the model trained only on primary tumor samples.