Supplemental figure legends 1-8, Supplementary tables 1-5, and Supplementary methods
Assessment of programmed death ligand 1 (PD-L1) expression by immunohistochemistry (IHC) has emerged as an important predictive biomarker across multiple tumor types. However, manual quantitation of PD-L1 positivity can be difficult and leads to substantial inter-observer variability. Although the development of artificial intelligence (AI) algorithms may mitigate some of the challenges associated with manual assessment and improve the accuracy of PD-L1 expression scoring, use of AI-based approaches to oncology biomarker scoring and drug development has been sparse, primarily due to the lack of large-scale clinical validation studies across multiple cohorts and tumor types. We developed AI-powered algorithms to evaluate PD-L1 expression on tumor cells by IHC and compared it with manual IHC scoring in urothelial carcinoma, non-small cell lung cancer, melanoma, and squamous cell carcinoma of the head and neck (prospectively determined during the phase II and III CheckMate clinical trials). 1,746 slides were retrospectively analyzed, the largest investigation of digital pathology algorithms on clinical trial datasets performed to date. AI-powered quantification of PD-L1 expression on tumor cells identified more PD-L1–positive samples compared with manual scoring at cutoffs of ≥1% and ≥5% in most tumor types. Additionally, similar improvements in response and survival were observed in patients identified as PD-L1–positive compared with PD-L1–negative using both AI-powered and manual methods, while improved associations with survival were observed in patients with certain tumor types identified as PD-L1–positive using AI-powered scoring only. Our study demonstrates the potential for implementation of digital pathology-based methods in future clinical practice to identify more patients who would benefit from treatment with immuno-oncology therapy compared with current guidelines using manual assessment.
BACKGROUND AND AIMS:Manual histological assessment is currently the accepted standard for diagnosing and monitoring disease progression in NASH, but is limited by variability in interpretation and insensitivity to change. Thus, there is a critical need for improved tools to assess liver pathology in order to risk stratify NASH patients and monitor treatment response.APPROACH AND RESULTS:Here, we describe a machine learning (ML)-based approach to liver histology assessment, which accurately characterizes disease severity and heterogeneity, and sensitively quantifies treatment response in NASH. We use samples from three randomized controlled trials to build and then validate deep convolutional neural networks to measure key histological features in NASH, including steatosis, inflammation, hepatocellular ballooning, and fibrosis. The ML-based predictions showed strong correlations with expert pathologists and were prognostic of progression to cirrhosis and liver-related clinical events. We developed a heterogeneity-sensitive metric of fibrosis response, the Deep Learning Treatment Assessment Liver Fibrosis score, which measured antifibrotic treatment effects that went undetected by manual pathological staging and was concordant with histological disease progression.CONCLUSIONS:Our ML method has shown reproducibility and sensitivity and was prognostic for disease progression, demonstrating the power of ML to advance our understanding of disease heterogeneity in NASH, risk stratify affected patients, and facilitate the development of therapies.
Abstract BackgroundHomologous recombination deficiency (HRD), originally described in tumors from patients with germline mutations in BRCA1/2 genes, renders cells sensitive to poly-ADP ribose polymerase inhibitors (PARPi) (1), but can be caused by mutations in other genes and is prevalent across multiple cancer types (2). HRD status is of clinical interest because it can indicate patient eligibility for treatment with PARPi. Currently, HRD status is determined by sequencing to identify BRCA mutations or genomic instability, but this has a high rate of failure (3). In this research study, we apply a deep-learning based computational approach to directly infer HRD status from digitized images of hematoxylin and eosin (H&E) stained histology samples in breast cancer tumors. MethodsDigitized whole slide images (WSI) of 931 H&E stained, formalin-fixed and paraffin-embedded (FFPE) breast adenocarcinoma (BRCA) tumor biopsies from the cancer genome atlas (TCGA) were used to train machine learning (ML) models to identify patients that are HRD based on human-interpretable features (HIFs) and end-to-end (E2E) modeling. To train the models, samples were split into training and validation sets designated either HRD or homologous recombination proficient (HRP) based on a previously generated aggregate HRD score (calculated from regions of loss of heterozygosity, large scale genomic instability, and telomeric allelic imbalance) by genomic analysis of the PanCancerAtlas (2). We applied an untuned HRD score threshold of 45 to assign class labels resulting in 142/931 (15.3%) HRD cases. Board certified pathologists (N=93) annotated tissue regions and cellular foci on the PathAI research platform yielding 65,477 annotations. ML models based on convolutional neural networks were trained to recognize breast cancer cells, lymphocytes, macrophages, plasma cells, fibroblasts, and tissue compartments including cancer epithelium, cancer stroma and necrosis within the H&E stained breast cancer samples. Two pipelines constructed H&E histology-based classifiers of HRD status. A weakly-supervised “end-to-end” model using ResNets extracted features from small image patches with an attention module to aggregate across patches and directly predict HRD status. The HIF-based approach used the tissue segmentation and cell identification classifiers to quantify histological features in the WSI. From the labeled images, we extracted 600 HIFs that capture complex relationships between cell and tissue types. HIFs and patient clinical covariates were applied as input to a Sparse Group Lasso model to predict the HRD status of the associated patients.ResultsML models predicted HRD status from H&E stained WSI. The area under the receiver operating characteristics curve (AUROC) was 0.87 for the HIF model and 0.80 for the E2E model. Both classifiers achieved high sensitivity for HRD status (0.86) with more moderate precision (F1 score HIF: 0.80 and E2E: 0.72). Our HIF with clinical covariates model revealed morphological features that were significantly associated with HRD compared with HRP. HRD samples were enriched for areas of necrosis, stromal fibroblasts, and tumor infiltrating lymphocytes (p< 0.001, Mann-Whitney U test). Conclusions Computational models built with the PathAI research platform identified HRD positive patients directly from routinely collected H&E stained WSIs and identified a histological basis for how mutational signatures impact the tumor microenvironment. Disclaimer: The PathAI platform and HRD model are not intended for diagnostic purposes. 1 Farmer et al., 2005. Nature 14;434(7035):917-212 Knijnenburg et al., 2018. Cell Rep 23, 239–2543 Hoppe et al., 2018. JNCI 110(7): djy0854 Coudray et al., 2018 Nat Med 24:1559-15675 Kather et al., 2019. Nat Med 25: 1054–1056pages Citation Format: Amaro Taylor-Weiner, Aryan Pedawi, Wan Fung Chui, James Diao, Jason Wang, Victoria Mountain, Benjamin Glass, Hunter Elliott, Ilan Wapinski, Michael Montalto, Aditya Khosla, Andrew H. Beck. Deep-learning based prediction of homologous recombination deficiency (hrd) status from histological features in breast cancer; a research study [abstract]. In: Proceedings of the 2020 San Antonio Breast Cancer Virtual Symposium; 2020 Dec 8-11; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2021;81(4 Suppl):Abstract nr PD6-04.
Computational methods have made substantial progress in improving the accuracy and throughput of pathology workflows for diagnostic, prognostic, and genomic prediction. Still, lack of interpretability remains a significant barrier to clinical integration. We present an approach for predicting clinically-relevant molecular phenotypes from whole-slide histopathology images using human-interpretable image features (HIFs). Our method leverages >1.6 million annotations from board-certified pathologists across >5700 samples to train deep learning models for cell and tissue classification that can exhaustively map whole-slide images at two and four micron-resolution. Cell- and tissue-type model outputs are combined into 607 HIFs that quantify specific and biologically-relevant characteristics across five cancer types. We demonstrate that these HIFs correlate with well-known markers of the tumor microenvironment and can predict diverse molecular signatures (AUROC 0.601-0.864), including expression of four immune checkpoint proteins and homologous recombination deficiency, with performance comparable to 'black-box' methods. Our HIF-based approach provides a comprehensive, quantitative, and interpretable window into the composition and spatial architecture of the tumor microenvironment.
HER2 levels that will meaningfully identify patients who benefit from the HER2 antibody drug conjugate T-DxD are under investigation. We developed a quantitative machine learning (ML) method to measure HER2 expression and evaluated our approach for predicting patient outcomes. T-DXd demonstrated clinical activity in advanced-stage breast cancer (BC) patients in the DS8201-A-J101 trial. Digitized pathology images stained for HER2 from 154 patients were added to the PathAI research platform (PathAI; Boston, MA). The ML model was trained on 91,413 annotations created by 87 pathologists to identify HER2 positivity in BC cells, immune cells (IC, macrophages and lymphocytes), and tissue compartments in the HER2 stained BC samples. ML derived patient-level features were clustered and selected, with false discovery rate control, for predicting patient outcomes and HER2 status. Thresholds for feature-based patient selection were identified using cross validation optimizing for hazard ratio (HR) between patients with feature values above and below the threshold. 149 patients with BC treated by T-DXd were included in this study. Patients selected by high manual HER2 ASCO/CAP achieved an ORR of 53% vs. 41% (p=0.29). AI feature measuring complete membranous HER2 positive BC cells with increased cytoplasmic HER2 staining identified more patients than manual HER2 (86% vs. 82%) and patients selected by this achieved an ORR of 55% vs. 24% (p<0.009). A novel biomarker measuring the ratio of IC density in the BC stroma vs. BC epithelium achieved an ORR of 57% vs. 38% (p<0.027). A composite biomarker using both features achieved an ORR of 59% vs. 37% (p<0.011). Similar results were found with PFS. ML models identified tissue compartments and HER2 positivity in BC cells. ML-based feature combining IC density with complete membranous/increased cytoplasmic HER2 staining in BC cells achieved a comparable ORR (59% vs. 53%) to manual HER2 scoring. These results show the potential of ML models to select patients for HER2 therapy beyond manual HER2 scoring.
Abstract Background: Approximately 15-25% of patients with atypical ductal hyperplasia (ADH) diagnosed on breast core needle biopsy (CNB) are upgraded to ductal carcinoma in situ (DCIS) or invasive carcinoma (IC) on surgical excision. The reproducible identification of patients with ADH on CNB who are more likely to have upgrades at excision remains elusive. We hypothesized that a machine learning approach could be utilized to train models to recognize ADH on digitized pathology images and to identify cases of ADH more likely to be upgraded to DCIS or IC at excision. The purpose of this study was to determine the accuracy of the machine learning approach to identify ADH. Methods: 726 digitized images of CNB slides derived from 306 cases with a diagnosis of ADH between 11/2004-3/2018 were included in this study. Independent histologic review by two breast pathologists identified slides with and without ADH from each case. 39 board certified pathologists with experience in evaluation of breast biopsies were employed for tissue region annotation on the PathAI research platform (not intended for diagnostic purposes), yielding 14,118 tissue region annotations. Region annotations included ADH, ADH stroma, flat epithelial atypia (FEA), lobular neoplasia (LN), calcifications (Ca), columnar cell change/hyperplasia, sclerosing adenosis, papilloma, normal terminal duct lobular units and other non-atypical breast tissue regions. These annotations were used to train a convolutional neural network (CNN) with 35 layers and approximately 9 million parameters to identify ADH. The data were split into training and testing sets, representing 61.1% and 38.9% of the data respectively. The distribution of cases, images with ADH and cases with upgrade were balanced between the training and testing sets. Results: CNB specimens were assigned labels of “ADH” or “No ADH” based on histologic assessment. AI models were able to predict the diagnosis of ADH with 85% sensitivity (144 of 168 images within the test set) and 69% specificity (78 of 113 images within the test set). The slide-level area under the receiver operator curve (ROC) for this model was 0.84. Conclusions: A deep learning-based classifier showed strong performance for the identification of ADH from whole slide images of H&E stained breast CNBs. With further development, this approach may improve the reproducibility and standardization of the diagnosis of ADH. Future analyses will focus on determining if morphologic features of ADH extracted by the deep learning system can be used to predict upgrade to DCIS and IC. This approach may help stratify patients with ADH on CNB into those who require surgical excision and those who can be followed with active surveillance. Citation Format: Jennifer K. Kerner, Allison Cleary, Suyog Jain, Harsha Pokkalla, Benjamin Glass, Sam Grossmith, Maya Harary, Elizabeth Mittendorf, Andrew H. Beck, Aditya Khosla, Stuart J. Schnitt, Ilan Wapinski, Tari King. Artificial intelligence powered predictive analysis of atypical ductal hyperplasia from digitized pathology images [abstract]. In: Proceedings of the 2019 San Antonio Breast Cancer Symposium; 2019 Dec 10-14; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2020;80(4 Suppl):Abstract nr P5-02-02.
While computational methods have made substantial progress in improving the accuracy and throughput of pathology workflows for diagnostic, prognostic, and genomic prediction, lack of interpretability remains a significant barrier to clinical integration. In this study, we present a novel approach for predicting clinically-relevant molecular phenotypes from histopathology whole-slide images (WSIs) using human-interpretable image features (HIFs). Our method leverages >1.6 million annotations from board-certified pathologists across >5,700 WSIs to train deep learning models for high-resolution tissue classification and cell detection across entire WSIs in five cancer types. Combining cell- and tissue-type models enables computation of 607 HIFs that comprehensively capture specific and biologically-relevant characteristics of multiple tumors. We demonstrate that these HIFs correlate with well-known markers of the tumor microenvironment (TME) and can predict diverse molecular signatures, including immune checkpoint protein expression and homologous recombination deficiency (HRD). Our HIF-based approach provides a novel, quantitative, and interpretable window into the composition and spatial architecture of the TME.
3130 Background: IMpower150 is a phase 3 study measuring the effect of carboplatin and paclitaxel (CP) combined with atezolizumab (A) and/or bevacizumab (B) in patients with advanced nonsquamous NSCLC, testing the hypothesis that anti-PD-L1 therapy may be enhanced by the blockade of VEGF. Here, we apply a machine-learning based approach to quantify the tumor micro-environment (TME) and vasculature and identify associations with clinical outcome in IMpower150. Methods: Digitized H&E images were registered onto the PathAI research platform (n=1027). Over 200K annotations from 90 pathologists were used to train convolutional neural networks (CNNs) that classify human-interpretable features (HIFs) of cells and tissue structures from images. Blood vessel compression (BVC) indices were calculated using the long versus short axes for each predicted blood vessel. HIFs were clustered to reduce redundancy, and selected features were associated with progression free survival (PFS) within each arm (ABCP, ACP, and BCP) using Cox proportional hazard models. Results: We used the trained CNNs to generate 4,534 features summarizing each patient’s histopathology and TME. After association with survival and correction for multiple comparisons we identified clusters that were significantly associated with survival in at least one arm. Among patients receiving treatments that target PD-L1 (ABCP and ACP), high lymphocyte to fibroblast ratio (LFR) was associated with improved PFS (HR=0.64 (0.51, 0.81), p < 0.001) and showed no significant association with PFS among patients treated with BCP alone (HR=1.13 (0.85, 1.51), p=0.4). Among BCP treated patients, a higher average BVC within the tumor tissue was associated with improved PFS (HR=0.67 (0.50,0.90), p=0.01) and worse PFS among patients treated with ACP (HR=1.50 (1.10,2.06), p=0.009). Conclusions: We developed a deep learning-based assay for quantifying pathology features of the TME and vasculature from H&E images. Application of this system to Impower150 identified an association between high LFR and improved PFS among patients receiving PD-L1 targeting therapy, and between low BVC and improved PFS among patients receiving BCP. These findings support the importance of the TME and vasculature in determining response to PD-L1 and VEGF-targeting therapies.
Abstract The success of targeted or immune therapies is often hampered by the emergence of resistance and/or clinical benefit in only a subset of patients. We hypothesized that combining targeted therapy with immune modulation would show enhanced antitumor responses. Here, we explored the combination potential of erdafitinib, a fibroblast growth factor receptor (FGFR) inhibitor under clinical development, with PD-1 blockade in an autochthonous FGFR2K660N/p53mut lung cancer mouse model. Erdafitinib monotherapy treatment resulted in substantial tumor control but no significant survival benefit. Although anti–PD-1 alone was ineffective, the erdafitinib and anti–PD-1 combination induced significant tumor regression and improved survival. For both erdafitinib monotherapy and combination treatments, tumor control was accompanied by tumor-intrinsic, FGFR pathway inhibition, increased T-cell infiltration, decreased regulatory T cells, and downregulation of PD-L1 expression on tumor cells. These effects were not observed in a KRASG12C-mutant genetically engineered mouse model, which is insensitive to FGFR inhibition, indicating that the immune changes mediated by erdafitinib may be initiated as a consequence of tumor cell killing. A decreased fraction of tumor-associated macrophages also occurred but only in combination-treated tumors. Treatment with erdafitinib decreased T-cell receptor (TCR) clonality, reflecting a broadening of the TCR repertoire induced by tumor cell death, whereas combination with anti–PD-1 led to increased TCR clonality, suggesting a more focused antitumor T-cell response. Our results showed that the combination of erdafitinib and anti–PD-1 drives expansion of T-cell clones and immunologic changes in the tumor microenvironment to support enhanced antitumor immunity and survival.
The rise of multi-million-item dataset initiatives has enabled data-hungry machine learning algorithms to reach near-human semantic classification performance at tasks such as visual object and scene recognition. Here we describe the Places Database, a repository of 10 million scene photographs, labeled with scene semantic categories, comprising a large and diverse list of the types of environments encountered in the world. Using the state-of-the-art Convolutional Neural Networks (CNNs), we provide scene classification CNNs (Places-CNNs) as baselines, that significantly outperform the previous approaches. Visualization of the CNNs trained on Places shows that object detectors emerge as an intermediate representation of scene classification. With its high-coverage and high-diversity of exemplars, the Places Database along with the Places-CNNs offer a novel resource to guide future progress on scene recognition problems.
We propose a general framework called Network Dissection for quantifying the interpretability of latent representations of CNNs by evaluating the alignment between individual hidden units and a set of semantic concepts. Given any CNN model, the proposed method draws on a broad data set of visual concepts to score the semantics of hidden units at each intermediate convolutional layer. The units with semantics are given labels across a range of objects, parts, scenes, textures, materials, and colors. We use the proposed method to test the hypothesis that interpretability of units is equivalent to random linear combinations of units, then we apply our method to compare the latent representations of various networks when trained to solve different supervised and self-supervised training tasks. We further analyze the effect of training iterations, compare networks trained with different initializations, examine the impact of network depth and width, and measure the effect of dropout and batch normalization on the interpretability of deep visual representations. We demonstrate that the proposed method can shed light on characteristics of CNN models and training methods that go beyond measurements of their discriminative power.