Abstract Introduction: Lung adenocarcinoma (LUAD) is the most common non-small cell lung cancer, especially in never-smokers. Environmental pollution, particularly airborne particulates (PM2.5/PM5), is a critical factor driving lung cancer initiation [1]. However, reliable individual-level pollution exposure quantification is challenging. We recently introduced the lung pollutant index (LPI), an AI-derived metric quantifying tissue-resident pollutant burden from pathology images. However, pathology-based LPI is invasive and unsuitable for large-scale application. We propose a machine-learning framework to predict CT-based LPI (CT-LPI) from chest CT, enabling non-invasive pollutant burden quantification for individuals undergoing chest imaging, including high-risk smokers and never-smokers with incidental nodules. Methods: We retrospectively investigated 153 LUAD patients who received preoperative lung CT at MD Anderson Cancer Center (IRB 2023-0114) with surgical pathology. LPI computed from digitalized H&E images [2] categorized cohorts into LPI-high (n=61) and LPI-low (n=92). Region-wise radiomics features (793 per region) were extracted from 5-mm peritumoral ring, normal lung, and whole lung, including first-order, texture, shape, Laplacian-of-Gaussian, wavelet, and habitat features. We incorporated COPD-associated radiomic markers to enhance biological interpretability. We built a multi-regional ensemble framework with features selected using mutual information and Elastic-Net. Region-specific classifiers (Ridge Logistic, Gradient Boosting, CatBoost) were trained, with final predictions via weighted-average ensemble [3]. Results: In 5-fold cross-validation, the multi-regional ensemble achieved AUC 0.719 and ACC 0.687, improving AUC by 0.048 over the best single-regional model and outperforming simple averaging (0.710), LogisticNet (0.664), and ElasticNet (0.661). Incorporating COPD-associated markers improved AUC to 0.724. Conclusion: We developed CT-LPI, predicting tissue-resident pollutant burden from routine CT using complementary multi-regional features. This framework non-invasively quantifies individual exposure to environmental carcinogens implicated in lung cancer initiation. By enabling scalable monitoring of pollution-driven biological alterations, CT-LPI may support environmental exposure assessment and personalized risk stratification. Next steps include external validation, application to screening datasets, and integration with other biomarkers. [1] Hill W, et al. Nature. 2023. [2] Pan et al., Nature Cancer, under review. [3] Shaheen A, et al. Front Neurosci. 2022. Citation Format: Yutong Li, Xiaoxi Pan, Abishek Balachandra, Chingyi Young, Maria Esther Salvatierra, Carmen Behrens, Luisa Maren Solis Soto, Yinyin Yuan, Chengyue Wu, . Multi-regional CT-based radiomics fusion predicts pathological pollutant index associated with lung adenocarcinoma [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2773.
Whole slide images (WSIs) enable weakly supervised prognostic modeling via multiple instance learning (MIL). Spatial transcriptomics (ST) preserves in situ gene expression, providing a spatial molecular context that complements morphology. As paired WSI-ST cohorts scale to population level, leveraging their complementary spatial signals for prognosis becomes crucial; however, principled cross-modal fusion strategies remain limited for this paradigm. To this end, we introduce PathoSpatial, an interpretable end-to-end framework integrating co-registered WSIs and ST to learn spatially informed prognostic representations. PathoSpatial uses task-guided prototype learning within a multi-level experts architecture, adaptively orchestrating unsupervised within-modality discovery with supervised cross-modal aggregation. By design, PathoSpatial substantially strengthens interpretability while maintaining discriminative ability. We evaluate PathoSpatial on a triple-negative breast cancer cohort with paired ST and WSIs. PathoSpatial delivers strong and consistent performance across five survival endpoints, achieving superior or comparable performance to leading unimodal and multimodal methods. PathoSpatial inherently enables post-hoc prototype interpretation and molecular risk decomposition, providing quantitative, biologically grounded explanations, highlighting candidate prognostic factors. We present PathoSpatial as a proof-of-concept for scalable and interpretable multimodal learning for spatial omics-pathology fusion.
BACKGROUND:Cancer is a systemic disease with most deaths attributed to metastatic burden. Primary and metastatic tumors, albeit at different anatomic locations, are interconnected through multiple biological processes. Pre-clinical and clinical observations of growth acceleration of metastases after surgery, or abscopal effects outside the radiation field are widely reported, yet reliably triggering favorable and avoiding unfavorable systemic responses remains an unmet clinical need. Understanding local and systemic tumor interaction dynamics will help guide future treatments. METHODS:We analyze the data of multiple in vivo tumor models. We formalize the systemic interplay of tumors as mathematical differential equation and calibrate parameters for each cell line and mouse type. Using model selection metrics, we identify classic tumor growth models with a novel shared carrying capacity parsimoniously describe the pan-cancer experimental data. RESULTS:Shared systemic carrying capacity, metastatic spread potential, and metastatic growth rates differ across tested cell lines and mouse strains. Bi-directional concomitant systemic interconnectivity explains the observed metastatic explosion after primary tumor surgery. DISCUSSION:Future investigations should reproduce this analysis in clinical settings and evaluate whether this shared carrying capacity model could help stratify patients at risk of metastatic disease below clinical detectability and inform strategies to control oligometastatic cancer.
Xenium, a new spatial transcriptomics platform, enables subcellular-resolution profiling of complex tumor tissues. Despite the rich morphological information in histology images, extracting robust cell-level features and integrating them with spatial transcriptomics data remains a critical challenge. We introduce CellSymphony, a flexible multimodal framework that leverages foundation model-derived embeddings from both Xenium transcriptomic profiles and histology images at true single-cell resolution. By learning joint representations that fuse spatial gene expression with morphological context, CellSymphony achieves accurate cell type annotation and uncovers distinct microenvironmental niches across three cancer types. This work highlights the potential of foundation models and multimodal fusion for deciphering the physiological and phenotypic orchestration of cells within complex tissue ecosystems.
Background:Improved cancer risk stratification is needed to differentiate high-risk individuals with Barrett's esophagus (BE) from low-risk populations to reduce overtreatment and improve outcome. The evolution of BE towards adenocarcinoma is likely driven by a combination of genomic and microenvironmental factors, yet existing predictive models rarely integrate both using routine specimens. Method:We developed BEACON (Barrett Esophagus DNA content Abnormality and immune ecology for Cancer Outcome), a spatially aware framework predicting DNA content abnormalities and characterizing immune spatial ecology from routine histopathology. First, using 777 BE biopsies with flow cytometry-based DNA content data scanned at two institutions, we trained and tested DACOR (DNA content abnormality recognition), a multi-instance learning model that predicts DNA content abnormalities from histopathology. Next, complementary models for cell classification and tissue segmentation enabled spatial immune ecology metric computations. Lastly, a logistic regression model integrated molecular immune ecological features and epithelial morphology for cancer risk stratification. Results:DACOR achieved 0.825 AUC in the test cohort for DNA content abnormality prediction. DNA content abnormal regions exhibited increased lymphoplasma cellular inflammation versus normal regions (p=0.006). Patients classified as DNA content abnormal by DACOR demonstrated increased cancer progression (p=0.0001). Among patients with DNA content abnormality, cancer progressors exhibited increased plasma cell clustering adjacent to abnormal epithelium compared to non-progressors. The integrated risk classification model stratified DNA content abnormal patients into high- and low-risk groups with 0.817 AUC. Conclusion:BEACON spatially integrates molecular abnormality with immune spatial ecology to stratify BE patients by cancer progression risk using routine pathology images. This scalable, explainable approach could improve clinical decision-making and reduce unnecessary surveillance in low-risk patients.
Background: The tumor immune microenvironment is crucial in shaping the response to neoadjuvant chemotherapy in triple-negative breast cancer (TNBC). Stromal tumor-infiltrating lymphocytes (sTILs) and Ki-67 are important biomarkers to predict treatment response in TNBC (Abuhadra et al., 2023). However, manual assessment of sTILs faces challenges due to variability in interpretation across different observers (Van Bockstal et al. 2021). To obtain a consistent predictive model for treatment response with chemotherapy in TNBC, we utilized artificial intelligence (AI) methods for enhancement. Material & methods: We collected 408 hematoxylin and eosin (H&E)-stained images from ARTEMIS (NCT02276443) arm 1 pretreatment core biopsies, divided into discovery and validation cohorts as per a previous work (Abuhadra et al., 2023). We also used the same criteria to evaluate pathologic complete response (pCR), Ki-67, and sTILs. H&E images exhibiting metaplasia and giant cells were excluded, resulting in 201 slides for the discovery cohort and 193 slides for the validation cohort, with one slide per case. To calculate the TILs score, we applied AI pipelines on H&E images to identify epithelial, lymphoid, stromal, and other cells (AbdulJabbar et al., 2020), and to segment tissue into tumor, stroma, parenchyma, necrosis/hemorrhage, and adipose tissue (unpublished). To ensure the accuracy of cell classification, we combined the two pipelines to refine cell recognition. Specifically, cell types identified outside the tumor area but misclassified as epithelial cells underwent a secondary prediction. AI-derived TILs (AI-TILs) was calculated as the proportion of identified lymphoid cells within the tumor and stromal tissues, excluding other cell types. Results: To evaluate the AI pipelines, 20 images from discovery cohort were manually annotated by three pathologists, resulting in 8010 cells. We achieved an average balanced accuracy of 91.2% across the identified cell types. In the discovery cohort (n=201; pCR rate 42%, 85/201), AI-TILs was notably associated with manual sTILs (Spearman’s rho=0.49, P<0.001). In a multivariable model consisting of AI-TILs and Ki-67 to predict pCR, with AI-TILs cutoff (0.173, range: 0.03-0.73) established through recursive partitioning analysis, we found that higher AI-TILs levels were significantly associated with pCR (Odds ratio, 4.57, 95% CI=1.63-12.80, P=0.004). Manual sTILs showed a similar pattern (Odds ratio, 3.71, 95% CI=1.97-6.99, P<0.001) in a multivariable model with Ki-67. When combining Ki-67, sTILs and AI-TILs, the significant association between AI-TILs and pCR was retained (Odds ratio: 3.01, 95% CI=1.04-8.72, P=0.042). Using the same model, we achieved an AUC of 0.74 and a precision of 0.68, marginally improving the model consisting of Ki-67 and sTILs (AUC 0.71, precision 0.68). In the validation cohort (n=193; pCR rate 42%, 82/193), the AI-TILs score performed consistently, maintaining correlation with sTILs (rho=0.39, P<0.001). By employing the model constructed in the discovery dataset with Ki-67, sTILs, and AI-TILs to predict pCR, we achieved an AUC of 0.69 and a precision of 0.67, comparable to the model with Ki-67 and sTILs alone (AUC 0.68, precision 0.66). Calibration plots and Hosmer-Lemeshow test indicated that the model including AI-TILs aligned better with the actual data compared to the model without AI-TILs. Conclusion: This study demonstrates that AI-TILs from baseline biopsies can be a promising biomarker to enhance the prediction for neoadjuvant chemotherapy response in TNBC, alongside existing predictors such as manual sTILs, underscoring the potential clinical value of AI-TILs due to its reproducibility and objectivity. Citation Format: Xiaoxi Pan, Caner Ercan, Zhongya Wang, Roland Bassett Jr., Karina B Pinao Gonzales, Clinton Yam, Lei Huo, Yinyin Yuan. Al-Derived Tumor-Infiltrating Lymphocytes Enhance the Model with Baseline Stromal Tumor-Infiltrating Lymphocytes and Ki-67 in Predicting Pathologic Complete Response in an Early-Stage Triple-Negative Breast Cancer Prospective Trial [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2024; 2024 Dec 10-13; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(12 Suppl):Abstract nr PS1-07.
The abundance of tumor-infiltrating lymphocytes (TILs) has demonstrated prognostic value in breast cancer recurrence. However, manual TIL assessment is both time-consuming and prone to inter-observer variability. This study aimed to evaluate the performance of two AI-based TIL scoring pipelines, emphasizing the added value of incorporating tissue-based context insights for automated TILs scoring. We compared two in-house AI-based TILs scoring pipelines using the Translational Breast Cancer Research Consortium DCIS cohort. The first pipeline employs single-cell classification to identify epithelial, lymphocyte, stromal, and other cell types, followed by spatial analysis of immune populations relative to epithelial cells through clustering (AI-TILs-cluster). The second pipeline builds on this foundation by integrating a tumor microenvironment segmentation model as context guidance to compute the proportion of lymphoid cells within the epithelium and stromal compartments (AI-TILs-seg). In a cohort of 45 retrospectively collected hematoxylin and eosin (HE)-stained whole-slide tissue sections from 38 patients with available pathologists manual TILs, we evaluated the correlation of both scores and the median manual TIL score across the pathologists. Additionally, we assessed the prognostic value of the AI-derived TILs in 232 patients (530 HE slides) with up to 228 follow-up months (median: 78 months) by examining the association of per-patient AI-TILs with time to invasive breast cancer recurrence. This analysis was conducted using a multivariate Cox Proportional Hazards model that included patient age, tumor grade, DCIS lesion size, and estrogen receptors (ER) and progesterone receptors (PR) status. AI-TILs-cluster was not significantly correlated with median pathologist TILs, AI-TILs-seg, however, demonstrated a moderate correlation (Spearman’s rho = 0.43, p = 0.003). This finding was expected as the automatic scores were not designed to replicate the manual TILs. AI-TILs-seg was associated with an increased risk of invasive breast cancer (Hazard Ratio = 1.067, 95% Confidence Interval: 1.017-1.119, p= 0.008) independent of other variables. In comparison, the multivariate analysis including AI-TILs-cluster did not show significance in predicting the risk of invasive breast cancer. These findings suggest that incorporating contextual tumor microenvironment segmentation can enhance AI-derived TIL scoring, leading to better alignment with pathologist assessments and improved clinical relevance. Future work will focus on validating this scoring approach in larger cohorts and improving correlation with pathologist TILs. Sara Ranjbar, Xiaoxi Pan, Karina Pinao, Caner Ercan, Roberto Salgado, Hugo M. Horlings, Allison Hall, Lorraine M. King, Carlo Maley, E Shelly Hwang, Yinyin Yuan, Simon Castillo, Pingjun Chen. Al-derived tumor-infiltrating lymphocytes predicts risk of invasive breast cancer recurrence in ductal carcinoma in situ (DCIS) [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 6270.
Barrett’s esophagus (BE) is a precancerous condition characterized by the metaplastic replacement of esophageal epithelium in response to gastroesophageal reflux. Although the progression rate to invasive carcinoma is low, identifying high-risk patients remains a critical unmet need, especially as BE prevalence rises. While molecular and immune microenvironment changes associated with BE progression are documented (Nowicki-Osuch 2021), there is a gap in reliable pathology-level indicators to assess cancer risk. This study aims to uncover cancer-relevant changes in BE using computational pathology approaches. We analyzed a unique dataset comprising 742 esophageal biopsy slides from 138 patients, and with DNA content abnormalities, including aneuploid (n=125 slides from 26 patients), determined by flow cytometry. Clinical follow-up data for cancer progression was available for 117 of these patients. Using computational pathology models, we trained a multiple-instance learning framework for aneuploid prediction (Lu 2021, Ercan 2024) on 45 slides, and a single-cell detection model based on the ACformer architecture (Huang 2023), fine-tuned on 13, 343 cells. Spatial immune cell distribution within the epithelium was quantified using the Morisita-Horn index (MHI), which measures spatial overlap between immune cells and BE cells. The MIL-based aneuploid prediction model achieved a balanced accuracy of 74.3% and an AUC of 0.81 on the test set (n= 388), while the single-cell detection & classification model reached accuracies of 73.7% for detection and 98.3% for classification (in epithelium, lymphocyte, plasma cell, eosinophil, neutrophil and stromal cell classes). Immune cell MHI was significantly elevated in aneuploid biopsies (p < 10e-4, n=125). Aneuploid was found to strongly correlate with cancer progression (p < 10e-5, n=33). In aneuploid samples, immune cell infiltration analysis revealed that increased immune cell abundance and higher plasma cell MHI scores were significantly associated (p < 10e-4, < 10e-4, respectively) with cancer progression in the vicinity of aneuploid BE epithelium, consistent with the observations in (MK Strasser 2023). This study demonstrates the potential of Patho-Omic tools to predict aneuploid in BE biopsies and identify cancer progression-associated features directly from routine pathology slides. By combining slide-level aneuploid prediction with single-cell detection, we highlight the role of immune cell spatial dynamics in cancer progression. These findings lay the groundwork for stratifying BE patients by progression risk through biopsy image analysis. Caner Ercan, Xiaoxi Pan, Thomas G. Paulson, Matthew D. Stachler, Carlo C. Maley, Yinyin Yuan. Path-Omics: Elevated plasma cell infiltration into aneuploid epithelium predicts Barrett’s esphagus progression to adenocarcinoma [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7509.
High-plex digital spatial profiling (DSP) enables mapping protein expression onto tissue context to study heterogeneity of tumor microenvironment (TME). However, current techniques rely on manual annotations to select representative regions for profiling, introducing subjectivity and uncertainty. We present a deep-learning approach to predict spatial proteomics associated with immune response and proliferation in carcinomas from tissue space to high-throughput protein expression analysis across the full sections while preserving spatial context. Our study utilized a diverse dataset of 62 carcinomas and adenocarcinomas cases, with major types including Papillary Urothelial Carcinoma (25) and Colorectal Adenocarcinoma (5) among others, totaling 721 regions of interest (ROIs). Using morphology markers (panCK, CD45) for mIF images and 12 proteins for profiling, we conducted a two-phase study. First, YOLOv8 identified epithelial (panCK+) and immune cells (CD45+) from mIF images and correlated their abundance with protein expression. Second, we developed a multi-instance learning framework using Swin Transformer to predict protein expression from mIF images, processing 256×256 pixel adjacent patches (average 1, 947 per slide, total 120, 729) with attention mechanisms. The dataset was split into training (47), validation (9), and test (6) sets. We observed strong correlations between cell abundance from mIF images and protein expression, (panCK: rho=0.624, p<0.001; CD45: rho=0.676, p<0.001), validating the effectiveness of the YOLOv8 model in cell recognition. The multi-instance model demonstrated promising prediction performance, marginally improving the correlations for panCK (rho=0.701, p<0.001) and CD45 (rho=0.699, p<0.001) when compared with YOLOv8. For markers included in the protein profiling DSP panel but not for the morphology, our model captured the signals for immune checkpoint and proliferation markers: PD-L1 (rho=0.741, p<0.001), Ki-67 (rho=0.730, p<0.001). Moderate correlations were observed for stromal and immune cell markers (FAP-alpha: rho=0.569, p<0.001; CD163: rho=0.465, p=1.45e-03; CD3: rho=0.409, p=2.02e-02; CD4: rho=0.381, p=7.46e-03; FOXP3: rho=0.344, p=3.24e-02) while a weak correlation for CD8 (rho=0.210, p=4.83e-02). These findings suggest that cellular morphology and spatially encoded information in the TME can be learned by our model to predict immune checkpoint expression and proliferation. Our integration of mIF images with spatial proteomics enables high-throughput protein expression prediction from limited markers, potentially extending to H&E images. This could reduce sequential biopsies, enabling real-time treatment monitoring. The model's PD-L1 prediction capabilities provide insights into tumor-immune interactions to guide immunotherapy decisions. Yasin Shokrollahi, Tanishq Gautam, Alejandra Serrano, Simon P. Castillo, Pingjun Chen, Karina Pinao, Maria Esther Salvatierra, B. Leticia Rodriguez, Patient Mosaic Team, Luisa M. Solis Soto, Yinyin Yuan, Xiaoxi Pan. Artificial intelligence for predicting spatial proteomics using high-plex digital spatial profiling in carcinomas [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2428.
The tumor microenvironment (TME) plays a key role in lung cancer progression. Spatial transcriptomics offers insights into tumor heterogeneity but relies on integration with histological features to decode tumor transcriptomic programs, which is challenging due to TME complexity. To overcome this, we developed AI models that accurately identify 10 tissue types from histology images, streamlining TME analysis. We used public annotations to train AI models, DeepLabV3+ and Segformer, for TME segmentation. The training set included hematoxylin&eosin-stained slides from breast and lung cancers in TCGA, combined with newly curated annotations from TCGA and four internal slides of surgically resected tumors, totaling 3403 patches (breast: 3110; lung: 293) at 0.44 μm/pixel. The testing set included 2000 patches from lung cancer slides at 1 μm/pixel. The models were trained to classify 10 tissue types: tumor, stroma, inflammatory, necrosis/hemorrhage, adipose, bronchi epithelium, vessel, macrophage area, alveoli, and muscle. For comparison, predictions of adipose, vessel, and muscle were grouped into stroma, while inflammatory and macrophage areas were grouped as immune. Model performance was evaluated using the Dice score. We further validated the models using spot-level manual annotations from Visium data for four lung cancer samples. Pathologists labeled spots as tumor, stroma, immune aggregates, macrophage, bronchi epithelium, vessel, and alveoli using Loupe Browser. The type with the highest number of predicted pixels by AI within a spot-based patch was assigned as the predicted label. Concordance was assessed using Fleiss's kappa. At pixel level, Segformer outperformed DeepLabV3+, achieving an average Dice score of 0.856 and 0.809. At spot level, kappa index between pathologist and Segformer was 0.28-0.66, showing fair (n=2) to good (n=1) agreement, consistent with the concordance between pathologist and DeepLabV3+ (0.34-0.722). Kappa index between Segformer and DeepLabV3+ showed good (n=3) to perfect (n=1) agreements (0.677-0.842). Interestingly, the first annotated sample showed good agreement between the pathologist and AI models (0.66-0.722), whereas the last annotated sample demonstrated only fair agreement (0.28-0.343). Notably, this discrepancy was not observed when comparing Segformer and DeepLabV3+ directly, implying a high reproducibility for AI models. In terms of efficiency, the pathologist required about five hours in theory to annotate a slide containing around 12000 spots, while AI models, running on an A100 GPU, completed the segmentation in about 20 minutes. Our study shows the effectiveness and efficiency of AI models in recognizing the TME from histology images in lung cancer. These models streamline the integration of tissue histopathological features with Visium assay and facilitate the discovery of image-derived biomarkers. Xiaoxi Pan, Maria E. Salvatierra, Caner Ercan, Lakshmi Kakarala, Wei Lu, Ou Shi, Idania C. Lubo Julio, Ignacio I. Wistuba, Luisa M. Solis Soto, Yinyin Yuan. TMEseg: Connecting histopathology with spatial transcriptomics through tumor microenvironment segmentation for lung cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2426.
Barrett’s esophagus (BE) is the sole precursor to esophageal adenocarcinoma (EAC), and is an opportunity for developing biomarkers for cancer risk assessment. DNA content abnormalities, including aneuploidy, have been implicated in the progression to EAC in BE patients, but molecular assays require valuable tissue for its detection. We propose utilizing images from routine histology to detect ploidy status using deep learning. Employing a weakly supervised deep learning approach, multi-instance learning (MIL), we trained a model to predict ploidy using hematoxylin and eosin-stained whole slide images of endoscopic biopsies and flow cytometry results. The study introduces a novel data augmentation method for MIL, sequentially altering features from original and augmented images during training loops. This method improved the average area under curve (AUC) from 0.43, 0.64 and 0.81 for ResNet50, DenseNet121 and REMEDIS foundation model, respectively (training without any augmentation), to 0.61, 0.87 and 0.91 with the proposed augmentation strategy. The top-performing model, using REMEDIS foundation model as the backbone, achieved 0.93 AUC and 0.83 balanced accuracy to predict aneuploidy in the test cohort biopsies (n=279). Across all the patients (n=123), predicted aneuploidy status was correlated with progression to EAC (p=6.55e-06), similar to correlation with ploidy status based on flow cytometry results (p=2.84e-7). Supporting the findings, histologic nuclear features typically associated with dysplasia and DNA content abnormalities such as enlarged, hyperchromatic nuclei and loss of nuclear polarity, were seen in the samples called abnormal compared to the control diploid samples. In conclusion, our deep learning model efficiently predicts aneuploidy, a mechanism that has been shown to underpin BE progression to EAC. This method, preserving precious biopsy tissues, complements routine histology, offering potential for identifying individuals at high risk of progression through molecular-based advancements.
Abstract Background Barrett’s esophagus (BE) is the only known precursor of esophageal adenocarcinoma (EAC) and there is a need for biomarkers in BE for risk stratification for cancer progression. Aneuploidy has been suggested as a factor in the development, initiation or progression of EAC in patients with BE, and it has been found predictive for EAC progression (Hadjinicolaou, et al. 2020, Sikkema, et al. 2009). However, current assays for detecting aneuploidy require valuable tissue, hence validation studies are limited. We hypothesise that routine histology images can be used to detect ploidy status. We aim to develop a sensitive, accurate, and easy-to-use deep learning tool for this purpose. Methods We used a weakly supervised deep learning-based approach called clustering-constrained-attention multiple-instance learning (CLAM) (Lu et al., 2021) to detect aneuploidy on hematoxylin and eosin-stained whole slide images of endoscopical biopsies from BE, for which the ploidy status had been determined by flow cytometry. To benchmark the model’s performance, we also trained a traditional fully supervised algorithm, ResNet50 (He et al., 2016), with the same dataset. The models were trained on 388 slides (51 aneuploid) from the Seattle Barrett’s Esophagus Annotated Resource (BEAR) and then applied to an independent test cohort of BEAR patients, consisting of 279 slides (36 aneuploid). Results The multi-attention-branch CLAM model achieved AUC of 0.82 and 76.4% balanced accuracy for aneuploidy on the internal test subset (10% of the cohort). On the independent test dataset, the model achieved AUC of 0.85 and 79.3% balanced accuracy, while performance of the fully supervised model was AUC of 0.65 and 33.1 balanced accuracy. In a challenging subset for image-based diagnosis, including 253 samples (29 aneuploid) without dysplasia or atypical mitosis (a major feature of aneuploidy samples) the model achieved a 67.6% accuracy and an AUC of 0.83. Both the flow results (p=2.84e-7) and the model's predictions (p=2.18e-3) revealed a correlation between progression to EAC and ploidy status. Conclusion We developed a deep learning model for predicting aneuploidy in BE biopsies. The weakly-supervised approach can perform much better compared to the traditional fully-supervised model. Aneuploidy is correlated with cancer progression and stands as a promising candidate for the prediction of cancer progression in BE patients. Although atypical mitosis is the primary histological indicator of aneuploidy, our model can still make aneuploid predictions in the absence of atypical mitosis, as it doesn't rely solely on a single parameter. Our model is efficient, capable of processing hundreds of samples resource-effectively, and thus an ideal adjunct to standard histologic evaluation. This classifier may facilitate molecular-based improvements in identifying individuals at high risk of progression. Citation Format: Caner Ercan, Xiaoxi Pan, Thomas G. Paulson, Matthew D. Stachler, Carlo C. Maley, Yinyin Yuan. A novel approach to Barrett's esophagus risk stratification: Whole-slide image analysis for aneuploidy detection [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 6184.
Selecting regions of interest (ROIs) in whole-slide histology images (WSIs) is a crucial step for spatial molecular profiling. As a general practice, pathologists manually select ROIs within each WSI based on morphological tumor markers to guide spatial profiling, which can be inconsistent and subjective. To enhance reproducibility and avoid inter-pathologist variability, we introduce a novel immune-guided end-to-end pipeline to automate the ROI selection in multiplex immunofluorescence (mIF) WSIs stained with three cell markers (Syto13, CD45, PanCK). First, we estimate immune infiltration (CD45 ^+ expression) scores at the grid level in each WSI. Then, we incorporate the Pathology Language and Image Pre-Training (PLIP) foundational model to extract features from each grid and further select a subset of grids representative of the whole slide that comparatively matches pathologists’ assessment. Further, we implement state-of-the-art detection models for ROI detection in each grid, incorporating learning from pathologists’ ROI selection. Our study shows a significant correlation between our automated method and pathologists’ ROI selection across five different types of carcinomas, as evidenced by a significant Spearman’s correlation coefficient (> 0.785, p < 0.001), substantial inter-rater agreement (Cohen’s κ > 0.671), and the ability to replicate the ROI selection made by independent pathologists with excellent average performance (0.968 precision and 0.991 mean average precision at a 0.5 intersection-over-union). By minimizing manual intervention, our solution provides a flexible framework that potentially adapts to various markers, thus enhancing the efficiency and accuracy of digital pathology analyses.
Accurate cell detection in multiplex immunofluorescence (mIF) is crucial for quantifying and analyzing the spatial distribution of complex cellular patterns within the tumor microenvironment. Despite its importance, cell detection in mIF is challenging, primarily due to difficulties obtaining comprehensive annotations. To address the challenge of limited and unevenly distributed annotations, we introduced a streamlined semi-supervised approach that effectively leveraged partially pathologist-annotated single-cell data in multiplexed images across different cancer types. We assessed three leading object detection models, Faster R-CNN, YOLOv5s, and YOLOv8s, with partially annotated data, selecting YOLOv8s for optimal performance. This model was subsequently used to generate pseudo labels, which enriched our dataset by adding more detected labels than the original partially annotated data, thus increasing its generalization and the comprehensiveness of cell detection. By fine-tuning the detector on the original dataset and the generated pseudo labels, we tested the refined model on five distinct cancer types using fully annotated data by pathologists. Our model achieved an average precision of 90.42%, recall of 85.09%, and an F1 Score of 84.75%, underscoring our semi-supervised model's robustness and effectiveness. This study contributes to analyzing multiplexed images from different cancer types at cellular resolution by introducing sophisticated object detection methodologies and setting a novel approach to effectively navigate the constraints of limited annotated data with semi-supervised learning.
The introduction of the International Association for the Study of Lung Cancer grading system has furthered interest in histopathological grading for risk stratification in lung adenocarcinoma. Complex morphology and high intratumoral heterogeneity present challenges to pathologists, prompting the development of artificial intelligence (AI) methods. Here we developed ANORAK (pyrAmid pooliNg crOss stReam Attention networK), encoding multiresolution inputs with an attention mechanism, to delineate growth patterns from hematoxylin and eosin-stained slides. In 1,372 lung adenocarcinomas across four independent cohorts, AI-based grading was prognostic of disease-free survival, and further assisted pathologists by consistently improving prognostication in stage I tumors. Tumors with discrepant patterns between AI and pathologists had notably higher intratumoral heterogeneity. Furthermore, ANORAK facilitates the morphological and spatial assessment of the acinar pattern, capturing acinus variations with pattern transition. Collectively, our AI method enabled the precision quantification and morphology investigation of growth patterns, reflecting intratumoral histological transitions in lung adenocarcinoma.
Histologic growth patterns are associated with patient prognosis, thus recognized as an important part of the WHO classification in lung adenocarcinoma (Travis et al. 2015, Moreira et al., 2020). The wide spectrum of growth patterns proves challenging for reproducible and quantitative scoring. Currently, scoring is based on manual identification of the predominant pattern and percentages of patterns in routine diagnostic slides. The lack of an automated method also limits our ability to investigate the immune microenvironment of growth patterns. To overcome the above challenges, we present a deep learning method, Pyramid Stream Networks, to precisely segment growth patterns at pixel level. Unlike existing methods, the proposed method captures different spatial scales of the histology information by novel attention strategies at different learning stages. This problem-oriented design yields precise boundaries for each pattern, enabling the investigation of growth pattern heterogeneity, and the relationship with tumor microenvironment components. Experiments were conducted on 49 haematoxylin and eosin whole slide images (WSIs) from TRACERx 100 cohort (AbdulJabbar et al., 2020). Each WSI was sparsely annotated by 3 senior pathologists. A total of 2968 annotated patches were split into 5 folds for cross validation. We compared our method with two state-of-the-art methods applied in semantic segmentation, attention U-net (Oktay et al. 2018) and DeepLabV3+ (Chen et al. 2018). When evaluated at patch level, our method outperformed the better comparison method, DeepLabV3+, by 3.43% and 2.99% in pixel-wise Dice and overall precision (OP) (Dice: 60.34% vs. 56,91%, OP: 65.43% vs. 62.44%). When applied to WSIs, the model correctly predicted the predominant pattern for 38 out of 49 samples, achieving an accuracy of 77.55%. Interestingly, in the 11 discordant cases, 10 showed high intra-tumor heterogeneity of growth patterns, measured by Shannon diversity, highlighting the impact of intra-tumor heterogeneity on growth pattern assessment. Additionally, we combined the identified growth patterns with lymphocytic distribution measured in (AbdulJabbar et al., 2020) and revealed a significantly increased immune infiltration in proximity to the solid pattern as compared to others, which is in line with previous findings (Tavernari et al., 2021). In summary, by leveraging image-analysis and artificial intelligence techniques, we propose a new method for precise growth pattern segmentation from routine histology samples of lung adenocarcinoma. It provides quantitative and reproducible scores of growth patterns, which can be developed into a decision support system for pathologists and clinicians. Furthermore, through pattern-specific spatial mapping, it enables future studies of intra-tumor heterogeneity, such as the preferential infiltration of lymphocyte subsets adjacent to diverse growth patterns. Citation Format: Xiaoxi Pan, Hanyun Zhang, Anca-Ioana Grapa, Khalid AbdulJabbar, Shan E. Ahmed Raza, HO KWAN ALVIN CHEUNG, Takahiro Karasaki, John Le Quesne, David A. Moore, Charles Swanton, Yinyin Yuan. Precise segmentation of growth patterns in TRACERx lung adenocarcinoma [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 5055.