Background Recent advances in computational pathology enables AI-assisted diagnosis and risk stratification of breast cancer. This advance in technology will reduce the inconsistent reporting of breast cancer grading using Nottingham Histologic grading system. This study evaluated the implementation of the DeepGrade model for breast cancer grading using H&E-stained slides from breast cancer patients in Ethiopia. Objective To assess the accuracy, specificity, and sensitivity of the DeepGrade model in distinguishing between grade 1 and grade 3 tumours. Additionally, the study aimed to explore the model's ability to further classify Nottingham Histologic Grade 2 tumours into two distinct risk categories. Methods A retrospective analysis was conducted using data from 200 tumour samples from the Department of Pathology, Tikur Anbessa Specialized Hospital, Addis Ababa, Ethiopia. The performance of the DeepGrade model was compared with three pathologists' diagnosis using metrics such as specificity, sensitivity, area under the curve and agreement level. Results The DeepGrade model reached a 100% specificity for low-grade tumours and an 82.05% sensitivity for high-grade tumours, with an AUC of 0.914 and an agreement level of 86.79%. Our findings illustrated the model's strong agreement with the pathologist, and a Kappa coefficient of 0.71 (95% CI: 0.51-0.90). Conclusion The study showed the potential significance of DeepGrade model utilization in enhancing breast cancer grading practices in resource-limited settings. By adopting this model, a consistent and standardized grading system for breast cancer can be established, significantly enhancing the effectiveness of breast cancer management.
Pathology foundation models (PFMs) have recently emerged as powerful pretrained encoders for computational pathology, enabling transfer learning across a wide range of downstream tasks. However, systematic comparisons of these models for clinically meaningful prediction problems remain limited, especially in the context of survival prediction under external validation. In this study, we benchmark widely used and recently proposed PFMs for breast cancer survival prediction from whole-slide histopathology images. Using a standardized pipeline based on patch-level feature extraction and a unified survival modeling framework, we evaluate model representations across three independent clinical cohorts comprising more than 5,400 patients with long-term follow-up. Models are trained on one cohort and evaluated on two independent external cohorts, enabling a rigorous assessment of cross-dataset generalization. Overall, H-optimus-1 achieves the strongest survival prediction performance. More broadly, we observe consistent generational improvements across model families, with second-generation PFMs outperforming their first-generation counterparts. However, absolute performance differences between many recent PFMs remain modest, suggesting diminishing returns from further scaling of pretraining data or model size alone. Notably, the compact distilled model H0-mini slightly outperforms its larger teacher model H-optimus-0, despite using fewer than 8
Estrogen receptor-positive (ER+), HER2-negative (HER2-) breast cancer (BC) is the most common BC subtype. Patients with this subtype rarely achieve pathological complete response (pCR) after neoadjuvant chemotherapy (NACT), and the population most likely to benefit remains unclear. This highlights the need to identify reliable markers of response, particularly in the context of emerging chemoimmunotherapy strategies. We retrospectively included 415 patients diagnosed with ER+/HER2- BC between 2016 and 2020 from three Danish pathology departments. We analysed associations between standard clinicopathological variables and computationally assessed stromal tumor-infiltrating lymphocytes (sTILs) with response rates and long-term outcomes. To define clinically useful cut-offs, we performed threshold analysis for ER-expression and sTIL levels. The pCR rate in the study population was 6.3
Pathology foundation models (PFMs) have become central to computational pathology, aiming to offer general encoders for feature extraction from whole-slide images (WSIs). Despite strong benchmark performance, PFM robustness to real-world technical domain shifts, such as variability from whole-slide scanner devices, remains poorly understood. We systematically evaluated the robustness of 14 PFMs to scanner-induced variability, including state-of-the-art models, earlier self-supervised models, and a baseline trained on natural images. Using a multiscanner dataset of 384 breast cancer WSIs scanned on five devices, we isolated scanner effects independently from biological and laboratory confounders. Robustness is assessed via complementary unsupervised embedding analyses and a set of clinicopathological supervised prediction tasks. Our results demonstrate that current PFMs are not invariant to scanner-induced domain shifts. Most models encode pronounced scanner-specific variability in their embedding spaces. While AUC often remains stable, this masks a critical failure mode: scanner variability systematically alters the embedding space and impacts calibration of downstream model predictions, resulting in scanner-dependent bias that can impact reliability in clinical use cases. We further show that robustness is not a simple function of training data scale, model size, or model recency. None of the models provided reliable robustness against scanner-induced variability. While the models trained on the most diverse data, here represented by vision-language models, appear to have an advantage with respect to robustness, they underperformed on downstream supervised tasks. We conclude that development and evaluation of PFMs requires moving beyond accuracy-centric benchmarks toward explicit evaluation and optimisation of embedding stability and calibration under realistic acquisition variability.
In triple-negative and HER2-positive breast cancer, tumor-infiltrating lymphocytes (TILs) are established predictive and prognostic biomarkers, but their role in luminal subtypes remains unclear. We investigated the prognostic significance of TILs using AI-based image analysis in 2298 luminal breast cancers from the DBCG99c cohort, classified by PAM50. Stromal (sTIL) and intraepithelial (iTIL) densities were quantified automatically using a commercial platform. Higher stromal TIL infiltration was associated with improved overall survival, particularly within the first 5 years (heterogeneity p = 0.01). In multivariable models, higher sTILs independently predicted lower risk of distant or any recurrence (sHR = 0.95, 95% CI 0.91-0.99) and improved survival (HR = 0.91, 95% CI 0.85-0.97 early; HR = 0.99, 95% CI 0.97-1.00 late). Intraepithelial TILs were not prognostic. No significant interactions were observed by molecular subtype, nodal status, or grade, although a nonsignificant trend toward stronger iTIL effects appeared in luminal B tumors. AI-based TIL quantification thus provides independent prognostic information in high-risk luminal breast cancer.
555 Background: Stratipath Breast is an AI-based prognostic medical device for risk stratification of early-stage breast cancer using routine H&E-stained histopathology whole slide images (WSIs). AI-based analyses of histopathology slides offer an alternative to costly and logistically demanding genomic assays. In this study, the prognostic performance of Stratipath Breast was validated in the TAILORx trial (NCT00310180). Methods: WSIs from patients enrolled in TAILORx were analyzed by Stratipath Breast. After histopathology quality control and availability of clinical endpoints and WSIs, 5,519 patients were included. Stratipath Breast binary risk category, multi-level risk group, and continuous risk score were evaluated. The prognostic performance was analyzed by Kaplan-Meier statistic and log rank test, as well as concordance index (C-index). Multivariable Cox Proportional Hazard (PH) model adjusting for age, tumor size, histologic subtype and continuous Oncotype DX recurrence score (RS), both without and with histologic grade, was used to assess independent prognostic value. Recurrence-free interval (RFI) and distant recurrence-free interval (DRFI) were evaluated. Results: 53.9% (2,975/5,519) of patients were classified as Stratipath low risk and 46.1% (2,544/5,519) as high risk. A significant association of Stratipath risk category and group with RFI and DRFI was confirmed (p < 0.05). In multivariable Cox PH analyses, Stratipath high-risk category was an independent prognostic factor associated with worse outcomes (RFI HR = 1.54, 95%CI:1.29-1.84, p < 0.05; DRFI HR = 1.61, 95%CI: 1.30-1.99, p < 0.05). Stratipath risk category remained significant in multivariable analysis when histologic grade was included, whereas grade did not. RS remained prognostic significant in multivariable analyses together with Stratipath Breast, indicating potentially complementary prognostic information. C-index improved with inclusion of Stratipath Breast together with clinical variables and RS, exceeding that of grade. Conclusions: In the TAILORx trial Stratipath Breast was found to provide significant prognostic value. Stratipath Breast also provided independent prognostic value in multivariable analysis adjusting for standard clinicopathologic factors and RS. These findings support the clinical relevance of AI-driven morphology-based biomarkers for breast cancer risk stratification.
The prognostic performance of histologic grade in breast cancer is robust, but evidence for its clinical validity in the neoadjuvant setting is limited. Therefore, we evaluated grade in neoadjuvant-treated breast cancer to investigate associations with overall survival (OS) in the postneoadjuvant setting. In a multicentric neoadjuvant cohort (n = 507; diagnosed 2009-2018), we examined grade in preoperative biopsies and subsequent resected specimens and compared with controls of primary operated patients (n = 297). Survival analysis for the neoadjuvant cohort related to OS was estimated, with subgroup analysis for surrogate subtypes, using the Kaplan-Meier method and log-rank test. Multivariable Cox regression models were performed to calculate hazard ratios (HR) adjusted for established clinicopathological factors. A decrease in tumor grade between preoperative biopsies and resected specimens was more frequently observed in the neoadjuvant cohort (29.8%) compared with the nontreated control group (5.7%). Patients with high-grade tumors had a considerably worse prognosis compared with low-grade tumors in both biopsies and resected specimens (P values < .001). In subgroup analysis, we found that grade had prognostic value for the ER+/HER2- subtype (P value < .001). In multivariable analysis, grade in resection specimens remained an independent prognostic marker, related to OS (HR, 2.09; 95% CI, 1.30-3.35; P = .002), whereas grade in biopsies did not (HR, 1.40; 95% CI, 0.89-2.19; P = .14). This study shows that histologic tumor grade is associated with patient outcomes after neoadjuvant treatment. Postneoadjuvant grade should be considered a prognostic factor of use in therapeutic decision-making.
Emerging evidence indicates that estrogen receptor-low (ER-low)/human epidermal growth factor receptor 2 negative (HER2-) breast cancer (BC) may more closely resemble ER-negative (ER-zero, < 1
AI-based models for analysis of histopathology whole slide images (WSIs) are now common. However, image quality, particularly unsharp areas of WSIs, impacts model performance. In this study we investigate the impact of blur on deep learning models for WSI analysis. We propose a mixture of experts (MoE) strategy that mitigates the impact of unsharp areas in WSIs on classification performance by combining predictions from multiple expert models trained on data with varying levels of blur. The study included hematoxylin and eosin (H E) stained WSIs from 2093 breast cancer patients. Classification of histological grades 1 and 3 was used as a primary benchmarking case, and prediction of immunohistochemistry (IHC) markers (ER, PR, HER2) from H E as a secondary case. The proposed MoE strategy to improve robustness against blur was evaluated in both a deep CNN model (CNN_CLAM and MoE-CNN_CLAM) and a Vision Transformer-based histopathology foundation model (UNI_CLAM and MoE-UNI_CLAM). For each architecture, a baseline model was trained on sharp images, and multiple expert models were trained on tiles with added Gaussian blur at different levels. Model performance (area under the ROC curve) was evaluated under multiple levels of uniform blur, as well as in several simulated scenarios with a mixture of blur levels within the WSIs. Baseline model performance degraded with increasing blur for all evaluated architectures. Individual expert models trained on data with simulated Gaussian blur performed better on unsharp images compared to baseline models. The proposed MoE consistently outperformed its respective baseline models in simulation scenarios with various degrees of blur within WSIs. MoE-CNN_CLAM outperformed the baseline CNN_CLAM under moderate (AUC: 0.868 vs. 0.702) and mixed blur conditions (AUC: 0.890 vs. 0.875). MoE-UNI_CLAM outperformed the baseline UNI_CLAM model in both moderate (AUC: 0.950 vs. 0.928) and mixed blur conditions (AUC: 0.944 vs. 0.931). Unsharp image areas are common in WSIs and impact prediction performance. The proposed MoE strategy provided equal or substantially improved prediction performance under all evaluated test scenarios. The proposed methodology has the potential to increase quality and reliability of AI-based pathology models in both research and clinical applications.
Tumor-infiltrating lymphocytes (TILs) are a predictive and prognostic biomarker in triple-negative (TNBC) and HER2 + breast cancer (BC). This study applies artificial intelligence (AI) to evaluate their value in a multi-institutional cohort of TNBC and HER2 + BC patients treated with neoadjuvant chemotherapy (NACT). A supervised deep learning pipeline was developed to analyze hematoxylin and eosin-stained whole-slide images from a discovery cohort of 273 patients and a validation cohort of 245 BC patients. AI quantified stromal TILs percentage, stromal TILs density, and intraepithelial TILs density. Associations between AI-derived TILs metrics, clinicopathological characteristics, and patient outcomes were assessed. AI-based scores were highly correlated with pathologists’ scores (Spearman R = 0.61–0.77, p-val < .001). Higher AI-assessed TILs levels were significantly associated with better NACT response, and both stromal and intraepithelial TILs were strong and independent predictors of pathological complete response in TNBC and HER2 + subtypes. Furthermore, patients with higher TILs had longer disease-free survival and overall survival in the discovery cohort and TNBC subtype, but not in HER2 + BC. This study supports AI-driven TILs quantification as a predictive and prognostic tool in BC patients receiving NACT. AI-derived stromal and intraepithelial TILs densities are independent predictors of response, highlighting their potential for integration into digital pathology workflows for risk stratification.
BACKGROUND:Breast cancer prognostication is crucial for treatment decisions, and the Nottingham Histologic Grade (NHG) system is widely used. However, NHG suffers from interobserver variability, and its division into three risk groups leaves the intermediate group (comprising ∼50 % of patients) overrepresented, making individualized treatment planning challenging as prognosis within this group differ widely. OBJECTIVES:This study aimed to validate the prognostic value of Stratipath's low and high-risk categories and five risk groups and compare NHG performance with the Stratipath deep-learning-based model. METHODS:We analyzed clinical data from 2466 postmenopausal, ER+/HER2-breast cancer patients who did not receive chemotherapy according to guidelines at that time. The NHG and Stratipath models were compared using concordance index and hazard ratios (HR) for distant recurrence (DR), with time to any recurrence (TR) and overall survival (OS) as secondary endpoints. RESULTS:The Stratipath five-risk group model showed similar performance to the NHG-system in predicting DR (c-index 0.71 vs. 0.72). HR for DR for Stratipath risk groups 2, 3, 4, and 5 were 1.91 (95 % CI: 1.17-3.13), 2.63 (95 % CI: 1.63-4.24), 3.18 (95 % CI: 2.00-5.07), and 3.25 (95 % CI: 2.00-5.28), respectively (p < 0.0001). In the NHG 2 subgroup, Stratipath Breast retained prognostic value for DR (HR for groups 3-5 vs. group 1: 1.73-1.85; p = 0.05), with a c-index of 0.71. CONCLUSIONS:The Stratipath AI model performs similarly to the NHG system. Further prospective validation of the clinical benefits of differentiating Stratipath risk groups 2 and 3 in treatment strategies would be valuable.
Background With new emerging technologies for diagnostics and treatment for breast cancer, there is a demand for updated breast cancer costs based on current clinical practice. The objectives of this study were to estimate recent societal costs of breast cancer in Sweden and provide population-based patient-level cost estimates for health economic evaluations. Methods This prevalence-based cost-of-illness study was based on 2019 data linking multiple Swedish national registers. The analysis employed a societal perspective considering direct health care, informal care, and productivity losses. Total costs were estimated using a bottom-up micro-costing approach. Direct costs per patient-year were also estimated by subgroups, including age group, breast cancer subtype, breast cancer stage at diagnosis, and disease state defined by metastatic status. Findings 82,960 breast cancer patients diagnosed since 2008 were alive by the end of 2019. The annual societal cost of breast cancer in Sweden was €632 million, where the direct health care, informal care, and productivity losses accounted for 37%, 5%, and 57%, respectively. The cost per capita was €61. Costs of direct health care, including inpatient/outpatient care and prescribed drugs, varied by subgroups, where younger age, higher stage, and more adverse subtypes were associated with higher costs per patient-year. Patients with a diagnosis of de novo metastatic cancer incurred the highest mean cost per patient-year. Conclusion Breast cancer represents a large economic burden in Sweden. The mean cost estimates per patient-year are informative to future health economic evaluations for breast cancer screening and treatment.
Targeted monotherapies for cancer often fail due to inherent or acquired drug resistance. By aiming at multiple targets simultaneously, drug combinations can produce synergistic interactions that increase drug effectiveness and reduce resistance. Computational models based on the integration of omics data have been used to identify synergistic combinations, but predicting drug synergy remains a challenge. Here, we introduce Drug synergy Interaction Prediction (DIPx), an algorithm for personalized prediction of drug synergy based on biologically motivated tumor- and drug-specific pathway activation scores (PASs). We trained and validated DIPx in the AstraZeneca-Sanger (AZS) DREAM Challenge human cell-line dataset using two separate test sets: Test Set 1 comprised the combinations already present in the training set, while Test Set 2 contained combinations absent from the training set, thus indicating the model’s ability to handle novel combinations. The Spearman’s correlation coefficients between predicted and observed drug synergy were 0.50 (95% CI: 0.47–0.53) in Test Set 1 and 0.26 (95% CI: 0.22–0.30) in Test Set 2, compared to 0.38 (95% CI: 0.34–0.42) and 0.18 (95% CI: 0.16–0.20), respectively, for the best performing method in the Challenge. We show evidence that higher synergy is associated with higher functional interaction between the drug targets, and this functional interaction information is captured by PAS. We illustrate the use of PAS to provide a potential biological explanation in terms of activated pathways that mediate the synergistic effects of combined drugs. In summary, DIPx can be a useful tool for personalized prediction of drug synergy and exploration of activated pathways related to the effects of combined drugs.
Introduction: Stratipath Breast is a CE-IVD marked AI-based solution for prognostic risk stratification of breast cancer patients into high- and low-risk groups, using haematoxylin and eosin (H&E)-stained histopathology whole slide images (WSIs). In this retrospective validation study, we assess the prognostic performance of Stratipath Breast in independent breast cancer cases. Material and methods: This study included patients (N=2719) diagnosed with primary breast cancer at two healthcare locations in Sweden. The patients were stratified into low- and high-risk groups by Stratipath Breast using H&E stained WSIs from the surgically resected tumours. The prognostic performance was evaluated using time-to-event analysis by multivariable Cox Proportional Hazards analysis with progression-free survival (PFS) as the primary endpoint. Further, we evaluated the prognostic performance of the continuous slide score from the Stratipath Breast by stratifying patients into 5-level risk groups based on the 5-equally sized bins of the continuous slide score defined by quantiles. Results: In the clinically relevant ER+/HER2-oestrogen receptor (ER)+/human epidermal growth factor receptor 2 (HER2)- patient subgroup, the estimated Hazard Ratio (HR) associated with PFS between low- and high-risk groups was 2.76 (95% CI: 1.63-4.66, p-value < 0.001) after adjusting for established risk factors. In the ER+/HER2- Nottingham histological grade (NHG) 2 (intermediate risk) subgroup, the HR was 2.20 (95% CI: 1.22-3.98, p-value = 0.009) between low- and high-risk groups. For 5-level patient risk stratification based on the continuous slide score, we observed the adjusted HR for PFS of 3.88 (95% CI: 1.43-10.52, p-value = 0.008) comparing between the lowest- and highest-risk group in ER+/HER2- patient subgroup. Conclusion: The results indicate an independent prognostic value of Stratipath Breast in both the general breast cancer population, in the clinically relevant ER+/HER2- and subgroup and the NHG2/ER+/HER2- subgroups. Improved image-based risk stratification of intermediate-risk ER+/HER2- breast cancers provides information relevant for treatment decisions of adjuvant chemotherapy and has the potential to reduce both under and over-treatment, shorten lead times, and reduce costs compared to molecular diagnostics. Citation Format: Abhinav Sharma, Sandy Kang Lövgren, Kajsa Ledesma Eriksson, Yinxi Wang, Stephanie Robertson, Johan Hartman, Mattias Rantalainen. Validation of an AI-based solution for breast cancer risk stratification using routine digital histopathology images [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2024; 2024 Dec 10-13; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(12 Suppl):Abstract nr P4-03-28.
Deep learning enables the modelling of high-resolution histopathology whole-slide images (WSI). Weakly supervised learning of tile-level data is typically applied for tasks where labels only exist on the patient or WSI level (e.g. patient outcomes or histological grading). In the weakly supervised learning context, there is a need for a methodology that facilitates the identification of the precise spatial regions in WSI that drive the prediction of the slide label. Such information is also needed for any further spatial interpretation of predictions from such models. We propose a novel method, Wsi rEgion sElection aPproach (WEEP), for model interpretation. It provides a principled yet straightforward way to establish the spatial area of WSI required for assigning a particular prediction label. We demonstrate WEEP on a binary classification task in the area of breast cancer computational pathology. WEEP facilitates the identification of spatial regions in WSI that are driving the decision making of a particular weakly supervised learning model, which can be further visualised and analysed to provide spatial interpretability of the model. The method is easy to implement, is directly connected to the model-based decision process, and offers information relevant to both research and diagnostic applications.
The potential of artificial intelligence (AI) in digital pathology is limited by technical inconsistencies in the production of whole slide images (WSIs), leading to degraded AI performance and posing a challenge for widespread clinical application as fine-tuning algorithms for each new site is impractical. Changes in the imaging workflow can also lead to compromised diagnoses and patient safety risks. We evaluated whether physical color calibration of scanners can standardize WSI appearance and enable robust AI performance. We employed a color calibration slide in four different laboratories and evaluated its impact on the performance of an AI system for prostate cancer diagnosis on 1,161 WSIs. Color standardization resulted in consistently improved AI model calibration and significant improvements in Gleason grading performance. The study demonstrates that physical color calibration provides a potential solution to the variation introduced by different scanners, making AI-based cancer diagnostics more reliable and applicable in clinical settings.
Introduction Histopathological evaluation of prostate biopsies using the Gleason scoring system is critical for prostate cancer diagnosis and treatment selection. However, grading variability among pathologists can lead to inconsistent assessments, risking inappropriate treatment. Similar challenges complicate the assessment of other prognostic features like cribriform cancer morphology and perineural invasion. Many pathology departments are also facing an increasingly unsustainable workload due to rising prostate cancer incidence and a decreasing pathologist workforce coinciding with increasing requirements for more complex assessments and reporting. Digital pathology and artificial intelligence (AI) algorithms for analysing whole slide images show promise in improving the accuracy and efficiency of histopathological assessments. Studies have demonstrated AI’s capability to diagnose and grade prostate cancer comparably to expert pathologists. However, external validations on diverse data sets have been limited and often show reduced performance. Historically, there have been no well-established guidelines for AI study designs and validation methods. Diagnostic assessments of AI systems often lack preregistered protocols and rigorous external cohort sampling, essential for reliable evidence of their safety and accuracy.Methods and analysis This study protocol covers the retrospective validation of an AI system for prostate biopsy assessment. The primary objective of the study is to develop a high-performing and robust AI model for diagnosis and Gleason scoring of prostate cancer in core needle biopsies, and at scale evaluate whether it can generalise to fully external data from independent patients, pathology laboratories and digitalisation platforms. The secondary objectives cover AI performance in estimating cancer extent and detecting cribriform prostate cancer and perineural invasion. This protocol outlines the steps for data collection, predefined partitioning of data cohorts for AI model training and validation, model development and predetermined statistical analyses, ensuring systematic development and comprehensive validation of the system. The protocol adheres to Transparent Reporting of a multivariable prediction model of Individual Prognosis Or Diagnosis+AI (TRIPOD+AI), Protocol Items for External Cohort Evaluation of a Deep Learning System in Cancer Diagnostics (PIECES), Checklist for AI in Medical Imaging (CLAIM) and other relevant best practices.Ethics and dissemination Data collection and usage were approved by the respective ethical review boards of each participating clinical laboratory, and centralised anonymised data handling was approved by the Swedish Ethical Review Authority. The study will be conducted in agreement with the Helsinki Declaration. The findings will be disseminated in peer-reviewed publications (open access).
AI-based models for histopathology whole slide image (WSI) analysis are increasingly common, but unsharp or blurred areas within WSI can significantly reduce prediction performance. In this study, we investigated the effect of image blur on deep learning models and introduced a mixture of experts (MoE) strategy that combines predictions from multiple expert models trained on data with varying blur levels. Using H E-stained WSIs from 2,093 breast cancer patients, we benchmarked performance on grade classification and IHC biomarker prediction with both CNN- (CNN_CLAM and MoE-CNN_CLAM) and Vision Transformer-based (UNI_CLAM and MoE-UNI_CLAM) models. Our results show that baseline models' performance consistently decreased with increasing blur, but expert models trained on blurred tiles and especially our proposed MoE approach substantially improved performance, and outperformed baseline models in a range of simulated scenarios. MoE-CNN_CLAM outperformed the baseline CNN_CLAM under moderate (AUC: 0.868 vs. 0.702) and mixed blur conditions (AUC: 0.890 vs. 0.875). MoE-UNI_CLAM outperformed the baseline UNI_CLAM model in both moderate (AUC: 0.950 vs. 0.928) and mixed blur conditions (AUC: 0.944 vs. 0.931). This MoE method has the potential to enhance the reliability of AI-based pathology models under variable image quality, supporting broader application in both research and clinical settings.