Ensuring fairness and explainability is essential for the development of ethical, reliable, and effective AI systems in healthcare. Bias in AI models can contribute to disparities in clinical outcomes, challenging equity in medical decision-making. Content-Based Image Retrieval (CBIR) offers interpretable, visual tools to support diagnostic processes; however, these tools remain susceptible to biases inherent in the data. This study investigates covariate bias arising from differences in scanning devices within Foundation Models (FMs) used for CBIR in histopathology. We introduce a unique dataset comprising spatially co-registered images derived from the same histopathology slides scanned using two distinct scanners. This design enables a targeted analysis of scanner-induced variability in FM representations. Among the FMs assessed, Vision Transformer (ViT)-based architectures such as UNI, Virchow2, and GigaPath, demonstrated top performance and best generalization properties across scanners.
Triple negative breast cancer (TNBC) is the most aggressive breast cancer subtype with the poorest patient outcomes. The composition and active immune pathways of the immune microenvironment play a critical role in tumor progression. TNBC is frequently characterized by a high and heterogeneous immune infiltration in the tumor stroma. The degree of immune infiltration impacts current management: neoadjuvant chemotherapy and anti-PD1 agent pembrolizumab.1, 2, 3 Previously, our lab studied a treatment-naive cohort of 192 patients diagnosed with TNBC and found that patients with higher B cell infiltration in the stromal regions had better outcomes (Kanwar and Balde et. al, Cancer Res 2021)4. While previous studies employed low-plex methods of analysis to study archival formalin-fixed paraffin-embedded tissue samples, the emergence of high-plex in situ protein detection methods like the GeoMx Digital Spatial profiler (DSP) has drastically enhanced the resolution of the tumor microenvironment that can be assessed.5 Using the GeoMx DSP, we are quantifying a panel of 40 immune proteins, to expand our understanding of the immune microenvironment of the pre-treatment cohort previously studied. The protein panel enables the detection of multiple immune cell types and subtypes, certain immune pathways like T cell activation and exhaustion, and the expression of immunotherapeutic targets. Preliminary results show that patients with high CD163 expression, associated with immunosuppressive macrophages had poorer outcomes (p= 0.035). The immune microenvironment for larger tumors expressed higher T cell activation markers (p=0.016). While the overall heterogeneity in immune cells did not have an impact on patient outcomes, a high heterogeneity in T cell activation markers was significantly associated with patient outcomes (p=0.029) as well as tumor size (p=0.024). We have characterized the stromal immune cell infiltrates and their spatial heterogeneity, finding immune markers that may significantly impact patient outcomes. Our findings aid in the identification of therapeutic targets prevalent in the TNBC stroma as well as potentially informing therapy decisions based on the tumor immune microenvironment composition. 1. Won KA, Spruck C. (2020). Int J Oncol, 57(6), 1245-1261. DOI:10.3892/ijo.2020.5135 2. Marra A et al. (2020). NPJ Breast Cancer, 6, 54. DOI:10.1038/s41523-020-00197-2 3. Obidiro O et al. (2023). Pharmaceutics, 15(7), 1796. DOI:10.3390/pharmaceutics15071796 4. Kanwar N et al. (2021). Cancer Res, 81(24), 6196-6206. DOI:10.1158/0008-5472.CAN-21-1079 5. Bergholtz H et al. (2021). Cancers, 13(17), 4456. DOI:10.3390/cancers13174456 Prerana Sensharma, Huidan Zuo, Melanie Dawe, Megan Hopkins, Zeynep Baskurt, Osvaldo Espin-Garcia, Philippe L. Bedard, Melanie Spears, Susan J. Done. A spatial profiling approach to evaluating the prognostic impact of heterogeneity in the triple-negative breast cancer immune microenvironment [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2010.
Fairness and explainability are essential pillars for the development of ethical, trustworthy, and effective AI systems in healthcare. Biases in AI models, particularly those arising from data acquisition, can lead to disparities in clinical outcomes and compromise equity in medical decision-making. While Content-Based Image Retrieval (CBIR) systems offer interpretable visual context to support diagnostic reasoning, they remain vulnerable to dataset-induced biases. In this work, we present the first systematic investigation of covariate bias related to scanning devices in the context of Foundation Models (FMs) for CBIR applications in histopathology, a critical yet previously unexplored dimension in the literature. To facilitate this study, we curated a novel dataset to isolate scanner variability, comprising spatially co-registered image patches obtained from the same histology slides, each scanned using two different whole slide imaging devices. This design ensures confounding factors are held consistent across paired patches. Through a series of experiments across two distinct datasets, we demonstrate that the retrieval performance of state-of-the-art FMs is dependent on scanners, revealing a pronounced lack of robustness to scanner variability. Among the FMs evaluated, iBOT-Path, UNI, Virchow2, and Prov-GigaPath achieved the most stable performance across scanners and datasets. In contrast, KimiaNet, Virchow, HibouB, and HIPT have poor performance stability, reflecting lower generalizability. This study provides the first empirical evidence of scanner-induced covariate bias in CBIR applications using FMs, highlighting an important and previously unaddressed challenge in the deployment of trustworthy AI systems for computational pathology.
BACKGROUND:Emerging evidence indicates that tumour innervation promotes cancer progression via a non-canonical TLR7 signalling pathway. However, its impact across breast cancer subtypes, patient populations, associated molecular pathways, and oncogenic drivers remains poorly defined. METHODS:We analysed TLR7 signature scores in human breast cancer across multiple datasets and evaluated their associations with prognosis, clinical outcomes, TNBC subtypes, metastasis, molecular signatures, oncogenic signalling, and pathological complete response. RESULTS:We demonstrate that the TLR7score signature is significantly elevated in triple-negative breast cancer (TNBC) - the most aggressive breast cancer subtype-compared with ER⁺ disease. Within TNBC, high TLR7 signalling characterises basal- and mesenchymal-like tumours relative to the luminal androgen receptor (LAR) subtype. Across multiple breast cancer cohorts, including TNBC, TLR7score alone does not uniformly predict prognosis, as both high- and low-scoring tumours are associated with reduced survival. Using sequential cut-off analysis in seven independent clinical cohorts, we show that both TLR7score-high (e.g. , SCAN-B:HR = 4.7, P = 0.01) and TLR7score-low (HR = 3.37, P = 0.038) tumours are associated with unfavourable outcomes relative to intermediate-score tumours. TLR7score-high lesions are enriched for cell proliferation, neuronal, and mast cell-related pathways, as well as RB1 and TP53 loss and elevated E2F, PI3K, MET, and MYC signalling. In contrast, TLR7score-low tumours show increased ER signalling and are enriched for T cell-associated but not neuronal pathways, delineating innervated versus non-innervated TNBC phenotypes. Moreover, TLR7score correlates with pathological complete response (pCR) in a treatment-dependent manner. CONCLUSIONS:Collectively, these findings suggest that TNBC progression involves both TLR7-dependent and TLR7-independent mechanisms and that TLR7score may enable patient stratification for distinct therapeutic strategies.
Automated mitosis detection aims to reduce the time and variability associated with manual mitotic counting, but remains challenged by scarce annotations, high morphological diversity, and domain shift across scanners, stains, organs, and species. We propose a unified teacher–student framework for the MIDOG 2025 Challenge to perform mitosis segmentation (Track 1) and atypical mitosis classification (Track 2) within a single model. The method uses a U-Net backbone with contrastive and domain-adversarial learning for robust feature learning. A frozen teacher generates pseudo-masks for normal nuclei, mitotic figures, and hard negatives, guiding the student through semi-supervised consistency loss. For Track 2, a lightweight multi-scale classifier operates on encoder feature maps to categorize normal and atypical mitosis. Across preliminary and final tests, our method achieved F _1 -scores of 0.7660 and 0.6598 for Track 1, and balanced accuracies of 0.8414 and 0.8721 for Track 2, demonstrating the effectiveness of integrating segmentation and classification within a domain-generalized teacher–student framework.
Deep learning has proven capable of automating key aspects of histopathologic analysis. However, its context-specific nature and continued reliance on large expert-annotated training datasets hinders the development of a critical mass of applications to garner widespread adoption in clinical/research workflows. Here, we present an online collaborative platform that streamlines tissue image annotation to promote the development and sharing of custom computer vision models for PHenotyping And Regional Analysis Of Histology (PHARAOH; https://www.pathologyreports.ai/ ). Specifically, PHARAOH uses a weakly supervised, human-in-the-loop learning framework whereby patch-level image features are leveraged to organize large swaths of tissue into morphologically-uniform clusters for batched annotation by human experts. By providing cluster-level labels on only a handful of cases, we show how custom PHARAOH models can be developed efficiently and used to guide the quantification of cellular features that correlate with molecular, pathologic and patient outcome data. Moreover, by using our PHARAOH pipeline, we showcase how correlation of cohort-level cytoarchitectural features with accompanying biological and outcome data can help systematically devise interpretable morphometric models of disease. Both the custom model design and feature extraction pipelines are amenable to crowdsourcing, positioning PHARAOH to become a fully scalable, systems-level solution for the expansion, generalization and cataloging of computational pathology applications. Faust, Chen, and colleagues present PHARAOH, a collaborative computational pathology platform that allows histologists to quickly develop custom labelled image datasets to train and catalogue a variety of machine learning models for histopathological analysis.
The emergence of large foundation models (FMs) in histopathology, trained on extensive image datasets using high-performance graphics processing unit (GPU) clusters, has demonstrated significant potential in advancing computational pathology. FMs have potential to overcome the domain gap between training and testing datasets, which creates more translation opportunities. However, the reliance on vast computational resources and large-scale data often limits accessibility and widespread adoption of FMs. To address this limitation, we present HistoLite, a lightweight self-supervised learning framework designed to enable domain-invariant representation learning in histopathology. HistoLite utilizes customizable auto-encoders within a self-supervised learning paradigm that learns generalized and transferable features in an efficient manner. We evaluated the proposed framework using breast Whole Slide Images (WSIs) and benchmarked performance with state-of-the-art FMs for domain generalization. A novel dataset was curated that is of the same tissue slides, scanned by two different scanning platforms, which allows for specific analysis of covariate shifts due to scanner bias. Aspects evaluated include the difference in embeddings across scanners using novel representation shift metrics, including a robustness index, and accuracy, which looks at performance on downstream tasks. The top performing models were UNI, Virchow2 and Prov-GigaPath, likely due to large model sizes and training datasets. In general, most FMs were found to be susceptible to scanner-bias, as shown by differences in embeddings and drop in performance on the held-out scanner. This has significant implications for real-world deployment of FMs in histopathology. HistoLite offered low representation shift in embeddings, the lowest performance drop on out-of-domain data with modest classification accuracy, indicating the smaller model may exhibit a tradeoff between accuracy and generalization.
Foundation models (FMs) have introduced new opportunities for zero-shot generalization in downstream tasks through techniques such as linear probing and feature extraction. However, a systematic evaluation of their fairness regarding their sensitivity to covariate bias remains lacking, which is critical for clinical translation. In this study, we address this gap by constructing a unique dataset comprising identical glass slides digitized using two scanners, enabling a controlled simulation of covariate bias in data distributions. The same tissue regions (patches) are extracted from Whole Slide Images acquired on each scanner and processed through FMs to obtain zero-shot feature representations. We define and quantify ‘representation shift’, as the difference in feature vectors across scanners, and assess using metrics such as mean squared error, Kullback–Leibler divergence, and a novel clustering-based Calinski–Harabasz index. Our results demonstrate that FMs exhibit significant scanner-dependent variability in feature representations, highlighting a key generalization limitation. This sensitivity to acquisition device introduces the risk of unequal or unfair performance; an important consideration for the safe and fair deployment of FMs in clinical settings.
A Ki67 proliferation index (PI) of ≥20% determines high-risk status and guides treatment decisions in HR+/HER2− early breast cancer. The Ki67 PI, defined as the ratio of Ki67+ tumor cells to total malignant cells, is labor-intensive and subjective to score. AI support is necessary for routine and reliable clinical use. This is the first large-scale, international study on AI-aided Ki67 PI risk assessments in breast cancer. Ninety pathologists scored the continuous PI from 0% to 100% for 10 breast cancer tissue microarrays (TMA) with and without AI support. The AI tool provided a continuous PI and a Ki67+/- tumor nuclei overlay. The AI tool was developed and validated independently of the pathologists and data used in this study (PMID: 38218973). Using gold-standard manual counts, two pathologists established the ground-truth (GT) PI scores for each TMA, ranging from 7% to 28%. The pathologists’ PI scores were classified into low (<20%) and high risk (≥20%). The risk classification accuracy, sensitivity and specificity relative to the GT were obtained for manual and AI-aided scores, yielding 900 paired classification metrics (90 pathologists × 10 TMAs). Table 1 summarizes the classification metrics showing improvements with AI support for all demographics. The AI tool achieved 100% risk classification accuracy. Manual misclassification and AI-aided correction rates are also summarized in Table 1. The overall manual misclassification rate was 26%, and AI support successfully corrected 95% of these cases. The McNemar’s test, a nonparametric statistical test for paired binary data, assessed statistical differences in risk classification between manual and AI-aided scoring, an asterisk denotes significance (p<0.05). AI support can improve the clinical reliability of PI scoring for risk classification and treatment decisions. AI integration into routine practice can ensure consistent Ki67 use, improving drug response predictions and patient outcomes. Amanda Dy, Ngoc-Nhu Jennifer Nguyen, Melanie Dawe, Dimitrios Androutsos, Susan Done, April Khademi. Transforming breast cancer treatment decisions: AI tools improve Ki67 risk assessments with 90 pathologists [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7442.
[This corrects the article DOI: 10.1016/j.jpi.2022.100002.].
Kikuchi-Fujimoto disease (KFD, histiocytic necrotizing lymphadenitis) is a rare, benign disease in which the presenting clinical and radiological features often result in misclassification as a malignant process. We present the first report of concurrent KFD in the draining lymph nodes of a malignant phyllodes tumour, adding to the growing number of reports of KFD occurring in the context of malignancy, further compounding the existing diagnostic difficulties. An increased awareness of this condition with consideration for inclusion in the differential diagnosis of lymphadenopathy is required for improved diagnosis of this under-recognized entity.
Visualization of cancer during breast conserving surgery (BCS) remains challenging; the BCS reoperation rate is reported to be 20-70 n = 17 patients have received 20mg/kg bodyweight (BW) 5-ALA orally 2-4 h before imaging to facilitate the accumulation of PpIX within tumour cells. Tissue types were identified based on their colour appearance. Breast tumours in sectioned lumpectomies appeared red, which contrasted against the green connective tissues and orange-brown adipose tissues. In addition, ductal carcinoma in situ (DCIS) that was missed during intraoperative standard of care was identified at the surgical margin at <1 mm depth. In addition, artifacts due to the surgical drape, illumination, and blood within the surgical cavity were discovered. This study has demonstrated the detection of a grossly occult positive margin intraoperatively. Artifacts from imaging within the surgical cavity have been identified, and potential mitigations have been proposed. ClinicalTrials.gov Identifier: NCT01837225 (Trial start date is September 2010. It was registered to ClinicalTrials.gov retrospectively on April 23, 2013, then later updated on April 9, 2020, to reflect the introduction of the new imaging device.)
Ki-67 is a nuclear protein associated with proliferation, and a strong potential biomarker in breast cancer, but is not routinely measured in current clinical management due to a lack of standardization. Digital image analysis (DIA) is a promising technology that could allow high throughput analysis and standardization. There is a dearth of data on the clinical reliability as well as intra- and inter-algorithmic variability of different DIA methods. In this study, we scored and compared a set of breast cancer cases in which manually counted Ki-67 has already been demonstrated to have prognostic value (n=278) to five DIA methods; namely, Aperio ePathology, Definiens Tissue Studio, Qupath, an unsupervised IHC color histogram (IHCCH) algorithm and a deep learning pipeline piNET. The piNET system achieved high agreement (ICC: 0.850) and correlation (R= 0.85) with the reference score. The Qupath algorithm exhibited a high degree of reproducibility between all rater instances (ICC: 0.889). Although piNET performed well against absolute manual counts, none of the tested DIA methods classified common Ki-67 cutoffs with high agreement or reached the clinically relevant Cohen’s kappa of at least 0.8. The highest agreement achieved was Cohen’s kappa statistic of 0.73 for cutoffs 20% and 25% by the piNET system. The main contributors to inter-algorithmic variation and poor cutoff characterization included heterogeneous tumor biology, varying algorithm implementation, and setting assignments. It appears that image segmentation is the primary explanation for semi-automated intra-algorithmic variation, which involves significant manual intervention to correct. Automated pipelines such as piNET may be crucial in developing robust and reproducible unbiased DIA approaches to accurately quantify Ki-67 for clinical diagnosis in the future.
The Ki-67 proliferation index (PI) guides treatment decisions in breast cancer but suffers from poor inter-rater reproducibility. Although AI tools have been designed for Ki-67 assessment, their impact on pathologists' work remains understudied. 90 international pathologists were recruited to assess the Ki-67 PI of ten breast cancer tissue microarrays with and without AI. Accuracy, agreement, and turnaround time with and without AI were compared. Pathologists’ perspectives on AI were collected. Using AI led to a significant decrease in PI error (2.1% with AI vs. 5.9% without AI, p < 0.001), better inter-rater agreement (ICC: 0.70 vs. 0.92; Krippendorff’s α: 0.63 vs. 0.89; Fleiss’ Kappa: 0.40 vs. 0.86), and an 11.9% overall median reduction in turnaround time. Most pathologists (84%) found the AI reliable. For Ki-67 assessments, 76% of respondents believed AI enhances accuracy, 82% said it improves consistency, and 83% trust it will improve efficiency. This study highlights AI's potential to standardize Ki-67 scoring, especially between 5 and 30% PI—a range with low PI agreement. This could pave the way for a universally accepted PI score to guide treatment decisions, emphasizing the promising role of AI integration into pathologist workflows.
Primary Mucosa-associated lymphoid tissue (MALT) lymphoma is a rare diagnosis in the breast, and clinical diagnosis based on radiological features is often challenging. This study aimed to evaluate the clinicopathological, and radiological characteristics of the patients diagnosed with primary breast MALT lymphoma. This study examined 18 cases of primary MALT lymphoma of the breast diagnosed at a single tertiary center between January 2002 to December 2020. Medical charts, radiological imaging and original pathology slides were reviewed for each case. All cases were female (gender assigned at birth) and presented with a palpable mass or an incidental imaging finding. Imaging presentation ranged from mammographic asymmetries, circumscribed masses, and ultrasound masses lacking suspicious features. Seventeen cases were biopsied under ultrasound; one received a diagnostic excision biopsy. Microscopic examination of the breast specimens demonstrated atypical small lymphocyte infiltration with plasmacytoid differentiation and rare lymphoepithelial lesions. Immunohistochemistry was performed in all cases and established the diagnosis. Most patients were treated with radiotherapy, and only three were treated with chemotherapy. The median follow-up period was 4 years and 7.5 months, and all patients were alive at the last follow-up. Primary MALT breast lymphomas are usually indolent and non-systemic, and local radiotherapy may effectively alleviate local symptoms. Radiological findings show overlap with benign morphological features, which can delay the diagnosis of this unusual etiology. Although further studies involving a larger cohort could help establish the clinical and radiological characteristics of primary breast MALT lymphomas, pathology remains the primary method of diagnosis. University Health Network Ethics Committee (CAPCR/UHN REB number 19–5844), retrospectively registered.
Raw cell counts of FISH gene probes for (1) primary breast tumors, (2) metastatic sites, (3) pre- and post-treatment [Hungarian sample], and (4) TNBC cases.
Supplementary Figure 3: Association of chemokine expression with a T cell-inflamed phenotype was validated in a second independent PDAC patient cohort