
Distinguishing alcohol-associated steatohepatitis (ASH) from metabolic dysfunction-associated steatohepatitis (MASH) is challenging given the absence of pathognomonic differentiators. We developed and evaluated a weakly supervised deep learning model, as a single-institution proof-of-concept study, to test whether routine hematoxylin and eosin (H&E) whole-slide images (WSIs) of liver biopsies contain sufficient morphological information to predict steatohepatitis etiology at the slide level.A retrospective cohort of 1147 WSIs was assembled (train set: 1007, holdout test set: 140). Models were trained using 5-fold patient-level cross-validation at ×20 and ×40 magnification. Model interpretability was assessed through attention-based clustering with blinded pathologist review of high-attention patches.The ×20 model achieved a mean area under the receiver operating characteristic of 0.86 ± 0.01 on the test set, with balanced accuracy of 0.80. The ×40 model performed comparably. Specificity was high (0.93 at ×20), with an ASH sensitivity of 0.67 at ×20. Fibrosis-stratified showed preserved performance at advanced fibrosis (stage ≥3) with balanced accuracy of 82.5% and ASH sensitivity of 77.4%. Attention-based clustering localized ASH-enriched regions to active injury patterns including Mallory–Denk bodies, neutrophilic inflammation, cholestatic change, and pericellular fibrosis, whereas MASH-enriched regions showed steatosis with lower inflammatory activity.Weakly supervised deep learning applied to routine liver H&E WSIs can discriminate ASH from MASH with performance preserved at advanced fibrosis. The comparable performance of ×20 and ×40 magnification, and the alignment of model attention with established histological features, support the use of routine morphology as a decision-support input in cases with incomplete or conflicting clinical histories, particularly when the etiological distinction has the greatest implications for management.
Whole-slide imaging (WSI) is gradually being adopted for primary diagnosis in many pathology subspecialties. The benefits of WSI for the pathology workflow are numerous and include the use of artificial intelligence (AI) tools to support WSI analysis. However, digitization of hematopathology workflows has remained a challenge, in part due to the complexity of the hematopathology workflow and the unique characteristics of blood and bone marrow smears. Although AI applications to support hematopathology tasks, such as nucleated differential counts, are well-described, general implementation of these tools has been limited due to the lack of widely available WSIs. Consequently, there is an unmet need for new and innovative solutions to WSI in hematopathology. This review focuses on WSI of hematopathology cytological specimens, particularly, peripheral blood and bone marrow smears, and summarizes unique challenges and AI solutions to support WSI analysis in this area. We also discuss future opportunities for hematopathologist-driven innovation in AI and WSI in hematopathology.
Digital pathology increasingly seeks to extract quantitative vascular and microenvironmental features from routine histology images, but thin-section hematoxylin and eosin (H&E) slides can fragment vessels and obscure their spatial organization, constraining downstream computational analysis. Slide-free fluorescence-imitating brightfield imaging (FIBI) produces histology-like images directly from fresh or fixed tissue or from already prepared paraffin blocks within minutes and can better preserve apparent microvascular continuity in breast cancer and other specimens. In this exploratory digital pathology study, we acquired paired FIBI and H&E images from paraffin-embedded breast tissue blocks, manually segmented blood vessels in tumor and tumor-adjacent stroma, and quantified vascular architecture using both standard geometry-based metrics and topology-derived descriptors of inter-vessel arrangement, including a persistent-homology-based “vessel spacing” metric that captures multiscale clustering. Here, geometry refers to properties of individual vessels (size, length, and branching), whereas topology summarizes how vessels are arranged as a network, including how closely or loosely they cluster. FIBI images exhibited an easily appreciable increase in vascular information compared with matched H&E slides, with vessels appearing more continuous and more clearly resolved in both geometry (e.g., larger area, more branched, etc.) and network topology (e.g., decreased inter-vessel separation). Motivated by this qualitative impression of increased informational content, the goal of this study was to quantitatively assess and validate these differences by contrasting complementary geometry and topology-derived metrics of vessels in pairs of H&E and FIBI images. In particular, persistent homology-derived topology metrics provided information that was not linearly explained by standard geometric descriptors. Together, these findings demonstrate a practical digital pathology pipeline that integrates slide-free FIBI acquisition, vessel annotation, geometry- and topology-based feature extraction, and suggest that FIBI-derived vascular signatures, particularly when coupled with automated vessel segmentation, may provide informative input for future computational and artificial intelligence-based pathology models, and potentially clinical applications.
Accurate stratification of Hodgkin lymphoma (HL) by immunologic/histological subtypes and Epstein-Barr virus (EBV) status is essential for epidemiological and translational research, yet large-scale testing is impractical and expensive. Digital pathology models that utilize routinely used hematoxylin and eosin (H&E) whole-slide images (WSIs) could close this gap. We developed and validated a hierarchical Vision Transformer pipeline that aggregates cell-, patch-, and region-level context to predict EBV status and the three most prevalent immunological/histological HL subtypes: nodular sclerosis (NS), mixed cellularity (MC), and nodular lymphocyte-predominant HL (NLPHL)—from H&E-stained WSIs, and additionally evaluated a standard attention-based multiple-instance learning (ABMIL) baseline for direct architectural comparison. The development pool comprised 1643 HL cases (1952 WSIs) from 18 Danish hospitals and was used for hospital-preserving 5-fold cross-validation; external validation was performed on an independent hold-out cohort of 458 cases (532 WSIs) from five hold-out hospitals. For subtype prediction, analyses were restricted to the 1560 cases belonging to NS, MC, or NLPHL. On the external EBV cohort (N=458) the hierarchical pipeline achieved an area under the receiver operating characteristic curve (ROC–AUC) of 0.73 (95% confidence interval (CI) 0.68–0.77), precision–recall (PR)–AUC 0.57 (95% CI 0.49–0.66) with recall (sensitivity) 0.74 (95% CI 0.67–0.80) and macro-F1 score 0.60 (95% CI 0.54, 0.65). For 3-class subtype prediction on 359 external cases, discrimination reached ROC–AUC 0.84 (95% CI 0.80–0.88) and PR–AUC 0.63 (95% CI 0.56–0.71) with a macro-F1 of 0.56 (95% CI 0.48, 0.64) and macro-recall 0.55 (95% CI 0.47, 0.63); residual errors were dominated by NS-MC confusions. An ABMIL baseline using the same patch embeddings achieved ROC–AUC 0.76 (95% CI 0.72–0.81) for EBV and 0.89 (95% CI 0.86–0.92) subtype prediction, outperforming the hierarchical model on both tasks. This multicenter study shows that both hierarchical and attention-based architectures can determine EBV status and major HL subtypes directly from routine H&E slides with externally validated performance across hospitals, whereas the finding that the simpler baseline outperformed the hierarchical model suggests that strong foundation-model embeddings combined with attention-based pooling may reduce the need for explicit multi-scale modelling in cohorts of this size.
The display is a vitally important component in the digital imaging chain, being placed at the final step at the interface to the pathologist. In medical imaging fields such as radiology, there are numerous guidelines and standards relating to the display. However, in digital pathology, there are relatively few recommendations for displays, and those that are available can quickly become obsolete in this fast-moving field. There is general disagreement over whether consumer-grade displays are adequate for digital pathology reporting, or whether medical-grade displays are required, as is the practice in radiology. However, irrespective of the type of display chosen, it is essential that the device can ensure a reproducible, reliable, optimal appearance for diagnostic assessment. This is particularly important to support the ongoing clinical implementation of digital pathology and engender confidence in this relatively new and disruptive technology from the perspective of medical professionals, patients, and the public. Here, we describe and examine the existing guidance on displays in digital pathology and the current landscape of devices. We present tools and methodologies to aid display selection, propose minimum recommendations for pathology display specifications, and present a framework for design of a display quality assurance method.
Background Cytology specimens are essential for lung cancer diagnosis, yet distinguishing benign from malignant cells and subtyping adenocarcinoma, squamous cell carcinoma, and small-cell lung cancer remains challenging due to limited cellular material. Methods We systematically searched seven databases for studies evaluating artificial intelligence-based lung cancer classification in cytology. After screening, 32 studies were included, of which 20 contributed 33 analyzable two-by-two data rows for diagnostic test accuracy meta-analysis. Outcomes were pooled using bivariate logit random-effects models. Summary receiver operating characteristic curves, Fagan nomograms, Deeks' test for publication bias, and quality assessments using QUADAS-2 and artificial intelligence-specific tools were performed. Results For benign versus malignant diagnosis, 18 data rows encompassing 3464 observations yielded a pooled sensitivity of 89.3% (95% confidence interval, 84.1–92.9), a specificity of 88.5% (95% confidence interval, 83.2–92.3), and an area under the curve of 0.952. At a median prevalence of 58.2%, the post-test probability of malignancy after a positive artificial intelligence result was 91.2%, whereas a negative result yielded a post-test probability of 14.6%. For one-versus-rest subtyping among adenocarcinoma, squamous cell carcinoma, and small-cell lung cancer, 14 data rows comprising 2687 observations demonstrated a pooled sensitivity of 77.5%, specificity of 89.5%, and area under the curve of 0.901. Small-cell lung cancer showed the strongest performance (specificity 93.2%, area under the curve 0.949), whereas squamous cell carcinoma had lower sensitivity (67.0%). Substantial heterogeneity was observed across studies. Deeks' test suggested potential small-study effects, and the risk of bias was high in all included studies. Conclusion Artificial intelligence-based cytology classification achieves high diagnostic accuracy for benign-malignant discrimination and shows promising but less mature performance for histological subtyping. Prospective external validation, standardized reporting, and clinically integrated human–artificial intelligence studies are required before routine clinical deployment.
The Dutch Nationwide Pathology Databank (Palga) was established in 1971 and achieved nationwide coverage of all Dutch pathology departments in 1991. It serves as the central repository for pathology data across the Netherlands. Palga manages a nationwide IT infrastructure enabling the collection, structuring, and secure sharing of pathology data from all Dutch pathology departments. This unique infrastructure supports diagnostics in patient care by enabling pathologists to access nationwide patient pathology histories and by supporting standardized, structured reporting, particularly in population-based screening programs and oncological pathology. It also serves as a valuable resource for secondary data use, supporting scientific research and healthcare policy decision-making.This article provides an overview of the Palga infrastructure, with a focus on its embedded scientific database for secondary use. It describes how pathology data are collected, structured, and shared, and explains how data are securely processed in compliance with current legislation, with particular emphasis on safe secondary use.Secure secondary use of pathology data is expected to increasingly advance evidence-based healthcare research for the benefit of society. Palga's nationwide infrastructure provides substantial added value for patient care, health research, and healthcare quality improvement, while enabling responsible and socially valuable secondary data use through collaboration and robust legal and ethical governance.
Introduction:Digital pathology continues to transform the daily routine of pathology, in terms of the increasingly automated laboratory and in the diagnostic paradigm through the adoption of artificial intelligence (AI) tools to support diagnosis-computational pathology. The reliability and performance of these tools depend on the whole-slide image (WSI) quality being guaranteed a priori. Pre-analytical quality control step that underpins this guarantee, and artifact detection remains largely qualitative and is frequently overlooked in routine digital pathology. This operational feasibility study evaluated whether an adaptation of GrandQC, an open-source AI tool, enables automated, quantitative artifact assessment of a complete single-day biopsy workload from a high-throughput digital pathology laboratory, analyzed retrospectively. Material and methods:A random biopsies day of 2025 at Centro de Anatomia Patológica Germano de Sousa (CAPGS) was selected as a sample to test the performance of GrandQC on the WSI generated (n = 544) in order to simulate the daily workflow. A script was created to quantify the pixels corresponding to the type of artifact automatically, creating an Excel file for registering and statistical analysis. Results:Analysis took a median of 24 s per WSI, detecting a median of 1.46% of tissue area with some type of artifact. Dark Spots and blurring areas were the most representative detected artifacts. Conclusion:GrandQC is a valuable tool in the quantitative quality control of biopsies tissue, allowing quick evaluation, signaling types of artifacts, and identifying cases that need to be reviewed before being handed over to the pathologist allowing the recognition of opportunities to improve laboratory histology quality and precision medicine.
Programmed death ligand-1 (PD-L1) expression is a key biomarker for identifying non-small cell lung cancer (NSCLC) patients eligible for immunotherapy, but its immunohistochemistry assessment is subject to high interobserver variability. This study aimed to develop and validate an automated system based on artificial intelligence (AI) to detect and classify PD-L1-positive cells in digital lung biopsies from Latin American patients with NSCLC. A total of 141 biopsies were digitized, and deep learning models were trained for tissue segmentation and analysis. YOLOv8 was used for cell detection and ResNet50 for classification into four categories (tumor or non-tumor cells, PD-L1 positive or negative). The tumor proportion score (TPS) was calculated using these classifications. The YOLOv8 model achieved 76% precision and a mAP@0.5 of 63.7% in cell detection, whereas ResNet50 reached 85.2% sensitivity in identifying PD-L1-positive tumor cells. The TPS estimated by the system demonstrated 97% concordance with manual evaluation by an expert pathologist. In addition, spatial PD-L1 maps revealed intratumoral heterogeneity, supporting its relevance for immunotherapy decisions. These findings show that automated PD-L1 quantification using AI is feasible and accurate, improves diagnostic reproducibility, and supports personalized treatment decisions, especially in resource-limited settings.
Medical imaging data are an essential resource for research and teaching; however, regulations such as the Health Insurance Portability and Accountability Act require careful consideration of protected health information (PHI) contained within images and their metadata. De-identification, the removal of PHI to prevent subject re-identification, is often essential for regulatory compliance. In pathology, whole-slide imaging (WSI) files-high-resolution scans of entire glass slides-are essential for research and machine learning model training. Currently, many scanner vendors use proprietary WSI formats; some vendors offer de-identification tools, which are often manual and tedious to use. Alternatively, physical slides may be re-scanned with obfuscated labels, though this is resource intensive. Here, we describe an institutional-scale automated WSI de-identification pipeline, which utilizes an informatics-based approach to convert clinical WSI from 11 proprietary formats into de-identified WSI in multiple open formats. This approach minimizes manual processes and additional slide-scanning hardware, storage, and personnel. De-identification requests are submitted through a zero-footprint web portal following the Fast Healthcare Interoperability Resources ServiceRequest data model. De-identification is verified via a human-in-the-loop review, and archives are distributed using cloud-based storage. Since deployment in November 2024, our pipeline has de-identified 819 cases comprising 4322 WSI files generated by 7 scanner models in 4 formats. We demonstrate that the conversion process is predictable, linearly scalable, and reliable across de-identification request sizes (238× variation), image sizes (120× variation), and capture format. This pipeline eliminates the need for manual de-identification or re-scanning while achieving high throughput and reliability at an institutional scale.
Accurate nuclei detection and classification in hematoxylin and eosin (H&E) whole-slide images (WSIs) is a key task in computational pathology, particularly for quantitative analysis of the tumor microenvironment. However, this task remains highly challenging due to variations in nuclei morphology, staining procedures, scanners, organs, magnifications, and WSI artifacts. In addition, many existing pipelines rely on computationally demanding architectures and post-processing procedures, making gigapixel WSI analysis time-consuming. In this work, CellPrior-Net (CP-Net) is proposed, an efficient nuclei detection and classification pipeline that utilizes a lightweight convolutional neural network architecture and hematoxylin (H) channel as prior information to enhance nuclei-aware feature learning. Extensive benchmarking was conducted against state-of-the-art pipelines on eight public and private datasets (total:~10.4 M nuclei) obtained from different organs, scanners, magnifications, and clinical centers. Experimental results demonstrate that CP-Net achieves comparable performance while significantly reducing inference time. Furthermore, CellQuant-Net was introduced—an end-to-end nuclei quantification pipeline—that integrates a quality assessment model to exclude regions with artifacts, followed by CP-Net cell detection and classification. The pipeline is publicly available on GitHub, and provides a potentially efficient and scalable framework for downstream computational pathology applications.
Background:Artificial intelligence (AI)-enabled software is increasingly integrated into digital and computational pathology, driving new regulatory and quality management requirements that extend beyond traditional laboratory practice and classical medical-device oversight. Practical, experience-based guidance on balancing development agility with global regulatory readiness remains limited for biomarker developers and translational pathologists navigating the convergence of AI governance, cybersecurity, and regulated clinical deployment. Methods:We reviewed our multi-year regulatory, quality, information security management program, and software lifecycle artifacts associated with AI-based diagnostic software development across multiple jurisdictions. Documentation, change control processes, and internal coordination mechanisms were analyzed to identify structural patterns supporting parallel progress in regulatory submissions, product releases, and assurance infrastructure. Results:Our assessment consistently showed four transferable operational determinants to maintain iterative development while preserving regulatory readiness across jurisdictions: (1) Regulatory submissions and approvals (e.g., IVDR, FDA clearance); (2) product releases and lifecycle control; (3) quality and assurance infrastructure, including quality management system (QMS) certifications (e.g., ISO 13485, MDSAP); and (4) cybersecurity and information security certifications (e.g., ISO 27001, HITRUST, and C5). Together, these determinants enabled coordination of regulatory, release, and quality milestones in parallel, reducing friction at later submission stages and supporting synchronized readiness across jurisdictions. Conclusions:This technical note presents a transferable framework for managing AI-based pathology software development in regulated environments. High-quality deployment requires more than model performance alone; it depends on technical, regulatory, and operational maturity across many dimensions, including change control, documentation, security-aligned quality systems, post-market surveillance, technical support, and workflow integration. The presented framework is directly relevant to computational scientists, pathologists, and laboratory/medical directors tasked with evaluating AI systems by supporting informed evaluation and adoption decisions, including procurement considerations, in clinical and biopharma settings.
Tumor cellularity (TC) is a key histopathological measure that impacts the quality of molecular testing and the choice of treatment. Estimation of TC is usually manual, by visual assessment by pathologists, which makes it prone to inter-observer variability. Advancements in the application of artificial intelligence (AI) in digital pathology facilitate TC evaluation and scale up existing pathology workflows. In practical terms, these AI systems analyze digitized tissue slides computationally, producing objective TC scores that can support pathologists in determining whether a specimen contains sufficient tumor material for molecular testing. This review presents a compilation of AI-driven methodologies for TC estimation by summarizing publicly available datasets and benchmarking resources, examining methodological advancements in automated TC assessment, and delineating the research and commercial tools accessible for TC quantification. It also discusses the challenges that hinder clinical translation and the necessity for transparent and interpretable outputs that align with the pathologists' reasoning.
Background:Digital pathology and artificial intelligence (AI) are transforming cancer diagnostics worldwide, yet their implementation in sub-Saharan Africa remains largely undocumented. The region faces a critical shortage of pathologists while bearing an increasing cancer burden, making AI-assisted diagnostics particularly relevant. However, practical deployment in resource-constrained environments raises unique technical, logistical, and educational challenges that differ substantially from those encountered in high-income settings. Objective:To systematically document the technical, logistical, and practical challenges encountered during the implementation of QuPath, an open-source digital pathology platform, for breast cancer immunohistochemical (IHC) biomarker assessment at a reference pathology lab in Cameroon, and to propose actionable solutions for similar resource-limited settings. Methods:We conducted a prospective implementation study at the Centre Pasteur du Cameroun, Yaoundé, involving 39 cases of invasive breast carcinoma with IHC for estrogen receptor, progesterone receptor, Ki67, and HER2. We documented all phases of the digital pathology workflow: pre-analytical slide preparation, slide digitization (performed remotely at Erasme University Hospital, Brussels), image transfer and storage, QuPath algorithm training, and automated analysis on a consumer-grade laptop (4 GB RAM, 500 GB storage). Challenges were categorized into four domains: hardware and infrastructure constraints, pre-analytical and scanning issues, software training and optimization, and human factors including the learning curve. Results:Of 130 IHC slides, 18 (13.8%) required re-scanning due to detection failures or blurred images despite pre-scanning quality control. The absence of a local scanner necessitated international slide shipment, adding 8-12 weeks of delay and logistical complexity. Processing on a 4 GB RAM consumer laptop averaged 20 min per case (range: 5-60 min), with frequent application freezes on large tissue sections. Image files averaged 1.5 GB each at ×40 magnification, rapidly exhausting the 500 GB local storage capacity. The QuPath random tree classifier required manual annotation of representative tumor, stromal, and lymphoid regions on all 39 hematoxylin and eosin-stained cases before deployment on IHC slides, representing approximately 15-20 h of pathologist time. Exploratory concordance analysis suggested clinically meaningful agreement with expert pathologist scoring for three of the four biomarkers assessed, with detailed analytical validation reported separately. Conclusions:Implementing open-source digital pathology in sub-Saharan Africa is feasible but requires strategic planning around infrastructure, logistics, and training. We propose a practical framework addressing minimum hardware requirements, quality control protocols adapted to tropical environments, and a structured training program for pathologists. Our experience demonstrates that despite significant constraints, AI-assisted biomarker assessment can be successfully deployed in resource-limited settings, offering a pathway to improved diagnostic standardization where it is most needed.
Histopathology classification of breast cancer is still a case of difficulty in sorting out the complex morphology of the tissues, low inter-class variability, and manual biopsy analysis which is time consuming, costly, and inter-observer-dependent. Convolutional neural network-based methods have been shown to perform well for breast cancer imaging diagnosis, but they frequently lack the ability to capture the long-range spatial dependencies and lack interpretability for clinical decision-making. To tackle these problems, this study introduces an interpretable vision transformer ensemble framework TransBreast-Net for breast cancer histopathology classification. The proposed framework is based on transfer learning from large data augmentation and strong preprocessing on BreakHis and ICIAR datasets. The architectures of three transformers, namely CaiT S24 224, DeiT Small Patch16_224, and Swin Small Patch4_ Window7_224, are used to extract the complementary representations of local and global features. However, ensemble methods with Swin + DeiT for binary classification, and Swin + CaiT for multi-class classification are designed to enhance classification robustness and generalization. Experimental results show good performance with 99.35% accuracy in binary classification (benign vs. malignant) and 97% accuracy in multi-class classification (benign, in situ, invasive, and normal). In addition, visual explanations are embedded using Gradient-weighted Class Activation Mapping to enhance interpretability by identifying diagnostically relevant tissue regions that are important in the model decision-making process. The proposed TransBreast-Net framework exhibits high classification accuracy, good robustness, and has potential clinical applications for artificial intelligence diagnosis in breast cancer.
Fungal infections are an increasing risk to human health. They can pose a threat to life and cause a variety of health problems. The traditional diagnosis of fungal infections is challenging due to several reasons, such as the lack of clinical mycologists, costly procedures, a high time commitment, and the need for accuracy and specificity requirements. However, early fungal infections detection is essential for effective treatment. In this work, an explainable fine-tuned ResNet34 model for fungi classification is proposed by integrating transfer learning with a learnable threshold-based ReLULeaky activation function to enrich feature representation and classification performance. To improve feature extraction and convergence, the proposed learnable threshold approach dynamically adjusts activation levels during backpropagation. Our proposed fine-tuned ReLULeaky-ResNet34 method outperforms many tests, achieving the best accuracy (95.39%), F1-score (96%), and precision (97%). In addition, the model achieves a 99.40% area under the curve score, ensuring robust classification performance. The study highlights the efficacy of adaptive thresholds by methodically comparing the current and proposed activation functions. Interpretability confirms that the model focuses on biologically significant morphological features. These results demonstrate that our fine-tuned ReLULeaky-ResNet34 model outperforms for accurate and faster fungi classification.
Accurate pathology billing is crucial for financial sustainability, process optimization, and regulatory compliance. Current procedural terminology (CPT) codes 88300–88309 denote increasing complexity in gross and microscopic examination of surgical specimens, and correct assignment is critical to prevent misbilling, yet reports contain interpretive information subject to ambiguous assignment. Although machine learning methods have been developed for CPT code prediction, most fail to address challenges of integrating heterogeneous, variably structured reports across institutions with differing documentation practices. This study aims to advance machine learning approaches to leverage large-scale, diverse pathology report corpora and assess their adaptability across institutions. Transformer-based encoder models were initially fine-tuned on 59,923 cases from Dartmouth-Hitchcock Medical Center and further fine-tuned on 174,045 cases from Cedars-Sinai Medical Center, where final performance in CPT code prediction was evaluated. This sequential fine-tuning strategy was compared against models trained solely on the secondary institution's data (single-stage) and against traditional bag-of-words machine learning approaches (Naive Bayes, Random Forest, XGBoost). The best-performing model, a sequentially fine-tuned SciBERT-Longformer, achieved a macro-F1 of 0.8912. Sequential fine-tuning consistently yielded improvements in F1 for rare CPT codes for which there exists limited training data. As validation, SHAP analyses were used to interpret model predictions and verify alignment between influential keywords and established CPT code terminology. Model predictions were ultimately shown to be efficient, explainable, and grounded in clinical knowledge, further supporting potential for adaptation to real-world settings. Sequential fine-tuning on varied pathology corpora improved coding performance, highlighting the potential of institution-specific transformer models to streamline billing, reduce administrative burden, and enhance reimbursement accuracy.
Whereas anti-PD-(L)1 therapies are widely used in cancer treatment, only a subset of patients achieve long-term survival. Companion diagnostics based on PD-L1 expression have limited predictive power for these therapies, motivating development of additional predictive markers. We demonstrate a digital pathology (DP) method to quantify the density and epithelial infiltration of cytotoxic T cells near the tumor epithelial-stromal interface (ESI) using pancytokeratin-CD8 immunohistochemistry slides. We used POPLAR (n = 188), a phase 2 clinical trial of atezolizumab vs. docetaxel in previously treated non-small cell lung cancer, to train a machine learning outcome prediction model, generating a novel score, the Digital Assessment of Cytotoxic T cell Infiltration (DACTI). DACTI correlated with bulk RNAseq signatures of immune infiltration, providing verification that the DP measurements captured genuine aspects of the tumor microenvironment. A greater concentration of CD8+ T cells at the ESI and greater infiltration of those cells into the tumor epithelium were associated with benefit from atezolizumab, but not docetaxel. We validated DACTI in OAK (n = 879), a randomized phase 3 trial with treatment arms identical to POPLAR. Atezolizumab-treated DACTI-high patients in the validation set had longer overall survival than docetaxel-treated patients (n = 337, hazard ratio [HR] = 0.67, 95% confidence interval [CI]: 0.53-0.86; treatment interaction Cox model p-value = 0.028), including in patients with PD-L1 negative tumors (n = 61, HR = 0.47, 95% CI:0.27-0.84). This difference between arms was not observed in the DACTI-low group (n = 518, HR = 0.94, 95% CI: 0.78-1.13). These results suggest that automated analysis of pathology images may be able to direct immunotherapy treatments to patients who will most benefit.