Advances in digital pathology, image analysis, and artificial intelligence (AI) are rapidly transforming how pathologists and researchers interact with tissue samples and enable the development of diagnostic tools that harness high-resolution whole-slide images; these advances are in turn creating new opportunities for research, education, and routine clinical care globally. Liver disease is no exception, and digital pathology and AI have many applications in the diagnosis of liver cancer and liver diseases and in the assessment and management of transplantation. Although quantitative image analysis techniques have been applied to liver disease in research settings for over 50 years, recent improvements in image resolution, data storage, and the availability of advanced AI methods such as deep learning have driven multiple exciting developments. In this Review, we summarise the advancements in digital pathology, image analysis, and AI in liver disease. Key challenges such as access to and the logistics of using digital solutions, quality issues, and appropriate guidance in research and clinical use are reviewed, along with potential solutions to these challenges in the context of liver pathology and liver disease. Digital technologies are well established in liver pathology research, and access in clinical practice is increasing, with potential to address current laboratory challenges. Further evaluation is required to assess real-world effectiveness, clinical safety, and implementation of AI tools in liver pathology.
The current melanoma staging system predicts 74% of the variance in survival, with prognostic biomarkers subject to high levels of inter-observer variation. This work assesses whether a previously developed convolutional neural network (CNN) for invasive melanoma segmentation in whole slide images (WSIs) may reveal new insights into melanoma morphology and patient prognosis. This paper uses Cox proportional multivariate regression analyses to evaluate the ability of the CNN outputs to predict patient survival across 745 WSIs from 5 data sources. Five objective histomorphological parameters of tumour size and shape that are independently associated with overall and melanoma-specific survival were created from the CNN: tumour area(log) (HR 1.48 CI 1.30-1.68, p < 0.001), tumour perimeter(log) (HR 1.86 CI 1.48-2.32, p < 0.001), major axis length(log) (HR 1.88 CI 1.42-2.48, p < 0.001), Nodularity Index(log) (HR 1.77 CI 1.28-2.43, p < 0.001) and digital Breslow thickness(log) (HR 2.04, CI 1.63-2.54, p < 0.001). These results indicate that melanoma segmentation of the entire lesion within a WSI may be used to predict patient outcome. Moreover, this technology can be used to make new morphological discoveries to provide information not currently contained within our staging system (e.g. Nodularity Index), as well as provide objectivity and automation of current biomarkers (e.g. digital Breslow thickness). Further work is required to validate this initial discovery and evaluation.
Variation in digital whole slide images can be introduced across entire the digital pathology pathway, but one source of variation is the WSI scanner itself. We measured the range of variation across 9 different makes and models of WSI scanner using commercial and bespoke test objects. Slides were prepared representing a range of contrast, colour and resolution test signals. Differences in contrast were seen in both a shift of the maximum and minimum contrast levels and the introduction of non-linearity in some scanners. These contrast differences were also reflected in H&E images with additional shifts seen within the eosin colour space for some of the scanners. The resolution test pattern showed a wide variation in the ability of scanners to accurately reproduce the smaller details. This study describes the image variability of nine WSI scanners and highlights the importance of independent quality assessment and control.
This study systematically reviews and evaluates published research on machine learning models that integrate histopathology whole slide images and high-throughput -omic data to predict overall survival in cancer. A comprehensive search of PubMed, EMBASE, and Cochrane CENTRAL was conducted through August 12, 2024, with citation screening for additional studies. Eligible studies applied machine learning or deep learning methods to multimodal data combining pathology images and -omics. Data extraction followed the CHARMS checklist, and risk of bias was assessed using the PROBAST + AI tool. Narrative synthesis was conducted in line with PRISMA 2020 guidelines. Forty-eight studies published since 2017 met inclusion criteria, spanning 19 cancer types. All relied on The Cancer Genome Atlas dataset. Modelling approaches included regularised Cox regression (n = 4), classical machine learning (n = 13), and deep learning (n = 31). Reported concordance indices ranged from 0.550 to 0.857, with most multimodal models outperforming unimodal counterparts. However, all studies were assessed as having high or unclear risk of bias—most often due to limited external validation, insufficient reporting, and minimal assessment of clinical utility. This review highlights a rapidly evolving yet methodologically underdeveloped field. While model performance is promising, improvements in data standardisation, reporting practices, and real-world contextualisation are critical for clinical translation. This work was funded by the National Pathology Imaging Cooperative (NPIC), supported by UK Research and Innovation (Project no. 104687).
IntroductionMelanoma is the most aggressive form of skin cancer, with around 40–50% of cases harbouring BRAF mutations that can be effectively targeted to improve patient outcomes. Typically, BRAF mutations are identified through DNA sequencing, which can be costly and contribute significantly to turn-around-times. Predicting BRAF status directly from haematoxylin and eosin (H&E) whole-slide images (WSIs) could offer a faster and more accessible alternative within digital pathology workflows.MethodsIn this study, we evaluated six pathology foundation models for WSI-based BRAF mutation prediction using an attention-based multiple instance learning (ABMIL) framework. Training and testing models with five-fold cross-validation, using 1,109 WSIs from 745 patients and performing external validation with TCGA and VisioMel cohorts.ResultsModels which utilised H-Optimus-1 features achieved the highest AUROC values on the internal test set (0.80 (0.72, 0.88)) and VisioMel (0.79 (0.75, 0.83) external test set, while UNI2-h features demonstrated superior sensitivity and higher TCGA performance (AUROC: 0.76 (0.66, 0.86)). Subgroup analyses indicated that model performance was impacted by histopathological features such as ulceration and tissue area, while t-SNE plots showed that high-attention foundation model features tended to cluster based on the dataset the features were from, rather than biological signal.DiscussionOverall, foundation models combined with weakly supervised learning provide competitive performance for BRAF status prediction in melanoma WSIs, but our findings highlight the need for diverse melanoma datasets and the importance of evaluating model performance beyond a single metric when developing H&E-based computational biomarkers.
Background Histopathological assessment of lymph nodes is essential in breast cancer care. Artificial intelligence (AI) has shown promise in supporting this process, yet successful clinical adoption requires attention to human-AI interaction and user interface design. This study explored the user experience of an AI-based clinical decision support tool assisting pathologists in diagnosing breast cancer metastases. Methods We conducted a within-subjects mixed-methods study in which ten pathologists evaluated two AI-based clinical decision support user interface prototypes for breast cancer metastasis detection in sentinel nodes. ANNO displayed threshold based on-image annotations with minimal automation, whereas RAVS presented four likelihood regions of interest (ROI) per slide with automated measurement and case summaries. Quantitative measures included review time, diagnostic outcomes, and System Usability Scale (SUS) scores. Qualitative data were gathered via semi-structured interviews and analysed thematically. Results ANNO achieved higher usability scores (mean SUS 90.3) and was associated with a sense of control and resembled current diagnostic workflow, whereas RAVS (mean SUS 80.0) was preferred by participants valuing rapid navigation and automated measurements. Interview findings highlighted trade-offs: ANNO preserved diagnostic overview and autonomy, while RAVS improved navigational efficiency but reduced whole-slide context. ROI presentation strategy seemed to influence preference. Trust in AI was generally cautious and linked to needs for transparency and long-term clinical exposure. Conclusion Both interfaces were perceived as potentially clinically valuable but differed in their usability and workflow integration. Our exploratory findings indicates that for future development of AI-based decision aiding tools for this diagnostic task, a combination of high human control with targeted automation is advisable, combined with visual clarity, and an effective ROI presentation strategy.
Variation in H&E staining is often a confounder when comparing data in multi-site studies, and impacts on the generalisability of AI models. In this paper, slides containing serial sections of control liver tissue were sent out to eight clinical laboratories in England to stain over a two-week period. H&E staining optical density was quantified from the WSIs to measure the level of stain variation present within, and between, the eight laboratories. The laboratories used similar staining instruments and reagents, but their H&E stain durations varied. Intra-laboratory consistency was high (9-12% variation) but inter-laboratory variation was higher (maximum median difference of 31%) and the eight laboratories were statistically significantly different (p < 0.001). In addition, multiple colour-normalisation algorithms were applied to the image data using a range of target images. An automated nuclear count was then applied as an indicator of the impact of colour normalisation on the dataset and compared with pathologist manual 'ground truth' count. From three representative labs, the manual 'ground truth' nuclear counts ranged between 129 and 249, and the automated nuclear counts ranged between 152 and 621. Colour normalisation did improve inter-laboratory automated count variation (QCD) from 10% to 5%, but the results varied, particularly depending upon target image selection. The results from this work suggest that colour normalisation may improve consistency between slides but has the potential to be a confounder in its own right.
AIMS:In the end-to-end digital pathology workflow, variability can be introduced at each step, resulting in differences in the final image dataset. The effectiveness of quality control processes at each step of the workflow will impact the extent and relevance of this variability. METHODS:To assess the maturity of whole slide imaging (WSI) quality processes for the whole digital pathology workflow, we conducted an online questionnaire across 19 digitally active members of the Bigpicture consortium. RESULTS:A key finding was that a lower proportion of centres are implementing rigorous quality processes and checks processes at the post-scanning steps of the WSI workflow, such as 'digital reporting and display' (44%) and computational analysis (34%), when compared with pre-scanning steps such as 'pre-staining' (72%) and 'staining' (77%). CONCLUSIONS:This information allows us to identify priorities for quality improvement of the overall WSI workflow.
BACKGROUND:Liver biopsy assessment by pathologists remains the gold standard for diagnosing metabolic dysfunction-associated steatotic liver disease (MASLD). Current automated image analysis tools for patient risk stratification are often proprietary or not applicable to whole slide images (WSIs). Here, we introduce "Liver-Quant," an open-source Python package for quantifying steatosis and fibrosis in liver WSIs. METHOD:Liver-Quant leverages colour and morphological features to measure Steatosis Proportionate Area (SPA) and Collagen Proportionate Area (CPA). We evaluated the method using an internal dataset of 414 WSIs from adult patients (Leeds Teaching Hospitals NHS Trust, 2016-2022) and an external public dataset (109 WSIs). Semi-quantitative scores were extracted from pathological reports. The Spearman rank coefficient (ρ) assessed correlations between computed SPA/CPA and pathologist scores. RESULTS:Steatosis quantification showed a substantial correlation (ρ = 0.92), while fibrosis quantification yielded a moderate correlation (ρ = 0.51). We further investigated the impact of three staining dyes (Van Gieson (VG), Picro Sirius Red (PSR), and Masson's Trichrome (MTC)) on fibrosis quantification (n = 18). Stain normalisation yielded excellent agreement in CPA measurements across all three stains. Without normalisation, PSR achieved the strongest correlation with human scores (ρ = 0.9) followed by VG (ρ = 0.8) and MTC (ρ = 0.59). Finally, we explored the impact of apparent magnification on SPA and CPA. High-resolution images (0.25 or 0.50 μm per pixel (MPP)) were necessary for accurate SPA measurement, while lower resolution (10 MPP) sufficed for CPA measurements. CONCLUSIONS:Liver-Quant offers an open-source solution for rapid and precise MASLD quantification in WSIs applicable to multiple histological stains.
Hematoxylin and eosin (H&E) staining accounts for over 80% of slides stained worldwide. Although routinely used, there are high levels of variation between labs due to different staining methods. Staining is a pivotal part of slide preparation, but quality control is largely subjective, with overall clinical assurance provided by external quality assessment (EQA) services, underpinned by expert assessment. Digital pathology offers the potential to provide objective quantification of stain, through color analysis, to augment EQA assessment.This large-scale study evaluated H&E staining in 247 international labs participating in the UK NEQAS CPT EQA programme. Tissue sections were circulated to each lab to stain using their routine H&E staining protocol. The slides were reviewed by independent expert UK NEQAS CPT assessors, and quantitative digital analysis was conducted, comprising of H&E color deconvolution and color difference determination (ΔE).Most labs (69%) achieved an EQA score indicating good or excellent staining, with high inter-observer concordance to support this (92.5% within one mark of each other). H&E color difference, ΔE, showed 60% of labs were within 2 ΔE of the mean, which is considered as only perceptible through close observation. There was little correlation found between H&E intensity and assessor score, however, the H&E intensity ratio indicated a trend with assessor score suggesting there may be an optimal stain relationship that should be investigated further.The presented hybrid analysis combines expert analysis with objective data. This has the potential to inform upon optimal tissue staining and allows us to consider quantitative standards of H&E staining in pathology practice.
Digital Pathology has provided a platform to use Artificial Intelligence (AI) to assist pathologists with diagnosis and reporting. An AI tool is being developed that analyzes digital Hematoxylin and Eosin (stained tissue) images associated with a skin cancer case and pre-populates a report with required parameters. The aim of this AI pathology assistant is to save pathologist time and increase reporting efficiency. This study assessed ease of use and acceptability of a first iteration of the AI tool. Twelve pathologists were recruited across seven UK hospitals and participated in a think-aloud evaluation, completing a pathology report using the novel tool, after which they participated in a brief interview. The think-aloud identified several issues that can inform tool development to improve ease of use. AI performance (inaccuracy populating report items) constrained assessment of tool acceptability and added tasks to the reporting process. This finding emphasizes the importance of AI accuracy (1) for assessing if and how such tools can be integrated into clinician's workflow to increase efficiency, and (2) for cultivating clinician trust in tool performance to support adoption in practice.
Multimodal machine learning integrating histopathology and molecular data shows promise for cancer prognostication. We systematically reviewed studies combining whole slide images (WSIs) and high-throughput omics to predict overall survival. Searches of EMBASE, PubMed, and Cochrane CENTRAL (12/08/2024), plus citation screening, identified eligible studies. Data extraction used CHARMS; bias was assessed with PROBAST+AI; synthesis followed SWiM and PRISMA 2020. Protocol: PROSPERO (CRD42024594745). Forty-eight studies (all since 2017) across 19 cancer types met criteria; all used The Cancer Genome Atlas. Approaches included regularised Cox regression (n=4), classical ML (n=13), and deep learning (n=31). Reported c-indices ranged 0.550-0.857; multimodal models typically outperformed unimodal ones. However, all studies showed unclear/high bias, limited external validation, and little focus on clinical utility. Multimodal WSI-omics survival prediction is a fast-growing field with promising results but needs improved methodological rigor, broader datasets, and clinical evaluation. Funded by NPIC, Leeds Teaching Hospitals NHS Trust, UK (Project 104687), supported by UKRI Industrial Strategy Challenge Fund.
Purpose: There are many radiological datasets for breast cancer, some which have supported the development of AI medical devices for breast cancer screening and image classification. This review aims to identify mammography datasets (including digitised screen film mammography, 2D digital mammography and digital breast tomosynthesis) used in the development of AI technologies and present their characteristics, including their transparency of documentation, content, populations included and accessibility. Materials and methods: MEDLINE and Google Dataset searches identified studies describing AI technology development and referencing breast imaging datasets up to June 2024. The characteristics of each dataset are summarised. In particular, the accompanying documentation was reviewed with a focus on diversity and inclusion of populations represented within each dataset. Results: 254 datasets were referenced in the literature search, 190 were privately held, 36 had barriers which prevented access, and 28 were accessible. Most datasets originated from Europe, East Asia and North America. There was poor reporting of individuals' attributes: 32 (12 %) datasets reported race or ethnicity; 76 (30 %) reported female/male categories with only one dataset explicitly defining whether these categories represented sex or gender attributes. Conclusion: Through this review, we demonstrate gaps in the data landscape for mammography, highlighting poor representation globally. To ensure datasets in breast imaging have maximum utility for researchers, their characteristics should be documented and limitations of datasets, such as their representativeness of populations and settings, should inform scientific efforts to translate data-driven insights into technologies and discoveries.
CONTEXT.—:The current melanoma staging system does not account for 26% of the variance seen in melanoma-specific survival, therefore our ability to predict patient outcome is not fully elucidated. Morphology may be of greater significance than in other solid tumors, with Breslow thickness remaining the strongest prognostic indicator despite being subject to high levels of interobserver variation. The application of convolutional neural networks to whole slide images affords objective morphologic metrics, which may reveal new insights into patient prognosis. OBJECTIVE.—:To develop and evaluate a convolutional neural network for invasive cutaneous melanoma detection in whole slide images for the generation of objective prognostic biomarkers based on tumor morphology. DESIGN.—:One thousand sixty-eight whole slide images containing cutaneous melanoma from 5 data sets were used in the initial development and evaluation of the convolutional neural network. A 2-class tumor segmentation network with a fully convolutional architecture was trained using sparse annotations. The network was evaluated at per-pixel and per-tumor levels as compared to manual annotation, as well as variation across 3 scanning platforms. RESULTS.—:The convolutional neural network located conventional cutaneous invasive melanoma tissue with an average per-pixel sensitivity and specificity of 97.59% and 99.86%, respectively, across the 5 test sets. There were high levels of concordance between the tumor dimensions generated by the model as compared to manual annotation, and between the tumor dimensions generated by the model across 3 scanning platforms. CONCLUSIONS.—:We have developed a convolutional neural network that accurately detects invasive cutaneous conventional melanoma in whole slide images from multiple data sources. Future work should assess the use of this network to generate metrics for survival prediction.
Introduction: Digital pathology is an important resource in modern pathology education, with many examples of large databases of educational cases now available online. However, there remains a lack of standardization in retrieval and categorization of cases for training within and between institutions. Methods: Over 1600 teaching terms applicable to histopathology were developed and mapped to corresponding SNOMED CT terms to create a large directory of teaching cases. This was then integrated as a pre-defined list into a clinical PACS system for a national digital pathology project covering multiple hospitals and pathology training programs. Results: This resource allows easy allocation of teaching term labels to cases with educational value. A substantial catalog of educational cases has been generated already, with ongoing efforts to expand this with cases from routine clinical practice. The catalog is fully searchable by term for use in training and examinations. Conclusions: A large directory of digital pathology teaching cases was developed with associated corresponding SNOMED CT terms. The directory of terms will be shared with other healthcare providers to expand its use and utility.