Background:Deep learning (DL) can extract predictive and prognostic biomarkers from histopathology whole slide images, but its interpretability remains elusive. Methods:We introduce MoPaDi (Morphing histoPathology Diffusion), a framework for generating counterfactual explanations for histopathology that reveal which morphological features drive deep learning classifier predictions. MoPaDi combines self-supervised diffusion autoencoders with task-specific classifiers to manipulate images and flip predictions by altering visual features. To address weakly supervised scenarios common in pathology, it integrates multiple instance learning. We validate MoPaDi across four tasks: tissue type, cancer subtypes, slide origin, and a biomarker (microsatellite instability) classification. Counterfactuals are evaluated through pathologist studies and quantitative analysis. Results:MoPaDi achieves high image reconstruction quality (MS-SSIM 0.966-0.992) and good classification performance (AUCs 0.76-0.98). In a blinded observer study, two pathologists misclassified between 26.7% and 63.3% of synthetic images as real across all tasks, indicating that MoPaDi-generated images often exhibit high perceptual realism. Furthermore, counterfactual images revealed interpretable, pathology-consistent morphological changes recognizable by experts. Conclusion:MoPaDi is a practical and extensible framework for counterfactual explanations generation in computational pathology. It enables interpretable, model-specific insight into what morphological changes drive classification outcomes, improving interpretability in clinical deep learning models.
Background:Genomic data is essential for clinical decision-making in precision oncology. Bioinformatic algorithms are widely used to analyze next-generation sequencing (NGS) data, but they face two major challenges. First, these pipelines are highly complex, involving multiple steps and the integration of various tools. Second, they generate features that are human-interpretable but often result in information loss by focusing only on predefined genetic properties. This limitation restricts the full potential of NGS data in biomarker extraction and slows the discovery of new biomarkers in precision oncology. Methods:We propose an end-to-end deep learning (DL) approach for analyzing NGS data. Specifically, we developed a multiple instance learning DL framework that integrates somatic mutation sequences to predict two compound biomarkers: microsatellite instability (MSI) and homologous recombination deficiency (HRD). To achieve this, we utilized data from 3,184 cancer patients obtained from two public databases: The Cancer Genome Atlas (TCGA) and the Clinical Proteome Tumor Analysis Consortium (CPTAC). Results:Our proposed deep learning method demonstrated high accuracy in identifying clinically relevant biomarkers. For predicting MSI status, the model achieved an accuracy of 0.98, a sensitivity of 0.95, and a specificity of 1.00 on an external validation cohort. For predicting HRD status, the model achieved an accuracy of 0.80, a sensitivity of 0.75, and a specificity of 0.86. Furthermore, the deep learning approach significantly outperformed traditional machine learning methods in both tasks (MSI accuracy, p-value = 5.11×10-18; HRD accuracy, p-value = 1.07×10-10). Using explainability techniques, we demonstrated that the model's predictions are based on biologically meaningful features, aligning with key DNA damage repair mutation signatures. Conclusion:We demonstrate that deep learning can identify patterns in unfiltered somatic mutations without the need for manual feature extraction. This approach enhances the detection of actionable targets and paves the way for developing NGS-based biomarkers using minimally processed data.
Artificial intelligence (AI) methods enable humans to analyse large amounts of data, which would otherwise not be feasibly quantifiable. This is especially true for unstructured visual and textual data, which can contain invaluable insights into disease. The hepatology research landscape is complex and has generated large amounts of data to be mined. Many open questions can potentially be addressed with existing data through AI methods. However, the field of AI is sometimes obscured by hype cycles and imprecise terminologies. This can conceal the fact that numerous hepatology research groups already use AI methods in their scientific studies. In this review article, we aim to assess the contemporaneous use of AI methods in hepatology in Europe. To achieve this, we systematically surveyed all scientific contributions presented at the EASL Congress 2024. Out of 1,857 accepted abstracts (1,712 posters and 145 oral presentations), 6 presentations (∼4%) and 69 posters (∼4%) utilised AI methods. Of these, 55 posters were included in this review, while the others were excluded due to missing posters or incomplete methodologies. Finally, we summarise current academic trends in the use of AI methods and outline future directions, providing guidance for scientific stakeholders in the field of hepatology.
Background and aims: Hepatocellular carcinoma (HCC) is a highly fatal tumor, for which early detection and risk stratification is crucial, yet remains challenging. We aimed to develop an interpretable machine-learning framework for HCC risk stratification based on routinely collected clinical data. Methods: We leverage data obtained from over 900,000 individuals and 983 cases of HCC across two large-scale population-based cohorts: the UK Biobank study and the "All Of Us Research Program". For all of these patients, clinical data from timepoints years before diagnosis of HCC was available. We integrate data modalities including demographics, electronic health records, lifestyle, routine blood tests, genomics and metabolomics to offer a unique, multi-modal perspective on HCC risk. Results: Our random-forest-based model significantly outperforms all publicly available state-of-the-art risk-scores, with an AUROC of 0.88 both for internal and external test sets. We demonstrate robustness of our model across ethnic subgroups, a major advance over previous models with variable performance by ethnicity. Further, we perform extensive feature-importance analysis, showcasing our approach as an interpretable framework. We provide all model weights and an open-source web calculator to facilitate further validation of our model. Conclusion: Our study presents a robust and interpretable machine-learning framework for HCC risk stratification, which offers the potential to improve early detection and could ultimately reduce disease burden through targeted interventions. ### Competing Interest Statement JNK declares consulting services for Bioptimus, France; Owkin, France; DoMore Diagnostics, Norway; Panakeia, UK; AstraZeneca, UK; Mindpeak, Germany; and MultiplexDx, Slovakia. Furthermore, he holds shares in StratifAI GmbH, Germany, Synagen GmbH, Germany, and has received a research grant by GSK, and has received honoraria by AstraZeneca, Bayer, Daiichi Sankyo, Eisai, Janssen, Merck, MSD, BMS, Roche, Pfizer and Fresenius. TB has served on advisory boards for AdvanzPharma/Intercept Pharmaceuticals, SOBI, Novartis, and Gilead, and has received speaker fees from Falk Foundation, CSL Behring, Norgine, Intercept, Abbvie, Gilead, Merck, and Gore. OSMEN holds shares in StratifAI GmbH, Germany. Apichat Kaewdech re-ceived research grants or support from Roche, Roche Diagnostics, and Abbott Laboratories, and honoraria from Roche, Roche Diagnostics, Abbott Laboratories, and Esai. ### Clinical Protocols [https://github.com/schneiderlabac/hcc\_u\_soon][1] ### Funding Statement JC is supported by the Mildred-Scheel-Postdoktorandenprogramm of the German Cancer Aid (grant #70115730). JNK is supported by the German Federal Ministry of Health (DEEP LIVER, ZMVI1-2520DAT111), the Max-Eder-Programme of the German Cancer Aid (grant #70113864), the German Federal Ministry of Education and Research (PEARL, 01KD2104C; CAMINO, 01EO2101; SWAG, 01KD2215A; TRANSFORM LIVER, 031L0312A; TANGERINE, 01KT2302 through ERA-NET Transcan), the German Academic Exchange Service (SECAI, 57616814), the German Federal Joint Committee (Transplant.KI, 01VSF21048) the European Union's Horizon Europe and innovation programme (ODELIA, 101057091; GENIAL, 101096312) and the National Institute for Health and Care Research (NIHR, NIHR213331) Leeds Biomedical Research Centre. The views expressed are those of the author(s) and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care. DT is supported by the German Federal Ministry of Education and Research (SWAG, 01KD2215A; TRANSFORM LIVER), the European Union's Horizon Europe and innovation programme (ODELIA, 101057091). TL was funded by the German Cancer Aid (Deutsche Krebshilfe-DECADE 70115166), the Federal Ministry of Education and Research (BMBF - TRANSFORM LIVER 031L0312B) and the Federal Ministry of Health (BMG - DEEP LIVER 2520DAT111). TB is supported by the German Research Foundation (SFB1382 Project ID 403224013/B07). C.V.S is supported by a grant from the Interdisciplinary Centre for Clinical Research within the faculty of Medicine at the RWTH Aachen University (PTD 1-13/IA 532313), the Junior Principal Investigator Fellowship program of RWTH Aachen Excellence strategy and the NRW Rueckkehr Programme of the Ministry of Culture and Science of the German State of North Rhine-Westphalia. K.M.S is supported by the Federal Ministry of Education and Research (BMBF) and the Ministry of Culture and Science of the German State of North Rhine-Westphalia under the Excellence strategy of the federal government and the Laender as well as the NRW Rueckkehr Programme of the Ministry of Culture and Science of the German State of North Rhine-Westphalia. C.V.S and K.M.S are supported by the CRC 1382 project A11 and B09 funded by Deutsche Forschungsgesellschaft (DFG, German Research Foundation) - Project-ID 403224013 - FB 1382". D.Y.Z. is supported by the National Heart, Lung, and Blood Institute of the National Institute of Health under award number F30HL172382. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: UK Biobank data, including NMR metabolomics, are publicly available to bona fide researchers upon application at http://www.ukbiobank.ac.uk/using-the-resource/. Detailed information on predictors and endpoints used in this study is presented in Supplementary Tables 1-25. This study used data from the All of Us Research Pro-gram's Controlled Tier Dataset v7, available to authorized users on the Researcher Workbench. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes UK Biobank data, including NMR metabolomics, are publicly available to bona fide researchers upon application at http://www.ukbiobank.ac.uk/using-the-resource/. Detailed information on predictors and endpoints used in this study is presented in Supplementary Tables 1-25. This study used data from the All of Us Research Program's Controlled Tier Dataset v7, available to authorized users on the Researcher Workbench. [1]: https://github.com/schneiderlabac/hcc_u_soon
Liver cancer has high incidence and mortality globally. Artificial intelligence (AI) has advanced rapidly, influencing cancer care. AI systems are already approved for clinical use in some tumour types (for example, colorectal cancer screening). Crucially, research demonstrates that AI can analyse histopathology, radiology and natural language in liver cancer, and can replace manual tasks and access hidden information in routinely available clinical data. However, for liver cancer, few of these applications have translated into large-scale clinical trials or clinically approved products. Here, we advocate for the incorporation of AI in all stages of liver cancer management. We present a taxonomy of AI approaches in liver cancer, highlighting areas with academic and commercial potential, and outline a policy for AI-based liver cancer management, including interdisciplinary training of researchers, clinicians and patients. The potential of AI in liver cancer is immense, but effort is required to ensure that AI can fulfil expectations.
Latent diffusion models (LDMs) have emerged as a state-of-the-art image generation method, outperforming previous Generative Adversarial Networks (GANs) in terms of training stability and image quality. In computational pathology, generative models are valuable for data sharing and data augmentation. However, the impact of LDM-generated images on histopathology tasks compared to traditional GANs has not been systematically studied.We trained three LDMs and a styleGAN2 model on histology tiles from nine colorectal cancer (CRC) tissue classes. The LDMs include 1) a fine-tuned version of stable diffusion v1.4, 2) a Kullback-Leibler (KL)-autoencoder (KLF8-DM), and 3) a vector quantized (VQ)-autoencoder deploying LDM (VQF8-DM). We assessed image quality through expert ratings, dimensional reduction methods, distribution similarity measures, and their impact on training a multiclass tissue classifier. Additionally, we investigated image memorization in the KLF8-DM and styleGAN2 models.All models provided a high image quality, with the KLF8-DM achieving the best Frechet Inception Distance (FID) and expert rating scores for complex tissue classes. For simpler classes, the VQF8-DM and styleGAN2 models performed better. Image memorization was negligible for both styleGAN2 and KLF8-DM models. Classifiers trained on a mix of KLF8-DM generated and real images achieved a 4% improvement in overall classification accuracy, highlighting the usefulness of these images for dataset augmentation.Our systematic study of generative methods showed that KLF8-DM produces the highest quality images with negligible image memorization. The higher classifier performance in the generatively augmented dataset suggests that this augmentation technique can be employed to enhance histopathology classifiers for various tasks.
Detecting misleading patterns in automated diagnostic assistance systems, such as those powered by Artificial Intelligence, is critical to ensuring their reliability, particularly in healthcare. Current techniques for evaluating deep learning models cannot visualize confounding factors at a diagnostic level. Here, we propose a self-conditioned diffusion model termed DiffChest and train it on a dataset of 515,704 chest radiographs from 194,956 patients from multiple healthcare centers in the United States and Europe. DiffChest explains classifications on a patient-specific level and visualizes the confounding factors that may mislead the model. We found high inter-reader agreement when evaluating DiffChest's capability to identify treatment-related confounders, with Fleiss' Kappa values of 0.8 or higher across most imaging findings. Confounders were accurately captured with 11.1% to 100% prevalence rates. Furthermore, our pretraining process optimized the model to capture the most relevant information from the input radiographs. DiffChest achieved excellent diagnostic accuracy when diagnosing 11 chest conditions, such as pleural effusion and cardiac insufficiency, and at least sufficient diagnostic accuracy for the remaining conditions. Our findings highlight the potential of pretraining based on diffusion models in medical image classification, specifically in providing insights into confounding factors and model robustness.