
Advances in digital pathology, image analysis, and artificial intelligence (AI) are rapidly transforming how pathologists and researchers interact with tissue samples and enable the development of diagnostic tools that harness high-resolution whole-slide images; these advances are in turn creating new opportunities for research, education, and routine clinical care globally. Liver disease is no exception, and digital pathology and AI have many applications in the diagnosis of liver cancer and liver diseases and in the assessment and management of transplantation. Although quantitative image analysis techniques have been applied to liver disease in research settings for over 50 years, recent improvements in image resolution, data storage, and the availability of advanced AI methods such as deep learning have driven multiple exciting developments. In this Review, we summarise the advancements in digital pathology, image analysis, and AI in liver disease. Key challenges such as access to and the logistics of using digital solutions, quality issues, and appropriate guidance in research and clinical use are reviewed, along with potential solutions to these challenges in the context of liver pathology and liver disease. Digital technologies are well established in liver pathology research, and access in clinical practice is increasing, with potential to address current laboratory challenges. Further evaluation is required to assess real-world effectiveness, clinical safety, and implementation of AI tools in liver pathology.
Electroencephalography is the most commonly used diagnostic tool for epilepsy. However, interpreting electroencephalograms (EEGs) requires expertise that is not widely available. Advances in digital technology and wearables have enabled large-scale EEG recording, generating vast amounts of data that cannot be managed through traditional visual interpretation by experts. Artificial intelligence (AI) has the potential to augment human expertise and reduce workloads. The application of artificial neural networks in analysing clinical EEG recordings has led to major breakthroughs, bringing AI-based EEG interpretation closer to clinical implementation. In this Review, we summarise the most important research and development results in this field from a clinical perspective. We provide an overview of AI applications in spike and seizure detection; analysis of data from wearable electroencephalographs, patients who are critically ill, and epilepsy surgery; and the automated interpretation of clinical EEGs.
Accurate prediction of treatment response and selection of optimal treatments remain challenging in epilepsy management. With no reliable surrogate biomarkers for treatment response, the current process of selecting an antiseizure medication remains largely a trial-and-error approach. Other non-pharmacological treatment options, such as epilepsy surgery, are viable alternatives for patients with drug-resistant epilepsy. Statistical and machine learning techniques have been used to predict seizure outcomes associated with antiseizure medications and epilepsy surgery. Recent breakthroughs in deep learning have unveiled new pathways and opportunities, potentially revolutionising personalised treatment selection in health care. In this Review, we explore a broad range of studies that have used various statistical and machine learning methodologies, with particular emphasis on state-of-the-art deep learning techniques to predict the outcomes of both pharmaceutical and surgical treatments for epilepsy. We also review potential future research trajectories and address the inherent challenges of incorporating machine learning into the clinical management of epilepsy.
BACKGROUND:Colonoscopy is a frequently performed procedure to diagnose and treat colorectal cancer. However, endoscopist competence varies considerably. Adenoma detection rate (ADR) is the most established indicator of colonoscopy quality. We aimed to evaluate the effect of a computer-aided quality (CAQ) feedback and monitoring system on the ADR of endoscopists. METHODS:We conducted a cluster-randomised, controlled trial with stepped-wedge randomisation among patients aged 50-74 years with a positive faecal immunochemical test (FIT) result across three university hospitals in Denmark (Hilleroed University Hospital [site A], Bispebjerg University Hospital [site B], and Herlev University Hospital [site C]). The unit of randomisation was defined as the hospital site, resulting in three clusters. Patients received standard-of-care colonoscopies by endoscopists during the control period. During the intervention period, endoscopists received CAQ feedback according to the Copenhagen Colonoscopy Concept (CCC) after each colonoscopy on the following quality metrics: colonoscopy retraction score (retraction distance, tip retraction, and efficiency), cecal intubation rate, intubation time, number of polyps detected, withdrawal time, and distance to cecum. The primary outcome measure was ADR, analysed at individual colonoscopy level. The secondary outcome measures were polyp detection rate (PDR), adenomas per colonoscopy, polyps per colonoscopy, adenocarcinoma detection rate (ACDR), and withdrawal time. This trial is registered at ClinicalTrials.gov (NCT04862793) and is complete. FINDINGS:Between Feb 23, 2022, and April 19, 2024, 1109 patients in the intervention group (site A: 710, site B: 65, and site C: 334 patients) and 1001 patients in the control group (site A: 446, site B: 127, and site C: 428 patients) were enrolled and included in the final analysis (935 [44·3%] patients were female; 1175 [55·7%] were male). 17 endoscopists (site A: four endoscopists, site B: three endoscopists, and site C: ten endoscopists) participated and acted as their own controls. The intervention period had a higher ADR (539 [48·6%] vs 434 [43·4%]; odds ratio 1·24, 95% CI 1·02-1·52; p=0·032), PDR (54·2% vs 50·7%; 1·22, 1·00-1·49; p=0·048), ACDR (8·9% vs 6·5%; 1·51, 1·06-2·15; p=0·027), and withdrawal time (960 s vs 915 s; incident rate ratio (IRR; 1·11, 95% CI 1·06-1·17; p=0·003) than the control period. Per colonoscopy, there was no significant difference between the intervention period and control period for mean number of detected adenomas (1·00 [SD 1·58] vs 0·85 [1·34]; IRR 1·15, 95% CI 0·99-1·33; p=0·069), tubular adenomas (0·93 [1·51] vs 0·80 [1·33]; 1·15, 0·99-1·35; p=0·078), tubulovillous adenomas (0·02 [0·16] vs 0·02 [0·16]; 1·06, 0·54-2·06; p=0·86), sessile serrated adenomas (0·04 [0·48] vs 0·03 [0·17]; 1·11, 0·48-2·60; p=0·79), and polyps (1·36 [2·07] vs 1·27 [2·03]; 1·10, 0·95-1·27; p=0·22). INTERPRETATION:Implementation of the CCC CAQ-feedback system improved ADR, PDR, and ACDR, and prolonged the withdrawal time. Future studies should assess whether this implementation leads to a decrease in post-colonoscopy colorectal cancer and explore the effect in centres that do not use FIT-based screening. FUNDING:EU Horizon, Danish Cancer Society, Capital Region of Denmark, Danish Cancer Research Foundation, Vissing Foundation, Aase and Ejnar Danielsens Foundation, and Ambu.
BACKGROUND:Automated electrocardiogram (ECG) assessment tools to assist clinicians in diagnosis have improved substantially over the past decade; however, the full diagnostic and predictive potential of ECGs is limited when using traditional machine learning approaches because of an over-reliance on task-specific labels. We aimed to develop a large-scale ECG foundation model, ECG-CLIP, with improved generalisability and clinical applicability for ECG analysis. METHODS:This study used data from the Scripps Health GE MUSE system, a large-scale, retrospective dataset containing more than 1·7 million ECGs and paired clinician-overread ECG reports from 542 288 patients, collected in-clinic between Jan 15, 2008, and Jan 15, 2019. ECG-CLIP was pretrained in a two-stage pretraining framework consisting of masked ECG reconstruction and contrastive learning combining ECG data with textual annotations. MIMIC-IV, an independent, external dataset with more than 800 000 ECGs, was used to evaluate ECG-CLIP in cardiovascular disease detection (acute myocardial infarction, cardiac amyloidosis, and hypertrophic cardiomyopathy), cardiovascular disease prediction (atrial fibrillation from normal sinus rhythm), and adverse health outcome prediction (30-day emergency department and post-surgical mortality and 3-year onset of chronic kidney disease and type 2 diabetes) tasks. We compared the performance of ECG-CLIP against that of two supervised baseline models and one non-ECG foundation model, as well as three ECG signal-only foundation models: ECGFounder, DeepECG-SL, and DeepECG-SSL. Performance was evaluated using a five-fold cross-validation framework and reported as the area under the receiver operating characteristic curve (AUC). We assessed label efficiency by varying the number of positive labels in the training set of each downstream task. FINDINGS:ECG-CLIP consistently showed superior performance across all detection and prediction tasks compared with the supervised baseline and non-ECG foundation models, as well as a high label efficiency, reaching the same AUC in cardiovascular disease detection tasks as the best-performing comparator model with an average of 90·8% less training data. Compared with ECG signal-only foundation models, ECG-CLIP showed superior performance in low-label settings. When only ten positive labels were present in the dataset, ECG-CLIP significantly outperformed the runner-up model on tasks including acute myocardial infarction detection (AUC 0·910 [95% CI 0·903-0·916] vs 0·884 [0·876-0·892]), cardiac amyloidosis detection (0·790 [0·767-0·812] vs 0·777 [0·754-0·799]), hypertrophic cardiomyopathy detection (0·772 [0·751-0·793] vs 0·754 [0·733-0·774]), and atrial fibrillation prediction (0·777 [0·773-0·781] vs 0·765 [0·761-0·769]). The model has good performance in the detection of anterior and inferior acute myocardial infarction when using single ECG leads, outperforming supervised baseline models when using leads typically considered less informative. INTERPRETATION:ECG-CLIP is a generalisable foundation model that can learn clinically meaningful representations from ECGs and paired reports for diagnostic and prediction tasks. This approach enables accurate, data-efficient, and interpretable risk predictions across diverse clinical tasks, supporting scalable deployment in real-world and resource-limited settings. FUNDING:The VoLo Foundation and the National Center for Advancing Translational Sciences at the National Institutes of Health.
Despite decades of health-care digitalisation efforts worldwide, health information systems remain highly fragmented, with multiple vendor-specific silos that communicate through incomplete solutions. This fragmentation prevents the creation of real-time, lifelong patient health records and becomes increasingly problematic as demand grows for person-centred care, data-driven clinical practice, and greater patient involvement in health-care decisions. To address these challenges and establish a foundation for a nationwide electronic health record (EHR), the Spanish Ministry of Health commissioned a steering committee to develop recommendations based on a comprehensive national consensus. The committee conducted a Delphi study comprising 45 items across four domains, which was distributed to 220 experts from June 23, 2023, to Sept 26, 2023. With a response rate of 69·1% (152/220), the study achieved consensus in a single round, with all items reaching the pre-established threshold of greater than or equal to 70% agreement (scores 7-9 on a 9-point Likert scale), and consensus ranging from 118 (77·6%) to 151 (99·3%) of 152 responses (44 items ≥80%). The resulting recommendations were externally validated by an international advisory board, which assessed their consistency and alignment with global best practices and standards. The final set included 20 recommendations across four domains: justification of need (2 items), functional characteristics (7 items), technical characteristics (6 items), and governance (5 items). These recommendations provide a roadmap for developing a robust, integrated national health information system centred on a standardised, longitudinal EHR. The proposed approach moves beyond generic calls for interoperability by embedding clinical knowledge into open, standardised EHR architectures through ontology-driven semantic integration, supported by federated governance and citizen-controlled data use. This roadmap equips Spain to implement a longitudinal, knowledge-driven national record while providing a scalable model for other countries transforming fragmented health information systems.
Artificial intelligence (AI)-based prediction models, including risk scoring systems and decision support systems, are being increasingly adopted in health care. Addressing AI fairness is essential to fighting health disparities and ensuring equitable model performance and patient outcomes. However, numerous and conflicting definitions of fairness complicate this effort. In this Viewpoint, we aim to support the transition of AI fairness from theory to practice using appropriate fairness metrics. We assess the relation of 27 fairness definitions identified in the literature to the model's intended use, type of decision influenced, and ethical principles of distributive justice. Because of limitations in some notions of fairness, we argue that clinical utility, performance-based metrics (such as area under the receiver operating characteristic curve), calibration, and statistical parity are the most relevant group-based metrics for medical applications. Through two use cases, we show that different metrics might be applicable depending on the intended use and ethical framework. Our approach provides practical guidance for fair AI development, helping AI developers and assessors to evaluate model fairness and understand the effects of bias mitigation strategies, thereby supporting equitable AI-based implementations.
BACKGROUND:Identifying patients with glaucoma who are at risk of rapid disease progression is crucial to preventing vision loss. We aimed to develop and externally validate G-PROG, a deep learning model that predicts 2-5-year glaucoma progression from baseline colour fundus photographs (CFPs). METHODS:G-PROG was trained and validated on data from a single centre (UZ Leuven, Leuven, Belgium); the other datasets (Brussels, Belgium; Liège, Belgium; Tampere, Finland; Mainz, Germany; and Hangzhou, China) served as external test sets. Across six glaucoma departments, we analysed 161 827 fundus images from 127 962 visits (13 913 patients), totalling 128 021 eye-years of follow-up. Progression was defined by the G-RISK slope, calculated via within-eye linear regression on longitudinal G-RISK predictions over follow-up intervals of 2-5 years. G-RISK is a previously validated deep learning model that quantifies glaucomatous optic nerve damage from CFPs. We trained 20 G-PROG configurations with varying inclusion criteria applied to the number of visits, image quality, time between visits, and G-RISK at baseline. Performance was evaluated using the area under the receiver operating characteristic curve (AUC), the coefficient of determination (R2), and explained variance score (EVS). G-RISK slope as a progression biomarker was validated against the visual field mean deviation (MD) slope and average retinal nerve fibre layer thickness (RNFL) slope. FINDINGS:Significant AUC values were obtained in 18 out of 20 model configurations, with internal validation reaching a maximum AUC of 0·98 (95% CI 0·97-1·00) across follow-up intervals (2-5 years). In glaucomatous eyes with a baseline G-RISK exceeding 0·6, the maximum AUC was 0·92 (0·85-0·98). For external validation, the predictions from the eight top-performing configurations (selected based on positive R2 and minimal discrepancy between R2 and EVS in internal validation) were averaged. Maximum AUC values ranged from 0·74 to 0·86 across the five test datasets. G-RISK slope showed significant agreement with established progression markers, with maximum AUCs of 0·82 for MD slope and 1·00 for average RNFL slope. INTERPRETATION:Externally validated across five international cohorts, G-PROG predicts 2-5-year glaucoma progression from baseline CFPs. Prospective evaluation is warranted to assess whether G-PROG can improve risk stratification and resource allocation in glaucoma care. FUNDING:This work was funded and supported by grants from the National Medical Research Council, National Research Foundation Singapore, National Health Innovation Centre Singapore, SingHealth and Duke-NUS, Duke-NUS, the Singapore Eye Research Institute and Nanyang Technological University and the Singapore Eye Research Institute, the Competitive Research Funding of the Pirkanmaa Wellbeing Services County, the LUX-Foundation for Glaucoma Research, state funding for university-level health research at Tampere University Hospital, Wellbeing Services County of Pirkanmaa, the Tampere University Hospital Support Foundation, and the Belgian Ophthalmology Cooperation in Clinical Sciences initiative hosted by the Funds for Research in Ophthalmology.
BACKGROUND:Non-communicable diseases are the leading cause of death globally. Smartphone apps can offer benefits for individuals, health-care professionals, and governments in the prevention and management of such conditions. We aimed to systematically evaluate the effectiveness of app-based interventions in improving the outcomes of non-communicable diseases and in modifying their metabolic and behavioural risk factors. METHODS:For this umbrella review and meta-analysis, we searched eight databases (Embase, Epistemonikos, IEEE Xplore Digital Library, APA PsycInfo via Ovid, PubMed, Scopus, Web of Science Core Collection, and Cochrane Central Register of Controlled Trials) for systematic reviews with meta-analysis published between Jan 1, 2013, and Jan 10, 2024, with no restrictions by geographical location or language. Additional studies were located through citation chaining and searching of reference lists. Eligible studies reviewed randomised controlled trials or controlled studies focused on adults (aged ≥18 years) with or at risk of non-communicable diseases and the use of app-based interventions for managing or improving the outcomes of these diseases and related health and risk factors, both metabolic and behavioural. Two investigators (EK and MFV) used COVIDENCE software to screen abstracts and full texts and to subsequently extract data from eligible studies. In case of missing data, authors of the relevant articles were contacted for unreported data or additional details. Effect sizes were measured as the standardised mean difference (SMD) and were aggregated through meta-analyses. 95% CIs for each review were synthesised using a random-effects model and prediction intervals were based on a t distribution. When more than ten reviews were available, publication bias was assessed visually through the inspection of funnel plots and by Egger's test; if bias was suspected, a trim-and-fill analysis was applied to estimate a revised effect size. The quality of the included reviews was evaluated with the AMSTAR 2 checklist, the certainty of evidence for each outcome was assessed using the GRADE framework and the Ioannidis criteria, and heterogeneity was measured by calculating I2 values. This study was registered with PROSPERO, CRD42023426735. FINDINGS:Of 6951 unique records identified by our searches, 383 underwent full-text review and 78 systematic reviews with meta-analysis, covering 31 outcomes, were included in the study. These reviews covered 496 primary studies and involved a total of 177 373 participants. App-based interventions were found to be effective in lowering diastolic blood pressure (SMD -0·414 [95% CI -0·606 to -0·221], I2=93%), systolic blood pressure (-0·444 [-0·689 to -0·199], I2=96%), glycated haemoglobin (-0·587 [-0·715 to -0·460], I2=91%), fasting blood glucose concentration (-1·189 [-1·605 to -0·774], I2=93%), 2 h postprandial glucose concentration (-1·229 [-1·609 to -0·848], I2=95%), anxiety (-0·215 [-0·407 to -0·023], I2=92%), depression (-0·097 [-0·176 to -0·019], I2=77%), stress (-0·336 [-0·528 to -0·143], I2=87%), bodyweight (-0·427 [-0·594 to -0·260], I2=91%), BMI (-0·265 [-0·522 to -0·007], I2=92%), waist circumference (-0·310 [-0·464 to -0·156], I2=75%), and sedentary time (-0·600 [-1·121 to -0·079], I2=22%). Additionally, apps significantly improve diet quality (0·551 [0·261-0·842], I2=91%), exercise capacity (0·259 [0·119-0·398], I2=17%), moderate-to-vigorous physical activity (0·240 [0·025-0·456], I2=75%), number of steps taken daily (0·489 [0·209-0·770], I2=83%), multiple physical activity outcomes (0·466 [0·243-0·689], I2=90%), mindfulness (0·293 [0·177-0·409], I2=77%), wellbeing (0·186 [0·065-0·307], I2=74%), quality of life (0·227 [0·083-0·371], I2=80%), and medication adherence (0·688 [0·410-0·965], I2=88%) when compared with control groups. However, no significant effect was found on cardiovascular mortality; HDL, LDL, total cholesterol, or triglyceride concentrations; body fat; distress; fruit and vegetable intake; hospitalisation; or smoking abstinence. Of the 78 systematic reviews, only one (1%) was rated as being of high quality, with three (4%) of moderate quality, 17 (22%) of low quality, and 57 (73%) of critically low quality. According to Ioannidis criteria, five outcomes were categorised as having highly suggestive (class II) evidence, with eight outcomes having suggestive evidence (class III), eleven outcomes having weak evidence, and seven outcomes categorised as non-significant. INTERPRETATION:Our analyses provide evidence that app-based interventions support significant improvements in multiple outcomes of and risk factors for non-communicable diseases, including cardiovascular diseases, glucose control, mental health outcomes, physical activity, and quality of life. The integration of app-based interventions into health-care systems should be prioritised to enhance patient care and health outcomes. FUNDING:None.
BACKGROUND:Current artificial intelligence (AI) models for medical imaging predominantly focus on a single imaging modality and a single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training these models typically requires large, well labelled datasets, which are costly and labour intensive to prepare. We aimed to train and evaluate an AI model that can interpret diverse imaging modalities across specialties while maintaining robust performance within each modality. METHODS:We developed Multimodal, Multi-Disease Medical Imaging Foundation Model (MerMED-FM), a multi-specialty model trained using self-supervised learning and a memory module. MerMED-FM was pretrained on publicly sourced, unlabelled medical images from 12 specialties and seven imaging modalities: chest x-rays, CT, ultrasound, histopathology, colour fundus photography (CFP), optical coherence tomography (OCT), and dermatoscopy. After pretraining, the model was fine-tuned, validated, and evaluated for the diagnosis of a range of diseases on 26 public datasets and five private datasets comprising radiology, histopathology, and ophthalmology images. MerMED-FM was compared against a general-domain vision foundation model, various specialist single-modality foundation models, and a multispecialty foundation model. Models were fine-tuned using 10%, 30%, 50%, and 100% of data, with primary comparative analyses conducted using a 10% label fraction. The primary outcome was the area under the receiver operating characteristic curve (AUROC), which was summarised by imaging modality. FINDINGS:MerMED-FM was trained on around 3·3 million images from 53 publicly available, unlabelled datasets, comprising 713 931 chest x-rays, 292 353 CT slices, 389 885 ultrasound frames, 1 017 712 pathology patches, 333 099 CFP images, 176 719 OCT slices, and 401 059 dermatoscopy images. Strong performance was achieved across all modalities at a label fraction of only 10%, with mean AUROC values of 0·844 for chest x-rays, 0·906 for CT, 0·818 for ultrasound, 0·908 for histopathology, 0·810 for CFP, 0·962 for OCT, and 0·827 for dermatoscopy. INTERPRETATION:MerMED-FM has the potential to be a highly adaptable, versatile, cross-specialty foundation model that enables robust interpretation of medical imaging across diverse medical disciplines. FUNDING:National Medical Research Council, Singapore and the Agency for Science, Technology and Research, Singapore.
BACKGROUND:The European Organisation for Research and Treatment of Cancer's STRASS trial, the only completed randomised study of preoperative radiotherapy in retroperitoneal sarcoma, showed no overall benefit. Its subgroup analysis has been interpreted as supporting radiotherapy for all patients with liposarcoma, whereas the STREXIT extension has been interpreted as supporting radiotherapy for well differentiated liposarcoma and low-grade or intermediate-grade dedifferentiated liposarcoma. We aimed to identify subsets of patients who might benefit from preoperative radiotherapy and to quantify this benefit in terms of abdominal recurrence. METHODS:In this artificial intelligence (AI)-based reanalysis of the STRASS dataset, we trained a random survival forest model on all 266 randomly assigned patients (radiotherapy plus surgery vs surgery alone) to predict 5-year abdominal recurrence-free survival under both treatment options, based on pretreatment variables. These predictions were used to fit an optimal policy tree (OPT) that partitions patients into nodes by predicted abdominal recurrence-free survival benefit from radiotherapy. We then compared outcomes in OPT-defined subgroups, STRASS subgroups, and STREXIT subgroups through Kaplan-Meier curves, and Fine-Gray competing-risks models within STRASS. FINDINGS:The OPT partitioned the cohort into seven subgroups; three subgroups (152 of 266 patients) were predicted to benefit from radiotherapy, and for two of these subgroups the benefit was statistically significant: patients with well differentiated liposarcoma aged 60 years or younger and patients with dedifferentiated liposarcoma who underwent curative-intent surgery. In these two subgroups combined, 3-year abdominal recurrence-free survival was 79% (95% CI 70-89) with radiotherapy versus 58% (48-72) without (hazard ratio [HR] 0·40 [95% CI 0·22-0·71], p=0·0016), and the cumulative incidence of abdominal recurrence was significantly lower with radiotherapy (17% [95% CI 9-27] vs 33% [22-45] without radiotherapy; HR 0·40, p=0·0090). Inverse probability of censoring weight-adjusted 5-year abdominal recurrence-free survival estimates showed similar absolute gains (25·6 percentage points). A Cox model found a significant radiotherapy-age interaction (p=0·012). By contrast, STRASS-defined and STREXIT-defined radiotherapy subgroups did not show significant abdominal recurrence-free survival improvement when re-evaluated within STRASS. INTERPRETATION:The AI-guided partition of STRASS identified younger patients with well differentiated liposarcoma and patients with dedifferentiated liposarcoma and curative-intent surgery as subgroups in which radiotherapy appears to meaningfully improve abdominal recurrence-free survival, whereas radiotherapy strategies for all well differentiated liposarcoma and low-grade or intermediate-grade dedifferentiated liposarcoma were not supported by the randomised controlled trial data. These findings argue for a more selective use of preoperative radiotherapy in retroperitoneal sarcoma and provide a concrete basis for focused future trials, which, if successful, could substantiate these findings before these strategies become standards of care. FUNDING:National Cancer Institute and Memorial Sloan Kettering Cancer Center.