Carotid artery stenosis (CAS) is a significant risk factor for ischemic stroke and a marker of systemic vascular burden. Hemodynamically significant CAS may induce microvascular changes in the retinal circulation, detectable through near-infrared (NIR) reflectance imaging. This study explored the potential of machine learning to identify CAS from NIR fundus images. We used a dataset of 1483 NIR images from 469 patients, including 121 with confirmed CAS and 348 controls. The data were split into training and testing sets (70:30). Two models were developed: (1) a logistic regression model using quantitative vascular biomarkers computed by a specialized deep learning segmentation model, and (2) a deep learning model trained on raw NIR images. Both models achieved an area under the receiver operating characteristic curve of 0.70. These results highlight the potential of machine learning for detecting CAS in NIR fundus imaging, though further validation in larger cohorts is needed.
While circadian phenotypes of paroxysmal atrial fibrillation (AF) have been observed in patient cohorts, differences in parameter definition and methodological limitations have left the existence and clinical relevance of distinct AF circadian phenotypes unclear. We hypothesize that paroxysmal AF comprises multiple circadian phenotypes (‘chronophenotypes’), each associated with different survival outcomes and population characteristics. We analyzed 24 h Holter recordings and clinical data from 58 995 examinations collected in 20 primary care facilities in Israel. AF episodes were detected using ArNet2 , a deep learning model for AF detection trained on over 51 000 h of Holter recordings. Unsupervised hierarchical clustering was then applied to identify different chronophenotypes of AF based on the short-term AF burden. Each chronophenotype was further characterized by demographic, clinical, and treatment differences, including survival outcomes. The analysis resulted in the following distinct chronophenotypes: Nocturnal-to-Morning (12 AM–10 AM), Evening-to-Early Morning (3 PM-4 AM), and Daytime (9 AM-7 PM). These chronophenotypes were associated with different AF burden and clinical outcomes. The Evening-to-Early Morning chronophenotype was associated with a higher AF burden ( $p \lt $ 0.001) than the other paroxysmal AF chronophenotypes. However, the Nocturnal-to-Morning and Daytime chronophenotypes were associated with a higher risk of mortality; Daytime was also associated with a higher risk of heart failure. The study introduces a new approach to discovering chronophenotypes from Holter recordings. The results support the existence of distinct AF chronophenotypes, which are associated with significantly different average AF burden ( $p \lt $ 0.001) and clinical outcomes. The open-source implementation is publicly available at https://github.com/aim-lab/AFtoolkit .
Excessive daytime sleepiness (EDS) refers to a physiological state where individuals have difficulty remaining alert during the day. Managing EDS is particularly challenging to study and treat due to its multifaceted nature. Assessment methods include both subjective and objective approaches. Subjective evaluation often relies on simple, widely accepted, and widely used questionnaires; however, these tools are inherently limited by self-reporting bias. Objective assessment, on the other hand, primarily involves two well-known and reliable tests, but these are costly, time-consuming, and impractical for use outside of sleep units. Therefore, developing an objective tool that can quickly and accurately detect a decline in alertness, while remaining reliable, easy to use, and affordable, is of critical importance for sleep clinicians, safety organizations, and researchers. According to PRISMA guidelines, we did a systematic analysis of 95 studies that used photoplethysmography (PPG) for assessing EDS, drowsiness, and/or fatigue during the last 15 years (2010-2025). With advances in wearable technology, particularly through PPG and artificial intelligence, achieving this goal may be attainable. The next essential step is rigorous validation against established gold-standard tests to ensure the tool meets scientific and clinical standards for widespread adoption.
Open-source datasets of digital fundus images annotated for diabetic retinopathy (DR) staging often suffer from inconsistent or inaccurate labeling but manually relabeling them at scale is impractical. A key need is a semi-supervised methodology which automatically flags images most likely to be mislabeled for expert review. We present DRStageNet2, our model improvement to DRStageNet, which results exclusively from our data refinement strategy with an expert-in-the-loop who reannotates images selectively in a data-driven manner across multiple rounds of model retraining and inference. Through this process, the expert reviewed a total of 3984 images among the 91 984 images used (2.8%). The expert modified the labels a total of 2592 (65%) of the reviewed images. DRStageNet2 improved out-of-domain generalization performance across five external validation sets. DRStageNet2’s multiclass accuracy of International Clinical Diabetic Retinopathy Scale scores 0–4 improved across all datasets with a mean increase of 3.6%. Its quadratic-weighted Cohen kappa score (Q-kappa) increased from an initial average of 0.89 (range: 0.86–0.91) to 0.91 (range: 0.89–0.93) with improvements of up to 6.2% in one dataset. DRStageNet2 then, achieves a Q-kappa score at the upper end of previously reported inter-rater agreement between retinal specialists.
Objective. Photoplethysmography, a non-invasive optical technique that measures changes in blood volume in the microvascular bed of tissue, offers a promising approach for monitoring physiological changes during sleep. This study evaluates differential photoplethysmography signal patterns that can distinguish between apneas vs hypopneas, which are key features of sleep-related breathing disorders. Approach. We analyzed data from 263 severe (apnea hypopnea index ⩾30) obstructive sleep apnea patients, using recordings from the Multi-Ethnic Study of Atherosclerosis. Over 57 000 respiratory events occurring during stage N2 sleep were included. A machine learning model was trained on 89 features derived from the photoplethysmography signal, using the pyPPG toolbox, to classify: apneas vs hypopneas in the supine and lateral sleep posture, and posture-specific differences for each respiratory event type. Main results. Results showed that photoplethysmography signal characteristics significantly differed between apneas vs hypopneas. The model achieved an area under the receiver operation characteristic curve of 0.80 in the lateral posture and 0.83 in the supine posture. However, classification performance was low when distinguishing between apneas and hypopneas in the lateral vs the supine position with an area under the receiver operation characteristic curve of 0.62 for apneas and 0.64 for hypopneas. The discriminative signal features were consistent across different periods of the night. Significance. These findings indicate that photoplethysmography can detect meaningful differences in sleep-related breathing events and support its potential as a foundation for wearable diagnostic and monitoring tools that are personalized, accessible, and cost-effective.
A major challenge in translating medical AI systems into clinical practice is their limited generalization. In the field of physiological time series analysis, we propose a fine-tuning framework that leverages multiple small annotated datasets from diverse domains to improve out-of-distribution generalization performance (OOD-GP). Through an ablation study, we demonstrate the performance of our framework by evaluating the role of incorporating a greater number of independent datasets for tine-tuning to improve OOD-GP. Our experiments involve thirteen publicly available electrocardiogram and electroencephalogram datasets across four distinct tasks. In addition, we develop a method to measure the alignment of the latent space of target domains. We use this method to interpret our results, suggesting that multi-source domain training facilitates the learning of robust cross-domain features while minimizing learning of shortcut features. To support further research, we provide reproducible source code, establishing a framework and benchmark for studies on OOD-GP [URL provided upon publication.]
Introduction: Premature Ventricular Contractions (PVCs) are common cardiac arrhythmias originating from the ventricles. Accurate detection remains challenging due to variability in electrocardiogram (ECG) waveforms caused by differences in lead placement, recording conditions, and population demographics. Methods: We developed uPVC-Net, a universal deep learning model to detect PVCs from any single-lead ECG recordings. The model is developed on four independent ECG datasets comprising a total of 8.3 million beats collected from Holter monitors and a modern wearable ECG patch. uPVC-Net employs a custom architecture and a multi-source, multi-lead training strategy. For each experiment, one dataset is held out to evaluate out-of-distribution (OOD) generalization. Results: uPVC-Net achieved an AUC between 97.8% and 99.1% on the held-out datasets. Notably, performance on wearable single-lead ECG data reached an AUC of 99.1%. Conclusion: uPVC-Net exhibits strong generalization across diverse lead configurations and populations, highlighting its potential for robust, real-world clinical deployment.
Glaucomatous optic neuropathy (GON), affecting an estimated 64.3 million people globally, causes irreversible vision loss when not detected early. Traditional diagnosis requires time-consuming ophthalmic examinations by specialists. Recent deep learning models for automating GON detection from colour fundus photographs (CFP) have shown promise but often suffer from limited generalizability across different ethnicities, disease groups and examination settings. To address these limitations, we introduce GONet, a robust deep learning model developed using seven independent datasets, including over 119 000 CFPs with gold-standard annotations and from patients of diverse geographic backgrounds. GONet consists of a DINOv2 pre-trained self-supervised vision transformer fine-tuned using a multisource domain strategy. GONet demonstrated high out-of-distribution generalizability, with an AUC of 0.88-0.99 in target domains. GONet performance was similar or superior to state-of-the-art works and the cup-to-disc ratio, by up to 18.4%. GONet is available via Lirot.ai (www.aimlab-technion.com/lirot-ai). We also contribute a new dataset consisting of 747 CFPs with GON labels as open access, available at https://doi.org/10.13026/pdxv-m215.
Heart failure (HF) affects 11.8% of adults aged 65 and older, reducing quality of life and longevity. Preventing HF can reduce morbidity and mortality. We hypothesized that artificial intelligence (AI) applied to 24-hour single-lead electrocardiogram (ECG) data could predict the risk of HF within five years. To research this, the Technion-Leumit Holter ECG (TLHE) dataset, including 69,663 recordings from 47,729 patients, collected over 20 years, was used. Our deep learning model, DeepHHF, trained on 24-hour ECG recordings, achieved an area under the receiver operating characteristic curve of 0.80 that outperformed a model using 30-second segments and a clinical score. High-risk individuals identified by DeepHHF had a two-fold chance of hospitalization or death incidents. Explainability analysis showed DeepHHF focused on arrhythmias and heart abnormalities. This study highlights the feasibility of deep learning to model 24-hour continuous ECG data, capturing paroxysmal events essential for reliable risk prediction. Artificial intelligence applied to single-lead Holter ECG is non-invasive, inexpensive, and widely accessible, making it a promising tool for HF risk prediction.
Self-supervised learning (SSL) has enabled Vision Transformers (ViTs) to learn robust representations from large-scale natural image datasets, enhancing their generalization across domains. In retinal imaging, foundation models pretrained on either natural or ophthalmic data have shown promise, but the benefits of in-domain pretraining remain uncertain. To investigate this, we benchmark six SSL-pretrained ViTs on seven digital fundus image (DFI) datasets totaling 70,000 expert-annotated images for the task of moderate-to-late age-related macular degeneration (AMD) identification. Our results show that iBOT pretrained on natural images achieves the highest out-of-distribution generalization, with AUROCs of 0.80-0.97, outperforming domain-specific models, which achieved AUROCs of 0.78-0.96 and a baseline ViT-L with no pretraining, which achieved AUROCs of 0.68-0.91. These findings highlight the value of foundation models in improving AMD identification and challenge the assumption that in-domain pretraining is necessary. Furthermore, we release BRAMD, an open-access dataset (n=587) of DFIs with AMD labels from Brazil.
The past few years have witnessed a rapid proliferation of AI systems for the automated interpretation of digital fundus images (DFIs). Although these systems have achieved strong performance on retrospective datasets and several have recently obtained FDA approval, their comparative performance against human expert interpretation remains insufficiently characterized. The aim of this study was to compare the diagnostic performance of an AI system, Lirot.ai, with that of board-certified ophthalmologists in detecting referable eye diseases from single, color non-mydriatic DFIs. This study evaluated the detection of three vision threatening ophthalmic diseases: (1) referable diabetic retinopathy (rDR), (2) referable age-related macular degeneration (rAMD), and (3) glaucomatous optic neuropathy (GON). Primary endpoints included the sensitivity and specificity for detection of each referable eye disease by Lirot.ai versus human readers. Secondary endpoints include inter-reader agreement ( κ ), human-hours required to annotate, and subset analysis of detection performance on images marked ungradable by human readers. A total of 10 ophthalmologists were recruited for the study: 80% of whom were specialists and 80% of whom had ⩾6 years of clinical experience. Fleiss’ k among readers was 0.401 for GON detection, 0.663 for rAMD detection, and 0.719 for rDR detection. Lirot.ai’s sensitivities across GON, rAMD, and rDR detection were the maximum among readers while maintaining specificities within the observed range of reader performance. Lirot.ai detected GON with an area under the receiver operating curve (AUROC) of 0.860 (sensitivity and specificity of 88.2% and 74.3%, respectively), rAMD with an AUROC of 0.923 (sensitivity and specificity of 92.9% and 89.6%, respectively), and rDR with an AUROC of 0.954 (sensitivity and specificity of 91.7% and 94.0%, respectively) where sensitivities and specificities maximize Youden’s Index. Lirot.ai met or exceeded readers’ disease detection across GON, rDR, and rAMD as evident by readers’ sensitivity-specificity performance along or below the ROC curve. Lirot.ai demonstrated expert-level performance across three referable eye diseases and even surpassed human readers in GON detection, demonstrating the potential of AI to provide consistent, high-quality interpretation of fundus images. These results highlight how AI systems can enhance clinical decision-making and expand access to eye care by supporting large-scale, efficient screening pathways.
Objective.Large vessel occlusion (LVO) stroke presents a major challenge in clinical practice due to the potential for poor outcomes with delayed treatment. Treatment for LVO involves highly specialized care, in particular endovascular thrombectomy, and is available only at certain hospitals. Therefore, prehospital identification of LVO by emergency ambulance services, can be critical for triaging LVO stroke patients directly to a hospital with access to endovascular therapy. Clinical scores exist to help distinguish LVO from less severe strokes, but they are based on a series of examinations that can be time-consuming and may be impractical for patients with dementia or those who cannot follow commands due to their stroke. There is a need for a fast and reliable method to aid in the early identification of LVO. In this study, our objective was to assess the feasibility of using 30 s photoplethysmography (PPG) recording to assist in recognizing LVO stroke.Approach.A total of 88 patients, including 25 with LVO, 27 with stroke mimic (SM), and 36 non-LVO stroke patients (NL), were recorded at the Liverpool Hospital emergency department in Sydney, Australia. Demographics (age, sex), as well as morphological features and beating rate variability measures, were extracted from the PPG. A binary classification approach was employed to differentiate between LVO stroke and NL + SM (NL.SM). A 2:1 train-test split was stratified and repeated randomly across 100 iterations.Main results.The best model achieved a median test set area under the receiver operating characteristic curve of 0.77 (0.71-0.82).Significance.Our study demonstrates the potential of utilizing a 30 s PPG recording for identifying LVO stroke.
Optical Coherence Tomography (OCT) is essential in ophthalmology for cross-sectional imaging of the retina. Pretrained foundation models facilitate task-specific model development by enabling fine-tuning with limited labeled data. However, current foundation models rely on a single B-scan (usually the central slice), overlooking volumetric context. This research investigates video foundation models to capture full 3D retinal structure and improve diagnostic performance. V-JEPA, a state-of-the-art video foundation model, was benchmarked against retinal foundation models (RETFound, VisionFM) and a natural image foundation model (DINOv2). All were fine-tuned to detect Age-related Macular Degeneration or Glaucomatous Optic Neuropathy using five OCT datasets. V-JEPA consistently equaled or outperformed image-based models, achieving an average AUROC of 0.94 (0.80-0.99), versus 0.90 (0.76-0.98) for the best image model, a statistically significant improvement (p < 0.001). To our knowledge, this is the first application of transformer-based video models to volumetric OCT, highlighting their promise in 3D medical imaging.
Background: Sleep staging is critical for diagnosing sleep disorders. Traditional methods in clinical settings involve time-intensive scoring procedures. Recent advancements in data-driven algorithms using photoplethysmogram (PPG) time series have shown promise in automating sleep staging in adults. However, for children, algorithm development is hindered by the limited availability of datasets, with the Childhood Adenotonsillectomy Trial (CHAT) being the only substantial source, comprising recordings from children aged 5-10. This limitation constrains the evaluation of algorithmic generalization performance. Methods: We employed a deep learning model for sleep staging from PPG, initially trained using a large dataset of adult sleep recordings, and fine-tuned it on 80% of the CHAT dataset (CHAT-train) for the task of three-class sleep staging (wake, REM, non-REM). The resulting algorithm performance was compared to the same model architecture but trained from scratch on CHAT-train (benchmark). The algorithms are evaluated on the local test set, denoted CHAT-test, as well as on a newly introduced independent dataset. Results: Our deep learning algorithm achieved a Cohen's Kappa of 0.88 on CHAT-test (versus 0.65), and demonstrated generalization capabilities with a Kappa of 0.72 on the external Ichilov dataset for children above 5 years old (versus 0.64) and 0.64 for those below 5 (versus 0.53). Significance: This research establishes a new state-of-the-art performance for the task of sleep staging in children using raw PPG. The findings underscore the value of transfer learning from the adults to children domain. However, the reduced performance in children under 5 suggests the need for further research and additional datasets covering a broader pediatric age range to fully address generalization limitations.
The Hillel Yaffe Age Related Macular Degeneration (HYAMD) dataset is a longitudinal collection of 1,560 Digital Fundus Images (DFIs) from 325 patients examined at the Hillel Yaffe Medical Center (Hadera, Israel) between 2021 and 2024. The dataset includes an AMD cohort of 147 patients (aged 54-94) with varying stages of AMD and a control group of 190 diabetic retinopathy (DR) patients (aged 24-92). AMD diagnoses were based on comprehensive clinical ophthalmic evaluations, supported by Optical Coherence Tomography (OCT) and OCT angiography. Non-AMD DFIs were sourced from DR patients without concurrent AMD, diagnosed using macular OCT, fluorescein angiography, and widefield imaging. HYAMD provides gold-standard annotations, ensuring AMD labels were assigned following a full clinical assessment. Images were captured with a DRI OCT Triton (Topcon) camera, offering a 45 deg field of view and 1960 x 1934 pixel resolution. To the best of our knowledge, HYAMD is the first open-access retinal dataset from an Israeli sample, designed to support AMD identification using machine learning models.
Objective. sleep staging is essential for diagnosing sleep disorders and managing sleep health. Traditional methods require time-consuming manual scoring. Recent photoplethysmography (PPG)-based deep learning models perform well on local datasets but struggle with external generalization due to data drift.Approach. this study evaluates multi-source domain training for improving out-of-distribution generalization in four-class sleep staging (wake, light, deep, rapid eye movement) from raw PPG time-series. The trained deep learning model is denoted SleepPPG-Net2. Additionally, we examined the impact of demographic factors, ethnicity, and obstructive sleep apnea (OSA) on performance. SleepPPG-Net2 was benchmarked against two state-of-the-art models.Main results. SleepPPG-Net2 outperformed benchmark models, improving generalization performance (Cohen's kappa) by up to 21%. Performance disparities were observed in relation to age, sex, and OSA severity.Significance. SleepPPG-Net2 enhances PPG-based sleep staging and provides insights into demographic and clinical influences on model performance.
Objective. Diabetic retinopathy (DR) is a serious diabetes complication that can lead to vision loss, making timely identification crucial. Existing data-driven algorithms for DR staging from digital fundus images (DFIs) often struggle with generalization due to distribution shifts between training and target domains.Approach. To address this, DRStageNet, a deep learning model, was developed using six public and independent datasets with 91 984 DFIs from diverse demographics. Five pretrained self-supervised vision transformers (ViTs) were benchmarked, with the best further trained using a multi-source domain (MSD) fine-tuning strategy.Main results. DINOv2 showed a 27.4% improvement in L-Kappa versus other pretrained ViT. MSD fine-tuning improved performance in four of five target domains. The error analysis revealing 60% of errors due to incorrect labels, 77.5% of which were correctly classified by DRStageNet.Significance. We developed DRStageNet, a DL model for DR, designed to accurately stage the condition while addressing the challenge of generalizing performance across target domains. The model and explainability heatmaps are available atwww.aimlab-technion.com/lirot-ai.
Aims/Purpose: Deriving vascular features of retinal images is a proposed noninvasive method to assess vascular health. Although several studies have linked cardiovascular risk to retinal image features, they often rely on limited datasets or have used deep learning approaches with limited explainability. This research introduces an end‐to‐end method for analyzing the retinal vasculature in large retinal image datasets, applied for primary open angle glaucoma (POAG).Methods: 115,237 retinal images of the UZ Leuven Glaucoma Clinic were extracted. 4858 unique images remain of POAG patients (n = 3010) and controls (n = 1848) after image quality assessment, optic disk detection, region of interest definition, automated segmentation of arterioles and venules (LUNet algorithm), and parametrization of the vascular biomarkers (PVBM toolbox). Analysis of covariance and linear mixed models were used to adjust for age, sex and disc size.Results: Both arteriolar and venular diameter, area, length, tortuosity, branching angle, endpoints, intersection points and both mono‐ as multifractal dimension levels are independently lower in POAG patients and older patients (all p < 0.001), after adjustment for sex, disc size and multiple testing. The multivariate linear mixed models additionally show that the vascular features are significantly influenced by sex and disc size, which necessitates correct adjustment.Conclusions: This is the first time a fully automated end‐to‐end pipeline is published for the analysis of retinal vascular geometry in a large cohort (n = 4858). Given the independent similarities in retinal vascular geometry changes between older age and POAG prevalence the theory of pronounced vascular ageing in POAG patients is proposed. This novel approach enables quantitative analysis of eye vasculature and supports the study of its correlation with specific diseases, facilitating reproducible analysis of large datasets and providing explainability to deep learning approaches.