Background:A comprehensive transthoracic echocardiogram involves the assessment of over 70 parameters, placing a substantial burden on sonographers and physicians for manual annotation with considerable inter-observer variability. Prior open-source segmentation models have largely addressed 2D B-mode ventricular function, leaving a gap in the spectral Doppler and atrial measurements required for valvular and diastolic assessment such as velocity-time integral (VTI) and atrial chamber size. Methods:In this retrospective multi-cohort study, we developed EchoNet-Segmentation, comprehensive task-specific deep learning segmentation models for left and right atrial area and VTI Doppler measurements. Training used 186,712 sonographer-annotated images from 93,978 studies (56,855 patients) at Cedars-Sinai Medical Center (CSMC). Performance was evaluated on a held-out CSMC test set, a CSMC temporal split, an external Kaiser Permanente Northern California cohort, and the public MIMIC-Echo dataset. Findings:On the CSMC held-out test set, our AI models showed strong agreement with sonographer measurements, with R² of 0.817-0.882 and mean absolute error (MAE) of 1.13-3.80 cm for automated VTI measurements, and R² of 0.675-0.747 and MAE of 2.48-2.52 cm² for left and right atrial area segmentation. Performance was consistently confirmed on the CSMC temporal split (VTI: R² 0.606-0.866, atrial area: R² 0.694-0.705) and on the KPNC external cohort (VTI: R² 0.575-0.859, atrial area: R² 0.803-0.876), on the MIMIC-Echo dataset. Robustness was demonstrated on a different vendor's machines and across subgroups. EchoNet-Segmentation outperformed an open-source medical image foundation model with bounding-box, point prompt configurations on R², MAE, and Dice score on both held-out test dataset and MIMIC apical four-chamber data. Interpretation:EchoNet-Segmentation is the first open-source framework that delivers accurate, generalizable automated measurement across several key routine echocardiographic parameters, supporting end-to-end automation of clinically important echocardiographic assessments. Public release of model weights, code, and demonstration tools can facilitate reproducibility, research use and clinical deployment. Funding:Funding Statement: This work was supported by NIH NHLBI grants R00HL157421, R01HL173526, and R01HL173487 to D.O. Research in context:Evidence before this study: We searched PubMed and arXiv from database on April 1, 2026, for studies of deep learning-based segmentation of echocardiographic images, using the terms ("echocardiography" OR "echocardiogram") AND ("deep learning" OR "artificial intelligence") AND ("segmentation" OR "measurement"). Prior work has demonstrated automated segmentation of cardiac chambers and left ventricular ejection fraction estimation, and a small number of studies have reported deep learning models for velocity-time integral (VTI) or atrial size measurement. However, openly available models and code remain largely restricted to left ventricular structures, ejection fraction, and wall thickness, and commercial tools remain proprietary. To our knowledge, no open-source framework has comprehensively addressed VTI measurements across multiple Doppler views together with atrial chamber size in a single, reproducible toolkit, and existing models have not been systematically benchmarked against general-purpose medical-image foundation models on echocardiographic tasks.Added value of this study: We developed and validated EchoNet-Segmentation, a suite of task-specific deep learning models for several clinically important echocardiographic parameters: left and right atrial area and five VTI measurements (aortic valve, mitral valve, left ventricular outflow tract, right ventricular outflow tract, and pulmonary valve). The models were trained on the largest real-world collection of sonographer-annotated echocardiograms reported to date (186,712 images from 56,855 patients) in an academic center in the United States and showed strong agreement with sonographer measurements on a held-out internal test set, a temporal split cohort, an external cohort from a different health system, and a publicly available cohort recorded on a different vendor's ultrasound machines. EchoNet-Segmentation outperformed the publicly released medical-image foundation model (MedSAM2) on cardiac chamber segmentation across both internal and public dataset benchmarks. All model weights, training and inference code, demonstration tools, and the manual segmentation masks used for the public benchmark are openly released.Implications of all the available evidence: EchoNet-Segmentation enables end-to-end automation of routine transthoracic echocardiographic measurements with previously released open-source models. By openly releasing model weights, training code, and benchmark data, this work provides a reproducible foundation that the broader research and clinical community can build on, fine-tune for specific populations or imaging protocols, and integrate into clinical workflows. Prospective validation and randomized studies will be needed to define the impact of automated measurement on diagnostic accuracy, workflow efficiency, and clinical outcomes.
Sleep foundation models have recently demonstrated strong performance on in-domain polysomnography tasks, including sleep staging, apnea detection, and disease risk prediction. In this work, we investigate whether sleep biosignals can serve as an effective pretraining distribution for learning representations that transfer beyond sleep to adjacent domains. Following sleep foundation models, we perform sleep-only multimodal contrastive pretraining (with a leave-one-out objective) and evaluate transfer to non-sleep EEG and ECG, two well-benchmarked biosignal modalities with heterogeneous datasets and clinically meaningful downstream tasks. Across eight downstream tasks spanning multiple EEG and ECG datasets, sleep pretraining consistently improves performance relative to training from scratch. Moreover, on several tasks, we achieve performance competitive with or surpassing prior specialized state-of-the-art and foundation models.
BACKGROUND:Transthoracic echocardiography (TTE) is the most commonly performed cardiac imaging modality with over 30 million studies annually. Demand for timely expert interpretation continues to outpace capacity, creating diagnostic delays and inter-observer variability that impact patient care. Recent research has suggested computer vision artificial intelligence (AI) models can generate accurate preliminary comprehensive TTE reports, however, prospective evaluation is needed to determine whether AI-assisted TTE interpretation can improve clinician efficiency while preserving diagnostic accuracy. METHODS:AI ECHO INSIGHT is a prospective randomized blinded clinical trial conducted at Kaiser Permanente Northern California that will evaluate 1200 historical TTE studies (1000 consecutive unselected studies plus 200 with moderate or greater valvular disease) interpreted using three workflows: (1) AI-generated preliminary report finalized by a blinded cardiologist (AI-assisted); (2) cardiologist-generated preliminary report finalized by a blinded cardiologist (cardiologist-assisted); and (3) sonographer-generated preliminary report finalized by a blinded cardiologist (sonographer-assisted). The primary outcome is the rate of substantial change between preliminary and final reports, comparing the AI-assisted workflow to the pooled cardiologist-assisted and sonographer-assisted workflows. Secondary outcomes include cardiologist interpretation time for report finalization, superiority testing for diagnostic accuracy, and reporting consistency. CONCLUSION:AI ECHO INSIGHT is a prospective randomized blinded clinical trial evaluating the clinical impact of AI-assisted TTE interpretation on diagnostic accuracy, cardiologist efficiency, and reporting consistency in real-world echocardiography workflows. TRIAL REGISTRATION:ClinicalTrials.gov registration number NCT07229300.
Sleep is a fundamental biological process with broad implications for physical and mental health, yet its complex relationship with disease remains poorly understood. Polysomnography (PSG)-the gold standard for sleep analysis-captures rich physiological signals but is underutilized due to challenges in standardization, generalizability and multimodal integration. To address these challenges, we developed SleepFM, a multimodal sleep foundation model trained with a new contrastive learning approach that accommodates multiple PSG configurations. Trained on a curated dataset of over 585,000 hours of PSG recordings from approximately 65,000 participants across several cohorts, SleepFM produces latent sleep representations that capture the physiological and temporal structure of sleep and enable accurate prediction of future disease risk. From one night of sleep, SleepFM accurately predicts 130 conditions with a C-Index of at least 0.75 (Bonferroni-corrected P < 0.01), including all-cause mortality (C-Index, 0.84), dementia (0.85), myocardial infarction (0.81), heart failure (0.80), chronic kidney disease (0.79), stroke (0.78) and atrial fibrillation (0.78). Moreover, the model demonstrates strong transfer learning performance on a dataset from the Sleep Heart Health Study-a dataset that was excluded from pretraining-and performs competitively with specialized sleep-staging models such as U-Sleep and YASA on common sleep analysis tasks, achieving mean F1 scores of 0.70-0.78 for sleep staging and accuracies of 0.69 and 0.87 for classifying sleep apnea severity and presence. This work shows that foundation models can learn the language of sleep from multimodal sleep recordings, enabling scalable, label-efficient analysis and disease prediction.
Background and Aims:Accurate classification of mitral stenosis (MS) remains a significant clinical challenge. This study aimed to develop an artificial intelligence (AI) framework to automatically detect clinically significant MS from echocardiography. Methods:We developed EchoNet-MS, an open-source end-to-end integrated approach combining video based convolutional neural networks to assess MS severity and differentiate rheumatic etiology from echocardiography and validated its performance across four cohorts. Results:EchoNet-MS was trained and validated in total of 431,612 videos from 44,671 studies from three different healthcare system. Combining assessments from multiple echocardiographic videos, the model was trained on a Kaiser Permanente Northern California (KPNC) cohort of 8,677 studies from 7,576 patients with a range of MS severity. The model was validated on a KPNC held-out test cohort (N=1,623) and a temporally distinct cohort (N=19,206), as well as Stanford Healthcare (SHC) cohort (N=3,333) and Cedars-Sinai Medical Center (CSMC) cohort (N=72,909). EchoNet-MS achieved excellent discrimination of severe MS with AUC 0.937 [95% CI: 0.913 - 0.958] in the KPNC held-out cohort, 0.994 [0.986 - 0.999] in the temporally distinct cohort, 0.991 [0.986 - 0.995] in SHC, and 0.973 [0.958 - 0.987] in CSMC. The model achieved excellent performance in classifying both rheumatic or non-rheumatic MS with AUC ranging from 0.890 and 0.967. Conclusions:EchoNet-MS accurately assesses MS severity and etiology using information from multiple echocardiographic views. Its strong performance generalizes robustly to external cohorts and shows potential as an automated clinical decision support tool.
Background:Right ventricular (RV) function is an important predictor of morbidity and mortality in various cardiovascular conditions. Nevertheless, its echocardiographic assessment is challenging due to its complex anatomy and location in the chest, resulting in limited inter-observer reproducibility. Objectives:We aimed to develop a novel deep learning model - EchoNet-RV - to segment the RV in apical 4-chamber view (A4C) echocardiographic videos and estimate RV fractional area change (RVFAC). Methods:For training EchoNet-RV, 7,169 expert-annotated A4C echocardiographic videos were used. The model's performance was evaluated on a held-out internal test set of 1,320 A4C videos and two international external test sets of 3,107 and 1,077 A4C videos from two separate centers. Additionally, the associations between the predicted RVFAC values and the composite endpoint of heart failure hospitalization or all-cause death were also analyzed in the first external test set. Results:EchoNet-RV segmented the RV with Dice coefficients of 0.893 (0.891-0.895), 0.797 (0.796-0.798), and 0.788 (0.785-0.790) and predicted RVFAC with mean absolute errors of 5.795 (5.560-6.031), 5.830 (5.692-5.970), and 6.362 (6.064-6.660) percentage points in the held-out test set and the two external test sets, respectively. In 500 randomly selected videos from the external test sets, EchoNet-RV's prediction error was significantly lower than the inter-observer variability (p<0.001). Moreover, it identified RVFAC <35% with areas under the receiver operating characteristic curve of 0.859 (0.843-0.876), 0.725 (0.710-0.740), and 0.684 (0.653-0.713) in the three test sets. EchoNet-RV also outperformed two multi-task models, EchoPrime and PanEcho, in estimating RVFAC and identifying RV dysfunction in the external test sets. In the first external test set, predicted RVFAC values were inversely associated with the composite endpoint (adjusted HR: 0.948 [0.917-0.979], p<0.001), independent of age, sex, cardiovascular risk factors, and left ventricular systolic function. Conclusions:EchoNet-RV enables the rapid and automated assessment of RVFAC, with strong potential to become a valuable tool for the echocardiographic evaluation of RV function and disease surveillance.
Importance:Mobile phone-recorded echocardiogram videos are commonly used in point-of-care, telemedicine, and resource-limited workflows, but artificial intelligence models for left ventricular ejection fraction (LVEF) estimation have primarily been evaluated on native Digital Imaging and Communications in Medicine (DICOM) videos. Objective:To evaluate whether previously described artificial intelligence models for LVEF estimation retain performance when applied to mobile phone-recorded echocardiographic videos. Design:Multicenter model validation study comparing model-estimated LVEF with clinician-reported LVEF. Setting:Three medical centers: Kaiser Permanente Northern California, Beth Israel Deaconess Medical Center through MIMIC-IV-ECHO, and Cedars-Sinai Medical Center. Participants:Source studies with clinician-reported LVEF and apical 4-chamber or apical 2-chamber views, yielding 6209 phone-recorded videos from 2648 studies and 2611 patients. Exposures:Mobile phone recording of native echocardiographic videos and fine-tuning of pretrained models using mobile phone-recorded videos from the Kaiser Permanente Northern California training cohort. Main Outcomes and Measures:Mean absolute error in ejection fraction percentage points, R2 for continuous estimation, and area under the receiver operating characteristic curve for identifying ejection fraction greater than 50%. Results:The study included 6209 mobile phone-recorded echocardiographic videos from 2648 studies and 2611 patients; the weighted mean age was 68.4 years, and 1031 patients were male (39.5%). Without phone-video fine-tuning, the primary model achieved a mean absolute error of 7.00 percentage points, coefficient of determination of 0.49, and area under the receiver operating characteristic curve of 0.91 on phone-recorded videos; corresponding native DICOM performance was 6.08 percentage points, 0.60, and 0.93, respectively. On the 2396-video fine-tuning evaluation cohort, fine-tuning improved primary model performance to a mean absolute error of 6.96 percentage points, coefficient of determination of 0.61, and area under the receiver operating characteristic curve of 0.93. Fine-tuning the public EchoNet-Dynamic model improved performance from 9.36 percentage points, 0.37, and 0.84 to 7.86 percentage points, 0.50, and 0.89, respectively. Progressive central zoom preprocessing degraded model performance. Conclusions and Relevance:These findings suggest that artificial intelligence-assisted left ventricular ejection fraction estimation from mobile phone-recorded echocardiograms may be feasible when native image export is unavailable, although prospective evaluation is needed before clinical deployment.
Polysomnography (PSG), the gold standard test for sleep analysis, generates vast amounts of multimodal clinical data, presenting an opportunity to leverage self-supervised representation learning (SSRL) for pre-training foundation models to enhance sleep analysis. However, progress in sleep foundation models is hindered by two key limitations: (1) the lack of a shared dataset and benchmark with diverse tasks for training and evaluation, and (2) the absence of a systematic evaluation of SSRL approaches across sleep-related tasks. To address these gaps, we introduce Stanford Sleep Bench, a large-scale PSG dataset comprising 17,467 recordings totaling over 163,000 hours from a major sleep clinic, including 13 clinical disease prediction tasks alongside canonical sleep-related tasks such as sleep staging, apnea diagnosis, and age estimation. We systematically evaluate SSRL pre-training methods on Stanford Sleep Bench, assessing downstream performance across four tasks: sleep staging, apnea diagnosis, age estimation, and disease and mortality prediction. Our results show that multiple pretraining methods achieve comparable performance for sleep staging, apnea diagnosis, and age estimation. However, for mortality and disease prediction, contrastive learning significantly outperforms other approaches while also converging faster during pretraining. To facilitate reproducibility and advance sleep research, we will release Stanford Sleep Bench along with pretrained model weights, training pipelines, and evaluation code.
Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these videos are designed for human learners, often mixing medical imagery with narration, explanatory diagrams, and contextual framing. In this work, we investigate whether such pedagogically rich, yet non-standardized and heterogeneous videos can effectively teach general-domain vision-language models biomedical knowledge. To this end, we introduce OpenBiomedVi, a biomedical video instruction tuning dataset comprising 1031 hours of video-caption and Q/A pairs, curated through a multi-step human-in-the-loop pipeline. Diverse biomedical video datasets are rare, and OpenBiomedVid fills an important gap by providing instruction-style supervision grounded in real-world educational content. Surprisingly, despite the informal and heterogeneous nature of these videos, the fine-tuned Qwen-2-VL models exhibit substantial performance improvements across most benchmarks. The 2B model achieves gains of 98.7 text tasks. The 7B model shows improvements of 37.09 image tasks, with a slight degradation of 2.7 respective base models. To address the lack of standardized biomedical video evaluation datasets, we also introduce two new expert curated benchmarks, MIMICEchoQA and SurgeryVideoQA. On these benchmarks, the 2B model achieves gains of 99.1 respectively, demonstrating the models' ability to generalize and perform biomedical video understanding on cleaner and more standardized datasets than those seen during training. These results suggest that educational videos created for human learning offer a surprisingly effective training signal for biomedical VLMs.
BACKGROUND:Accurate measurement of echocardiographic parameters is crucial for the diagnosis of cardiovascular disease and tracking of change over time; however, manual assessment requires time-consuming effort and can be imprecise. Artificial intelligence has the potential to reduce clinician burden by automating the time-intensive task of comprehensive measurement of echocardiographic parameters. OBJECTIVES:The purpose of this study was to develop and validate open-sourced deep learning semantic segmentation models for the automated measurement of 18 anatomic and Doppler measurements in echocardiography. METHODS:We trained models for the automated measurement of echocardiography parameters using data sets between 2011 and 2023 from Cedars-Sinai Medical Center (CSMC). The outputs of segmentation models were compared with sonographer measurements from temporal split data from CSMC and an external data set from Stanford Healthcare (SHC) to access accuracy and precision. RESULTS:We used 877,983 echocardiographic measurements from 155,215 studies from CSMC to develop EchoNet-Measurements, an open-source deep learning model for echocardiographic annotation. The models demonstrated high accuracy when compared with sonographer measurements from held-out data from CSMC and an independent external validation data set from SHC. Measurements across all 9 B-mode and 9 Doppler measurements had high accuracy (mean coverage probability of 0.796 and 0.839 and mean relative difference of 0.120 and 0.096 on held-out test set from CSMC and external data set from SHC, respectively). When evaluated end-to-end on 2,103 temporally distinct studies at CSMC, EchoNet-Measurements had similar reasonable performance (mean coverage probability 0.803 and mean relative difference of 0.108). Performance was consistent across patient characteristics including age, sex, and atrial fibrillation, obesity status, and machine vendors. CONCLUSIONS:EchoNet-Measurements achieves high accuracy in automated echocardiographic quantification and potential for assisting the clinicians in the echocardiography workflow. This open-source model provides the foundation for future developments in artificial intelligence applied to echocardiography.
Echocardiography (echo), or cardiac ultrasound, is the most widely used imaging modality for cardiac form and function due to its relatively low cost, rapid acquisition time, and non-invasive nature. However, ultrasound acquisitions are often limited by artifacts and noise that hinder diagnostic interpretation in clinical settings. Existing methodologies for denoising echos consist solely of traditional filtering-based algorithms or deep learning methods developed on radio-frequency (RF) signals which prevents clinical applicability and scalability. To address these limitations, we introduce the first deep generative model capable of simulating ultrasound noise developed on B-mode data. Using this generative model, we develop a synthetic dataset of paired clean and noisy echo images to train a downstream model for real-world image denoising and demonstrate state-of-the-art performance in both internal and external experiments. In both held-out test sets, our method results in echo images with higher gCNR in comparison to noisy image counterparts and images derived from a comparable method which is consistent with provided visual comparisons. Our experiments showcase the potential of our method for future clinical use to improve the quality of echo acquisitions. To encourage further research into the field, we release our source code and model weights at https://github.com/echonet/image_quality.
Although the heart has complex three-dimensional (3D) anatomy, conventional medical imaging with cardiac ultrasound relies on a series of 2D videos showing individual cardiac structures. 3D echocardiography is a developing modality that now offers adequate image quality for clinical use, with potential to streamline acquisition and improve assessment of off-axis features. We propose an automated method to select standard 2D views from 3D cardiac ultrasound volumes, allowing physicians to interpret the data in their usual format while benefiting from the speed and usability of 3D scanning. Applying a deep learning view classifier and downstream heuristics based on anatomical landmarks together with heuristics provided by cardiologists, we reconstruct standard echocardiography views. This approach was validated by three cardiologists in blinded evaluation (96% accuracy in 1,600 videos from 2 hospitals). The downstream 2D videos were also validated in their ability to detect cardiac abnormalities using AI echocardiography models (EchoPrime and PanEcho) as well as ability to generate clinical-grade measurements of cardiac anatomy (EchoNet-Measurement). We demonstrated that the extracted 2D videos preserve spatial calibration and diagnostic features, allowing clinicians to obtain accurate real-world interpretations from 3D volumes. We release the code and a dataset of 29 3D echocardiography videos https://github.com/echonet/3d-echo .
The de-identification (deID) of protected health information (PHI) and personally identifiable information (PII) is a fundamental requirement for sharing medical images, particularly through public repositories, to ensure compliance with patient privacy laws. In addition, preservation of non-PHI metadata to inform and enable downstream development of imaging artificial intelligence (AI) is an important consideration in biomedical research. The goal of MIDI-B was to provide a standardized platform for benchmarking of DICOM image deID tools based on a set of rules conformant to the HIPAA Safe Harbor regulation, the DICOM Attribute Confidentiality Profiles, and best practices in preservation of research-critical metadata, as defined by The Cancer Imaging Archive (TCIA). The challenge employed a large, diverse, multi-center, and multi-modality set of real de-identified radiology images with synthetic PHI/PII inserted. The MIDI-B Challenge consisted of three phases: training, validation, and test. Eighty individuals registered for the challenge. In the training phase, we encouraged participants to tune their algorithms using their in-house or public data. The validation and test phases utilized the DICOM images containing synthetic identifiers (of 216 and 322 subjects, respectively). Ten teams successfully completed the test phase of the challenge. To measure success of a rule-based approach to image deID, scores were computed as the percentage of correct actions from the total number of required actions. The scores ranged from 97.91
Sleep is a fundamental biological process with profound implications for physical and mental health, yet our understanding of its complex patterns and their relationships to a broad spectrum of diseases remains limited. While polysomnography (PSG), the gold standard for sleep analysis, captures rich multimodal physiological data, analyzing these measurements has been challenging due to limited flexibility across recording environments, poor generalizability across cohorts, and difficulty in leveraging information from multiple signals simultaneously. To address this gap, we curated over 585,000 hours of high-quality sleep recordings from approximately 65,000 participants across multiple cohorts and developed SleepFM, a multimodal sleep foundation model trained with a novel contrastive learning approach, designed to accommodate any PSG montage. SleepFM produces informative sleep embeddings that enable predictions of future diseases. We systematically demonstrate that SleepFM embeddings can predict 130 future diseases, as modeled by Phecodes, with C-Index and AUROC of at least 0.75 on held-out participants (Bonferroni-corrected p < 0.01). This includes accurate predictions for death (C-Index: 0.84 [95% CI: 0.81-0.87]), heart failure (C-Index: 0.80 [95% CI: 0.77-0.83]), chronic kidney disease (C-Index: 0.79 [95% CI: 0.77-0.81]), dementia (C-Index: 0.85 [95% CI: 0.82-0.87]), stroke (C-Index: 0.78 [95% CI: 0.76-0.81]), atrial fibrillation (C-Index: 0.78 [95% CI: 0.75-0.81]), and myocardial infarction (C-Index: 0.81 [95% CI: 0.78-0.84]). The model's generalizability was further validated through strong performance on the Sleep Heart Health Study (SHHS), a dataset unseen during pre-training. Additionally, SleepFM demonstrates strong performance on traditional sleep analysis tasks, achieving competitive results in both sleep staging (mean F1 scores: 0.70-0.78) and sleep apnea diagnosis (AUROC: 0.90-0.94). Beyond these standard applications, our analysis reveals that specific sleep stages and physiological signals carry distinct predictive power for different diseases. This work demonstrates how foundation models can leverage sleep polysomnography data to uncover the extensive relationship between sleep physiology and future disease risk.
Background: Accurate assessment of mitral stenosis (MS) severity is critical to guide timely clinical management. Current evaluation relies on expert interpretation of B-mode and Doppler echocardiography, requiring integration of multiple views and skilled Doppler imaging. This study aimed to develop and validate a deep learning model for automated MS severity assessment using multi-view B-mode and color Doppler echocardiographic videos. Methods: We developed a two-stage framework for automated MS assessment. First, four video-based convolutional neural networks were trained to classify MS severity from distinct echocardiographic views: B-mode [parasternal long-axis (PLAX), apical three-chamber and five-chamber (AP)] and color Doppler [PLAX-color, and AP-color]. Next, outputs from the four models were integrated using a machine learning ensemble (HistGradientBoostingClassifier) to produce a study-level MS severity classification. Performance was assessed using area under the receiver operating characteristic curve (AUC) on held-out test set from Kaiser Permanente (KP) and Stanford Health Care (SHC). Results: The models were trained on 66,714 videos from 3,921 studies at KP, and evaluated on internal held-out test data (KP; 7,344 videos, 438 studies) and external test data (SHC; 29,181 videos, 1,988 studies). The multi-view ensemble model demonstrated strong performance, achieving a macro-AUC of 0.853 (95% CI: 0.828–0.877; Figure 1) on the KP test dataset. Generalizability was confirmed on the external SHC cohort with an AUC of 0.974 (95% CI: 0.968–0.981; Figure 2). Conclusion: This study confirmed the ability for multi-view deep learning models to assess MS severity. The model demonstrated accurate, generalizable performance and highlights the potential of AI-powered decision support tools in echocardiographic evaluation of MS.
Echocardiography is the most widely used cardiac imaging modality, capturing ultrasound video data to assess cardiac structure and function1. Artificial intelligence (AI) in echocardiography has the potential to streamline manual tasks and improve reproducibility and precision2. However, most echocardiography AI models are single-view, single-task systems that do not synthesize complementary information from multiple views captured during a full examination3,4, and thus lead to limited performance and scope of applications. To address this problem, we introduce EchoPrime, a multi-view, view-informed, video-based vision-language foundation model trained on over 12 million video-report pairs. EchoPrime uses contrastive learning to train a unified embedding model for all standard views in a comprehensive echocardiogram study with representation of both rare and common diseases and diagnoses. EchoPrime then utilizes view classification and a view-informed anatomical attention module to weight video-specific embeddings that accurately map the relationship between echocardiographic views and anatomical structures. With retrieval-augmented interpretation, EchoPrime integrates information from all echocardiogram videos in a comprehensive study and performs holistic clinical interpretation. In datasets from five international independent health-care systems, EchoPrime achieves state-of-the-art performance on 23 diverse benchmarks of cardiac form and function, surpassing the performance of both task-specific approaches and previous foundation models. Following rigorous clinical evaluation, EchoPrime can assist physicians in the automated preliminary assessment of comprehensive echocardiography.
Background and Aims:Accurate assessment of aortic stenosis (AS) requires integration of both structural and functional information characterized by visual traits as well as quantitation of gradients. Existing artificial intelligence (AI) models utilize solely either structural or functional information. Methods:We developed EchoNet-AS, an open-source end-to-end integrated approach combining video based convolutional neural networks to assess valve motion as well as segmentation models to automate the measurement of aortic valve peak velocity to classify AS severity. Results:EchoNet-AS was trained on 210,193 images from 16,076 studies from Kaiser Permanente Northern California (KPNC) and validated on 1,589 held-out test studies and a temporally distinct cohort of 19,206 studies. The final model was also externally validated on 2,415 studies from Stanford Healthcare (SHC) and 9,038 studies from Cedars-Sinai Medical Center (CSMC). Combining assessments from multiple echocardiographic videos and Doppler measurements, EchoNet-AS achieved excellent discrimination of severe AS with AUC 0.964 [95% CI: 0.952 - 0.973] in the KPNC held-out cohort and 0.985 [0.981 - 0.988] in the temporally distinct cohort, which was superior to models using single views or only Doppler measurements. The performance was consistently robust in distinct external cohorts with an AUC 0.985 [0.975 - 0.992] at SHC and 0.989 [0.986 - 0.992] at CSMC. Conclusions:EchoNet-AS synthesizes information from both B-mode videos and Doppler images to accurately assess AS severity. Its strong performance generalizes robustly to external validation cohorts and shows potential as an automated clinical decision support tool.
Sleep is a complex physiological process evaluated through various modalities recording electrical brain, cardiac, and respiratory activities. We curate a large polysomnography dataset from over 14,000 participants comprising over 100,000 hours of multi-modal sleep recordings. Leveraging this extensive dataset, we developed SleepFM, the first multi-modal foundation model for sleep analysis. We show that a novel leave-one-out approach for contrastive learning significantly improves downstream task performance compared to representations from standard pairwise contrastive learning. A logistic regression model trained on SleepFM's learned embeddings outperforms an end-to-end trained convolutional neural network (CNN) on sleep stage classification (macro AUROC 0.88 vs 0.72 and macro AUPRC 0.72 vs 0.48) and sleep disordered breathing detection (AUROC 0.85 vs 0.69 and AUPRC 0.77 vs 0.61). Notably, the learned embeddings achieve 48% top-1 average accuracy in retrieving modality clip pairs from 90,000 candidates. This work demonstrates the value of holistic multi-modal sleep modeling to fully capture the richness of sleep recordings. SleepFM is open source and available at https://anonymous.4open.science/r/sleepfm.
Echocardiography is the most widely used cardiac imaging modality, capturing ultrasound video data to assess cardiac structure and function. Artificial intelligence (AI) in echocardiography has the potential to streamline manual tasks and improve reproducibility and precision. However, most echocardiography AI models are single-view, single-task systems that do not synthesize complementary information from multiple views captured during a full exam, and thus lead to limited performance and scope of applications. To address this problem, we introduce EchoPrime, a multi-view, view-informed, video-based vision-language foundation model trained on over 12 million video-report pairs. EchoPrime uses contrastive learning to train a unified embedding model for all standard views in a comprehensive echocardiogram study with representation of both rare and common diseases and diagnoses. EchoPrime then utilizes view-classification and a view-informed anatomic attention model to weight video-specific interpretations that accurately maps the relationship between echocardiographic views and anatomical structures. With retrieval-augmented interpretation, EchoPrime integrates information from all echocardiogram videos in a comprehensive study and performs holistic comprehensive clinical echocardiography interpretation. In datasets from two independent healthcare systems, EchoPrime achieves state-of-the art performance on 23 diverse benchmarks of cardiac form and function, surpassing the performance of both task-specific approaches and prior foundation models. Following rigorous clinical evaluation, EchoPrime can assist physicians in the automated preliminary assessment of comprehensive echocardiography.