IntroductionVentriculoarterial coupling (VAC), traditionally assessed using the invasive “gold standard” Ea/Ees ratio (arterial elastance to end-systolic elastance), is clinically important but requires approximations that are impractical at the bedside. Oscillatory power fraction (OPF), defined as the ratio of oscillatory to total hydraulic power and derived from synchronized aortic flow–pressure measurements, has been proposed as a pragmatic VAC surrogate. This study aimed to clarify the possible correlation between VAC and OPF in pharmacologically induced hemodynamic extremes.MethodsIn 14 juvenile pigs challenged with vasoactive (phenylephrine, nitroprusside) and cardioactive drugs (dobutamine, esmolol) titrated toward protocolized targets (± 40% MAP, ± 50% dP/dt), we obtained a comprehensive dataset using intracardiac conductance catheters, intraaortic pressure probes, and ascending aortic flow probes. Linear mixed models with animal as a random effect compared baseline and drug-effect states. The coupling of the two responses was assessed from each animal’s baseline-to-effect change, corrected for measurement error estimated from the replicate recordings.ResultsEa/Ees changed by 10%–36% across the four interventions, whereas OPF changed by 1.5%–12%, and under phenylephrine and dobutamine, the change in OPF was negligible (|d| < 0.2, both p > 0.2) despite measurable changes in Ea/Ees. Esmolol was the only challenge under which both indices changed substantially in the same direction (Ea/Ees: + 35.4%, OPF: + 12.2%). Whether the magnitudes of the two responses were coupled could not be determined. Dobutamine markedly increased Ees, total power, and oscillatory power while preserving OPF despite a reduction in Ea/Ees, indicating improved VAC but stable pulsatile efficiency. Vasoactive drugs (phenylephrine, nitroprusside) substantially altered Ea/Ees without relevant OPF changes.ConclusionEa/Ees and OPF are not interchangeable. OPF responded far less strongly than Ea/Ees to the loading changes imposed here and did not resolve them under two of the four challenges. This is consistent with the two indices reflecting predominantly steady and pulsatile aspects of arterial load, although the present data do not demonstrate that separation directly. While OPF cannot replace Ea/Ees, power-derived indices offer practical adjuncts for characterizing ventriculoarterial interactions and the effects of selected drugs at the intensive care bedside.
PURPOSE OF REVIEW:To review recent advances in artificial intelligence (AI) for left ventricular (LV) strain echocardiography, with emphasis on studies published during the preceding 18 months, and to assess current evidence for measurement performance, workflow integration, emerging AI-based approaches, disease-specific applications and barriers to widespread clinical adoption. RECENT FINDINGS:The most mature evidence relates to automated global longitudinal strain (GLS) measurement, where recent studies show high feasibility, improved reproducibility and reduced operator dependence compared with conventional analysis. AI-based tools have also moved beyond retrospective post-processing towards real-time acquisition support and workflow integration, with evidence of shorter analysis time and more standardized image acquisition. Emerging approaches suggest a shift from single-task GLS automation towards multitask and phenotyping-oriented models in which strain is integrated with broader echocardiographic interpretation. Disease-specific applications are most studied in acute ischaemic heart disease and cardio-oncology, where AI-based strain has shown promise for risk stratification, surveillance and clinical triage. However, agreement with conventional methods remains imperfect and varies substantially across models, decision thresholds are uncertain, and prospective evidence for improved clinical outcomes is still limited. SUMMARY:AI for LV strain is evolving from single-task automated measurement towards broader support for reproducibility, workflow integration and clinical interpretation. Current priorities are robust external validation, transparent reporting of feasibility and unsuccessful analyses, and prospective evaluation of clinical utility and workflow impact in routine practice, with clinician oversight as a critical component of safe implementation.
Abstract In this proof‐of‐concept study, we evaluated the feasibility of non‐invasive estimation of cardiac power metrics—total power, steady power, oscillatory power, and oscillatory power fraction—and compared the prognostic and diagnostic value with established echocardiographic metrics. We prospectively included 29 patients (mean age 76 ± 13 years, 24% women) hospitalized with decompensated heart failure. Left ventricular outflow tract flow waveforms were derived from Doppler echocardiography and synchronized with a continuous arterial pressure waveform from a finger‐volume‐clamp device (INL382, Finapres Medical Systems B.V., Amsterdam, Netherlands). Total power was computed as the integral of the instantaneous pressure‐flow product per second, steady power as mean arterial pressure multiplied by mean flow, and oscillatory power as their difference. All measures were calculated from the same consecutive heartbeats covering three respiratory cycles. The feasibility of obtaining cardiac power metrics was 91%. Total, steady, and oscillatory power predicted short‐term all‐cause mortality (log rank p < 0.015), whereas left ventricular ejection fraction, global longitudinal strain, and myocardial work indices did not. These findings suggest that cardiac power metrics may be useful in risk stratification and warrant validation in larger cohorts.
Speckle tracking echocardiography (STE) is the clinical standard for myocardial strain estimation. Despite good performance on global strain (GLS), its accuracy for regional strain remains limited, even though this biomarker is highly relevant for early diagnosis and the characterization of subtle abnormalities. from clinical data. Deep learning is a promising alternative, but its development is constrained by the lack of reliable motion references. Existing solutions rely either on STE-derived labels or on simulations generated by physics-based models, but these synthetic sequences still have limited realism compared with clinical data.In this paper, we propose a novel simulation strategy that incorporates speckle decorrelation measures from real videos and uses an iterative refinement process to improve the motion realism in the simulations. We created an open-source photorealistic dataset of 1,478 videos with reference motion, which was used to train an echocardiographic motion estimation algorithm. The proposed method achieves unmatched performance on global and regional strain, notably reaching a GLS variability of 1.42
Myocardial point tracking (MPT) has recently emerged as a promising direction for motion estimation in echocardiography, driven by advances in general-purpose point tracking methods. However, myocardial motion fundamentally differs from motion encountered in natural videos, as it arises from physiologically constrained deformation that is spatially and temporally continuous throughout the cardiac cycle. Consequently, motion trajectories typically remain locally confined despite substantial tissue deformation. Motivated by these properties, we revisit the architectural design for MPT and find that coarse initialization in commonly used two-stage coarse-to-fine architectures may be unnecessary in this domain. In this work, we propose a fine-stage-only architecture, EchoTracker2, which enriches pixel-precise features with local spatiotemporal context and integrates them with long-range joint temporal reasoning for robust tracking. Experimental results across in-distribution, out-of-distribution (OOD), and public synthetic datasets show that our model improves position accuracy by 6.5% and reduces median trajectory error by 12.2% relative to a domain-specific state-of-the-art (SOTA) model. Compared to the best general-purpose point tracking method, the improvements are 2.0% and 5.3%, respectively. Moreover, EchoTracker2 shows better agreement with expert-derived global longitudinal strain (GLS) and enhances test-rest reproducibility. Source code will be available at: https://github.com/riponazad/ptecho.
Myocardial strain from echocardiography is a key biomarker for cardiac function. Recent deep learning methods show strong performance for myocardial motion tracking but often lack physiological constraints, leading to temporal drift across the cardiac cycle. Consequently, tracked points may not return to their relative initial positions at the end of each cardiac cycle, producing inaccurate strain estimates and even divergence in some cases. We propose a deep learning framework that compensates for drift during myocardial tracking. We extend a state-of-the-art echocardiographic tracking method (TAS-Net) with persistent memory tokens that share information across sliding windows over full cardiac cycles. A teacher-student fine-tuning strategy on real echocardiographic data then enforces physiologically consistent cyclic motion while preserving tracking accuracy. Experiments show reduced global and regional strain drift, improved agreement with clinical references, and better test-retest reproducibility, supporting more reliable myocardial strain estimation in clinical practice.
Aims:Mitral annular plane systolic excursion (MAPSE) is an accessible echocardiographic measure of left ventricular (LV) function. However, manual measurement methods are operator-dependent and time-consuming. We developed a multistep deep learning (DL) method for off-line and real-time fully automated MAPSE estimation, and aimed to assess agreement, reproducibility, time efficiency, and feasibility compared with standard manual measurements. Methods and results:The DL-based method was evaluated in two retrospective cohorts (n = 1775) and one prospective cohort (n = 51). Agreement between DL-MAPSE on B-mode images and experts' manual M-mode measurements was evaluated in all datasets. Evaluation of test-retest reproducibility, time efficiency using real-time analysis during acquisition, and agreement with cardiac magnetic resonance (CMR)-imaging were performed in subsets of the datasets. DL-MAPSE demonstrated good agreement with manual measurements, with bias 2.9 mm (95% CI 2.8-3.0 mm) and Pearson coefficient 0.81 (95% CI 0.79-0.84) in the primary dataset, and a lower bias of 1.0 mm against CMR-MAPSE compared with -2.1 mm using manual M-mode. Both DL and manual measurements showed good test-retest reproducibility (ICC 0.82 and 0.76, respectively). Real-time DL measurements reduced measurement and acquisition time by 51% (mean 1 min 50 s) per examination. The DL method demonstrated excellent feasibility (96%). Conclusion:This novel DL method for fully automated MAPSE demonstrated excellent feasibility, robust reproducibility, and good agreement with both manual M-mode and CMR-derived measurements. Automated DL-MAPSE could substantially reduce analysis time and enhance reproducibility, increasing its clinical value as a marker of LV systolic function.
Aims:Global longitudinal strain (GLS) is a sensitive measure of left ventricular (LV) function, though hampered by significant variability. Data on how echocardiographic recordings and measurements influence GLS and its variability are scarce. We aimed to study the consequences of variability in echocardiographic recordings and measurement procedures for GLS in a normal population. Methods and results:Healthy participants (n = 1412) from the Trøndelag Health Study were examined by echocardiography according to current recommendations. Four experienced operators read and re-read all GLS recordings (2824 measurement procedures). Characteristics of echocardiographic recordings and measurement procedures were extracted by deep learning (DL) and software metadata.LV apical endocardium was positioned at a mean distance of 26 mm from the top of the sector and 6 mm to the left of the central axis. Mean LV foreshortening ranged 1.2-2.1 mm in apical views and the recordings were well standardized with the preferred cut-planes for tilt and rotation. GLS (in percentage points) was lower with reduced image quality (0.8), longer LVs (0.2 per cm), wider regions of interest (0.8 per mm) during the measurement procedure, and longer distance (0.07 per mm) from the transducer to the LV apex, respectively. Moreover, GLS was 0.7 percentage points higher per mm apical foreshortening, and also higher when a more refined ROI initialization was used. Conclusion:Using novel DL methodology and available metadata, we found that indicators of image standardization and measurement procedures influenced GLS. Implementing such quality indicators during recording and measurement procedures may improve the efficacy of echocardiography and ease the interpretation of differences between studies.
Background:Echocardiographic image data accumulating in echo labs are a highly valuable but underutilized resource for cardiac imaging research. Despite the availability of large image databases, quantitative measurements required for clinical analysis and research remain limited. Retrospective manual measurements are highly time-consuming and susceptible to operator-related variability. Moreover, data curation and quality control metrics are needed to prepare real-world data for analysis. Methods:Deep learning-based image analysis can provide fully automated, rapid, and consistent extraction of measurements, given that the data have been properly curated. In this work, we develop an automated pipeline for data curation of a large echo database of 14 326 exams from 9678 patients and evaluate automated measurements of left ventricular ejection fraction (LVEF) and left atrial volume index (LAVI) as a use case. Results:In validation subsample of 1763 subjects with varying image quality and cardiac diseases and 1488 healthy subjects, the pipeline output was compared with manual measurements. Bland-Altman analysis revealed a bias [standard deviation (SD)] of -1.8% (7.6%) for LVEF and 3.3 mL/m² (8.1 mL/m²) for LAVI and demonstrated robust performance for varying image quality and pathological conditions. Additionally, in the large part of the database of 9678 exams without clinical measurements, the automated data curation and measurement quality control resulted in 79% measured data with high confidence. Conclusion:This work highlights the potential of deep learning-based automated measurements in echocardiography for data mining in large real-world databases, paving the way for advancements in cardiac imaging research and diagnostics.
BACKGROUND:Left ventricular (LV) global longitudinal strain (GLS) offers advantages over LV ejection fraction, including improved diagnostic sensitivity, reproducibility, and prognostic value. However, current semiautomatic analyses are time-consuming and operator dependent, impeding widespread adoption of GLS in routine clinical practice. OBJECTIVES:We aimed to assess the feasibility, precision, and time efficiency of GLS measurements using a deep learning (DL) platform that performs real-time GLS analysis during image acquisition and incorporates DL tools to support standardization, to evaluate whether DL-assisted acquisitions can enhance image quality metrics relevant to strain analyses. METHODS:A DL platform was developed for fully automated real-time GLS analysis, including tools that detect and alert the operator to foreshortening or baseline drift. In this controlled prospective study, 50 patients (mean age, 56 years; 64% male) were included. Two image sets were acquired by different operators using the DL platform and a conventional workflow, and GLS and image quality were compared. RESULTS:Overall feasibility of DL-based GLS measurements was 94%. Absolute GLS was 14.8 ± 3.2 using the DL platform workflow and 16.2 ± 3.3 with manual reference measurements, with a bias of -1.3 and limits of agreement ranging from -3.5 to 0.8. Correlation was excellent (r = 0.94). Images acquired with the DL platform showed significantly less baseline drift and borderline improved territorial strain agreement than the reference acquisition. The median time obtaining GLS with the DL platform was reduced by 57% compared to the conventional workflow, from 4 minutes and 48 seconds to 2 minutes and 4 seconds. CONCLUSION:The DL platform for fully automated real-time GLS measurements was feasible, precise, and time efficient. Real-time DL-based feedback allows operators to optimize images during acquisition, thus improving quality metrics relevant to GLS analyses. Implementing this method in clinical practice could streamline workflow and improve efficiency in the echocardiographic laboratory.
Abstract Background Accurate measurements of left ventricular (LV) wall thickness and chamber dimensions in the parasternal long-axis view (PLAX) are essential in the standard echocardiographic examination. Manual assessments introduce variability, and conducting the numerous measurements required for a comprehensive echocardiographic assessment is time-consuming. Therefore, there is potential for improving workflow. Purpose To develop a novel deep learning (DL)-based method for fully automated, real-time measurements of LV wall thickness and chamber dimensions while scanning the patient, and to validate its performance in an unselected prospective cohort. Methods PLAX recordings from 207 subjects were annotated to train DL networks for cardiac segmentation, timing of end-systole and end-diastole, measurements of LV wall thickness and chamber dimensions. The application presents results from three consecutive cardiac cycles in real-time to the operator for review (Figure 1). Agreement and correlation with reference measurements were evaluated, and the combined acquisition and measurement times for both the novel- and reference methods were recorded. Results Feasibility for real-time measurements in PLAX was 98%, with one patient excluded due to poor acoustic window. Agreement with reference measurements was good for interventricular septal diameter (IVSd) (bias - 0.7 mm, limits of agreement (LoA) -4.38 to 2.94 mm), posterior wall thickness (PWd) (bias 0.8 mm, LoA -2.32 to 3.84 mm) and left ventricular internal diastolic diameter (LVIDd) (bias 0.0 mm, LoA -7.0 to 6.9 mm, Figure 2A). Strong correlations were found for IVSd (R=0.84), PWd (R=0.66) and LVIDd (R=0.85), all significant (p <0.001). The average time for acquiring PLAX images was 53 ± 22 seconds, with an additional 42 ± 11 seconds for reference analysis, totaling 94 ± 27 seconds to complete reference PLAX measurements. In contrast, the real-time application took only 59 ± 34 seconds to obtain the same measurements (Figure 2B), resulting in a mean time reduction of 36 seconds (95% CI 24-47 seconds). Conclusion This novel DL-based method for fully automated, real-time analysis of key metrics in PLAX is time-efficient, feasible, and yields measurements in good agreement with manual reference. This method has the potential to substantially improve efficiency and precision of PLAX assessment in clinical practice.
OBJECTIVE:Quantitative measurements remain under-used in echocardiography due to measurement variability and time consumption. This study aimed to develop a deep learning (DL) pipeline for the automatic extraction of clinically relevant quantitative measurements from the parasternal short-axis (PSAX) view. METHODS:We trained an nnU-Net model to segment the left ventricle (LV) lumen and myocardium in PSAX images. Based on end-diastole (ED) and end-systole (ES) frames in the echocardiograms, we calculated the LV lumen area, LV fractional area change, mean wall thickness (MWT) and global circumferential strain. Segmentations and measurements were validated through comparison with two manual observers. RESULTS:Compared with manual references, the nnU-Net model achieved a high grade of similarity for the LV lumen and myocardium, with Dice coefficients of 0.93±0.04 and 0.84±0.08, and 95th percentile Hausdorff distances of 3.2±1.9mm and 3.4±1.7mm, respectively. The Dice coefficients and 95th percentile Hausdorff distances were on par or better for the evaluation dataset. DL-based measurements were in line with interobserver precision and variability. Automatic timing of ED and ES frames based on echocardiograms and LV lumen area produced similar results to using manual timing by experts. The subject-level feasibility was 90.4%. Furthermore, DL-based measurement of MWT at ED differentiated subjects with and without hypertension (p<0.001). CONCLUSION:Our DL-based measurement pipeline for PSAX achieved performance comparable to expert annotators, positioning it as a possible substitute for tedious manual measurements. DL-derived MWT was able to differentiate hypertensive from non-hypertensive subjects, which indicates the potential clinical utility of fully automated PSAX measurements.
Abstract Background Accurate quantification of left ventricular (LV) systolic function is crucial in echocardiography. LV Ejection Fraction (LV EF) and LV Global Longitudinal Strain (LV GLS), rely on clear delineation of the endocardial border, limiting their utility in patients with suboptimal acoustic windows. In contrast, Mitral Annular Plane Systolic Excursion (MAPSE) requires only the visualization of the mitral annulus with its strong acoustic reflections, offering near-perfect feasibility. However, conventional M-mode MAPSE is angle-dependent and prone to errors due to out-of-line movement of the mitral annulus. Purpose To validate and test a novel optical-flow-based deep learning (DL) application for fully automated B-mode MAPSE measurements in real-time during scanning (DL-MAPSE), assessing its agreement with conventional M-mode MAPSE, and correlation with LV GLS and LV EF. Methods Deep learning networks were trained to recognize apical views, determine the timing of end-systole and end-diastole, identify landmarks, estimate motion, and track B-mode MAPSE in real-time (Figure 1). This approach was tested in a prospective cohort of 52 patients. M-mode MAPSE and 2DS LV GLS measurements were obtained using semi-automatic software (EchoPAC version 204, GE Healthcare). Apical 4-chamber septal and lateral MAPSE values were acquired and averaged with both methods. Agreement and correlation were assessed using Bland-Altman plots and Pearson correlation coefficients. Results The feasibility for DL-MAPSE was 98%. We observed a method-specific bias, with lower MAPSE values for DL-MAPSE in the apical 4-chamber view (bias -2.95 mm, LoA -7.56 to 1.66, Figure 2A). There was a moderate correlation between methods (r = 0.60, 95% CI 0.40–0.75). DL-MAPSE and GLS showed a moderate-to-strong correlation (r = 0.69, 95% CI 0.51–0.81, figure 2B), which was stronger than that between GLS and M-mode MAPSE (r = 0.57, 95% CI 0.34–0.73). However, the difference in correlation strength was not statistically significant after Fishers’s Z-test (p = 0.33). MAPSE and LV EF had weak-to-moderate correlations (DL-MAPSE: r = 0.38, M-Mode MAPSE: r = 0.42, p = 0.80). For a subgroup of patients, acquisition and measurement times for DL-MAPSE and M-mode MAPSE were compared. M-mode image acquisition (septal and lateral-focused) took an average of 36 ± 10 seconds, with analysis requiring 37 ± 8 seconds, totaling a mean of 73 ± 13 seconds to acquire M-mode MAPSE values. In contrast, the acquisition and analysis time of DL-MAPSE was 33 ± 15 seconds, resulting in a mean time reduction of 40 seconds (95% CI 32-48 seconds). Conclusion Fully automated DL-based B-mode MAPSE measurements in real-time are feasible, efficient, and with strong correlation to LV GLS. By addressing the potential pitfalls of incorrect angling and out-of-line movement, this method offers the potential for more reliable MAPSE measurements, enhancing its utility in clinical practice. Agreement and correlation
Abstract Background Heart failure (HF) is a complex and debilitating condition. The pathophysiology of HF is related to the dynamic interplay governing the energy transfer from the heart to the vascular bed. This energy transfer can be estimated by calculating the heart's external power, measured in watts, which is the product of flow and pressure divided by time. Although power measurements can potentially improve hemodynamic assessment in patients with heart failure, clinical estimation of cardiovascular energy transfer is rarely used due to cumbersome analyses and dependence on invasive measurements. We propose a novel method based on echocardiography and continuous non-invasive blood pressure measurement to estimate cardiovascular energy transfer. Aim To test whether non-invasive estimates of cardiovascular energy transfer can differentiate between HF patients with reduced and preserved left ventricular ejection fraction (LVEF). Methods We included patients hospitalized with decompensated HF not optimally treated in the study. All patients underwent an echocardiographic examination using a GE HealthCare Vivid E95 scanner. We used the biplane Simpson method to calculate LVEF and the pulsed wave Doppler velocity spectrum in the left ventricular outflow tract to estimate flow. Continuous non-invasive blood pressure measurements were obtained using a Finapres finger cuff device calibrated with the brachial blood pressure. We developed an application for tracing the velocity spectrum, synchronizing flow and pressure curves, and calculating the power curves as the product of continuous synchronized flow and pressure curves. The patients were categorized into two groups based on whether LVEF was < 40% or ≥ 40%. Results We observed that cardiovascular energy transfer, estimated as external power, differed between patients with LVEF <40% vs. ≥40%. The study cohort comprised twenty-eight patients (25% women), averaging 76 years (SD 11). Eighteen patients had LVEF <40%, with a mean LVEF of 27% (SD = 9%), mean velocity time index (VTI) 11.7 cm (SD = 2.94), and) and mean systolic blood pressure 116 mmHg (SD = 19). Ten patients had LVEF ≥40%, with a mean LVEF of 47% (SD = 6 %), VTI 16.5 cm (SD = 5.50), and mean systolic blood pressure 143 mmHg (SD = 20). External power was significantly different between patients with LVEF <40% vs. ≥40%, 0.81 W (SD = 0.21 W) vs. 1.43 W (SD = 0.41 W), respectively, p<0.0001 (Fig. 1). We observed a modest but significant correlation between LVEF and power (R = 0.28, P-value < 0.005). Conclusion Cardiovascular energy transfer, estimated non-invasively as external power, was significantly lower in acutely decompensated heart failure patients with LVEF <40% compared to those with LVEF ≥ 40%. More extensive studies should explore the possible value of estimating cardiovascular energy transfer in stratification and therapeutic decisions in patients with HF.
Abstract Background Two-dimensional echocardiographic strain measurements have limited reproducibility, hampering the ability to detect subtle changes in repeated examinations. This is partly related image acquisition, including apical foreshortening and suboptimal transducer rotation and/or tilt, resulting in non-standardized imaging planes. However, the impact of image acquisitions on measured global longitudinal strain (GLS) values has not been quantified in large studies due to the labor-intensive task of measuring apical foreshortening and characterizing the imaging plane. Recent advances in echocardiographic image analysis based on deep learning (DL) may shed new light on this problem by enabling measurements of apical foreshortening and estimation of transducer angulation without human intervention. Purpose To quantify the effect of apical foreshortening and transducer rotation/tilt on GLS values. Methods To isolate the effect of foreshortening and transducer rotation/tilt on GLS values, 1395 patients without cardiac pathology were included in the analysis. For each patient, GLS was measured in the three apical views (four chamber, two chamber, long-axis) using commercial software. The same recordings were automatically analyzed using DL to obtain indicators of apical foreshortening (apex displacement in the long axis and short axis directions) and of the transducer rotation and tilt relative to the heart. Correlations between strain values and the foreshortening/rotation/tilt indicators were calculated. When significant correlation was detected, a regression analysis was performed to quantify the absolute change in GLS that can be related to foreshortening and transducer rotation/tilt. The first and last 5% quantiles of each DL indicator were excluded from these analyses. Results For each view, GLS was significantly correlated (p<0.05) with at least one of the foreshortening or transducer angulation indicators (Fig. 1). The regression analyses suggested that apical foreshortening and suboptimal transducer rotation/tilt can cause GLS variations up to 1.3 percentage points, representing a relative change of 6.5%. Conclusions Apical foreshortening and transducer rotation/tilt quantified using automatic DL methods were related to GLS values in a large dataset. The magnitude of the relationship suggests that apical foreshortening and transducer rotation/tilt can limit the usefulness of GLS measurements, especially in the context of repeated examinations where detecting small relative changes is important. Overall, this retrospective analysis confirms the importance of minimizing apical foreshortening and of standardizing transducer angulation to reduce minimal detectable changes in GLS and improve its reproducibility.
Background Automated measurements in cardiac imaging with the use of deep learning (DL) is a highly active area of research and innovation. However, some concerns challenge the translation of DL methods from research to clinical implementation. Objectives The authors evaluated 3 challenges for cardiac measurements by DL using left ventricular ejection fraction (LVEF) for management of heart failure and discuss mitigation strategies. Methods Using 3 different populations (N = 3,538), automated LVEF measurements were obtained with the use of supervised end-to-end learning and analyzed in terms of HF management. Three common challenges related to evaluation metrics, training data, and model generalization were studied. Results For the evaluation challenge, the authors identified significant unreliability of the AUC when applied to dichotomized heart failure diagnosis. Specifically, AUC varied from 0.71 to 0.98 owing solely to changes in population characteristics. For the training data challenge, model performance could be enhanced even after reducing the number of training subjects by 40%. For the generalization challenge, a performance degradation was observed compared with internal data when testing the model on external data. Integrating medical imaging domain knowledge in the DL framework effectively helped to recover performance and improve generalizability. Conclusions Both training data and generalization aspects challenge the performance of DL algorithms for automated cardiac measurements. In addition, evaluation metrics challenge the ability to detect underperforming algorithms. By considering evaluation metrics and training data distribution, and incorporating imaging domain knowledge, the design and evaluation of DL models can be improved, leading to more robust models, improved interpretation, and easier comparison across data sets. These findings may guide researchers and clinicians in implementing DL models for cardiovascular imaging.