BACKGROUND:Dyslipidemia, distinctly characterized by hypertriglyceridemia and impaired high-density lipoprotein cholesterol (HDL-C) metabolism in patients with chronic kidney disease (CKD), is associated with kidney disease progression. However, evidence on the renal outcomes of Triglyceride (TG)-targeting lipid-lowering therapies (LLTs), such as fibrates and niacin, often prescribed to patients with non-dialysis-dependent (NDD) CKD, is insufficient and equivocal. METHODS:We designed this retrospective cohort study using a target trial emulation framework and leveraged data from a nationwide cohort of 3,562,882 US Veterans to identify incident CKD patients initiating de novo LLTs. We compared eGFR decline from baseline and the odds of rapid eGFR decline in fibrate and niacin initiators versus statin initiators using the mixed-effects model-predicted 1-year, 3-year, and 10-year eGFR slopes, adjusted for baseline covariates. We also compared the risk of major adverse kidney events [MAKE] (composite of a sustained ≥57% decrease in eGFR from baseline, sustained eGFR<15 mL/min/1.73 m2, incident end-stage renal disease, and all-cause death) using Cox proportional hazards models adjusted for baseline covariates during follow-up of up to 10 years. RESULTS:Among 68,073 patients with incident CKD, 62,866 initiated de novo statin, 3,202 fibrate, or 2,005 niacin therapy. Compared to statin use, de novo fibrate initiators had slower eGFR decline (difference in 10-year slope: 0.26 mL/min/1.73 m2/year [95% CI, 0.11, 0.40], P<0.001), whereas the eGFR decline was not statistically significant in the niacin vs. statin comparison. The risk of MAKE was not significantly different in fibrate (hazard ratio, 0.94 [95% CI, 0.88, 1.01], P=0.09) or niacin (0.96 [0.89, 1.04], P=0.31) users after a median follow-up of 3.69 years. CONCLUSIONS:Among patients with NDD-CKD, unlike niacin, long-term fibrate use (compared to statin use) was associated with slower eGFR decline, suggesting renoprotective effects of fibrates. Further studies are warranted to investigate whether fibrate use is beneficial for preventing other outcomes, such as adverse cardiovascular events.
BACKGROUND:In patients with chronic kidney disease (CKD), lower and subfunctional high-density lipoprotein cholesterol (HDL-C) is associated with poor cardiovascular outcomes. Notwithstanding such poor outcomes, the primary therapeutic target in patients with CKD is low-density lipoprotein cholesterol (LDL-C), and the comparative effectiveness of commonly used lipid-lowering therapies (LLTs) in changing HDL-C levels in patients with non-dialysis-dependent CKD (NDD-CKD) remains unclear. METHODS:In this retrospective cohort study, using a target trial emulation framework, we examined a nationwide cohort of 3,562,882 US Veterans with normal kidney function enrolled between October 2004 and September 2006 and identified 247,270 incident CKD patients eligible for de novo LLT exposure, occurring during longitudinal follow-up until October 2019. We defined de novo LLT initiation using pharmacy dispensation data and followed patients for up to 1 year. We compared the intraindividual slopes of HDL-C levels in de novo fibrates and niacin users with those in statin users, using mixed-effects models adjusted for baseline and time-varying covariates. We also compared the odds of having a clinically meaningful (>10% from baseline) increase in HDL-C following LLT initiation. RESULTS:A total of 38,223 patients with incident CKD initiated de novo LLT (statin [n = 35,284], fibrate [n = 1,805], and niacin [n = 1,134]). The mean (standard deviation) age was 67.3 (10.5) years; 95.0% were men, and 20.6% were Black. Compared to statin users, the multivariable-adjusted annualized intraindividual increase of HDL-C was significantly higher following fibrate {1.15 mg/dL/year (95% confidence interval [CI]: 0.43, 1.87); p = 0.002} and niacin monotherapy (2.51 mg/dL/year [95% CI: 1.62, 3.41]; p < 0.001). Furthermore, niacin (OR: 1.37 [95% CI: 1.07, 1.75]; p = 0.012) was more likely than statins to provide a clinically meaningful elevation in HDL-C. Our findings were consistent in several sensitivity analyses. CONCLUSION:Among patients with NDD-CKD, de novo prescriptions of fibrates or niacin are associated with a greater increase in HDL-C levels compared to statins. Further studies are warranted to investigate whether such differences have meaningful effects on clinical outcomes.
The apple does not fall far from the tree is an old idiom that encapsulates a key concept: being related extends beyond merely sharing genetic material. It often implies sharing a common environment, including culture, language, dietary habits, and geographical location. In this study, we show that the analysis of genetic relatedness can serve as an indicator of health conditions by capturing the combined influences of genetic inheritance and shared environmental factors. We mapped the genomic data and electronic health records from 13k individuals in the Integrative Genomics Biorepository cohort, to neighborhood-level geographic data and integrated census-based environmental metrics. We used an identity-by-descent (IBD) based hierarchical community detection algorithm to identify four main communities closely aligning with continental ancestry, and seventeen subcommunities. We found uneven exposure to elevated environmental stressors across subcommunities, and we were able to identify subcommuinities at high risk of respiratory and dermatological conditions, demonstrating the potential of our framework and its possible application to public health interventions. Interestingly, for conditions such as congenital disorders, the most important differences were detected between subcommunities within the same community, indicating the relevance of considering a fine-grained population structure. These findings show how genetic relatedness, even at distant levels, can reflect shared environments and social determinants of health, providing a framework for understanding health disparities. ### Competing Interest Statement The authors have declared no competing interest.
Rationale: Rare pulmonary diseases (RPDs) in children, such as childhood interstitial lung disease (chILD) and genetic forms of bronchiectasis (including cystic fibrosis [CF] and primary ciliary dyskinesia [PCD]), are difficult to diagnose and often misdiagnosed as asthma, leading to delays in appropriate treatment. Early diagnosis through newborn genetic screening has dramatically increased CF life expectancy, highlighting the importance of early RPD diagnosis for early intervention and preservation of pulmonary function. Anecdotal evidence suggests that RPD are often initially misdiagnosed as severe or difficult-to-treat asthma, and misdiagnosing RPD as asthma delays proper care and exposes children to harmful and costly asthma treatments. However, the extent of this problem in clinical practice has not been explored. Objective: This study aims to estimate the prevalence of undiagnosed RPDs in children with severe asthma, testing the hypothesis that RPDs are more common in severe compared to non-severe asthma cases. Methods: We leveraged data from the Genomic Information Commons (GIC), a network of six academic children's hospitals with access to electronic health records (EHRs), biosamples, and genomic data. Using GIC's PIC-SURE query tool, we analyzed EHRs from 14,907,602 patients for ICD-10 codes related to asthma, severe asthma, bronchiectasis, and chILD. Chi-square tests were used to assess RPD prevalence differences between severe and non-severe asthma patients. For 6,019 asthma patients with exome sequence data, we further analyzed rare genetic variants associated with RPD. Results: Of 441,864 patients diagnosed with asthma across GIC hospitals, 10,753 had severe asthma (2.4%), 11,388 had bronchiectasis (2.6%), and 7,834 had chILD (1.8%). Among patients with severe asthma, RPD prevalence was 5.9 times higher than in those with non-severe asthma (4.7% vs. 0.8%, p < E-16), with notable increases for both bronchiectasis (3.5% vs. 0.5%, 6.8-fold increase) and chILD (1.2% vs. 0.3%, 4.3-fold-increase). Among patients with exome sequence data, those with severe asthma showed a 4.1-fold increase in RPD-related functional genetic variants compared to non-severe asthma cases (17.3% vs. 4.2%, p < 0.00001). Filaggrin (FLG) gene variants, linked to severe asthma, were present in ∼1/3 of severe cases. Excluding FLG variants, enrichment remained significant, especially in genes causing PCD. Conclusion: This large-scale analysis suggests that 1 in every 21 patients with severe asthma carry an RPD diagnosis. Though reliant on ICD-10 codes, our observations are further supported by a 4-fold increase in RPD-causing loss-of-function genetic variants, underscoring the need for heightened clinical awareness of these disorders in patients with severe asthma.
Rationale & Objectives:Extreme ambient temperatures have been associated with a higher risk of acute and chronic health outcomes, including cardiovascular, respiratory, infectious, and kidney diseases. However, there is a lack of synthesized comprehensive evidence regarding the association of ambient temperature and renal colic in existing literature. Study Design:Systematic review and meta-analysis of epidemiological studies. Setting & Population:Population of any geographic areas regardless of their age, sex/gender, ethnicity, or any other population characteristics. Selection Criteria for Studies:We conducted literature searches in PubMed, Scopus, CINAHL complete, Web of Sciences, and additional sources until June 4, 2024, following the Population-Exposure-Comparator-Outcome (PECO) framework. Exposure:Daily ambient temperature. Outcomes:Confirmed cases of renal colic, including its underlying causes, such as kidney stone/nephrolithiasis or urolithiasis, urinary tract infection, and pyelonephritis, etc. Data Extraction:Two investigators performed data extraction to ensure data consistency and quality. Analytical Approach:Random-effects meta-analyses using the DerSimonian and Laird method. Results:Of 982 initially retrieved articles, we included 26 articles in the systematic review, of which 23 were eligible for the heat effect meta-analysis and 7 were included in the cold effect meta-analysis. Despite high heterogeneity, the result showed a 2.4% higher risk of renal colic for 1°C higher daily ambient temperature (Cohen's d, 0.013 [95% CI, 0.010-0.015]; RR, 1.024 [95% CI, 1.020-1.028]; I 2 , 99.7%, P <0.001). For a 1 °C lower daily ambient temperature, the risk of renal colic was 1.5% (Cohen's d, 0.008 [95% CI: -0.000 to 0.016]; RR, 1.015 [95% CI, 1.000-1.029]; I 2 , 73.2%; P > 0.05), although not statistically significant. Limitations:Limitations of this systematic review include high heterogeneity and publication bias. Conclusions:Elevated daily ambient temperature is associated with the risk of renal colic, suggesting an adverse effect of high ambient temperature on kidney function. Registration:PROSPERO registration number: CRD420245555.
Background:Electronic health records contain inconsistently structured or free-text data, requiring efficient preprocessing to enable predictive health care models. Although artificial intelligence-driven natural language processing tools show promise for automating diagnosis classification, their comparative performance and clinical reliability require systematic evaluation. Objective:The aim of this study is to evaluate the performance of 4 large language models (GPT-3.5, GPT-4o, Llama 3.2, and Gemini 1.5) and BioBERT in classifying cancer diagnoses from structured and unstructured electronic health records data. Methods:We analyzed 762 unique diagnoses (326 International Classification of Diseases [ICD] code descriptions, 436 free-text entries) from 3456 records of patients with cancer. Models were tested on their ability to categorize diagnoses into 14 predefined categories. Two oncology experts validated classifications. Results:BioBERT achieved the highest weighted macro F1-score for ICD codes (84.2) and matched GPT-4o in ICD code accuracy (90.8). For free-text diagnoses, GPT-4o outperformed BioBERT in weighted macro F1-score (71.8 vs 61.5) and achieved slightly higher accuracy (81.9 vs 81.6). GPT-3.5, Gemini, and Llama showed lower overall performance on both formats. Common misclassification patterns included confusion between metastasis and central nervous system tumors, as well as errors involving ambiguous or overlapping clinical terminology. Conclusions:Although current performance levels appear sufficient for administrative and research use, reliable clinical applications will require standardized documentation practices alongside robust human oversight for high-stakes decision-making.
The Biorepository and Integrative Genomics (BIG) Initiative in Tennessee has developed a pioneering resource to address gaps in genomic research by linking genomic, phenotypic, and environmental data from a diverse Mid-South population, including underrepresented groups. We analyzed 13,152 exomes from BIG and found significant genetic diversity, with 50% of participants inferred to have non-European or several types of admixed ancestry. Ancestry within the BIG cohort is stratified, with distinct geographic and demographic patterns, as African ancestry is more common in urban areas, while European ancestry is more common in suburban regions. We observe ancestry-specific rates of novel genetic variants, which are enriched for functional or clinical relevance. Disease prevalence analysis linked ancestry and environmental factors, showing higher odds ratios for asthma and obesity in minority groups, particularly in the urban area. Finally, we observe discrepancies between self-reported race and genetic ancestry, with related individuals self-identifying in differing racial categories. These findings underscore the limitations of race as a biomedical variable. BIG has proven to be an effective model for community-centered precision medicine. We integrated genomics education, and fostered great trust among the contributing communities. Future goals include cohort expansion, and enhanced genomic analysis, to ensure equitable healthcare outcomes.
Adherence to scheduled radiation therapy (RT) is a key determinant of cancer treatment quality and outcomes. For this study, we developed an interpretable AI model to identify 1) patients at risk for multiple unplanned RT interruptions and 2) modifiable factors contributing to an elevated risk of RT interruption. We retrospectively analyzed clinical, socioeconomic, demographic, and behavioral data from 2,525 RT patients treated at the University of Tennessee Medical Center (UTMC) in Knoxville. The study cohort was dichotomized into patients with 0-1 unplanned RT interruptions (Class 0; n≈2000) and those missing >2 sessions (Class 1; n≈500). The dataset was partitioned into training, validation, and test sets (70:15:15 ratio), with class imbalance addressed in the training set by synthetic data generation via Tabular Variational Autoencoder. Twenty-seven candidate features were initially evaluated for multicollinearity using correlation matrices, heatmap visualization, and Variance Inflation Factor analysis. We applied feature selection methods (correlation-based techniques and causality-based approaches) to limit further modeling to the most predictive 15 core features. We compared XGBoost and Neural Networks-based classifiers, with each model undergoing hyperparameter optimization using Bayesian optimization methods. SHapley Additive exPlanations (SHAP) analysis was used to identify influential predictors. The final optimized XGBoost model provided an overall accuracy of 82% and AUC-ROC of 63% on the independent test set. All tested models yielded similar performance, confirming the consistent predictive value of our selected features despite class imbalance. SHAP analyses identified dominant predictive contributions from treatment factors (prescribed radiation dose per session), patient resources (insurance coverage, marital status, social vulnerability indices), and travel distance to the radiotherapy facility. Supplementary causal analysis employing total causal effect methods further corroborated the direct influence of all these features. Our results suggest that causal inference and explainable AI modeling can provide useful interrogative strategies to identify modifiable predictors of RT adherence. Further refinement of predictive decision-support tools may lead to automated approaches to match high-risk patients with personalized interventions (e.g. community-based care navigation and/or patient psychosocial support) in real-world clinical settings to overcome social barriers to RT access. Rezaur Rashid, Soheil Hashtarkhani, Parnian K. Rahimabad, Brianna M. White, Fekede A. Kumsa, Lokesh Chinthala, Janet A. Zink, Christopher L. Brett, Robert L. Davis, David L. Schwartz, Arash Shaban-Nejad. Machine Learning and Causal Inference-Based Predictive Risk Modeling of Unplanned Radiation Treatment Interruption [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A061.
Psychological distress impacts cancer treatment adherence, outcomes, and quality of life. While predictive models can identify patients at elevated risk, understanding the mechanistic drivers of distress will be essential to design personalized psychosocial interventions. In this study, we applied explainable machine learning to predict and interpret baseline distress levels in cancer patients preparing to undergo radiation therapy. We retrospectively analyzed data from 2,103 patients treated at the University of Tennessee Medical Center (UTMC) in Knoxville. 1,953 were formally screened by an ONC Distress Depression Score instrument prior to radiation therapy, with 1,759 providing a numeric distress rating (0–10). An XGBoost regression model was trained to predict patient distress scores using a broad set of variables, including age, body mass index (BMI), ICD-coded cancer diagnosis category, marital status, insurance type, smoking and alcohol use, and neighborhood-level socioeconomic indicators such as median household income, educational attainment, and social vulnerability. Hyperparameters were optimized via Bayesian search, and model performance was evaluated with cross-validation. To ensure interpretability, we employed Shapley Additive Explanations (SHAP) to quantify each feature’s contribution to the predicted distress and to visualize the direction and non-linearity of these effects. Distress levels were skewed toward the lower end of the scale, with 55.6% of patients reporting mild (<4), 20.0% moderate (4–6), and 15.9% severe (7–10) distress. The final model demonstrated moderate predictive accuracy, explaining approximately 59% of the variance in distress scores (R2 = 0.592, RMSE = 2.00, MAE = 1.67). SHAP analysis provided interpretable insights into how individual features influenced distress predictions. The top six contributors to predicted distress were younger age, lung cancer diagnosis, metastatic disease, current smoking, dual Medicare-Medicaid insurance coverage, and divorced marital status; all were associated with higher distress levels. These relationships highlighted both linear and non-linear effects, offering clinically meaningful explanations of patient-level risk. Our results suggest that explainable AI can contribute accurate predictions and interpretable insights into the contributors of psychological distress in cancer patients prior to radiation therapy. By combining XGBoost modeling with SHAP-based explanation, we uncovered potentially modifiable candidate drivers of distress. These findings support ongoing investigations to integrate interpretable machine learning strategies into clinical workflows to enhance preemptive, personalized supportive care interventions during radiation therapy planning. Soheil Hashtarkhani, Rezaur Rashid, Parnian K. Rahimabad, Fekede A. Kumsa, Brianna M. White, Lokesh Chinthala, Janet A. Zink, Christopher L. Brett, Robert L. Davis, David L. Schwartz, Arash Shaban-Nejad. Uncovering Social and Clinical Determinants of Baseline Distress Prior to Radiation Therapy: An Explainable Machine Learning Approach [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A062.
The Biorepository and Integrative Genomics (BIG) Initiative in Tennessee has developed a pioneering resource to address gaps in genomic research by linking genomic, phenotypic, and environmental data from a diverse Mid-South population, including underrepresented groups. We analyzed 13,152 exomes from BIG and found significant genetic diversity, with 50% of participants inferred to have non-European or several types of admixed ancestry. Ancestry within the BIG cohort is stratified, with distinct geographic and demographic patterns, as African ancestry is more common in urban areas, while European ancestry is more common in suburban regions. We observe ancestry-specific rates of novel genetic variants, which are enriched for functional or clinical relevance. Disease prevalence analysis linked ancestry and environmental factors, showing higher odds ratios for asthma and obesity in minority groups, particularly in the urban area. Finally, we observe discrepancies between self-reported race and genetic ancestry, with related individuals self-identifying in differing racial categories. These findings underscore the limitations of race as a biomedical variable. BIG has proven to be an effective model for community-centered precision medicine. We integrated genomics education, and fostered great trust among the contributing communities. Future goals include cohort expansion, and enhanced genomic analysis, to ensure equitable healthcare outcomes.
The goal of this study is to investigate the contributing factors resulting in asthma and related allergic disorders across different subpopulations, focusing on the impact of social determinants of health (SDOH). Using data from the All of Us Research Program, we employed Random Forest (RF) models to predict asthma risk and determine the most important variables among various SDOH factors studied. We developed race-specific models tailored to each racial group, as well as a global model encompassing all racial groups. Our analysis reveals no significant difference in predictive performance between the race-specific models and the race-combined model. However, SHAP (SHapley Additive exPlanations) analysis reveals that the race-specific models substantially differ in terms of the most important SDOH factors identified as the key influencing factors to the stratification. These findings emphasize the role of specific SDOH factors in shaping asthma outcomes across demographic groups and the importance of developing models that can avoid bias and highlight the relevant actionable insights to equitably advance the understanding of SDOH and health outcome associations for everyone.
10022 Background: Childhood cancer survivors (CCS) face an increased risk of developing cardiomyopathy during adult life due to late onset of treatment (e.g. anthracyclines and chest directed radiation) associated cardiotoxicity. Early identification of survivors at elevated risk using low cost and easily accessible data modalities can identify survivors in need of echocardiographic screening. Our goal is to utilize 12-lead electrocardiogram (ECG) develop and externally validate an artificial intelligence model (ECG-AI) for prediction of five-year risk for cardiomyopathy among CCS. Methods: We developed three deep learning models including a modified ResNet (a convolutional neural network architecture), an encoder-attention network, and a dual attention network to predict cardiomyopathy risk from 10 second 12-lead ECGs. We used the St Jude Lifetime Cohort Study (SJLIFE) for model building using a 60/20/20 patient-level split for training, validation, and holdout testing. We evaluated model performance for all cardiomyopathy grades (Common Terminology Criteria for Adverse Events) and specifically grade 3 (severe) cases. The final model was externally validated in the Dutch Childhood Cancer Survivor Study (DCCSS-LATER) cohort. For the DCCSS-LATER cohort, we evaluated accuracy of ECG-AI for the five-year cardiomyopathy risk prediction. Results: SJLIFE analytical cohort included 7,632 ECGs from 4,795 unique participants with no cardiomyopathy. 228 participants developed cardiomyopathy at least one year after index ECG date. Participants were 79% white, 11% Black, and 49% male with mean age at ECG of 33±10 years. In SJLIFE holdout, the encoder-attention model achieved the highest performance (area under the receiver operating characteristic curve (AUC) 0.75 for all grades and 0.84 for ≥grade 3 cases). The modified ResNet and dual attention models achieved AUCs of 0.69 and 0.72, respectively. DCCSS-LATER data included 749 ECGs with from 330 unique patients (48% male, age at ECG of 28±10 years). 22 patients developed cardiomyopathy at least one year after index ECG date. The encoder-attention model achieved an AUC of 0.74 for 5-year cardiomyopathy risk prediction. We note that the cardiomyopathy grading was not available in DCCSS-LATER, this study instead used a broader cardiomyopathy diagnosis information. Conclusions: ECG-AI analysis of standard 10 second 12-lead ECGs can identify childhood cancer survivors at risk for future cardiomyopathy with moderate to high accuracy depending on the cardiomyopathy severity. Future studies will focus on improving accuracy by incorporating clinical data such as B-type natriuretic peptides, left ventricular ejection fraction.
BACKGROUND:Preeclampsia is a hypertensive disorder in pregnancy known to increase the risk of mortality and other pregnancy-related issues, such as prematurity. Currently, there no known prophylactics or treatment options available for preeclampsia. More research is needed to better understand factors that increase preeclampsia risk. Vitamin D deficiency is consistently associated with developing preeclampsia. In addition to micronutrient deficiency, the presence of two fetal apolipoprotein L1 high-risk variants are also associated with preeclampsia risk. We hypothesized that a potential additive effect between high-risk apolipoprotein L1 genotype status and nutritional deficiencies would place individuals at a higher risk of developing preeclampsia. OBJECTIVE (S):The objective of this study was to determine the risk of developing preeclampsia in African American women with vitamin D deficiency and maternal/fetal high-risk apolipoprotein L1 genotype. STUDY DESIGN:This was a case-control study using a subset of 999 African American mother and infant pairs collected from the Conditions Affecting Neurocognitive Development and Learning in Early Childhood cohort in Memphis, TN. We performed multiple logistic regression to examine the association of preeclampsia with 2nd and 3rd trimester vitamin D concentrations. Concentrations were dichotomized into high or low categories. Vitamin D deficiency was defined as a concentration less than 20 ng/mL. Further analyses assessed whether maternal or fetal apolipoprotein genotype status modified the association between vitamin D association and preeclampsia. The reference group included individuals with both high vitamin D and low-risk apolipoprotein genotype. RESULTS:Pregnancies with low vitamin D in the 3rd trimester were at an increased risk for preeclampsia (odds ratio 2.10; 95 % confidence interval 1.09-4.12; P-value, 0.03). Risk for preeclampsia was greatest among pregnancies with fetal high-risk genotype and low vitamin D levels in the 2nd trimester (odds ratio, 2.79; 95 % confidence interval, 1.06-6.83; P-value, 0.03) and 3rd trimester (odds ratio 6.40; 95 % confidence interval 2.07-19.18; P-value, <0.01). CONCLUSION(S):Our significant findings suggest that the risk of preeclampsia associated with low vitamin D levels, especially during the 3rd trimester, is magnified by the presence of fetal high-risk apolipoprotein L1 genotype.
Background: There are delays on identification of people with low left ventricular ejection fraction (LVEF) and Heart Failure with preserved EF (HFpEF). There is a need for low cost and accessible tools to identify individuals who can benefit from more comprehensive exams such as echocardiogram (ECHO). Research Goal To develop electrocardiographic artificial intelligence (ECG-AI) models that can classify low EF and HFpEF. Methods: We developed an ECG-AI model using convolutional neural networks to classify 12-lead ECG’s into four categories: rEF( LVEF<40), mEF (40≤LVEF<50), HFpEF (clinical HF diagnosis with EF≥50) and controls with no HF diagnosis within 5-years of the index ECG. For rEF, mEF, and HFpEF categories, ECHO were used to determine LVEF. Patients with clinical HF diagnosis without ECHO data were excluded. In patients with ECHO, only available ECGs within 30-days of the ECHO study was utilized. The ECG-AI models were trained on ~80%, validated on ~10%, and tested on the remaining 10% hold out data from Wake Forest Baptist Health (WFBH) ECG repository and externally validated using data from the University of Tennessee Health Science Center (UTHSC). Results: The ECG-AI model was developed using 1,078,198 digital ECGs, 114,068 ECHO derived LVEF values, and EHR data from 165,243 patients (73% White, 19% Black, 52% female, with a mean age (SD) of 58(15) years). There were 32,962, 40,997, 11,037, and 993,202 ECGs from 8,555 rEF, 13,116 mEF, 3,704 HFpEF, and 139,868 control patients, respectively. The ECG-AI achieved a hold-out AUC of 0.76 (0.74-0.77) in classifying ECGs as HFpEF or not (Table 1). The UTHSC external validation cohort included 273 rEF, 167 mEF, 459 HFpEF, and 35,254 control patients, with 35% White, 62% Black, 60% female, mean age of 51(18) years. Our model identified HFpEF patients with an AUC of 0.75 (0.73-0.76) in the UTHSC cohort. Conclusion: ECG data alone result in moderate, moderately high, and very high accuracies in classifying HFpEF, mEF, and rEF, respectively. This could guide future models with incorporation of simple patient demographics and comorbidities to improve accuracy. Such ECG-AI models can assist with identifying who may benefit from further clinical evaluation and imaging studies.
Background: Fatal coronary heart disease (FCHD) affects ~650,000 people yearly in the US. Electrocardiographic artificial intelligence (ECG-AI) models can predict adverse coronary events, yet their application to FCHD is understudied. Objectives: The study aimed to develop ECG-AI models predicting FCHD risk from ECGs. Methods (Retrospective): Data from 10 s 12-lead ECGs and demographic/clinical data from University of Tennessee Health Science Center (UTHSC) were used for model development. Of this dataset, 80% was used for training and 20% as holdout. Data from Atrium Health Wake Forest Baptist (AHWFB) were used for external validation. We developed two separate convolutional neural network models using 12-lead and Lead I ECGs as inputs, and time-dependent Cox proportional hazard models using demographic/clinical data with ECG-AI outputs. Correlation of the predictions from the 12- and 1-lead ECG-AI models was assessed. Results: The UTHSC cohort included data from 50,132 patients with a mean age (SD) of 62.50 (14.80) years, of whom 53.4% were males and 48.5% African American. The AHWFB cohort included data from 2305 patients with a mean age (SD) of 63.04 (16.89) years, of whom 51.0% were males and 18.8% African American. The 12-lead and Lead I ECG-AI models resulted in validation AUCs of 0.84 and 0.85, respectively. The best overall model was the Cox model using simple demographics with Lead I ECG-AI output (D1-ECG-AI-Cox), with the following results: AUC = 0.87 (0.85–0.89), accuracy = 83%, sensitivity = 69%, specificity = 89%, negative predicted value (NPV) = 92% and positive predicted value (PPV) = 55% on the AHWFB validation cohort. For this, the 2-year FCHD risk prediction accuracy was AUC = 0.91 (0.90–0.92). The 12-lead versus Lead I ECG FCHD risk prediction showed strong correlation (R = 0.74). Conclusions: The 2-year FCHD risk can be predicted with high accuracy from single-lead ECGs, further improving when combined with demographic information.
Background: Tissue hypoxia and chronic anemia associated with sickle cell disease (SCD) leads to structural and physiological alterations in the heart. Early detection of heart failure (HF) in patients with SCD can assist with timely interventions, but current methods (e.g., echocardiogram and heart MRI) are not easily accessible in resource-deprived settings. The integration of artificial intelligence (AI)-powered tools utilizing low-cost ECG data to increase the power to detect more patients eligible for early treatment, thus improving patient outcomes, and needs to be validated. Hypothesis: We hypothesize that ECG-AI models developed to detect incident HF in the general population can detect HF in SCD patients. Methods/Approach: We previously developed an ECG-AI model employing convolutional neural networks to classify patients with HF using a large ECG-repository at Wake Forest Baptist Health (WFBH). This model was developed using 1,078,198 digital ECGs from 165,243 patients, 73% White, 19% Black, and 52% female individuals, with a mean age (SD) of 58 (15) years. The hold-out AUC of this previous model in distinguishing ECGs of HF patients from controls was 0.87. In this study, we externally validated this ECG-AI model using SCD patients’ data from the University of Tennessee Health Science Center (UTHSC). Additionally, a logistic regression (LR) model was constructed in the UTHSC cohort by incorporating other simple demographic variables with the outcome of ECG-AI model. Results/Data: The UTHSC external validation cohort included data from 2,107 SCD patients (188 HF and 1,919 SCD patients with no HF), 98% were Black, 72% were female, with a mean age of 39 (14) years. Despite demographic differences between the validation (more Blacks) and derivation cohorts (lower age), our ECG-AI model accurately identified HF with an AUC of 0.80 (0.77-0.82) in the UTHSC SCD cohort. When incorporating ECG-AI outcome (an ECG-based risk value between 0 and 1), age, sex, and race in a LR model, the AUC significantly improved (DeLong Test, p<0.01) to 0.84 (0.82-0.87) with a sensitivity of 0.76 (0.73-0.79) and specificity of 0.76 (0.74-0.79). Conclusion: The ECG-AI model can detect HF in SCD patients from ECG data alone with moderately high accuracy. The accuracy improves with incorporation of simple demographic data. Future studies will incorporate other clinical risk factors of HF and sickle cell genotype data for further improved accuracy allowing its use for surveillance.
Background Children from families with low socioeconomic status (SES), as determined by income, experience several negative outcomes, such as higher rates of newborn mortality and behavioral issues. Moreover, associations between DNA methylation and low income or poverty status are evident beginning at birth, suggesting prenatal influences on offspring development. Recent evidence suggests neighborhood opportunities may protect against some of the health consequences of living in low income households. The goal of this study was to assess whether neighborhood opportunities moderate associations between household income (HI) and neonate developmental maturity as measured with DNA methylation. Methods Umbilical cord blood DNA methylation data was available in 198 mother-neonate pairs from the larger CANDLE cohort. Gestational age acceleration was calculated using an epigenetic clock designed for neonates. Prenatal HI and neighborhood opportunities measured with the Childhood Opportunity Index (COI) were regressed on gestational age acceleration controlling for sex, race, and cellular composition. Results Higher HI was associated with higher gestational age acceleration (B = .145, t = 4.969, p = 1.56x10-6, 95% CI [.087, .202]). Contrary to expectation, an interaction emerged showing higher neighborhood educational opportunity was associated with lower gestational age acceleration at birth for neonates with mothers living in moderate to high HI (B = -.048, t = -2.08, p = .03, 95% CI [-.092, -.002]). Female neonates showed higher gestational age acceleration at birth compared to males. However, within males, being born into neighborhoods with higher social and economic opportunity was associated with higher gestational age acceleration. Conclusion Prenatal HI and neighborhood qualities may affect gestational age acceleration at birth. Therefore, policy makers should consider neighborhood qualities as one opportunity to mitigate prenatal developmental effects of HI.