BackgroundDepressive disorder, particularly major depressive disorder (MDD), significantly impact individuals and society. Traditional analysis methods often suffer from subjectivity and may not capture complex, non-linear relationships between risk factors. Machine learning (ML) offers a data-driven approach to predict and diagnose depression more accurately by analyzing large and complex datasets.MethodsThis study utilized data from the National Health and Nutrition Examination Survey (NHANES) 2013-2014 to predict depression using six supervised ML models: Logistic Regression, Random Forest, Naive Bayes, Support Vector Machine (SVM), Extreme Gradient Boost (XGBoost), and Light Gradient Boosting Machine (LightGBM). Depression was assessed using the Patient Health Questionnaire (PHQ-9), with a score of 10 or higher indicating moderate to severe depression. The dataset was split into training and testing sets (80% and 20%, respectively), and model performance was evaluated using accuracy, sensitivity, specificity, precision, AUC, and F1 score. SHAP (SHapley Additive exPlanations) values were used to identify the critical risk factors and interpret the contributions of each feature to the prediction.ResultsXGBoost was identified as the best-performing model, achieving the highest accuracy, sensitivity, specificity, precision, AUC, and F1 score. SHAP analysis highlighted the most significant predictors of depression: the ratio family income to poverty (PIR), sex, hypertension, serum cotinine and hydroxycotine, BMI, education level, glucose levels, age, marital status, and renal function (eGFR).ConclusionWe developed ML models to predict depression and utilized SHAP for interpretation. This approach identifies key factors associated with depression, encompassing socioeconomic, demographic, and health-related aspects.
Coronary heart disease (CHD) is a major cause of morbidity and mortality worldwide. Identifying key risk factors is essential for effective risk assessment and prevention. Machine learning (ML) offers advanced methods for analyzing complex datasets, revealing novel predictors of CHD beyond traditional models. This study aims to evaluate the contribution of various risk factors to CHD, focusing on both established and novel markers using machine learning techniques. The study recruited 7,672 participants aged 30 to 84 years from Suita City, Japan, between 1989 and 1999. Over an average of 15 years, participants were monitored for cardiovascular events. Five ML models—Random Forest (RF), XGBoost, Support Vector Machine (SVM), Logistic Regression (LR), and LightGBM—were used. The optimal model was identified based on accuracy, sensitivity, specificity, and AUC. SHapley Additive exPlanations (SHAP) were then employed to explore the contribution of various risk factors to CHD. RF achieved the highest AUC (95% CI) of 0.94 (0.93-0.96), outperforming LR, SVM, XGBoost, and LightGBM. SHAP on the best model identified the top CHD predictors. Intima-media thickness of common carotid artery (IMT_cMax) was identified as the strongest predictor of CHD, highlighting the importance of arterial health. Systolic and diastolic blood pressure, along with lipid profiles (non-HDL cholesterol, HDL cholesterol, and triglycerides), were closely associated with CHD incidence. eGFR underscored the link between renal function and CHD. Novel insights included the impact of lower calcium levels, systemic inflammation (elevated WBC counts), fructosamine levels, and obesity-related factor (body fat percentage). A protective effect in females indicated the need for sex-specific CHD management strategies. ML, particularly the RF model combined with SHAP, effectively identified key risk factors for CHD, including arterial health, blood pressure, lipid profiles, renal function, and novel markers. These findings support a multifactorial approach to CHD risk assessment.
The Shiga Epidemiological Study of Subclinical Atherosclerosis was conducted in Kusatsu City, Shiga, Japan, from 2006 to 2008. Participants were measured for LDL-p through nuclear magnetic resonance technology. 740 men participated in follow-up and underwent 1.5 T brain magnetic resonance angiography from 2012 to 2015. Participants were categorized as no-ICAS, and ICAS consisted of mild-ICAS (1 to < 50%) and severe-ICAS (≥ 50%) in any of the arteries examined. After exclusion criteria, 711 men left for analysis, we used multiple logistic regression to examine the association between lipid profiles and ICAS prevalence. Among the study participants, 205 individuals (28.8%) had ICAS, while 144 individuals (20.3%) demonstrated discordance between LDL-c and LDL-p levels. The discordance “low LDL-c–high LDL-p” group had the highest ICAS risk with an adjusted OR (95% CI) of 2.78 (1.55–5.00) in the reference of the concordance “low LDL-c–low LDL-p” group. This was followed by the concordance “high LDL-c–high LDL-p” group of 2.56 (1.69–3.85) and the discordance “high LDL-c–low LDL-p” group of 2.40 (1.29–4.46). These findings suggest that evaluating LDL-p levels alongside LDL-c may aid in identifying adults at a higher risk for ICAS.
We leveraged machine learning (ML) techniques, namely logistic regression (LR), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), and LightGBM to predict coronary heart disease (CHD) and identify the key risk factors involved. Based on the Suita study, 7672 men and women aged 30 to 84 years without cardiovascular disease were recruited from 1989 to 1999, in Suita City, Osaka, Japan. Over an average period of 15 years, participants were diligently monitored until the onset of their initial cardiovascular event or relocation. CHD diagnoses encompassed primary heart attacks, sudden death, or coronary artery disease with bypass surgery or intervention. RF achieved the highest AUC (95% CI) of 0.79 (0.70–0.87), outperforming LR, SVM, XGBoost, and LightGBM. Shapley Additive Explanations (SHAP) on the best model identified the top CHD predictors. Notably, systolic blood pressure, non-HDL-c, glucose levels, age, metabolic syndrome, HDL-c, estimated glomerular filtration rate, hypertension, elbow joint thickness, and diastolic blood pressure were key contributors. Remarkably, elbow joint thickness was identified as a previously unrecognized risk factor associated with CHD. These findings indicated that ML methods accurately predict incident CHD risk. Additionally, ML has identified new incident CHD risk variables.
Background: This population-based study investigated the potential of machine learning algorithms to predict stroke incidence and identify important risk factors. This study aimed to evaluate the accuracy of these algorithms in constructing a stroke prediction model. Methods: Participants from the Suita study were included, and baseline measurements were used to predict stroke outcomes over a 15-year follow-up period. In total, 7,389 participants and 51 variables were investigated, including demographics, medical history, medical imaging, laboratory data, and lifestyle habits. Initially, unsupervised K-prototype clustering was used to group participants based on their stroke risk. Subsequently, five supervised models (logistic regression, random forest, support vector machine, extreme gradient boosting, and light gradient boosted machine) were applied to predict the stroke outcomes. The Shapley Additive Explanations (SHAP) method determined the most critical variables. Results: Unsupervised clustering revealed significant differences in stroke incidence among the three identified risk clusters (9.1%, 6.6%, and 3.2%). These clusters were categorized into high-, medium-, and low-risk groups. Among the supervised models, the random forest algorithm demonstrated the best performance. The top ten most important variables for predicting stroke incidence were identified using the SHAP, with age being the most influential variable. Other significant risk markers included systolic blood pressure, hypertension, estimated glomerular filtration rate, metabolic syndrome, and blood sugar level. Additionally, elbow joint thickness and fructosamine, hemoglobin, and calcium levels were found to be potential predictors of stroke risk. Notably, the variables identified by the SHAP were consistent with those obtained from the unsupervised clustering approach in the high-risk group. Conclusion: Machine learning algorithms provide accurate predictions of stroke incidence and offer valuable insights into subclinical markers without the need for prior assumptions of causality. This study presents a data-driven machine-learning framework for stroke risk prediction and biomarker identification.
Stroke constitutes a significant public health concern due to its impact on mortality and morbidity. This study investigates the utility of machine learning algorithms in predicting stroke and identifying key risk factors using data from the Suita study, comprising 7,389 participants and 53 variables. Initially, unsupervised K-prototype clustering categorized participants into risk clusters, while five supervised models including Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), eXtreme Gradient Boosting (XGBoost), and Light Gradient Boosted Machine (Light-GBM) were employed to predict stroke outcomes. Stroke incidence disparities among identified risk clusters using the unsupervised K-prototype clustering method are substantial, according to the findings. Supervised learning, particularly RF was a preferable option because of the higher levels of performance metrics. The Shapley Additive Explanations (SHAP) method identified age, systolic blood pressure, hypertension, estimated glomerular filtration rate, metabolic syndrome, and blood glucose level as key predictors of stroke, aligning with findings from the unsupervised clustering approach in high-risk groups. Additionally, previously unidentified risk factors such as elbow joint thickness, fructosamine, hemoglobin, and calcium level demonstrate potential for stroke prediction. In conclusion, machine learning facilitated accurate stroke risk predictions and highlighted potential biomarkers, offering a data-driven framework for risk assessment and biomarker discovery.
BACKGROUND:Patients with severe dengue who develop severe respiratory failure requiring mechanical ventilation (MV) support have significantly increased mortality rates. This study aimed to develop a robust machine learning-based risk score to predict the need for MV in children with dengue shock syndrome (DSS) who developed acute respiratory failure. METHODS:This single-institution retrospective study was conducted at a tertiary pediatric hospital in Vietnam between 2013 and 2022. The primary outcome was severe respiratory failure requiring MV in the children with DSS. Key covariables were predetermined by the LASSO method, literature review, and clinical expertise, including age (< 5 years), female patients, early onset day of DSS (≤ day 4), large cumulative fluid infusion, higher colloid-to-crystalloid fluid infusion ratio, severe bleeding, severe transaminitis, low platelet counts (< 20 x 109/L), elevated hematocrit, and high vasoactive-inotropic score. These covariables were analyzed using supervised models, including Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), k-Nearest Neighbor (KNN), and eXtreme Gradient Boosting (XGBoost). Shapley Additive Explanations (SHAP) analysis was used to assess feature contribution. RESULTS:A total of 1,278 patients were included, with a median patient age of 8.1 years (IQR: 5.4-10.7). Among them, 170 patients (13.3%) with DSS required mechanical ventilation. A significantly higher fatality rate was observed in the MV group than that in the non-MV group (22.4% vs. 0.1%). The RF and SVM models showed the highest model discrimination. The SHAP model explained the significant predictors. Internal validation of the predictive model showed high consistency between the predicted and observed data, with a good slope calibration in training (test) sets 1.0 (0.934), and a low Brier score of 0.04. Complete-case analysis was used to construct the risk score. CONCLUSIONS:We developed a robust machine learning-based risk score to estimate the need for MV in hospitalized children with DSS.
Background and objectives: Paediatric dengue-associated acute liver failure (PALF) is a rare and fatal complication. To date, clinical data regarding the combination of therapeutic plasma exchange (TPE) and continuous renal replacement therapy (CRRT) for the treatment of dengue-associated PALF are limited.Methods: We conducted a single-center, retrospective study of all children with dengue-associated PALF admitted to the paediatric intensive care unit of Children Hospital No.2, Vietnam, who were treated with TPE+CRRT between January 2021 and March 2022. The main study outcomes were in-hospital survival, normalisation of hepatic function, and hepatic encephalopathy improvement.Results: Twelve patients aged from 06 to 12 years underwent TPE+CRRT procedures. Among them, three (25 %) patients died of severe sepsis and septic shock confirmed by Enterobacteriaceae spp. haemocultures (stable on maintenance treatment of COVID-19-associated MIS-C with low dose of oral steroids on hospital admission), acute respiratory distress syndrome (ARDS), and clinically apparent intracranial haemorrhage. Nine patients (75 %) survived. The paediatric mortality risk score improved significantly at discharge compared with PICU admission (P < 0.01). Markedly, all twelve patients were diagnosed with hepatoencephalopathy of grades III and IV on PICU admission. After the combined TPE+CRRT interventions, there were substantial improvements in liver transaminases levels, coagulation profiles, and metabolic biomarkers. Normal neurological functions were observed in nine alive patients at hospital discharge. Only one patient experienced an adverse event of slightly low blood pressure, which rapidly self-resolved.Interpretation and conclusions: Combined TPE+CRRT significantly improved survival outcome, neurological sta-tus, and rapid normalisation of liver functions in dengue-associated PALF.
OBJECTIVES: Pediatric acute liver failure (PALF) is a fatal complication in patients with severe dengue. To date, clinical data on the combination of therapeutic plasma exchange (TPE) and continuous renal replacement therapy (CRRT) for managing dengue-associated PALF concomitant with shock syndrome are limited. DESIGN: Retrospective cohort study (January 2013 to June 2022). PATIENTS: Thirty-four children. SETTING: PICU of tertiary Children’s Hospital No. 2 in Vietnam. INTERVENTIONS: We assessed a before-versus-after practice change at our center of using combined TPE and CRRT (2018 to 2022) versus CRRT alone (2013 to 2017) in managing children with dengue-associated acute liver failure and shock syndrome. Clinical and laboratory data were reviewed from PICU admission, before and 24 h after CRRT and TPE treatments. The main study outcomes were 28-day in-hospital mortality, hemodynamics, clinical hepatoencephalopathy, and liver function normalization. MEASUREMENTS AND MAIN RESULTS: A total of 34 children with a median age of 10 years (interquartile range: 7–11 yr) underwent standard-volume TPE and/or CRRT treatments. Combined TPE and CRRT ( n = 19), versus CRRT alone ( n = 15), was associated with lower proportion of mortality 7 of 19 (37%) versus 13 of 15 (87%), difference 50% (95% CI, 22–78; p < 0.01). Use of combined TPE and CRRT was associated with substantial advancements in clinical hepatoencephalopathy, liver transaminases, coagulation profiles, and blood lactate and ammonia levels (all p values < 0.001). CONCLUSIONS: In our experience of children with dengue-associated PALF and shock syndrome, combined use of TPE and CRRT, versus CRRT alone, is associated with better outcomes. Such combination intervention was associated with normalization of liver function, neurological status, and biochemistry. In our center we continue to use combined TPE and CRRT rather than CRRT alone.
Cardiovascular disease (CVD) is one of the primary causes of death around the world. This study aimed to identify risk factors associated with CVD mortality using data from the National Health and Nutrition Examination Survey (NHANES). We created three models focusing on dietary data, non-diet-related health data, and a combination of both. Machine learning (ML) models, particularly the random forest algorithm, demonstrated robust consistency across health, nutrition, and mixed categories in predicting death from CVD. Shapley additive explanation (SHAP) values showed age, systolic blood pressure, and several other health factors as crucial variables, while fiber, calcium, and vitamin E, among others, were significant nutritional variables. Our research emphasizes the importance of comprehensive health evaluation and dietary intake in predicting CVD mortality. The inclusion of nutrition variables improved the performance of our models, underscoring the utility of dietary intake in ML-based data analysis. Further investigation using large datasets with recurring dietary recalls is necessary to enhance the effectiveness and interpretability of such models.
BACKGROUND:Risk factors for atherosclerotic disease including dyslipidemia have been shown to be associated with aortic valve calcification (AVC). Nuclear magnetic resonance (NMR)-measured lipoprotein particles, low-density and high-density lipoprotein particles (LDL-p, HDL-p) in particular, have emerged as novel markers of atherosclerotic disease; however, whether NMR-measured particles are associated with AVC remains to be determined. This study aimed to examine the association between NMR-based lipoprotein particle measurements and standard lipids with AVC. The primary variables of interest were LDL-p (nmol/L), HDL-p (μmol/L), LDL-cholesterol, and HDL-cholesterol (both in mg/dL). METHODS AND RESULTS:A community-based random sample of Japanese men aged 40-79 years examined in 2006-2008, in Shiga, Japan was studied. Presence of AVC was defined as an Agatston score >0. Lipoprotein particles were measured using NMR spectroscopy. In the main analysis, multivariable-adjusted odds ratios (ORs) and 95% confidence intervals (95% CIs) for the prevalence of AVC across the higher quartiles of lipids in reference to the lowest ones were obtained. Of 874 participants analyzed, 153 men had AVC. Multivariable-adjusted ORs of prevalent AVC for the highest vs. the lowest quartile were significantly elevated for LDL-p (OR, 2.20; 95% CI: 1.23-3.93) and LDL-cholesterol (OR, 2.16; 95% CI: 1.23-3.78). In contrast, neither HDL-p nor HDL-cholesterol was associated with AVC. CONCLUSIONS:The association of prevalent AVC with NMR-based LDL-p was comparable to that with LDL-cholesterol.