Background: Coronary heart disease (CHD) is a major cause of morbidity and mortality worldwide. Identifying key risk factors is essential for effective risk assessment and prevention. A data-driven approach using machine learning (ML) offers advanced techniques to analyze complex, nonlinear, and high-dimensional datasets, uncovering novel predictors of CHD that go beyond the limitations of traditional models, which rely on predefined variables. Objective: This study aims to evaluate the contribution of various risk factors to CHD, focusing on both established and novel markers using ML techniques. Methods: The study recruited 7672 participants aged 30-84 years from Suita City, Japan, between 1989 and 1999. Over an average of 15 years, participants were monitored for cardiovascular events. A total of 7260 participants and 28 variables were included in the analysis after excluding individuals with missing outcome data and eliminating unnecessary variables. Five ML models-logistic regression, random forest (RF), support vector machine, Extreme Gradient Boosting, and Light Gradient-Boosting Machine-were applied for predicting CHD incidence. Model performance was evaluated using accuracy, sensitivity, specificity, precision, area under the curve, F1-score, calibration curves, observed-to-expected ratios, and decision curve analysis. Additionally, Shapley Additive Explanations (SHAPs) were used to interpret the prediction models and understand the contribution of various risk factors to CHD. Results: Among 7260 participants, 305 (4.2%) were diagnosed with CHD. The RF model demonstrated the highest performance, with an accuracy of 0.73 (95% CI 0.64-0.80), sensitivity of 0.74 (95% CI 0.62-0.84), specificity of 0.72 (95% CI 0.61-0.83), and an area under the curve of 0.73 (95% CI 0.65-0.80). RF also showed excellent calibration, with predicted probabilities closely aligning with observed outcomes, and provided substantial net benefit across a range of risk thresholds, as demonstrated by decision curve analysis. SHAP analysis elucidated key predictors of CHD, including the intima-media thickness (IMT_cMax) of the common carotid artery, blood pressure, lipid profiles (non-high-density lipoprotein cholesterol, high-density lipoprotein cholesterol, and triglycerides), and estimated glomerular filtration rate. Novel risk factors identified as significant contributors to CHD risk included lower calcium levels, elevated white blood cell counts, and body fat percentage. Furthermore, a protective effect was observed in women, suggesting the potential necessity for gender-specific risk assessment strategies in future cardiovascular health evaluations. Conclusions: We developed a model to predict CHD using ML and applied SHAP methods for interpretation. This approach highlights the multifactor nature of CHD risk evaluation, aiming to support health care professionals in identifying risk factors and formulating effective prevention strategies.
Recent research has highlighted the substantial impact of gut microbiome on various aspects of human health, such as obesity, inflammation, infectious diseases, and cancer. As a result, gut microbiota composition is increasingly recognized as a potential health indicator and biomarker for disease. Numerous factors, including lifestyle, diet, and physical fitness, are known to shape the composition of the human microbiome. However, a significant challenge in elucidating the relationships between these factors and the gut microbiome lies in needing a comprehensive database that integrates diverse human microbiome profiles with extensive sample metadata. To address this issue, we developed an extensive human microbiome database for healthy individuals. This initiative led to the establishment of the NIBN Japan Microbiome Database (NIBN JMD), one of the largest resources of its kind, encompassing up to 1,000 metadata points and more than 2,000 microbiome samples, including data from longitudinal studies. In this article, we describe the creation and features of NIBN JMD, detailing the data collection, processing, and database implementation. NIBN JMD is publicly accessible at https://jmd.nibn.go.jp/ .
Coronary heart disease (CHD) is a major cause of morbidity and mortality worldwide. Identifying key risk factors is essential for effective risk assessment and prevention. Machine learning (ML) offers advanced methods for analyzing complex datasets, revealing novel predictors of CHD beyond traditional models. This study aims to evaluate the contribution of various risk factors to CHD, focusing on both established and novel markers using machine learning techniques. The study recruited 7,672 participants aged 30 to 84 years from Suita City, Japan, between 1989 and 1999. Over an average of 15 years, participants were monitored for cardiovascular events. Five ML models—Random Forest (RF), XGBoost, Support Vector Machine (SVM), Logistic Regression (LR), and LightGBM—were used. The optimal model was identified based on accuracy, sensitivity, specificity, and AUC. SHapley Additive exPlanations (SHAP) were then employed to explore the contribution of various risk factors to CHD. RF achieved the highest AUC (95% CI) of 0.94 (0.93-0.96), outperforming LR, SVM, XGBoost, and LightGBM. SHAP on the best model identified the top CHD predictors. Intima-media thickness of common carotid artery (IMT_cMax) was identified as the strongest predictor of CHD, highlighting the importance of arterial health. Systolic and diastolic blood pressure, along with lipid profiles (non-HDL cholesterol, HDL cholesterol, and triglycerides), were closely associated with CHD incidence. eGFR underscored the link between renal function and CHD. Novel insights included the impact of lower calcium levels, systemic inflammation (elevated WBC counts), fructosamine levels, and obesity-related factor (body fat percentage). A protective effect in females indicated the need for sex-specific CHD management strategies. ML, particularly the RF model combined with SHAP, effectively identified key risk factors for CHD, including arterial health, blood pressure, lipid profiles, renal function, and novel markers. These findings support a multifactorial approach to CHD risk assessment.
We leveraged machine learning (ML) techniques, namely logistic regression (LR), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), and LightGBM to predict coronary heart disease (CHD) and identify the key risk factors involved. Based on the Suita study, 7672 men and women aged 30 to 84 years without cardiovascular disease were recruited from 1989 to 1999, in Suita City, Osaka, Japan. Over an average period of 15 years, participants were diligently monitored until the onset of their initial cardiovascular event or relocation. CHD diagnoses encompassed primary heart attacks, sudden death, or coronary artery disease with bypass surgery or intervention. RF achieved the highest AUC (95% CI) of 0.79 (0.70–0.87), outperforming LR, SVM, XGBoost, and LightGBM. Shapley Additive Explanations (SHAP) on the best model identified the top CHD predictors. Notably, systolic blood pressure, non-HDL-c, glucose levels, age, metabolic syndrome, HDL-c, estimated glomerular filtration rate, hypertension, elbow joint thickness, and diastolic blood pressure were key contributors. Remarkably, elbow joint thickness was identified as a previously unrecognized risk factor associated with CHD. These findings indicated that ML methods accurately predict incident CHD risk. Additionally, ML has identified new incident CHD risk variables.
Background: This population-based study investigated the potential of machine learning algorithms to predict stroke incidence and identify important risk factors. This study aimed to evaluate the accuracy of these algorithms in constructing a stroke prediction model. Methods: Participants from the Suita study were included, and baseline measurements were used to predict stroke outcomes over a 15-year follow-up period. In total, 7,389 participants and 51 variables were investigated, including demographics, medical history, medical imaging, laboratory data, and lifestyle habits. Initially, unsupervised K-prototype clustering was used to group participants based on their stroke risk. Subsequently, five supervised models (logistic regression, random forest, support vector machine, extreme gradient boosting, and light gradient boosted machine) were applied to predict the stroke outcomes. The Shapley Additive Explanations (SHAP) method determined the most critical variables. Results: Unsupervised clustering revealed significant differences in stroke incidence among the three identified risk clusters (9.1%, 6.6%, and 3.2%). These clusters were categorized into high-, medium-, and low-risk groups. Among the supervised models, the random forest algorithm demonstrated the best performance. The top ten most important variables for predicting stroke incidence were identified using the SHAP, with age being the most influential variable. Other significant risk markers included systolic blood pressure, hypertension, estimated glomerular filtration rate, metabolic syndrome, and blood sugar level. Additionally, elbow joint thickness and fructosamine, hemoglobin, and calcium levels were found to be potential predictors of stroke risk. Notably, the variables identified by the SHAP were consistent with those obtained from the unsupervised clustering approach in the high-risk group. Conclusion: Machine learning algorithms provide accurate predictions of stroke incidence and offer valuable insights into subclinical markers without the need for prior assumptions of causality. This study presents a data-driven machine-learning framework for stroke risk prediction and biomarker identification.
Stroke constitutes a significant public health concern due to its impact on mortality and morbidity. This study investigates the utility of machine learning algorithms in predicting stroke and identifying key risk factors using data from the Suita study, comprising 7,389 participants and 53 variables. Initially, unsupervised K-prototype clustering categorized participants into risk clusters, while five supervised models including Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), eXtreme Gradient Boosting (XGBoost), and Light Gradient Boosted Machine (Light-GBM) were employed to predict stroke outcomes. Stroke incidence disparities among identified risk clusters using the unsupervised K-prototype clustering method are substantial, according to the findings. Supervised learning, particularly RF was a preferable option because of the higher levels of performance metrics. The Shapley Additive Explanations (SHAP) method identified age, systolic blood pressure, hypertension, estimated glomerular filtration rate, metabolic syndrome, and blood glucose level as key predictors of stroke, aligning with findings from the unsupervised clustering approach in high-risk groups. Additionally, previously unidentified risk factors such as elbow joint thickness, fructosamine, hemoglobin, and calcium level demonstrate potential for stroke prediction. In conclusion, machine learning facilitated accurate stroke risk predictions and highlighted potential biomarkers, offering a data-driven framework for risk assessment and biomarker discovery.
A cross-sectional study involving 224 healthy Japanese adult females explored the relationship between ramen intake, gut microbiota diversity, and blood biochemistry. Using a stepwise regression model, ramen intake was inversely associated with gut microbiome alpha diversity after adjusting for related factors, including diets, Age, BMI, and stool habits (β = −0.018; r = −0.15 for Shannon index). The intake group of ramen was inversely associated with dietary nutrients and dietary fiber compared with the no-intake group of ramen. Sugar intake, Dorea as a short-chain fatty acid (SCFA)-producing gut microbiota, and γ-glutamyl transferase as a liver function marker were directly associated with ramen intake after adjustment for related factors including diets, gut microbiota, and blood chemistry using a stepwise logistic regression model, whereas Dorea is inconsistently less abundant in the ramen group. In conclusion, the increased ramen was associated with decreased gut bacterial diversity accompanying a perturbation of Dorea through the dietary nutrients, gut microbiota, and blood chemistry, while the methodological limitations existed in a cross-sectional study. People with frequent ramen eating habits need to take measures to consume various nutrients to maintain and improve their health, and dietary management can be applied to the dietary feature in ramen consumption.
Hookah, or waterpipe, is a tobacco smoking device that has gained popularity in the United States. A growing body of evidence demonstrates that waterpipe smoke (WPS) is associated with various adverse effects on human health, including infectious diseases, cancer, and cardiovascular diseases (CVDs), particularly thrombotic events. However, the molecular mechanisms through which WPS contributes to disease development remain unclear. In this study, we utilized an analytical approach based on the Comparative Toxicogenomics Database (CTD) to integrate chemical, gene, phenotype, and disease data to predict potential molecular mechanisms underlying the effects of WPS, based on its chemical and toxicant profile. Our analysis revealed that CVDs were among the top disease categories with regard to the number of curated interactions with WPS chemicals. We identified 5674 genes common between those modulated by WPS chemicals and traditional tobacco smoking. The CVDs with the most curated interactions with WPS chemicals were hypertension, atherosclerosis, and myocardial infarction, whereas "particulate matter", "heavy metals", and "nicotine" showed the highest number of curated interactions with CVDs. Our analysis predicted that the potential mechanisms un-derlying WPS-induced thrombotic diseases involve common phenotypes, such as inflammation, apoptosis, and cell proliferation, which are shared across all thrombotic diseases and the three aforementioned chemicals. In terms of enriched signaling pathways, we identified several, including chemokine and MAPK signaling, with particulate matter exhibiting the most statistically significant association with all 12 significant signaling pathways related to WPS chemicals. Collectively, our predictive comprehensive analysis provides evidence that WPS negatively impacts health and offers insights into the potential mechanisms through which it exerts these effects. This information should guide further research to explore and better understand the WPS and other tobacco product-related health consequences.
Introduction The prevalence of frailty is on the rise with the aging population and increasing life expectancy, which often is accompanied by comorbidities. Frailty can be effectively detected using Frailty index such as KCL index. Early detection of frailty allows applying measures that reduce the conversion rate to frail, and improve the quality of life in the frail people. Therefore, to facilitate the screening of frailty status at the primary care level, we suggest to produce a shorter version of the KCL questionnaire. Aim To understand the importance of KCL components in the decision making process for frailty and use machine learning approach to shorten the Questionnaire while maintaining reasonable accuracy, making it easier to screen for frailty in primary care. Methods We developed an automated framework of three steps: Feature importance determination using Shap values, testing models with Cross-validation with increased addition of selected features. Moreover, we validated the reliability of KCL to detect frailty by comparing the results of KCL criteria with the unsupervised clustering of the data. Results Our approach allowed us to identify the most important questions in the KCL questionnaire and demonstrate its performance using a short version with only four questions (4) Do you visit homes of friends?, (6) Are you able to go upstairs without using handrails or the wall for support? (10) Do you feel anxious about falling when you walk?, and (25) (In the past two weeks) Have you felt exhausted for no apparent reason?). We also showed that the data clustering corresponds well with the results of KCL criteria. Discussion and Conclusion While it is difficult to predict pre-frail status using shorter KCL questionnaire, it was shown to be fairly accurate in predicting frail status using only four questions.
Additional file 2. Table S1: WGCNA eigengenes.
The gut microbiota is closely related to good health; thus, there have been extensive efforts dedicated to improving health by controlling the gut microbial environment. Probiotics and prebiotics are being developed to support a healthier intestinal environment. However, much work remains to be performed to provide effective solutions to overcome individual differences in the gut microbial community. This study examined the importance of nutrients, other than dietary fiber, on the survival of gut bacteria in high-health-conscious populations. We found that vitamin B1, which is an essential nutrient for humans, had a significant effect on the survival and competition of bacteria in the symbiotic gut microbiota. In particular, sufficient dietary vitamin B1 intake affects the relative abundance of Ruminococcaceae, and these bacteria have proven to require dietary vitamin B1 because they lack the de novo vitamin B1 synthetic pathway. Moreover, we demonstrated that vitamin B1 is involved in the production of butyrate, along with the amount of acetate in the intestinal environment. We established the causality of possible associations and obtained mechanical insight, through in vivo murine experiments and in silico pathway analyses. These findings serve as a reference to support the development of methods to establish optimal intestinal environment conditions for healthy lifestyles.
Corona virus disease 2019 (COVID-19) increases the risk of cardiovascular occlusive/thrombotic events and is linked to poor outcomes. The underlying pathophysiological processes are complex, and remain poorly understood. To this end, platelets play important roles in regulating the cardiovascular system, including via contributions to coagulation and inflammation. There is ample evidence that circulating platelets are activated in COVID-19 patients, which is a primary driver of the observed thrombotic outcome. However, the comprehensive molecular basis of platelet activation in COVID-19 disease remains elusive, which warrants more investigation. Hence, we employed gene co-expression network analysis combined with pathways enrichment analysis to further investigate the aforementioned issues. Our study revealed three important gene clusters/modules that were closely related to COVID-19. These cluster of genes successfully identify COVID-19 cases, relative to healthy in a separate validation data set using machine learning, thereby validating our findings. Furthermore, enrichment analysis showed that these three modules were mostly related to platelet metabolism, protein translation, mitochondrial activity, and oxidative phosphorylation, as well as regulation of megakaryocyte differentiation, and apoptosis, suggesting a hyperactivation status of platelets in COVID-19. We identified the three hub genes from each of three key modules according to their intramodular connectivity value ranking, namely: COPE, CDC37, CAPNS1, AURKAIP1, LAMTOR2, GABARAP MT-ND1, MT-ND5, and MTRNR2L12. Collectively, our results offer a new and interesting insight into platelet involvement in COVID-19 disease at the molecular level, which might aid in defining new targets for treatment of COVID-19–induced thrombosis.
The gut microbiome is an important determinant in various diseases. Here we perform a cross-sectional study of Japanese adults and identify the Blautia genus, especially B. wexlerae , as a commensal bacterium that is inversely correlated with obesity and type 2 diabetes mellitus. Oral administration of B. wexlerae to mice induce metabolic changes and anti-inflammatory effects that decrease both high-fat diet–induced obesity and diabetes. The beneficial effects of B. wexlerae are correlated with unique amino-acid metabolism to produce S-adenosylmethionine, acetylcholine, and l -ornithine and carbohydrate metabolism resulting in the accumulation of amylopectin and production of succinate, lactate, and acetate, with simultaneous modification of the gut bacterial composition. These findings reveal unique regulatory pathways of host and microbial metabolism that may provide novel strategies in preventive and therapeutic approaches for metabolic disorders.
Non-small cell lung cancer (NSCLC) is the most prevalent form of lung cancer and a leading cause of cancer-related deaths worldwide. Using an integrative approach, we analyzed a publicly available merged NSCLC transcriptome dataset using machine learning, protein-protein interaction (PPI) networks and bayesian modeling to pinpoint key cellular factors and pathways likely to be involved with the onset and progression of NSCLC. First, we generated multiple prediction models using various machine learning classifiers to classify NSCLC and healthy cohorts. Our models achieved prediction accuracies ranging from 0.83 to 1.0, with XGBoost emerging as the best performer. Next, using functional enrichment analysis (and gene co-expression network analysis with WGCNA) of the machine learning feature-selected genes, we determined that genes involved in Rho GTPase signaling that modulate actin stability and cytoskeleton were likely to be crucial in NSCLC. We further assembled a PPI network for the feature-selected genes that was partitioned using Markov clustering to detect protein complexes functionally relevant to NSCLC. Finally, we modeled the perturbations in RhoGDI signaling using a bayesian network; our simulations suggest that aberrations in ARHGEF19 and/or RAC2 gene activities contributed to impaired MAPK signaling and disrupted actin and cytoskeleton organization and were arguably key contributors to the onset of tumorigenesis in NSCLC. We hypothesize that targeted measures to restore aberrant ARHGEF19 and/or RAC2 functions could conceivably rescue the cancerous phenotype in NSCLC. Our findings offer promising avenues for early predictive biomarker discovery, targeted therapeutic intervention and improved clinical outcomes in NSCLC.
Introduction Optimizing a protocol for 16S microbiome data analysis with QIIME2 is challenging for biologists with no programming experience. It involves a multi-step process and multiple parameters and options that must be tested and determined. Usually, researchers need to investigate various options, which requires running the analysis several times and comparing the results to decide the best set of tools and parameters for the data under investigation. Such an optimization process requires a substantial effort. In addition, it leads to several copies of the results with different sets of parameters, making the process relatively inefficient and difficult to reproduce. Results By combining the analysis strengths of QIIME2 with the flexibility in the definition of pipelines provided by Snakemake, here we introduce "Snaq," a snakemake pipeline that helps automate and optimize 16S data analysis using QIIME2. Snaq incorporates the definition of analysis rules with the intention of following a descriptive file name scheme, providing the functionality required to achieve faster protocol optimization, full pipeline automation, and handling data accumulation. Discussion Snaq offers a descriptive file naming system and automatically analyzes a data set by downloading and installing the required databases and classifiers, all through a single command-line instruction. It works natively on Linux, Mac. It works on Windows through containers and is potentially extendable by adding new rules. To make the pipeline versatile and easily adjustable, we adopted a convention of including all the critical parameter values inside a target file name and called this scheme descriptive target file naming. At the same time, other parameters are left as default. This means that Snakemake will parse the target file name and infer the sequence of steps and the parameter values needed. Then the target file will be created accordingly. Conclusion This pipeline will substantially reduce the efforts in sending commands and prevent the confusion caused by the accumulation of analysis results due to testing multiple parameters. Moreover, adding new rules can extend Snaq according to the users' needs. Github repository https://github.com/attayeb/snaq
Optimizing and automating a protocol for 16S microbiome data analysis with QIIME2 is a challenging task. It involves a multi-step process, and multiple parameters and options that need to be tested and determined. In this article, we describe Snaq, a snakemake pipeline that helps automate and optimize 16S data analysis using QIIME2. Snaq offers an informative file naming system and automatically performs the analysis of a data set by downloading and installing the required databases and classifiers, all through a single command-line instruction. It works natively on Linux and Mac and on Windows through the use of containers, and is potentially extendable by adding new rules. This pipeline will substantially reduce the efforts in sending commands and prevent the confusion caused by the accumulation of analysis results due to testing multiple parameters.
The gut microbiome is an important determinant in various diseases. Here we performed a cross-sectional study of Japanese adults and identified the Blautia genus, especially B. wexlerae , as a commensal bacterium that is inversely correlated with obesity and type 2 diabetes mellitus. Oral administration of B. wexlerae to mice induced metabolic changes and anti-inflammatory effects that decreased both high-fat diet–induced obesity and diabetes. The beneficial effects of B. wexlerae were mediated directly by unique amino-acid metabolism to produce S-adenosylmethionine, acetylcholine, and l-ornithine and indirectly by carbohydrate metabolism resulting in the accumulation of amylopectin and production of succinate, lactate, and acetate and simultaneous modification of the gut bacterial composition. These findings reveal unique regulatory pathways of host and microbial metabolism that may provide novel strategies in preventive and therapeutic approaches for metabolic disorders.
Dietary plant lignans are converted inside the gut to enterolignans enterodiol (ED) and enterolactone (EL), which have several biological functions, and health benefits. In this study, we characterized the gut microbiome composition associated with enterolignan production using data from a cross-sectional study in the Japanese population. We identified enterolignan producers by measuring ED and EL levels in subject’s serum using liquid chromatography-tandem mass spectrometry. Enterolignan producers show more abundant proportion of Ruminococcaceae and Lachnospiraceae than non-enterolignan producers. In particular, subjects with EL in their serum had a highly diverse gut microbiome that was rich in Ruminococcaceae and Rikenellaceae. Moreover, we built a random forest classification model to classify subjects to either EL producers or not using three characteristic bacteria. In conclusion, our analysis revealed the composition of gut microbiome that is associated with lignan metabolism. We also confirmed that it can be used to classify the microbiome ability to metabolize lignan using machine learning approach.
Remote health monitoring has become a necessity due to reduced healthcare access resulting from pandemic lockdowns and the increasing aging population. Electrocardiography (ECG) is the standard for cardiac monitoring and arrhythmia identification, but it is inconvenient for long-time remote monitoring. Recently, Magnetocardiography (MCG) sensors that operate at room temperature became available based on spintronic sensors. However, MCG analysis is affected by the low-frequency noise present at the sensors. In this paper, we present an artificial intelligence (AI)-aided multi-model pipeline combining two AI architectures, defined as model-M1 and model-M2, targeted for ultra-edge Internet of Things (IoT) sensors to simulate arrhythmia detection. Model-M1 is a denoising preprocessor based on a sliding-window assisted deep-learning (DL) model. We investigate various methods to achieve high accuracy with lightweight computation. Model-M2 is a lightweight DL model that analyzes denoised ECG output from model-M1 to identify arrhythmia. We use multiple publicly available clinically annotated datasets to evaluate our proposal. We find that denoising by model-M1 retains the features, which assist the model-M2 in achieving high classification accuracy, compared to using a conventional moving average filter. This AI pipeline architecture is promising for privacy-preserving ultra-edge medical sensing devices.