Current hypertension guidelines focus on blood pressure control, but incorporating end-organ imaging could improve understanding of disease manifestations. We undertook a systematic review to evaluate current task-level applications of artificial intelligence (AI) and computational approaches to imaging for hypertension identification, phenotyping, and outcome prediction. A systematic search was conducted across multiple databases up to end of December 2025. Retrieved studies were grouped by AI task, and a thematic qualitative analysis per-task was conducted to evaluate organ-specific findings, AI methodologies, and research gaps. For quantitative synthesis, the I2 statistic derived from Cochran's Q test was used to assess heterogeneity, and forest plots were generated to visualize effect sizes. The review was registered with PROSPERO (CRD42023427430). The search strategy yielded 48 studies. Thematic analysis categorized the studies into five major tasks, with the majority employing supervised learning for classification processes. Nearly half of the studies focused on the heart. However, paucity of studies performed multi-organ assessment, external validation, and phenotyping or predicting future risk. AI and computational approaches in imaging achieved an overall sensitivity of 0.84 [0.69-0.93] in identifying hypertension from normotension, highest with brain imaging. Sensitivity reached 0.92 [0.90-0.94] in discriminating hypertension from hypertrophic cardiomyopathy. Current research focusses primarily on hypertension prediction using single organ information. While results are promising, datasets remain small with limited external validation. There remains a need for discovery-oriented research to uncover disease heterogeneity, multi-organ phenotypes, and support personalized and targeted interventions.
Heart failure affects over 64 million individuals globally, contributing to elevated mortality rates and substantial healthcare costs. This study investigates the potential of retinal optical coherence tomography features combined with routine clinical variables as biomarkers for the detection of heart failure, exploring a potential avenue for improved risk assessment and screening support using explainable machine-learning tools. A comprehensive dataset of normal and heart failure patients' demographic and medical records including retinal measurements from both eyes was used. Among the machine learning models employed, the Extreme Gradient Boosting model demonstrated the best performance, achieving an accuracy of 73.31%, a precision of 71.81%, and an area under the receiver operating characteristic curve of 0.837. Explainability analyses further revealed that macular thickness metrics, particularly in the inner temporal subfield, inner nasal subfield, and outer superior subfields of the left eye, along with key clinical indicators such as age, body mass index, and glycated hemoglobin, were the most influential predictors of heart failure status. Local explanation methods also provided patient-level reasoning consistent with overall cohort patterns. To our knowledge, this is the first study to use an integrated, explainable approach incorporating bilateral retinal optical coherence tomography measurements with routine clinical indicators for heart failure detection, providing an interpretable and accessible alternative to black-box models while helping address the cost, invasiveness, and limited accessibility of existing heart failure diagnostic tools.
Congenital heart disease (CHD) is the most common birth defect and a major cause of infant morbidity and mortality worldwide. While echocardiography remains the diagnostic gold standard, its cost and reliance on expert interpretation limit availability in low- and middle-income countries (LMICs). This underscores the need for scalable, affordable, and robust approaches that can operate reliably under real-world signal degradation and low-quality recording conditions. In this study, we present a multimodal framework for pediatric CHD detection that integrates phonocardiograms (PCG), electrocardiograms (ECG), and clinical symptoms. Features from PCG and ECG were extracted using a pretrained foundation audio model, while key symptoms were selected using SHAP (SHapley Additive exPlanations) and encoded as structured embeddings. These modalities were combined in a fusion network with modality dropout to handle missing inputs, and mutual learning was applied to promote knowledge sharing across unimodal and multimodal branches. The framework was validated on a dataset of 751 pediatric patients in Bangladesh, comprising 3,435 synchronized PCG-ECG recordings with expert-confirmed labels. In 10-fold patient-wise cross validation, PCG alone achieved 90% accuracy, ECG alone 79%, and symptoms alone 73%. Combining PCG and ECG improved accuracy to 94%, while the full multimodal system reached 94.5% accuracy, 94.2% sensitivity, 95.6% specificity, and an AUROC of 96%. Importantly, even under degraded signal conditions when both PCG and ECG were of low quality, the multimodal framework maintained an accuracy of 88% and an AUROC of 87%, demonstrating superior performance compared to the unimodal models. These findings demonstrate that integrating PCG, ECG, and symptoms enables accurate, resilient CHD screening, offering a practical pathway for scalable early detection in LMICs.
Supernumerary Robotic Fingers (SRFs) have proven their capabilities in enhancing manipulation and dexterity in precision tasks enabling multitasking in healthy individuals, while also playing a compensatory role for those with motor impairments. However, the literature on SRF adaptation and embodiment remains scarce, particularly in terms of the neurophysiological basis underlying their introduction, use, and assimilation. This study employed a controlled-time-delay-stability (CTDS) framework to investigate physiological embodiment by capturing multimodal signals (encephalography (EEG), electrodermal-activity (EDA), photoplethysmography (PPG), and respiratory) in 30 adults performing three activities-of-daily-living. Tasks were completed across three phases: without SRF, with SRF immediately after attachment, and with SRF following individualized training. CTDS quantified pairwise strength and directionality of dynamic interactions while controlling for indirect effects. Network visualization provided a holistic view of brain-body reconfiguration with SRF use. Resting baselines showed no significant differences across phases, suggesting SRF attachment alone did not introduce sufficient stress to alter brain-body interactions. During tasks, brain-body interactions exhibited task-dependent modulation. This modulation was suppressed with initial SRF attachment but reappeared after training/familiarization, with post-training patterns becoming statistically indistinguishable from those without the SRF. These results indicate that SRFs can be effectively embodied into the body schema after targeted training. The transient plasticity observed in first-time users may serve as a biomarker for customizing SRF training and tracking neurorehabilitation progress. Extending prior findings of rapid cortical adaptation, this study shows that brain-body networks reconfigure in a task-specific manner toward assimilation following brief familiarization. Future work warrants clinical validation and translation.
Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal wearable computing. However, existing approaches do not operate as full-stack wearable processors, i.e., they do not simultaneously address task-specific classification performance, disentangled and interpretable representation learning, fusion, and generative modeling of highly heterogeneous multi-modal time series. To address this gap, we introduce Omni-modal Variational Decomposition Autoencoders (OmniDecVAEs), a framework that efficiently learns multi-purpose representations in a unified and scalable manner from arbitrarily many modalities. OmniDecVAEs extend DecVAEs by learning modality-conditioned time-frequency latent subspaces through a multi-view self-supervised decomposition loss and a shared asymmetric autoencoder (AE) architecture. Results on a challenging omni-modal human activity recognition (HAR) setting with up to thirty modalities, demonstrate the ability of OmniDecVAEs to learn full-stack wearable representations. When compared to transformer-based and VAE-based methods, OmniDecVAEs full-stack disentangled representation properties lead to accuracy improvements of 1.01
BACKGROUND:Hypertension induces structural and functional damage in multiple organs. Evidence of subclinical damage increases risk of vascular events and death but can be difficult to identify in the clinic. We developed a novel machine learning approach that quantifies current hypertension-associated multiorgan damage, mapping progression from health to advanced disease, in a pseudotemporal manner and predicts organ-specific disease progression trajectories. METHODS:We analyzed 566 multimodal imaging and nonimaging variables from 27 099 participants in the UK Biobank imaging substudy to develop a semisupervised contrastive trajectory inference (cTI) framework that models multiorgan alterations associated with hypertension exposure, including heart, brain, kidneys, vasculature, lungs, liver, and metabolic information. Model stability was validated through multiple internal validation steps, and external validity was tested on 5507 participants from the Atherosclerosis Risk in Communities study (ARIC). Clinical relevance was evaluated against existing risk scores and through ability to predict survival and incident multiorgan disease for up to 7 years, across both UK Biobank and ARIC. RESULTS:In the UK Biobank (mean age 63.27±7.48 years; 53.4% women) our global organ damage score (HyperScore) achieved an area under the curve of 0.964 (0.941-0.987) for identification of individuals with severe end-organ disease and robust stability in cross-validation with a mean root mean square error of 0.104±0.084. Survival odds differed significantly across HyperScore stages (P<0.001), whereas stratification by blood pressure was nonsignificant. We further revealed 6 hypertensive disease phenotypes (HyperTrajectory), characterized by predominant cardiac, lipoprotein, atherothrombosis, brain, cardiorenal, and liver features, respectively. External testing in ARIC confirmed stability of the model, with Jensen-Shannon distances as low as 0.10 for HyperScore distributions, without significant deviation in organ damage progression patterns (P>0.05) and consistent end-organ and outcome characteristics between ARIC and UK Biobank across HyperTrajectories. CONCLUSIONS:Machine learning-derived global organ damage scores are feasible in hypertension and enable identification of distinct hypertension-associated organ-disease phenotypes. New frameworks for hypertension assessment and monitoring using imaging to derive personalized risk assessment and phenotype-specific intervention may be achievable.
Understanding the structure of complex, nonstationary, high-dimensional time-evolving signals is a central challenge in scientific data analysis. In many domains, such as speech and biomedical signal processing, the ability to learn disentangled and interpretable representations is critical for uncovering latent generative mechanisms. Traditional approaches to unsupervised representation learning, including variational autoencoders (VAEs), often struggle to capture the temporal and spectral diversity inherent in such data. Here we introduce variational decomposition autoencoding (VDA), a framework that extends VAEs by incorporating a strong structural bias toward signal decomposition. VDA is instantiated through variational decomposition autoencoders (DecVAEs), i.e., encoder-only neural networks that combine a signal decomposition model, a contrastive self-supervised task, and variational prior approximation to learn multiple latent subspaces aligned with time-frequency characteristics. We demonstrate the effectiveness of DecVAEs on simulated data and three publicly available scientific datasets, spanning speech recognition, dysarthria severity evaluation, and emotional speech classification. Our results demonstrate that DecVAEs surpass state-of-the-art VAE-based methods in terms of disentanglement quality, generalization across tasks, and the interpretability of latent encodings. These findings suggest that decomposition-aware architectures can serve as robust tools for extracting structured representations from dynamic signals, with potential applications in clinical diagnostics, human-computer interaction, and adaptive neurotechnologies.
Obstructive sleep apnoea (OSA) and major depressive disorder (MDD) frequently co-occur and both have been associated with autonomic dysregulation. Whether the co-occurrence alters the short-scale dynamical organisation of cardiac control during the cyclic alternating pattern (CAP) of NREM sleep is unknown. We retrospectively analysed overnight polysomnography from 44 adults with OSA, 18 with comorbid MDD diagnosed by psychiatrist-administered Mini-International Neuropsychiatric Interview (MINI) version 5 (OSAD+) and 26 without MDD (OSAD−) and from 10 healthy controls drawn from an institutional CAP repository. All MDD patients were antidepressant-naive at the time of PSG. After delta-band-based CAP A-phase detection on the C4–A1 EEG, we extracted 22 HRV indices spanning time, frequency (Lomb–Scargle) and nonlinear domains from RR intervals occurring inside CAP windows, and compared groups with two-tailed Mann–Whitney U tests, Benjamini–Hochberg FDR-corrected, supplemented by stepwise logistic regression and receiver-operating characteristic (ROC) analysis. The OSAD+ and OSAD− groups did not differ in age, sex, BMI, AHI or CAP rate. Sample entropy (SampEn) during CAP was reduced in OSAD− compared with healthy controls (0.87 ± 0.05 vs 1.05 ± 0.02, p = 0.004) and was not reduced in OSAD+ (1.00 ± 0.05; OSAD+ vs control p = 0.29). The OSAD+ vs OSAD− difference reached nominal significance (raw p = 0.043; AUC = 0.69) but did not survive correction for multiple comparisons across the 22 features (Benjamini–Hochberg-adjusted p = 0.83). A multivariable logistic-regression model retaining SampEn and high-frequency power (HF) achieved an AUC of approximately 0.77 and is reported as hypothesis-generating; no other linear or nonlinear HRV index distinguished the groups. In OSA, the short-scale irregularity of CAP-aligned RR intervals is reduced relative to healthy controls in the absence of MDD and is restored towards control values in the presence of comorbid MDD. The finding is exploratory and requires replication in larger cohorts with prospective insomnia screening (Insomnia Severity Index) and stratification by depression subtype, but it identifies CAP-aligned SampEn as a candidate physiological marker of the depressive phenotype within OSA and motivates a re-framing of “complexity loss” interpretations as context- and microstructure-dependent.
PurposeElectrooculography (EOG) provides a noninvasive measure of eye movements linked to affective processing, yet it is mainly used for artifact correction of electroencephalography (EEG) signals rather than analyzed as a physiological signal in its own right. EEG–EOG coupling has therefore not been well-established. This study aimed to determine whether emotion-specific changes in arousal and valence are reflected in directional and frequency-specific interactions between EEG rhythms and EOG signals.MethodsThe DEAP dataset with 32 participants, where each viewed 40 1-min music videos and rated their arousal/valence, was used (1,280 samples). EEG from eight electrodes was filtered into theta, alpha, beta, and gamma frequency bands, while horizontal and vertical EOG were also preprocessed. EOG complexity was assessed using sample, fuzzy, and permutation entropy. EEG–EOG coupling was assessed with the controlled time delay stability (CTDS) framework, which evaluates stability of partial cross-correlation delays.ResultsEntropy analysis showed emotion-related differences in horizontal and vertical EOG complexity (p < 0.005). EEG–EOG coupling varied with emotion, with the strongest effects at sensorimotor and frontal sites, primarily within the gamma band. Directional EOG-to-EEG coupling predominating at frontal, sensorimotor, and occipital sites. Differences were most pronounced when arousal and valence varied independently or in opposite directions, with fewer effects during parallel shifts.ConclusionEmotional states are mirrored by frequency- and channel-specific shifts in EEG–EOG interactions, a core component of the affective behavioral network. These results clarify the directional dynamics linking eye movement and cortical activity, revealing a structured, context-sensitive neural architecture for affective processing.
Congenital heart disease (CHD) is a critical condition that demands early detection, particularly in infancy and childhood. This study presents a deep learning model designed to detect CHD using phonocardiogram (PCG) signals, with a focus on its application in global health. We evaluated our model on several datasets, including the primary dataset from Bangladesh, achieving a high accuracy of 94.1 demonstrated robust performance on the public PhysioNet Challenge 2022 and 2016 datasets, underscoring its generalizability to diverse populations and data sources. We assessed the performance of the algorithm for single and multiple auscultation sites on the chest, demonstrating that the model maintains over 85 able to achieve an accuracy of 80 cardiologists deemed non-diagnostic. This research suggests that an AI- driven digital stethoscope could serve as a cost-effective screening tool for CHD in resource-limited settings, enhancing clinical decision support and ultimately improving patient outcomes.
Chronic Kidney Disease (CKD) is a progressive condition that requires accurate diagnosis and staging for effective clinical management. Conventional CKD diagnosis relies on estimated Glomerular Filtration Rate (eGFR), a measure of kidney function derived from serum biomarkers such as serum creatinine (SCr) and cystatin C (SCysC). However, eGFR calculations may be inaccurate when applied to diverse patient populations. This study proposes a machine learning (ML) system that integrates regression-based eGFR estimation, metaheuristic optimization using the Grey Wolf Optimizer (GWO), and multi-class classification with various ML models to enhance CKD staging and classification. The model estimates eGFR using three established CKD Epidemiology Collaboration (CKD-EPI) equations incorporating SCr, SCysC, and their combined values. Regression models assess predictive performance, specifically Linear Regression (LR) and Support Vector Regression (SVR). SVR demonstrates superior performance compared to LR for CKD-EPISCr-SCysC achieved a root mean squared error (RMSE) of 3.03, a mean absolute percentage error (MAPE) of 2.97%, and a coefficient of determination (R2) score of 0.97. The application of GWO for hyperparameter tuning has resulted in a 37.3% reduction in root mean square error (RMSE), a 37.4% drop in mean absolute percentage error (MAPE), and a 2.06% improvement in R2 to improve the precision of prediction. Once the model fine-tunes the eGFR estimations, it feeds them into various algorithms for CKD stage classification, including Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). Among these, XGBoost achieves the highest classification accuracy of 97.76%, along with an F1-score of 97.45%, demonstrating its effectiveness in CKD staging. Shapley Additive Explanations (SHAP) provide global and local feature importance insights, enhancing clinical decision-making and model transparency. Future research will validate the model using more extensive and more diverse datasets. Additionally, it will incorporate extra clinical parameters, including biomarkers and genetic data, to enhance the precision of CKD risk prediction. This research enhances AI-driven nephrology by providing a scalable, interpretable, and highly accurate solution for diagnosing and managing CKD.
Peers’ conversation provides a domain of rich emotional information. The latter, apart from facial and gestural expressions, it is also naturally conveyed via peers’ speech, contributing to the establishment of a dynamic emotion climate (EC) during their conversational interaction. Recognition of EC could provide an additional source in understating peers’ social interaction and behavior on top of peers’ actual conversational content. Here, we propose a novel approach for speech-based EC recognition, namely AffECt, by combining peers’ complex affect dynamics (AD) with deep features extracted from speech signals using Temporary Convolutional Neural Networks (TCNNs). AffECt was tested and cross-validated on data drawn from there open datasets, i.e., K-EmoCon, IEMOCAP, and SEWA, in terms of EC arousal/valence level classification. The experimental results have shown that AffECt achieves EC classification accuracy up to 83.3% and 80.2% for arousal and valence, respectively, clearly surpassing the results reported in the literature, exhibiting robust performance across different languages. Moreover, there is a distinct improvement when the AD are combined with the TCNN, compared to the baseline deep learning approaches. These results demonstrate the effectiveness of AffECt in speech-based EC recognition, paving the way for many applications, e.g., in patients’ group therapy, negotiations, and emotion-aware mobile applications
Vision plays a fundamental role in the control of human locomotion, including walking gait. Given that side-dominance is associated with differences in motor control, the present study aimed to determine if patches obscuring half of the visual field affect left- and right-side dominant individuals’ gait kinematics and accompanying leg muscle activation differently. Healthy right- (n = 15, age = 28.2 ± 5.5 years) and left-side (n = 9, age = 27.9 ± 5.8 years) dominant participants performed 10 min of walking trials on a treadmill at a self-selected speed with 5 min of rest between three randomized trials, i.e., wearing clear glasses or glasses with left-or right half-field eye patching. In addition to a set of spatiotemporal and kinematic gait parameters, the average activity during the separated gait cycle phases, and the start and end of muscle activation in % of the gait cycle were calculated from five muscles in three muscle groups. Our results indicate that gait kinematics of left- and right-side dominant participants were similar both in their dominant and non-dominant legs, regardless of half-field eye patching condition. On the other hand, inter-group differences were found in selected kinematic variables. For instance, in addition to larger but less variable step width, our results suggest larger ankle and knee ROM in right- vs. left-sided participants. Furthermore, medial gastrocnemius and biceps femoris muscle activation showed selected differences at certain phases of the gait cycle between participants’ dominant and non-dominant legs. However, it was also unaffected by the half-field eye patching condition. Moreover, the endpoint of medial gastrocnemius activation was affected by side-dominance, i.e., its activation ended earlier in the non-dominant leg of right- as compared to left-side dominant participants. Our results suggest no major differences in walking gait kinematics and accompanying muscle activation between half-field eye patching conditions in healthy adults; nevertheless, side-dominance may affect biomechanical and neuromuscular control strategies during walking gait.
Phase coherence (2) between respiratory sinus arrhythmia (RSA) and respiration has emerged as a promising metric for assessing the role of autonomic nervous system (ANS) activity and slow wave sleep (SWS) activities in healthy subjects. This study aims to investigate how 2 and SWS activity differ between Obstructive Sleep Apnea (OSA) patients with and without major depressive disorder (MDD) during overnight sleep and explore whether the correlation between 2 and SWS activity exists among those OSA patients compared to healthy individuals. Overnight electroencephalogram (EEG), electrocardiograms (ECG), and breathing using plethysmography were recorded from 104 subjects, including 35 healthy individuals (control), 34 OSA subjects with MDD (OSAD+) and 35 OSA subjects without MDD (OSAD+). Slow wave activity was computed by the amplitude envelope of the EEG delta-wave (0.5-4 Hz). The interbeat intervals (RRI) and respiratory movement were derived from ECG. RRI and respiration were resampled at a frequency of 10 Hz, and the band passed filtered within the range of 0.1-0.4 Hz before the Hilbert transform was used to extract instantaneous phases of the RSA and respiration. From the analytical signal of the Hilbert transform, the phase coherence (2) and amplitude of RSA (ARSA) were quantified. Additionally, the heart rate variability (HRV) features were calculated. Our results showed that overnight 2 was significantly greater, while the Low Frequency (LF) and High Frequency (HF) components of the HRV were significantly lower in OSAD+ compared to OSAD-. In addition, overnight delta-wave activity was greater in OSADcompared to both OSAD+ and control groups. Using auto- and cross-correlation analyses, we found that overnight profiles of 2 and delta-wave were correlated only in healthy individuals compared to OSAD+ and OSAD-, indicating that sleep apnea may only have an impact on this cortical-cardiorespiratory correlation rather than depression. Our findings suggest that 2 and SWS activity appear to be biomarkers for assessing depression in OSA patients, whereas their correlation pattern may serve as a marker for only OSA. This could enhance diagnostic precision and provide valuable insights into the complex physiological mechanisms underlying the corambid of OSA and MDD.
Obstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) is a prevalent sleep disorder characterized by recurrent episodes of obstructed breathing due to the relaxation of muscles in the upper airway during sleep, often linked with neuromuscular and cardiovascular disorders. This study introduces a novel method using traditional machine learning classifiers and surface electromyography (SEMG) features extracted from motor units (MUs) decomposed from chin electromyography (EMG) signals to screen for OSA events in OSAHS subjects. SEMG features were extracted from individual MUs decomposed from chin EMG segments using a novel dataset. An apnea detection algorithm was designed to label these events for OSAHS subjects across sleep stages. Analysis of motor neuron firing patterns in OSAHS subjects revealed lower activation during OSA events and higher activation during non-OSA segments. Additionally, we evaluated the proposed system on a publicly available dataset, achieving a maximum accuracy of 72% for OSAHS subjects in the midlife phase age group (40- 59 years) and 72.5% for subjects in the severe phase of OSAHS using Support Vector Machines (SVM). The random forest (RF) classifier demonstrated robust performance, achieving 97% accuracy, 93.2% sensitivity, 100% specificity, 100% precision, a 96.48% F1-score, and an area under the curve (AUC) of 0.996. This system facilitates early differentiation between OSA and non-OSA events, enabling timely intervention in the mild apnea phase to prevent progression to severe OSAHS. Moreover, it offers a convenient alternative to conventional polysomnography (PSG), enhancing diagnostic accessibility and clinical management.
Emotions result from complex, dynamic interactions between the brain and the body. While the neural basis of emotional processing is well established, increasing evidence highlights the importance of its coordinated activity with peripheral systems, such as facial or shoulder muscles, in shaping and reflecting emotional experiences. This study investigated cortico-muscular synchronization as a mechanism of brain-body communication during emotionally evocative stimuli. EEG and EMG signals from 32 participants in the DEAP dataset were analyzed as they viewed music videos designed to elicit different valence-arousal combinations. Directional interactions between cortical rhythms and the zygomaticus major (zEMG) and trapezius (tEMG) muscles were quantified using a network-based framework combining controlled time-delay stability (cTDS) and partial cross-correlation. Emotion-related modulation was observed in both EEG-EMG networks but was stronger and more widespread in EEG-tEMG networks, with significant differences in bottom-up θ, β, and γ influences and localized top-down effects (p < 0.05). Most significant differences between emotion states involved the tEMG→ γ link at occipital and temporal sites (p = 0.008 and p = 0.007, respectively), and tEMG→ θ links at central sites (p = 0.003). EEG-zEMG connectivity showed comparable link strengths to those of EEG-tEMG but only minor emotion-dependent changes, with three significant effects (p < 0.05): γ →zEMG at C4 and bottom-up zEMG → θ and zEMG → α at C3. These findings suggest that affective modulation selectively reorganizes cortico-muscular networks, with more pronounced effects in postural muscles compared to expressive muscles. This agrees with neuropsychological findings of noticeable muscle tension in the trapezius muscle when stressed.
Accurate sleep stage scoring is vital for diagnosing a wide range of sleep disorders. However, traditional polysomnography (PSG) methods are time-consuming, expensive, and require complex setups, making automated sleep stage scoring an important advancement in sleep medicine. This study presents a framework for automating sleep stage scoring using surface electromyography (sEMG) features extracted from chin electromyography (chin EMG) muscle activity and traditional machine learning (ML) models. The analysis is based on data from the Sleep-EDFx database. Three sleep staging frameworks-4-stage, 5-stage, and 6-stage models-were evaluated, with the 6-stage model achieving the highest performance using the Random Forest (RF) model. A maximum testing accuracy of 82 % was obtained. F1-scores for the individual stages were 80.64 % for Awake, 84.87 % for REM, 95.94 % for N4, 92.5 % for N3, 62.9 % for N2, and 82.4 % for N1, showing notable performance even for the challenging N1 stage, which represents the sleep transition state. The physiological factors influencing model performance, particularly the difficulty in predicting the N2 stage, are explored in detail. Additionally, the proposed models were validated using the ISRUC-Sleep dataset, where a Support Vector Machine (SVM) model achieved a maximum accuracy of 70.5 % for control subjects aged 40-59 (middle adulthood) and 74.1 % for sleep-disordered subjects classified with normal weight based on Body Mass Index (BMI). These results demonstrate the framework's generalizability across datasets and varying subject demographics. Overall, this study marks a significant advancement in automated sleep scoring by leveraging chin EMG signals for accurate and efficient classification, offering the potential to streamline the diagnosis and management of sleep-related disorders in clinical settings.