Reliable interpretation of electrocardiograms (ECGs) requires precise identification of P, QRS, and T (PQRST) wave boundaries. However, it remains challenging due to noise, signal quality variability, and inherent morphological diversity particularly in recordings from children. This study systematically compares the performance of leading deep neural networks (DNN) and heuristic-based delineation algorithms on ambulatory single-lead ECG signals focusing on temporal accuracy. Experiments were conducted using the publicly available LUDB dataset and a private validation dataset comprising 21,759 annotated single-lead wave segments from 611 children recorded using KardiaMobile ECG sensor. DNN were first trained on the LUDB dataset and subsequently tested on the validation dataset. The delineation performance was assessed using Sensitivity (Se) and positive-predictive-value (P+) metrics. The best-performing heuristic based and DNN models reached Se and P+ of (98.9% vs 97.9%) for P, (99.8% vs 99.2%) for QRS, and (98.7% vs 95.9%) for T wave fiducials, respectively. The lowest standard-deviation (in ms) of wave onset/offset delineation was achieved by attention based 1DU-Net model; ±16.6/ ±16.3 for P-wave, ±14.0/ ±16.3 for QRS, and ±26.3/ ±18.8 for T-wave, respectively. The findings indicate that optimized heuristic models can perform comparably to complex DNN, highlighting their efficiency and suitability for real-time ECG delineation in digital health monitoring applications. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement The study is funded by KU Leuven, center for affordable healthcare with funding reference REF23123123. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study was conducted according to the guidelines of the Declaration of Helsinki and approved by the Ethics Committees of the UZ Leuven with reference No. B3222022001075, Soddo christian hospital No. SCH1941015. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
Automated food intake gesture detection is a critical task in the domain of dietary monitoring, with the potential to objectively and continuously record individuals' eating patterns, thereby improving the quality of life (QoL) for diverse populations. Wrist-worn inertial measurement units (IMUs), have been extensively explored for this task and have demonstrated promising performance. Recently, ambient-based contactless radar sensor, has also shown feasibility for intake gesture detection. This study aims to investigate whether the complementary features of wearable and contactless sensors can be effectively leveraged through multimodal learning to further enhance detection performance. Additionally, this study addresses a key challenge in multimodal learning, the reduced robustness when handling missing modalities. To this end, we propose a robust multimodal temporal convolutional network with cross-modal attention (MM-TCN-CMA) framework, designed to efficiently integrate information from IMU and radar sensors, improve detection performance, and maintain effectiveness when modality-incomplete data are encountered during the inference phase. A dataset comprising 52 meal sessions (3,050 eating gestures and 797 drinking gestures) collected from 52 participants is developed and made publicly available for validation. Experimental results demonstrate that the proposed fusion framework achieves a segmental F1-score improvement of 4.3% and 5.2% over unimodal-Radar and unimodal-IMU baselines, respectively. Under missing modality conditions, the framework still yields performance gains of 1.3% and 2.4% in the missing-Radar and missing-IMU scenarios, respectively. To the best of our knowledge, this is the first study to explore a robust multimodal learning framework combining IMU and radar. The proposed robust radar-IMU fusion framework holds potential for broader applications in other continuous, fine-grained human activity recognition (HAR) tasks.
Early detection of Rheumatic Heart Disease (RHD) is essential in reducing its associated mortality and late complications. In resource-limited settings, automated detection using low-cost electrocardiogram (ECG) sensors can enhance prevention efforts. However, its effectiveness as a potential RHD screening tool in at-risk populations remains unexplored. This study aimed to investigate the utility of machine learning for classifying RHD in a cohort screened for RHD using low-cost ECG devices. The ECGs were collected from 611 at-risk schoolchildren using KardiaMobile, where 47 were confirmed RHD and 564 were healthy. First, the ECG fiducial points were annotated using a publicly available prominence-based delineator. Then, temporal, frequency, wavelet, and visibility graph-based features were extracted from six-leads and fed to the XGBoost classifier. A 10-fold cross-validation was used at different prediction score thresholds to obtain target sensitivity (Se) for screening RHD. Single-lead evaluation on Lead-II showed an F1-score of 60.9%, a Se of 59.6% and a positive-predictive-value (PPV) of 62.2%. However, using multiple leads improved the results, with an F1-score of 62.8%, a Se of 59.6% and a PPV of 66.7%. The best model performance was achieved by adjusting the threshold to 0.6 with Se and PPV of 66% and 51%, respectively. Error analysis revealed that T-wave and STT changes, as well as non-rheumatic mitral valve cases were among the false positive cases. Machine learning can enhance early detection by leveraging relevant ECG features and adjustable target sensitivity based on screening priorities and resource capacity. Measurements can be obtained without chest contact, using only the fingers and knees, thereby enabling use by non-clinical staff. This approach provides a scalable and cost-effective solution for RHD screening in high-prevalence regions. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research was funded by KU Leuven with reference number B/22/032, and Group-T 5E fund with reference number REF23123123 under Leuven center for affordable health technology. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study was conducted according to the guidelines of the Declaration of Helsinki and approved by the Ethics Committee of the University hospital Leuven with reference number (B3222022001075) and the institutional review board of Soddo General Christian Hospital with reference number (SCH1941015). Clinical trial number is not applicable. Written informed consent was obtained from the patient(s) to publish the study results. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using synchronized egocentric video, wearable IMUs, and expert-annotated FOG labels collected from 13 PD participants in their homes, we evaluate frozen representations from pretrained ego-video and time-series foundation models, alongside an IMU-based TCN trained from scratch, under leave-one-subject-out evaluation. The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Although ego-video alone did not outperform IMU-based sensing, it showed above-chance discrimination, and qualitative analyses suggest that egocentric vision may capture FOG-relevant information independent of IMUs. Together, these results support the use of pretrained ego-video representations to add contextual information to wearable-sensor-based clinical motion understanding in daily living.
Regular physical activity preserves functional independence in older adults, yet care-home residents often miss out because personalized supervision is scarce. Autonomous, technology-supported exercise platforms could deliver such guidance without additional staff time-but only if sessions are automatically monitored for safety and quality. We therefore designed a deep learning (DL) system that (a) recognizes individual exercise types and (b) estimates joint angle trajectories from a standard video recording. These outputs are used to compute objective exercise performance metrics (EPMs) such as duration, repetition count, motion variability, and range of motion. Seven care-home residents (aged between 65-94 years) performed six common rehabilitation exercises in front of a single camera while wearing 17 inertial sensors (Xsens MVN Awinda) that provided ground-truth joint angles. Two-dimensional skeleton poses estimated from the video were fed into a temporal convolutional neural network to recognize the exercises and estimate three-dimensional joint angles. We evaluated exercise segmentation with F1@50 and angle regression with mean per-joint angular error (MPJAE) across nine trunk and lower-limb joints, using leave-one-subject-out cross-validation. Pearson correlations assessed agreement between estimated and ground-truth EPMs. The DL model achieved an F1@50 of 0.92 (± 0.04) for exercise recognition and an MPJAE of 7.7° (± 0.91) for joint angle estimation. The estimated EPMs aligned closely with ground truth, achieving correlation scores of 0.93 (95% CI [0.90, 0.95]) for duration, 0.86 (95% CI [0.80, 0.90]) for repetition count, and between 0.3 and 0.9 for motion variability and range of motion across exercises. The DL algorithm reliably estimates key exercise outcomes from a single video stream. This video-based monitoring pipeline could enable unsupervised, technology-supported exercise assessment in residential care homes while safeguarding session quality and safety. Future work will validate the approach in larger and more diverse cohorts.
Abstract Background Rheumatic heart disease (RHD) remains a major public health concern across low- and middle-income countries in the Global South. Early detection through community-based screening of asymptomatic individuals has been identified as a critical strategy for reducing the disease burden. Despite this, the absence of accessible, automated population screening tools continues to impede implementation at scale. This study investigates the screening potential of integrating electrocardiography (ECG) and phonocardiography (PCG) for the early detection of RHD in asymptomatic schoolchildren. Methods The dataset was obtained as part of an ambulatory screening initiative conducted across multiple school sites in rural areas of Ethiopia. It comprised ECG and PCG recordings from 611 asymptomatic schoolchildren aged 10 to 20 years. A comprehensive set of time–frequency, visibility graph and non-linear features were extracted from both signal modalities. These features were subsequently evaluated using machine learning models to assess their utility in the automated screening of early RHD. Results The best model achieved an average 10-folds cross-validation scores on sensitivity, positive-predictive-value and F1-score of 59.6%, 63.6% and 60.8%, respectively for multimodal ECG and PCG signals. Whereas separate evaluation of ECG showed an F1-score of 61.1% and PCG achieved 23.5%. Key features included the T-wave, the area under the QRS complex, and entropy measures derived from beat visibility graphs in the ECG. In addition, visibility graph features from multi-band S1 and S2 heart sound segments, along with MFCC coefficients from the PCG, were also relevant. However, PCG alone performed poorly and did not show improved results over the ECG features. Conclusion Although auscultation is key clinical diagnosis tool in symptomatic RHD, combined PCG with ECG features does not enhance asymptomatic RHD detection using the ECG modality alone.
Background:Chronic obstructive pulmonary disease (COPD) is a leading cause of morbidity and mortality worldwide, with frequent exacerbations of COPD (ECOPD) significantly impacting patient health and health care systems. Predicting ECOPD early would increase patients' quality of life and decrease the economic burden. The advancement of wearable technologies and Internet of Things (IoT) sensors has enabled continuous remote monitoring (RM), offering new opportunities for early ECOPD prediction. However, effectively leveraging wearable data requires robust artificial intelligence (AI) frameworks capable of processing heterogeneous physiological and environmental information. Objective:This systematic review aims to provide a comprehensive overview of both hardware and software solutions for predicting ECOPD using RM. From the reviewed literature, we first focus on key physiological and environmental variables essential for COPD monitoring that can be extracted from wearables and IoT sensors. Second, we describe the wearable and IoT devices currently deployed in COPD management. Finally, we review machine learning, including deep learning models, used for ECOPD prediction, discussing limitations for real-world implementation. By bridging AI-driven data processing with real-world sensor applications, this review aims to outline the current landscape, existing challenges, and future directions for developing effective RM solutions for ECOPD predictions. Methods:A comprehensive search was conducted following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines to identify studies using AI or machine learning techniques for predicting ECOPD in in-home contexts. Results:This review identified 26 studies that met the inclusion criteria. Twenty studies aimed at predicting or detecting exacerbations at the onset. The variables tracked most frequently were heart rate (n=9), peripheral oxygen saturation (n=9), and symptoms (n=8). Daily or weekly sampling was most common (n=14). Most studies (n=13) applied machine learning models-primarily random forest (n=5), CatBoost (n=2), decision trees (n=2), and support vector machines (n=2). Deep learning was used in 3 papers, while the remaining applied rule-based logics and probabilistic models. Wearables and IoT were used in only 6 out of 20 studies. Six papers analyzed changes in vital parameters during prodromal phases, defined as the period shortly before the onset of an exacerbation. Three studies collected data continuously, 2 daily, and 1 compared once-daily versus overnight monitoring; 4 of these 6 used wearable devices. Conclusions:Overall, current evidence highlights the potential of continuous monitoring of physiological and environmental variables for early ECOPD prediction, offering advantages over questionnaires or once-daily measurements. While wearables and IoT devices show promise, their use remains limited. Many studies rely on balanced datasets that do not mirror real-world exacerbation patterns and lack external validation across diverse populations. Future research should emphasize large-scale validation, integration of multimodal data, and translation of AI models into clinically feasible tools to enable timely intervention and improve COPD management.
Abstract Video annotation is the gold standard for assessing Freezing of Gait (FOG) in Parkinsonian disorders, but it is time-consuming. Deep learning (DL)-based assessment of FOG using inertial measurement units ameliorates this problem but poses challenges. Particularly, the large heterogeneity between patients and assessment methods potentially affects detection performance between independent cohorts. To evaluate heterogeneity effects, we developed a DL model on a local cohort (85 participants; 2043 trials) and validated it across six external cohorts (256 participants; 1058 trials). Model-expert agreement on the percentage-of-time-frozen was strong locally (ICC = 0.886 [0.79, 0.90]) but reduced in external cohorts (ICC = 0.562 ± 0.141). Fine-tuning the DL model with just 50 min of external cohort data improved the ICC to 0.732 ± 0.138, approaching the lower boundary of the inter-rater agreement between two clinical raters using video annotation (ICC = 0.73–0.99). Therefore, while unified standards are still being developed, we propose a human-in-the-loop workflow as an effective intermediary and present a proof-of-concept web-based platform for fine-tuning and expert review ( aidfog.be ).
Prolonged sedentary behavior has been a major public health concern strongly associated with low back pain (LBP), the leading cause of disability worldwide. Monitoring pelvic rotations (PRs), a key factor in postural control, is crucial for LBP prevention and management. However, PR detection remains underexplored, and publicly available datasets are lacking. This study comprehensively investigates the feasibility of using inertial measurement units (IMUs) and machine learning (ML) to recognize PRs. We introduce a novel public dataset comprising over 15 hours of high-resolution IMU data from thirty-nine participants with sedentary lifestyles, specifically designed for PR detection across multiple activities, including sitting, standing, sit-to-stand-to-sit (STSTS) transitions, walking, and targeted anterior/posterior PRs during sitting and standing. We benchmark six classical ML classifiers and identify XGBoost as the most effective model, achieving weighted F1 scores of 0.988 and 0.895 in 2-class and 8-class configurations, respectively. We further conduct a systematic sensor placement optimization, demonstrating that a two-sensor configuration (pelvis and left upper leg) reached performance comparable to that of the full-body 17-sensor system. In the 8-class classification task (sit, stand, STSTS, walk, and targeted anterior/posterior PR), this optimal combination achieved a weighted F1 score of 0.872, closely matching the result obtained using all sensors. Our work provides the community with a unique dataset, a robust baseline for PR recognition, and a practical minimal-sensor solution, thereby paving the way for real-world wearable systems aimed at reducing prolonged static postures and improving LBP management in daily life.
Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomotion mode recognition to adapt control strategies and provide appropriate assistance during daily activities. However, public benchmarks are typically collected from healthy adults, lack temporally precise labels necessary for detecting mode transitions, or focus on a limited set of tasks. To support development and evaluation under realistic clinical constraints and daily mobility demands, we introduce RevalExo, a functional daily-activity benchmark for inertial and visual locomotion mode recognition. RevalExo is built around a standardized, clinically and ecologically validated daily-activity protocol reflecting the cumulative everyday mobility demands in ageing and clinical populations. The benchmark includes 27 participants across three cohorts: older adults without mobility impairments, stroke survivors, and older adults with probable sarcopenia. The full cohort was recorded with lower-body IMUs, while synchronized egocentric video was collected for a clinically feasible subset of 13 participants. RevalExo provides 10.1 hours of frame-level annotations across 11 locomotion modes, including 5.1 hours of paired inertial–visual recordings. We benchmark three challenges: unimodal and multimodal locomotion mode recognition across multiple horizons, cross-population generalization from older adults without mobility impairments to clinical cohorts, and vision-guided knowledge transfer to IMU-only models. Results confirm consistent gains from fusing inertial and visual inputs but reveal a substantial gap between general recognition (∼93% F1) and recognition during transitions (∼68% F1), alongside persistent challenges in cross-population generalization and cross-modal transfer. We release RevalExo to stimulate further research on these open challenges.
Overweight and obesity have emerged as widespread societal challenges, frequently linked to unhealthy eating patterns. A promising approach to enhance dietary monitoring in everyday life involves automated detection of food intake gestures. This study introduces a skeleton based approach using a model that combines a dilated spatial-temporal graph convolutional network (ST-GCN) with a bidirectional long-short-term memory (BiLSTM) framework, as called ST-GCN-BiLSTM, to detect intake gestures. The skeleton-based method provides key benefits, including environmental robustness, reduced data dependency, and enhanced privacy preservation. Two datasets were employed for model validation. The OREBA dataset, which consists of laboratory-recorded videos, achieved segmental F1-scores of 86.18% and 74.84% for identifying eating and drinking gestures. Additionally, a self-collected dataset using smartphone recordings in more adaptable experimental conditions was evaluated with the model trained on OREBA, yielding F1scores of 85.40% and 67.80% for detecting eating and drinking gestures. The results not only confirm the feasibility of utilizing skeleton data for intake gesture detection but also highlight the robustness of the proposed approach in cross-dataset validationClinical relevance-This skeleton-based intake gesture detection system can serve as an automated annotation tool for nutritional experts, facilitating the analysis of participants' eating behaviors.
Subclinical rheumatic valvular disease is a significant yet underdiagnosed contributor to the global rheumatic heart disease (RHD) burden. Early detection through population screening is essential to prevent its progression to severe RHD. Rhythm changes and prolongations of PR and QTc intervals in the ECG are described in the advanced RHD cases. However, these parameters were not yet studied in subclinical disease. To investigate the potential of ECG biomarkers for screening RHD in asymptomatic schoolchildren. ECG tracings from 135 at-risk schoolchildren aged 10 to 20 years were selected from a cohort screened for RHD in schools. Confirmatory diagnoses were based on echocardiographic findings, where 89 (F=48, M=41) were healthy and 46 (F=27, M=19) were positive for RHD (24 borderline RHD and 22 definite RHD). Independent, blinded reviewers manually annotated the ECG and QTc, P-wave duration (PWd) and the ratio between the P-wave duration and PR-interval (Pw/PR) were analyzed. The mean age of the study cohort at diagnosis was 16.3 ± 2.7 years, and 55.6%(n=75) of the participants were females. Atrial fibrillation was seen in 8%(n=4), and prolonged PR interval in 2% (n=1) of RHD positive cases. Both QTc (p=0.004) and PWd (p=0.013) showed significant difference between healthy and RHD positive subjects. There was no difference of QTc by age, although a difference was noted by gender. QTc >423ms (p=0.008) predicted the presence of RHD with a sensitivity and specificity of 71.7% and 52.8%, respectively. Multivariate regression of QTc, PWd, and Pw/PR provided AUC of 72.7%. The QTc was increased consistently with severity across age-groups above and below 16 years. The QTc, PWd and Pw/PR might serve as non-invasive biomarkers for the screening of RHD in at-risk school children. Monitoring alterations in these markers at an early stage of RHD is crucial for enabling prompt management and follow-up. It is thus evident that ECG may still be utilized as a beneficial instrument for intermittent ambulatory RHD screening in resource-limited settings.
Freezing of gait (FOG) is a debilitating symptom of Parkinson’s disease (PD), characterized by an absence or reduction in forward movement of the legs despite the intention to walk. Detecting FOG during free-living conditions presents significant challenges, particularly when using only inertial measurement unit (IMU) data, as it must be distinguished from voluntary stopping events that also feature reduced forward movement. Influences from stress and anxiety, measurable through galvanic skin response (GSR) and electrocardiogram (ECG), may assist in distinguishing FOG from normal gait and stopping. However, no study has investigated the fusion of IMU, GSR, and ECG for FOG detection. Therefore, this study introduced two methods: a two-step approach that first identified reduced forward movement segments using a Transformer-based model with IMU data, followed by an XGBoost model classifying these segments as FOG or stopping using IMU, GSR, and ECG features; and an end-to-end approach employing a multi-stage temporal convolutional network to directly classify FOG and stopping segments from IMU, GSR, and ECG data. Results showed that the two-step approach with all data modalities achieved an average F1 score of 0.728 and F1@50 of 0.725, while the end-to-end approach scored 0.771 and 0.759, respectively. However, no significant difference was found compared to using only IMU data in both approaches (p-values: 0.466 to 0.887). In conclusion, adding physiological data did not provide a statistically significant benefit in distinguishing between FOG and stopping. The limitations may be specific to GSR and ECG data, and may not generalize to other physiological modalities.
Rheumatic heart disease (RHD) arises from untreated streptococcal throat infections caused by beta-hemolytic group A streptococci, leading to progressive damage to cardiac valves. While echocardiography is the gold standard for RHD diagnosis, its use in low-income countries is limited due to scarce resources and a lack of trained professionals. Automated RHD detection via echocardiography and phonocardiography data has shown promising, but the effectiveness of electrocardiogram (ECG) for detecting RHD in endemic regions at cardiac wards with limited resource remains uncertain. This study explores the viability of ECG as a cost-effective tool for RHD detection in cardiac wards. The study utilizes a dataset comprising single-lead ECG recordings from 124 confirmed RHD patients and 46 healthy controls collected at a major referral hospital in Ethiopia. Additionally, an extended-RHDECG dataset, which consists of age-matched ECGs from the Physikalisch-Technische Bundesanstalt (PTB-XL) dataset and RHD ECGs, was utilized. A single lead ECG segment of 10-second duration per patient was resampled at 250 Hz. Temporal and relative wavelet energy (RWE) features combined with Convolutional Neural Network (CNN) model features were employed for classification of prevalent cardiovascular diseases in the context of the Global South. A 5-folds cross-validation on RHDECG dataset using CNN model showed an average accuracy (mean ± std) of 88.6 ± 0.2
Jump monitoring for volleyball players during training or a match can be crucial to prevent injuries, yet the measurement requires considerable workload and cost using traditional methods such as video analysis. Also, existing methods do not provide accurate differentiation between different types of jumps. In this study, an unobtrusive system with a single inertial measurement unit (IMU) on the waist was proposed to recognize the types of volleyball jumps. A Multi-Layer Temporal Convolutional Network (MSTCN) was applied for sequence-to-sequence (seq-to-seq) classification without using the sliding window technique. The model was evaluated on volleyball players during a lab session with a fixed protocol of jumping and landing tasks, and during four volleyball training sessions, respectively. The MS-TCN model achieved better performance than a state-of-the-art deep learning model but with lower computational cost. In the lab sessions, most jump counts showed small differences between the predicted jumps and video- annotated jumps, with an overall count showing a Limit of Agreement (LoA) of 0.1 +/- 3.40 (r = 0.884). For comparison, the proposed algorithm showed slightly worse results than VERT (a commercial jumping assessment device) with a LoA of 0.1 +/- 2.08 (r = 0.955) but the differences were still within a comparable range. In the training sessions, the recognition of three types of jumps exhibited a mean difference from observation of less than 10 jumps: block, smash, and overhead serve. These results showed the potential of using a single IMU to recognize the types of volleyball jumps. The proposed architecture provided high resolution of recognition and required fewer parameters compared with state-of-the-art models.
The physical load of jumps plays a critical role in injury prevention for volleyball players. However, manual video analysis of jump activities is time-intensive and costly, requiring significant effort and expensive hardware setups. The advent of the inertial measurement unit (IMU) and machine learning algorithms offers a convenient and efficient alternative. Despite this, previous research has largely focused on either jump classification or physical load estimation, leaving a gap in integrated solutions. This study aims to present a pipeline to automatically detect jumps and predict heights using data from a waist-worn IMU. The pipeline leverages a Multi-Stage Temporal Convolutional Network (MS-TCN) to detect jump segments in time-series data and classify the specific jump category. Subsequently, jump heights are estimated using three downstream regression machine learning models based on the identified segments. Our method is verified on a dataset comprising 10 players and 337 jumps. Compared to the result of VERT in height estimation (R-squared = -1.53), a commercial device commonly used in jump landing tasks, our method not only accurately identifies jump activities and their specific types (F1-score = 0.90) but also demonstrates superior performance in height prediction (R-squared = 0.53). This integrated solution offers a promising tool for monitoring physical load and mitigating injury risk in volleyball players.
Background:Home-based neurophysiological monitoring is improving the assessment and management of neurological conditions such as epilepsy. Technologies such as electroencephalography (EEG), electromyography (EMG), and accelerometry are increasingly integrated into wearable systems for at-home use. Due to an increasing amount of data from long-term monitoring, machine learning algorithms assist in automated data analysis. However, ensuring device accuracy, signal quality, and user compliance remains crucial for clinical useability. Objective:This chapter explores advances and challenges in at-home neurophysiological monitoring, with a primary focus on EEG systems and their applications.Content: The discussion highlights the technological advances and the challenges associated with at-home monitoring. The focus will be on EEG systems, as well as a discussion of EMG in epilepsy. Next, we will provide an overview of the clinical applications for home-based monitoring of epilepsy and sleep disorders. Lastly, we will briefly discuss emerging topics within home-based monitoring in movement disorders and neurodegenerative disorders. Conclusion:Future advancements are expected with new generations of wearable systems capable of providing long-term monitoring with minimal maintenance. Beyond epilepsy and sleep disorders, home-based technologies are also being investigated in other neurological diseases including movement disorders and neurodegenerative diseases showing the expanding scope of home-based technologies in neurology.
The Otago Exercise Program (OEP) serves as a vital rehabilitation initiative for older adults, aiming to enhance their strength and balance, and consequently prevent falls. While Human Activity Recognition (HAR) systems have been widely employed in recognizing the activities of individuals, existing systems focus on the duration of macro activities (i.e. a sequence of repetitions of the same exercise), neglecting the ability to discern micro activities (i.e. the individual repetitions of the exercises), in the case of OEP. This study presents a novel multi-task machine learning approach aimed at bridging this gap in recognizing the micro activities of OEP. To manage the limited dataset size, our model utilizes a Transformer encoder for feature extraction, subsequently classified by a Temporal Convolutional Network (TCN). Simultaneously, the Transformer encoder is employed for masked self-supervised learning to reconstruct input signals. Results indicate that the masked unsupervised learning task enhances the performance of the supervised learning (classification task), as evidenced by f1-scores surpassing the clinically applicable threshold of 0.8. From the micro activities, two clinically relevant outcomes emerge: counting the number of repetitions of each exercise and calculating the velocity during chair rising. These outcomes enable the automatic monitoring of exercise intensity and difficulty in the daily lives of older adults.
The Otago Exercise Program (OEP) represents a crucial rehabilitation initiative tailored for older adults, aimed at enhancing balance and strength. Despite previous efforts utilizing wearable sensors for OEP recognition, existing studies have exhibited limitations in terms of accuracy and robustness. This study addresses these limitations by employing a single waist-mounted Inertial Measurement Unit (IMU) to recognize OEP exercises among community-dwelling older adults in their daily lives. A cohort of 36 older adults participated in laboratory settings, supplemented by an additional 7 older adults recruited for at-home assessments. The study proposes a Dual-Scale Multi-Stage Temporal Convolutional Network (DS-MS-TCN) designed for two-level sequence-to-sequence classification, incorporating them in one loss function. In the first stage, the model focuses on recognizing each repetition of the exercises (micro labels). Subsequent stages extend the recognition to encompass the complete range of exercises (macro labels). The DS-MS-TCN model surpasses existing state-of-the-art deep learning models, achieving f1-scores exceeding 80% and Intersection over Union (IoU) f1-scores surpassing 60% for all four exercises evaluated. Notably, the model outperforms the prior study utilizing the sliding window technique, eliminating the need for post-processing stages and window size tuning. To our knowledge, we are the first to present a novel perspective on enhancing Human Activity Recognition (HAR) systems through the recognition of each repetition of activities.
AI-assisted food intake monitoring systems have drawn considerable attention from researchers. To date, various approaches have been proposed to objectively and unobtrusively detect food intake activities by utilizing novel sensors and machine learning techniques. In the development of automated food intake monitoring systems, one crucial step is to evaluate the generated results from machine learning models. In this study, we illustrate the challenge arising from the inefficiency of traditional sliding-window-based evaluation in translating results into clinical indices (i.e. number of bites). Additionally, existing evaluation metrics only focus on detection performance (count the occurrence of eating gestures); however, the segmentation performance (temporal boundary of eating gesture) is missed, which is also a clinically meaningful index. Apart from the discussion of existing evaluation methods in food intake monitoring, we introduce the segment-wise evaluation scheme using the Intersection Over Union (IoU) as threshold to assess performance. This method facilitates the evaluation of both the detection and segmentation performance of eating activities. Two public food intake datasets are used in our case study to illustrate that the segment-wise method can yield more detailed information and a more comprehensive evaluation when compared to existing metrics. The proposed evaluation scheme has the potential to be applied to other human activity recognition (HAR) cases.