Large language models (LLMs) are revolutionizing healthcare by improving diagnosis, patient care, and decision support through interactive communication. More recently, they have been applied to analyzing physiological time-series like wearable data for health insight extraction. Existing methods embed raw numerical sequences directly into prompts, which exceeds token limits and increases computational costs. Additionally, some studies integrated features extracted from time-series in textual prompts or applied multimodal approaches. However, these methods often produce generic and unreliable outputs due to LLMs’ limited analytical rigor and inefficiency in interpreting continuous waveforms. In this paper, we develop an LLM-powered agent for physiological time-series analysis aimed to bridge the gap in integrating LLMs with well-established analytical tools. Built on the OpenCHA, an open-source LLM-powered framework, our agent powered by OpenAI’s GPT-3.5-turbo model features an orchestrator that integrates user interaction, data sources, and analytical tools to generate accurate health insights. To evaluate its effectiveness, we implement a case study on heart rate (HR) estimation from Photoplethysmogram (PPG) signals using a dataset of PPG and Electrocardiogram (ECG) recordings in a remote health monitoring study. The agent’s performance is benchmarked against OpenAI GPT-4o-mini and GPT-4o, with ECG serving as the gold standard for HR estimation. Results demonstrate that our agent significantly outperforms benchmark models by achieving lower error rates and more reliable HR estimations. The agent implementation is publicly available on GitHub1.
Wearable technology has expanded the applications of photoplethysmography (PPG) in remote health monitoring, enabling real-time measurement of various physiological parameters, such as heart rate (HR), heart rate variability (HRV), and respiration rate (RR). While existing studies mainly focus on individual parameters derived from PPG, they often overlook the shared characteristics among these physiological parameters. Multitask learning (MTL) offers a promising solution by training a single model to perform multiple related tasks, leveraging their interdependencies. However, the potential of MTL has not been thoroughly investigated in the context of PPG analysis. In this paper, we develop MTL approaches that exploit shared underlying characteristics across PPG-related tasks to improve the performance of PPG-based applications. We propose customized multitask deep learning models for two applications: (1) PPG quality assessment for HR and HRV features collected in free-living conditions and (2) simultaneous HR and RR estimation from PPG. Our models are evaluated on a PPG dataset collected from 46 subjects wearing smartwatches during their daily activities. Results demonstrate that the proposed MTL methods significantly outperform baseline single-task models, achieving higher accuracy in quality assessment and reduced error rates in HR and RR estimation.
Preterm birth (PTB) remains a global health concern, impacting neonatal mortality and lifelong health consequences. Traditional methods for estimating PTB rely on electronic health records or biomedical signals, limited to short-term assessments in clinical settings. Recent studies have leveraged wearable technologies for in-home maternal health monitoring, offering continuous assessment of maternal autonomic nervous system (ANS) activity and facilitating the exploration of PTB risk. In this paper, we conduct a longitudinal study to assess the risk of PTB by examining maternal ANS activity through heart rate (HR) and heart rate variability (HRV). To achieve this, we collect long-term raw photoplethysmogram (PPG) signals from 58 pregnant women (including seven preterm cases) from gestational weeks 12-15 to three months post-delivery using smartwatches in daily life settings. We employ a PPG processing pipeline to accurately extract HR and HRV, and an autoencoder machine learning model with SHAP analysis to generate explainable abnormality scores indicative of PTB risk. Our results reveal distinctive patterns in PTB abnormality scores during the second pregnancy trimester, indicating the potential for early PTB risk estimation. Moreover, we find that HR, average of interbeat intervals (AVNN), SD1SD2 ratio, and standard deviation of interbeat intervals (SDNN) emerge as significant PTB indicators.
Agents represent one of the most emerging applications of Large Language Models (LLMs) and Generative AI, with their effectiveness hinging on multimodal capabilities to navigate complex user environments. Conversational Health Agents (CHAs), a prime example of this, are redefining healthcare by offering nuanced support that transcends textual analysis to incorporate emotional intelligence. This paper introduces an LLM-based CHA engineered for rich, multimodal dialogue-especially in the realm of mental health support. It adeptly interprets and responds to users' emotional states by analyzing multimodal cues, thus delivering contextually aware and empathetically resonant verbal responses. Our implementation leverages the versatile openCHA framework, and our comprehensive evaluation involves neutral prompts expressed in diverse emotional tones: sadness, anger, and joy. We evaluate the consistency and repeatability of the planning capability of the proposed CHA. Furthermore, human evaluators critique the CHA's empathic delivery, with findings revealing a striking concordance between the CHA's outputs and evaluators' assessments. These results affirm the indispensable role of vocal (soon multimodal) emotion recognition in strengthening the empathetic connection built by CHAs, cementing their place at the forefront of interactive, compassionate digital health solutions.
Sleep quality is crucial to both mental and physical well-being. The COVID-19 pandemic, which has notably affected the population’s health worldwide, has been shown to deteriorate people’s sleep quality. Numerous studies have been conducted to evaluate the impact of the COVID-19 pandemic on sleep efficiency, investigating their relationships using correlation-based methods. These methods merely rely on learning spurious correlation rather than the causal relations among variables. Furthermore, they fail to pinpoint potential sources of bias and mediators and envision counterfactual scenarios, leading to a poor estimation. In this paper, we develop a Causal Machine Learning method, which encompasses causal discovery and causal inference components, to extract the causal relations between the COVID-19 pandemic (treatment variable) and sleep quality (outcome) and estimate the causal treatment effect, respectively. We conducted a wearable-based health monitoring study to collect data, including sleep quality, physical activity, and Heart Rate Variability (HRV) from college students before and after the COVID-19 lockdown in March 2020. Our causal discovery component generates a causal graph and pinpoints mediators in the causal model. We incorporate the strongly contributing mediators (i.e., HRV and physical activity) into our causal inference component to estimate the robust, accurate, and explainable causal effect of the pandemic on sleep quality. Finally, we validate our estimation via three refutation analysis techniques. Our experimental results indicate that the pandemic exacerbates college students’ sleep scores by 8%. Our validation results show significant p-values confirming our estimation.
Internet-of-Things-based systems have recently emerged, enabling long-term health monitoring systems for the daily activities of individuals. The data collected from such systems are multivariate and longitudinal, which call for tailored analysis techniques to extract the trends and abnormalities in the monitoring. Different methods in the literature have been proposed to identify trends in data. However, they do not include the time dependency and cannot distinguish changes in long-term health data. Moreover, their evaluations are limited to lab settings or short-term analysis. Long-term health monitoring applications require a modeling technique to merge the multisensory data into a meaningful indicator. In this paper, we propose a personalized neural network method to track changes and abnormalities in multivariate health data. Our proposed method leverages convolutional and graph attention layers to produce personalized scores indicating the abnormality level (i.e., deviations from the baseline) of users' data throughout the monitoring. We implement and evaluate the proposed method via a case study on long-term maternal health monitoring. Sleep and stress of pregnant women are remotely monitored using a smartwatch and a mobile application during pregnancy and 3-months postpartum. Our analysis includes 46 women. We build personalized sleep and stress models for each individual using the data from the beginning of the monitoring. Then, we compare the two groups by measuring the data variations. The abnormality scores produced by the proposed method are compared with the findings from the self-report questionnaire data collected in the monitoring and abnormality scores generated by an autoencoder method. The proposed method outperforms the baseline methods in exploring the changes between high-risk and low-risk pregnancy groups. The proposed method's scores also show correlations with the self-report data. Consequently, the results indicate that the proposed method effectively detects the abnormality in multivariate long-term health monitoring.
The rapid development of wearable technology has enabled remote photoplethysmography (PPG)-based health monitoring in everyday settings, offering real-time and continuous monitoring of cardiovascular parameters, such as heart rate (HR) and heart rate variability (HRV). However, PPG signals collected in daily life are prone to artifacts and noise, posing challenges to HR and HRV extraction. The existing HR and HRV extraction methods cannot effectively handle noisy PPG signals and ensure accurate results. Additionally, current Python packages were primarily designed for analyzing "clean" PPG signals, limiting their performance in handling artifacts and noise and resulting in unreliable HR and HRV measurements. In this paper, we propose a robust end-to-end PPG processing pipeline to reliably extract HR and HRV from PPG signals collected in free-living settings. The pipeline comprises three machine learning-based PPG analysis methods: signal quality assessment, reconstruction of noisy signal, and systolic peak detection. We assess the proposed PPG pipeline using a dataset including PPG and Electrocardiogram (ECG) signals recorded from 46 individuals by smartwatches. Our evaluation demonstrates the proposed pipeline’s superior performance compared to two established benchmark methods in terms of correlation and mean absolute error with ECG as the reference. We also provide the Python implementation of our pipeline for the research community to facilitate integration into their solutions.
Photoplethysmography (PPG) is a non-invasive technique used in wearable devices to measure vital signs (e.g., heart rate). The method is, however, highly susceptible to motion artifacts, which are inevitable in remote health monitoring. Noise reduces signal quality, leading to inaccurate decision-making. In addition, unreliable data collection and transmission waste a massive amount of energy on battery-powered devices. Studies in the literature have proposed PPG signal quality assessment (SQA) enabled by rule-based and machine learning (ML)-based methods. However, rule-based techniques were designed according to certain specifications, resulting in lower accuracy with unseen noise and artifacts. ML methods have mainly been developed to ensure high accuracy without considering execution time and device’s energy consumption. In this paper, we propose a lightweight and energy-efficient PPG SQA method enabled by a semi-supervised learning strategy for edge devices. We first extract a wide range of features from PPG and then select the best features in terms of accuracy and latency. Second, we train a one-class support vector machine model to classify PPG signals into “Reliable” and “Unreliable” classes. We evaluate the proposed method in terms of accuracy, execution time, and energy consumption on two embedded devices, in comparison to five state-of-the-art PPG SQA methods. The methods are assessed using a PPG dataset collected via smartwatches from 46 individuals in free-living conditions. The proposed method outperforms the other methods by achieving an accuracy of 0.97 and a false positive rate of 0.01. It also provides the lowest latency and energy consumption compared to other ML-based methods.
ECG signal is among medical signals used to diagnose heart problems. A large volume of medical signal’s data in telemedicine systems causes problems in storing and sending tasks. In the present paper, a recursive algorithm with backtracking approach is used for ECG signal compression. This recursive algorithm constructs a mathematical estimator function for each segment of the signal using genetic programming algorithm. When all estimator functions of different segments of the signal are determined and put together, a piecewise-defined function is constructed. This function is utilized to generate a reconstructed signal in the receiver. The compression result is a set of compressed strings representing the piecewise-defined function which is coded through a text compression method. In order to improve the compression results in this method, the input signal is smoothed. MIT-BIH arrhythmia database is employed to test and evaluate the proposed algorithm. The results of this algorithm include the average of compression ratio that equals 30.97 and the percent root-mean-square difference that is equal to 2.38%, suggesting its better efficiency in comparison with other state-of-the-art methods.
Telemedicine refers to a group of modern medical services that are provided on the platform of advanced telecommunication technologies. One of these services is the screening for heart diseases, which are the leading cause of mortality across the world. But the development of telemedicine systems for cardiac screening faces multiple challenges. One of these challenges is the large volume of ECG signals, which makes them difficult to store and transfer. Of the many algorithms proposed for the compression of ECG signals, most rely on the processing of data as discrete numerical values. The alternative approach followed in this study is to model the signal compression problem into a regression problem and then convert it into a text compression problem. Using this approach, the paper presents a new genetic programming based method for the compression of ECG signals. The proposed method starts with denoising and smoothing the ECG signal with discrete wavelet transform and then constructing its mathematical model with a genetic programming based algorithm. This model is a piecewise mathematical function where each sub-function models one part of the signal. Next, the model is converted to a character string and regular expressions are used to extract the function coefficients and encode the symbols contained in the string. Finally, the strings and coefficients are compressed using the 12W and arithmetic encoding methods, respectively. The efficiency of the algorithm is evaluated through compression ratio (CR), percent root-mean-square difference (PRD), root-mean-square-error (RMSE) and quality score (QS) on MIT-BIH Arrhythmia Database records. The evaluation results demonstrate the good performance of the proposed method in comparison with other state-of-the-art techniques. (C) 2019 Elsevier Ltd. All rights reserved.