IntroductionGenerative artificial intelligence (GenAI) is becoming an important tool in medical product development. A main component of this development includes annotating, summarizing, and extracting key insights from expert interviews to identify clinical pain points and curate device requirements. These tasks are time- and labor-intensive, resulting in increased administrative burden and reduced efficiency. As a result, researchers have developed large language models (LLMs) that can disseminate research and interview findings with reduced workload and improved productivity. This study explores the use of GenAI, specifically GPT-4o, to extract user functional and design requirements from medical professional interviews for the iterative development of an infant heart rate detector for neonatal resuscitation.MethodsA total of 29 healthcare practitioners were interviewed using a semistructured interview format. The interviews were recorded and transcribed. GPT-4o was used to extract user insights from the transcripts, and the results were compared with manual interviewer notes.ResultsA total of 26 h of interview data were collected. All interviewees validated the clinical need for a modality that enables quick and accurate heart rate (HR) measurement during neonatal resuscitation. A set of user requirements was extracted from the interviews and curated under the themes of ease of use, fast and accurate HR measurement, reusability, display, battery life, start-up time, and cost. Also, quantitative analyses of the interviewee’s years of experience, clinical settings, and specialties were conducted.DiscussionThese analyses were conducted using GPT-4o and compared with ground-truth manual annotations to determine the accuracy and reliability of GenAI in content extraction and summarization.ConclusionOverall, this study explored the user requirements identified through in-depth interviews for the development of a pediatric medical device. It also aimed to demonstrate the potential of GenAI in curating these design requirements, offering a framework for researchers and product designers to explore the use of LLMs in curating user requirements and design specifications for medical devices.
Background:Stress among health care workers (HCWs) contributes to burnout, workforce attrition, and adverse patient outcomes. Although virtual reality (VR), psychoeducation, ecological momentary assessments (EMAs), and wearables have independently shown promise in stress research, no integrated digital suite has combined controlled stress induction, intervention delivery, and longitudinal real-world monitoring in HCWs. Objective:This study aimed to evaluate the feasibility, engagement, and preliminary effectiveness of a multimodal Digital Health Monitoring and Intervention suite for Stress framework integrating VR simulation, psychoeducation, EMAs, and wearable biometrics. We examined (1) the impact of VR simulation and psychoeducation on stress outcomes and (2) associations between physiological and self-reported mental health outcomes. Methods:Ninety-nine nurses (mean age 33.7, SD 8.9 yr, 87% female) were enrolled in 2023. We conducted a single-arm prospective cohort study (NCT05923398). Using convenience sampling, participants were recruited from social media advertisements, flyers, and email notices distributed through professional listservs. Participants completed ≥2-week baseline monitoring, a single VR session (2 runs separated by a brief psychoeducation intervention), and 12-week follow-up. In-VR stress was assessed using the Subjective Units of Distress Scale (SUDS) and 4-item Moral Injury Outcome Scale (MIOS-4), with synchronous heart rate variability. Longitudinal outcomes included weekly and biweekly EMAs alongside 70 wearable-derived features. Paired t tests, aligned rank transform ANOVA, and Pearson correlations informed study objectives, with P values adjusted for multiple comparisons. Qualitative content analysis classified emotional responses during and after VR. Results:VR significantly increased subjective stress across checkpoints in both runs, with attenuation in Run B relative to Run A (all P<.001). No significant heart rate variability differences were observed between runs (P=.15). During VR, 92% (91/99) of participants felt stressed, 36% (36/99) reported anxiety or nervousness, and 51% (50/99)-78% (77/99) endorsed anger, guilt, shame, and/or betrayal. Most (59/99, 60%) HCWs returned to an emotional baseline post-VR, although 12% (12/99) reported lingering distress. Immediate reliable improvements in anger, guilt, shame, and/or betrayal occurred for 50% (50/99)-75% (74/99) of participants post intervention. Anxiety (mean -0.53, SD 2.34; P=.03) and stress (mean -3.05, SD 11.35; P=.01) decreased 2 weeks post intervention, but were not sustained at 12 weeks. Increased sleep restlessness was the only wearable feature showing significant changes (mean 2.46%, SD 5.43; Padj<.001). In-VR stress correlated with 12-week real-world stress (SUDS: r=0.57-0.58; MIOS-4: r=0.58-0.61; all P<.01). Data completion exceeded 90%, with 71% achieving full compliance. Conclusions:This study moves beyond single-tool interventions to demonstrate the feasibility and preliminary effectiveness of an integrated, multimodal stress platform within a single coordinated framework. This trial demonstrates high engagement, short-term symptom responsiveness, ecological validity, and emotional safety. The framework provides a scalable model for proactive stress identification, skills training, and implementation in high-risk occupational settings. Randomized controlled trials are needed to establish sustained efficacy and optimize deployment for real-world implementation.
Myasthenia gravis (MG) is a neuromuscular disorder that can precipitate serious and fatal complications, especially when respiratory muscles are affected. Current electrodiagnostic methods, such as Single Fiber Electromyography (SFEMG) and Repetitive Nerve Stimulation (RNS), have notable limitations. While SFEMG is sensitive to neuromuscular abnormalities rather than being specific for MG, it is also a semi-invasive procedure that requires supervision by specialist clinicians and an expensive clinical setup. Conversely, RNS is highly specific for MG, but its sensitivity decreases in cases of low severity. Prior studies, including our previous work, have demonstrated that the manifestation of MG in affecting eye movements can be indirectly quantified using electrooculography (EOG) signals. This non-invasive method could develop into a widely used and easily deployed screening method, especially for early detection of MG using readily obtainable biomarkers. With this goal in mind, our research team developed data collection protocols using a standardized eye movement protocol that could be implemented in a standard clinical setting to acquire EOG data. In this article, we present the experimental setup, test protocols, sample outcomes, and preliminary analysis approaches. Common artifacts encountered, as well as technical and logistical challenges associated with such a setup, are also discussed. 1. An approach to non-invasively detect and quantify MG using eye movement signals. 2. Data collection protocols for acquiring eye movement signals designed for MG quantification 3. Novel eye movement-based biomarkers in quantifying MG.
IntroductionSince the onset of the COVID-19 pandemic, extensive research has focused on developing non-invasive diagnostic approaches of respiratory syndrome using biomedical signals, particularly cough and speech audio. Time-frequency representations combined with Machine Learning models have shown potential in identifying acoustic biomarkers associated with respiratory conditions. Although many existing approaches demonstrate high performance, their use may be limited in resource-constrained environments due to processing or implementation demands.MethodsIn this study, we propose an end-to-end approach for COVID-19 inference based on compressed time-domain audio signals. The method combines temporal signal compression strategies - Downsampling (DS) and Compressive Sensing (CS) - with a Convolutional Neural Network (CNN) trained directly on the waveforms. This design eliminates the need for handcrafted features or spectrograms, aiming to reduce computational complexity while preserving classification performance.ResultsTo evaluate the proposed structure, we used data from two open-access datasets, one for coughing and one for speech. Experimental results, assessed using accuracy and F1-score metrics, indicate that CS outperformed DS in most scenarios, particularly under high compression rates (e.g., 200 Hz and 100 Hz).DiscussionThese findings support the use of compressed audio-based classification in real-world embedded and mobile health systems, where computational efficiency is essential.
BACKGROUND:EEG is widely used to identify neural markers, personalize treatments, and evaluate interventions. However, low signal-to-noise ratio and mixing of artifactual distorted interpretation. Traditional preprocessing approaches, including Independent Component Analysis (ICA), require extensive manual inspection and are laborious. This study aims to develop a human-in-the-loop preprocessing framework, in which AI agents assist experts in identifying, labeling, and iteratively refining artifact removal decisions. NEW METHOD:The proposed framework integrates a large language model (LLM)-driven decision-making agent that calls EEG analysis tools to manage complex signal mixtures. The system combines standard preprocessing with an iterative reasoning loop, where the agent interprets probabilistic outputs from multiple classifiers to decide which components to retain or reject, and whether rerun is needed. Each iteration is re-evaluated through a closed-loop policy. This structure enables adaptive artifact correction guided by interpretable reasoning steps and continuous feedback from both model ensemble and human expert, ensuring reproducibility, auditability, and improved preprocessing efficiency. RESULTS:Evaluations were conducted using both synthetic EEG data and expert-annotated empirical datasets to assess artifact detection, ICA classification accuracy, and reconstruction quality. COMPARISON WITH EXISTING METHODS:Across these tests, the AI agent system has consistently outperformed conventional preprocessing pipelines, achieving equivalent or better signal cleaning of artifactual components achieving a Person's r of 0.666 ± 0.188, and RMSE of 5 × 10-6 ± 1 × 10-6 relative to expert-labeled baselines. CONCLUSIONS:The AI agent framework streamlines preprocessing while maintaining expert oversight via a closed-loop policy that removes high-confidence artifacts, reconstructs the signal, and re-runs analysis with fixed hyperparameters.
Engagement between client and therapist is a critical determinant of therapeutic success. We propose a multi-dimensional natural language processing (NLP) framework that objectively classifies engagement quality in counseling sessions based on textual transcripts. Using 253 motivational interviewing transcripts (150 high-quality, 103 low-quality), we extracted 42 features across four domains: conversational dynamics, semantic similarity as topic alignment, sentiment classification, and question detection. Classifiers, including Random Forest (RF), Cat-Boost, and Support Vector Machines (SVM), were hyperparameter tuned and trained using a stratified 5-fold cross-validation and evaluated on a holdout test set. On balanced (non-augmented) data, RF achieved the highest classification accuracy (76.7 and SVM achieved the highest AUC (85.4 performance improved significantly: RF achieved up to 88.9 F1-score, and 94.6 93.6 future larger-scale applications. Feature contribution revealed conversational dynamics and semantic similarity between clients and therapists were among the top contributors, led by words uttered by the client (mean and standard deviation). The framework was robust across the original and augmented datasets and demonstrated consistent improvements in F1 scores and recall. While currently text-based, the framework supports future multimodal extensions (e.g., vocal tone, facial affect) for more holistic assessments. This work introduces a scalable, data-driven method for evaluating engagement quality of the therapy session, offering clinicians real-time feedback to enhance the quality of both virtual and in-person therapeutic interactions.
Multi-sensor fusion can improve daily stress monitoring. Methods: A wrist-worn device includes a system of the galvanic skin response (GSR), PPG-derived heart rate variability (HRV), skin temperature, and SpO2, paired with self-reported questionnaires. The device streams data to a mobile app over Bluetooth Low Energy and updates the UI within 1–2 s. The physiological features are captured within a fixed window around each questionnaire time and undergo a mid-level fusion; late fusion is also evaluated using self-reports. Results: Against a commercial reference device, the proposed system achieved a mean absolute error of 0.23 for SpO2 and 4.94 for BPM in a one-day benchmark session. The system was validated through a technical evaluation using representative inputs and simulated survey labels. The fusion model was evaluated using simulated physiological and survey data. Using a support vector machine algorithm, a mean squared error of 0.08 was achieved when predicting simulated stress labels. Temperature was shown to have the strongest correlation with simulated stress levels at −0.43, followed by heart rate variability (HRV) at 0.36, while SpO2 had a negligible correlation at 0.09 in the current dataset. Conclusion: The system integrates multi-sensing, on-device preprocessing, BLE transmission, and a clear fusion workflow that creates a useful predictive performance of daily stress monitoring.
Audio source counting is a fundamental task of audio scene analysis related to other audio tasks such as speaker diarization and sound event detection. It is also a relatively unexplored audio task that presents a complex challenge. In particular, source counting performance is poor when the source count range is large, limiting its potential applications. This paper presents a novel approach to improve upon audio source counting through multi-task learning. We present a first of its kind empirical study on the hierarchical nature of audio source counting, introducing the coarse source counting task and a hierarchical multi-task learning framework, in order to better understand and investigate the audio source counting task through several case study scenarios. We perform multi-task learning with a ResNet architecture and demonstrate improvements to audio source counting accuracy by up to a 6% increase from the previous best result on the SARdBScene dataset. We also perform multi-task learning of audio source counting and acoustic scene classification as a step forward for robust audio scene analysis. These experimental results show improvements of up to 6% in source counting accuracy over state-of-the-art baselines, particularly in high source count scenarios. Our findings highlight that multi-task learning not only enhances accuracy, but also improves efficiency by replacing multiple task-specific models with a single robust network.
Each year, around 10% of infants globally will require resuscitation at birth. Pediatricians can use stethoscope, electrocardiogram (ECG) or pulse oximetry to determine heart rate (HR) which is used to guide resuscitation steps. HR must be acquired accurately and quickly. However, current HR detection modalities are either inaccurate or too slow. This work offers a novel infant heart rate detector (iHRD) using single-lead dry electrode ECG that can display HR accurately within the first 10 seconds of initial contact. A research ethics board approved validation study is conducted on 50 healthy newborns comparing iHRD's HR with clinical HR monitors at a community hospital. 3-minute newborn single-lead ECGs and HR are recorded, and HR is annotated every 2 seconds. Statistical HR analysis is performed to ensure iHRD's feasibility and reliability. With 2741 HR datapoints, excluding outliers, the iHRD detected HR with 94.5% accuracy with time from contact to HR display under 10 seconds. Overall, the iHRD using dry electrode single-lead ECG showed good results in providing reliable HR quickly for neonatal resuscitation efforts.
With the proliferation of wearable healthcare devices and garments in the last decade, the necessity of the storage capacity of the acquired biomedical signals; particularly multi-lead electrocardiogram (MECG) signals, and the importance of securing users' personal information have increased significantly. However, existing MECG compression and steganography algorithms are insufficient to address these challenges effectively. This paper presents a discrete cosine transform, singular value decomposition, and American standard code for information interchange (ASCII) character encoding-based highly efficient quality-guaranteed steganographed MECG compression algorithm. The algorithm is tested on three publicly available MECG databases totalling 2.98 months, and its performance is assessed through both qualitative and quantitative measures. The algorithm attains a compression ratio that is much higher than that provided by other algorithms that are developed to compress the MECG signals only. The benefits of using the proposed algorithm are fivefolds: first, the clinical qualities of the reconstructed MECG signals can be controlled precisely, second, user's personal information is restored with no error, third, reconstruction error of the MECG signals is dependent neither on the size of the user's information nor on the steganography operation, fourth, the probability of guesstimating the security-key is close to zero, and fifth, high compression performance.
Thematic analysis provides valuable insights into participants' experiences through coding and theme development, but its resource-intensive nature limits its use in large healthcare studies. Large language models (LLMs) can analyze text at scale and identify key content automatically, potentially addressing these challenges. However, their application in mental health interviews needs comparison with traditional human analysis. This study evaluates out-of-the-box and knowledge-base LLM-based thematic analysis against traditional methods using transcripts from a stress-reduction trial with healthcare workers. OpenAI's GPT-4o model was used along with the Role, Instructions, Steps, End-Goal, Narrowing (RISEN) prompt engineering framework and compared to human analysis in Dedoose. Each approach developed codes, noted saturation points, applied codes to excerpts for a subset of participants (n = 20), and synthesized data into themes. Outputs and performance metrics were compared directly. LLMs using the RISEN framework developed deductive parent codes similar to human codes, but humans excelled in inductive child code development and theme synthesis. Knowledge-based LLMs reached coding saturation with fewer transcripts (10-15) than the out-of-the-box model (15-20) and humans (90-99). The out-of-the-box LLM identified a comparable number of excerpts to human researchers, showing strong inter-rater reliability (K = 0.84), though the knowledge-based LLM produced fewer excerpts. Human excerpts were longer and involved multiple codes per excerpt, while LLMs typically applied one code. Overall, LLM-based thematic analysis proved more cost-effective but lacked the depth of human analysis. LLMs can transform qualitative analysis in mental healthcare and clinical research when combined with human oversight to balance participant perspectives and research resources.
A significant proportion of patients with major depressive disorder do not achieve remission after two antidepressant trials and are considered to suffer from Treatment-Resistant Depression (TRD). Repetitive transcranial magnetic stimulation (rTMS) is an effective treatment for TRD. However, relapse rates among remitters within the first year post-treatment are significant, and there are no validated markers of relapse. Wearable devices have shown positive results for longitudinal monitoring of health metrics and may be a promising tool for an early detection of relapse following rTMS treatment. To evaluate the feasibility of a wearable device (Oura ring) to monitor individuals receiving rTMS treatment for depression and its utility to detect early signs of depressive relapse in a 6-month follow-up period. This single-arm pilot study will recruit 20 outpatients with a major depressive episode receiving rTMS at St. Joseph’s Healthcare Hamilton, Ontario, Canada. Participants will be required to wear an Oura ring throughout the treatment course and during the six-month follow up. Clinical assessments, including Montgomery-Åsberg Depression Rating Scale (MADRS), Patient Health Questionnaire-9 (PHQ-9), Generalized Anxiety Disorder-7 (GAD-7), Insomnia Severity Index (ISI), and World Health Organization-Five Well-Being Index (WHO-5), will be collected at baseline, treatment end, and 3- and 6-month follow-ups, alongside bi-weekly PHQ-9 and GAD-7 scores. The primary outcomes are feasibility metrics (i.e., recruitment, adherence, retention, missing data, usability). Secondary outcomes will include the utility of wearable-based data for predicting relapse. The study was funded in December 2025, and data collection will start following Ethics Approval. This study will provide initial evidence on the feasibility and utility of a wearable-based digital phenotyping in individuals receiving rTMS for TRD. Our findings will inform the design of future large-scale studies aimed at wearable-supported relapse prevention and precision monitoring in depression care.
This article has been withdrawn at the request of the editor due to an error in the publishing process. The Publisher apologizes for any inconvenience this may cause. The full Elsevier Policy on Article Withdrawal can be found at https://www.elsevier.com/about/policies-and-standards/article-withdrawal.
Therapeutic engagement between client and clinician is a key indicator in determining treatment outcomes for clients with mental health disorders. Quantifying this type of engagement provides an opportunity for the development of an engagement quantification framework for therapeutic efficacy, based on a number of data streams including, body movement and synchronicity, speech, and gestures to determine an individual's level of engagement. In this paper, we present a subset of such a framework through the quantification of engagement based on Facial Affect Recognition, Head Motion, and Natural Language Processing. We propose the use of semantic analysis, emotion dynamics and transitions, and head motion to describe a participant's attention over the consultation. For emotion dynamics and transitions we employ seven standard categorical emotions; for head motion we use acute and chronic head movement; and for semantic analysis we employ Robustly Optimized BERT Pretraining Approach. These features derive two engagement levels: low and high. We performed experiments on the AnnoMI dataset, which contains 133 therapeutic consultation videos for low and high quality motivational interviews, and compared the resulting engagement to the level of motivational interviewing. We achieved an 89.1% average accuracy for the Clinician model and an 81.1% average accuracy for the Client model using Gradient Boost as a classifier.
With the advancement of wearable healthcare devices in recent times, there is a growing potential to revolutionize the current strategies of the detection of cardiac anomalies such as myocardial infarction (MI), making it viable to monitor out-of-the-hospital patients using conventional as well as unconventional and pseudo-electrocardiogram (ECG) leads. However, most of the MI detection and localization techniques which are available in the literature to date are based on either all the 12 conventional ECG leads, or subsets thereof. A few single-lead ECG-based techniques are also there, but their accuracies are not satisfactory. On the contrary, this article demonstrates a machine learning model that classifies MI from a single ECG beat of any of the 15 ECG-leads, including the 12 conventional leads and three vectorcardiogram (VCG)-leads, as available in the publicly accessible Physikalisch-Technische Bundesanstalt (PTB) Diagnostic ECG Database. In this proposed technique, first, the R-peaks are detected, and the arrhythmic beats are identified and excluded using a Ramanujan filter bank-based periodicity estimation technique. Next, a few hand-crafted features are extracted from the detected ECG beats and a set of important features are identified. Finally, Extra Trees classifier-based machine learning models with optimized hyperparameters, obtained using the particle swarm optimization (PSO) technique, are developed utilizing these hand-crafted features for the detection and localization of MI. High $F1$ -score and Matthew's correlation coefficient (MCC) corroborate the reliability of the proposed algorithm in both intra- and inter-patient paradigms.
Electrical Impedance Tomography (EIT) is a noninvasive imaging technique that utilizes electrical voltage data to reconstruct cross-sectional images of the human body. EIT has diverse applications such as brain imaging and gesture recognition, making it a valuable tool for both medical diagnostics and human-computer interaction. One potential drawback of EIT is the power consumption and computational requirements for image reconstruction, which may limit real-time applications. This paper presents a novel approach that directly applies machine learning to raw EIT voltage data, bypassing the image reconstruction phase to achieve faster and highly accurate gesture classification. We introduce the Statistical Analysis, Information Theory, and Data-Driven (SID) pipeline, which analyzes raw EIT data using statistical techniques, information theory, and feature ranking, followed by classification with machine learning models. Two datasets were collected for this study: one for five gesture recognition (1193 samples) and another for distinguishing between arm flexion and extension (719 samples). The proposed SID pipeline was applied where the gesture recognition and arm flexion/extension feature set was reduced from 40 to 2 and 1, respectively, and XGBoost was used for classification and achieved an accuracy of 90.38% and 100.00%, respectively. The proposed method demonstrated high accuracy in both tasks while using a reduced feature set. This is in addition to by-passing the image reconstruction for EIT, further enhancing computational efficiency and reduced power consumption, highlighting the potential of this approach for real-time applications.
BackgroundDepression is a prevalent global mental health disorder with substantial individual and societal impact. Natural language processing (NLP), a branch of artificial intelligence, offers the potential for improving depression screening by extracting meaningful information from textual data, but there are challenges and ethical considerations. ObjectiveThis literature review aims to explore existing NLP methods for detecting depression, discuss successes and limitations, address ethical concerns, and highlight potential biases. MethodsA literature search was conducted using Semantic Scholar, PubMed, and Google Scholar to identify studies on depression screening using NLP. Keywords included “depression screening,” “depression detection,” and “natural language processing.” Studies were included if they discussed the application of NLP techniques for depression screening or detection. Studies were screened and selected for relevance, with data extracted and synthesized to identify common themes and gaps in the literature. ResultsNLP techniques, including sentiment analysis, linguistic markers, and deep learning models, offer practical tools for depression screening. Supervised and unsupervised machine learning models and large language models like transformers have demonstrated high accuracy in a variety of application domains. However, ethical concerns related to privacy, bias, interpretability, and lack of regulations to protect individuals arise. Furthermore, cultural and multilingual perspectives highlight the need for culturally sensitive models. ConclusionsNLP presents opportunities to enhance depression detection, but considerable challenges persist. Ethical concerns must be addressed, governance guidance is needed to mitigate risks, and cross-cultural perspectives must be integrated. Future directions include improving interpretability, personalization, and increased collaboration with domain experts, such as data scientists and machine learning engineers. NLP’s potential to enhance mental health care remains promising, depending on overcoming obstacles and continuing innovation.