Background. Co-occurrence of depressive and anxiety symptoms is common; however, the shortage of mental health professionals and high treatment costs means additions to in-person care are necessary to meet treatment needs. The present paper evaluates whether a web-based cognitive bias modification for interpretation (CBM-I) intervention intended for anxious adults—but which targets a shared cognitive mechanism of anxiety and depression—is superior to web-based psychoeducation at reducing co-occurring depressive symptoms. Methods. Latent growth curve modeling was used to assess for superiority of CBM-I over psychoeducation at reducing depressive symptoms among adults with elevated anxiety and depressive symptoms. Interventions were completed once weekly for five weeks with a 2-month follow-up assessment. Analyses were replicated in two separate datasets (Study 1: N = 1069; Study 2: N = 517) and employed multiple variations of CBM-I (Study 1: standard, added low-intensity coaching; Study 2: standard, shortened, added self-referential content, added psychoeducation). Results. Study 1 showed significant superiority of CBM-I over psychoeducation at post-intervention (between-group d = 0.31); this was not replicated in Study 2, but the pattern of effect sizes was similar (d = 0.21). Both interventions showed significant reductions in depressive symptoms at post-intervention (within-group ds from baseline = 0.59-0.92) that were maintained at 2-month follow-up (ds from baseline = 0.75-1.04). Conclusions. The current core CBM-I intervention, designed to target experiences of anxiety, appears preliminarily to be effective in reducing co-occurring depressive symptoms, though its effectiveness is not consistently superior to that of psychoeducation and needs replication in randomized trials.
Background Quality communication is foundational to palliative care, but methods to monitor and track communication performance in clinical settings are limited. CommSense is an AI-powered system created by our team and deployed on wearable devices to securely capture and process clinical conversations and provide actionable feedback to clinicians (1,2). Objective To deploy CommSense and assess clinical and technical acceptability and feasibility. Methods CommSense was deployed within a palliative care clinic and audio-recorded conversations between clinicians and patients (≥18 years, any diagnosis). Brief surveys were completed by patients and clinicians after each recorded conversation to assess perception of the quality of the conversation. Conversation transcripts were cleaned, de-identified and annotated by clinical experts to provide ground truth data for CommSense model comparison. Large Language Models (LLMs) evaluated task-specific language analytics, specifically assessing balanced accuracy and precision of detecting understanding, empathy, emotion, clarity, and presence (3). Internal (patient and clinician users of CommSense) and external stakeholders (policy makers, communication scientists, administrators, and industry experts) participated in semi-structured interviews to explore their perceptions of CommSense. Results Fifty-one (n=51) conversations were collected by CommSense from 36 (n=36) patients and 7 (n=7) palliative care clinicians between May – August 2025. External stakeholder participants (n=14) discussed system-level implementation considerations, including: importance of integrating communication technologies within existing health system infrastructure; engagement of champions throughout the implementation process; attention to data security and privacy; and intentionality with the design and delivery of communication performance feedback. Participants also emphasized the benefit of CommSense to support trainees. Conclusion AI-powered technology has tremendous potential to support healthcare communication, particularly by offering trainees a way to objectively assess communication skills and longitudinally track their progress in implementing communication skills in the actual (non-simulation lab) setting. Future work should focus on optimal ways to provide contextualized communication feedback to clinicians.
Social interactions are fundamental to well-being, yet automatically detecting them in daily life—particularly using wearables—remains underexplored. Most existing systems are evaluated in controlled settings, focus primarily on in-person interactions, or rely on restrictive assumptions (e.g., requiring multiple speakers within fixed temporal windows), limiting generalizability to real-world use. We present an on-watch interaction detection system designed to capture diverse interactions in naturalistic settings. A core component is a foreground speech detector trained on a public dataset. Evaluated on over 100,000 labeled foreground speech and background sound instances, the detector achieves a balanced accuracy of 85.51%, outperforming prior work by 5.11%. We evaluated the system in a real-world deployment (N=38), with over 900 hours of total smartwatch wear time. The system detected 1,691 interactions, 77.28% were confirmed via participant self-report, with durations ranging from under one minute to over one hour. Among correct detections, 81.45% were in-person, 15.7% virtual, and 1.85% hybrid. We further developed a 15-second window-level audio-only model that enables faster interaction prediction, achieving a balanced accuracy of 90.39% and a sensitivity of 91.01% on 33,698 labeled windows. These results demonstrate the feasibility of real-world interaction sensing and open the door to adaptive, context-aware systems responding to users’ dynamic social environments.
Clinical assessments for neuromuscular disorders, such as Spinal Muscular Atrophy (SMA) and Duchenne Muscular Dystrophy (DMD), continue to rely on subjective measures to monitor treatment response and disease progression. We introduce a novel method using wearable sensors to objectively assess motor function during daily activities in 19 patients with DMD, 9 with SMA, and 13 age-matched controls. Pediatric movement data is complex due to confounding factors such as limb length variations in growing children and variability in movement speed. Our approach uses Shape-based Principal Component Analysis to align movement trajectories and identify distinct kinematic patterns, including variations in motion speed and asymmetry. Both DMD and SMA cohorts have individuals with motor function on par with healthy controls. Notably, patients with SMA showed greater activation of the motion asymmetry pattern. We further combined projections on these principal components with partial least squares (PLS) to identify a covariation mode with a canonical correlation of r = 0.78 (95% CI: [0.34, 0.94]) with muscle fat infiltration, the Brooke score (a motor function score) and age-related degenerative changes, proposing a novel motor function index. This data-driven method has the potential to inform future home deployments with wearable devices, allowing better longitudinal tracking of treatment efficacy for children with neuromuscular disorders.
Large language models (LLMs) are transforming healthcare by advancing clinical decision support, patient care, and administrative efficiency. However, effectively and sustainably integrating LLMs into healthcare systems requires addressing participatory gaps that may hinder alignment with stakeholders' practical and ethical needs. This paper explores how participatory methods can be applied throughout the development lifecycle of LLM-enhanced health systems (LLMHS), arguing that: (1) participatory approaches are critical for engaging stakeholders in LLMHS development, and (2) LLM techniques can create novel participatory opportunities that reinforce stakeholder engagement while driving technical innovation in LLMHS. This dual perspective highlights the potential of LLMHS to align technical sophistication with real-world healthcare demands, paving the way for next-generation health systems.
Rates of stress and anxiety are alarmingly high in university communities, but most people do not receive treatment. Mobile health (mHealth) interventions show promise to improve psychological symptoms and increase access to interventions, but little is known about their effects in the moment. The present study evaluated the short-term impact of brief mHealth sessions to determine which intervention features are associated with the greatest momentary self-reported improvements. Participants (N = 100 undergraduate students, graduate students, and university staff members) completed brief training sessions 1–2 times daily of MASK, a new mobile application for the university community that uses Cognitive Bias Modification for Interpretations (CBM-I) to shift anxious thinking patterns. Training sessions varied based on stressor domain/topic selected and writing requirements, among other features. Linear mixed effects models were used to test whether stressor domain or writing requirements predict post-training: (1) momentary affect, (2) reappraisal self-efficacy, and (3) emotion regulation self-efficacy. Self-reported improvement in state affect, reappraisal self-efficacy, and emotion regulation self-efficacy occurred for six out of eight stressor domains. Additionally, training sessions requiring less (vs. more) writing were associated with greater positive changes in affect, but not reappraisal or emotion regulation self-efficacy. Stressor domain and writing requirements are associated with different in-the-moment cognitive and affective outcomes, pointing to the need to tailor mHealth programs to users’ specific needs and current stressors.
Social interactions are a fundamental part of daily life and play a critical role in well-being. As emerging technologies offer opportunities to unobtrusively monitor behavior, there is growing interest in using them to better understand social experiences. However, automatically detecting interactions-particularly via wearable devices-remains underexplored. Existing systems are often limited to controlled environments, constrained to in-person interactions, and rely on rigid assumptions such as the presence of two speakers within a fixed time window. These limitations reduce their generalizability to capture diverse real-world interactions. To address these challenges, we developed a real-time, on-watch system capable of detecting both in-person and virtual interactions. The system leverages transfer learning to detect foreground speech (FS) and infers interaction boundaries based upon FS and conversational cues like whispering. In a real-world evaluation involving 11 participants over a total of 38 days (Mean = 3.45 days, SD = 2.73), the system achieved an interaction detection accuracy of 73.18%. Follow-up with six participants indicated perfect recall for detecting interactions. These preliminary findings demonstrate the potential of our system to capture interactions in daily life-providing a foundation for applications such as personalized interventions targeting social anxiety.
Social anxiety, characterized by fear of social interactions and negative evaluation, is both pervasive and impairing. Passive sensing offers opportunities for timely detection and intervention, yet most prior work has emphasized trait-level anxiety, with limited success in predicting intra-day fluctuations in state anxiety. To address this gap, we studied 72 socially anxious students who used our smartwatch–smartphone system for an average of nine days. The smartwatch collected multimodal physiological and behavioral signals (e.g., heart rate, movement), while the smartphone delivered seven randomly timed ecological momentary assessments per day to capture state anxiety. These multimodal data formed the basis for anxiety modeling. We present a weakly supervised multimodal framework that trains a base model on a public dataset, then transfers learning and personalizes predictions via meta-learning. Our model outperformed baselines, achieving 72.1% balanced accuracy on our dataset and 66.9% on an external dataset, demonstrating potential to support just-in-time-adaptive interventions.
Effective patient-provider communication is crucial in clinical care, directly impacting patient outcomes and quality of life. Traditional evaluation methods, such as human ratings, patient feedback, and provider self-assessments, are often limited by high costs and scalability issues. Although existing natural language processing (NLP) techniques show promise, they struggle with the nuances of clinical communication and require sensitive clinical data for training, reducing their effectiveness in real-world applications. Emerging large language models (LLMs) offer a new approach to assessing complex communication metrics, with the potential to advance the field through integration into passive sensing and just-in-time intervention systems. This study explores LLMs as evaluators of palliative care communication quality, leveraging their linguistic, in-context learning, and reasoning capabilities. Specifically, using simulated scripts crafted and labeled by healthcare professionals, we test proprietary models (e.g., GPT-4) and fine-tune open-source LLMs (e.g., LLaMA2) with a synthetic dataset generated by GPT-4 to evaluate clinical conversations, to identify key metrics such as ‘understanding’ and ‘empathy’. Our findings demonstrated LLMs’ superior performance in evaluating clinical communication, providing actionable feedback with reasoning, and demonstrating the feasibility and practical viability of developing in-house LLMs. This research highlights LLMs’ potential to enhance patient-provider interactions and lays the groundwork for downstream steps in developing LLM-empowered clinical health systems.
In HIV care, a strong rapport between patient and provider is essential for strengthening trust, enhancing therapy adherence, and ultimately leading to improved health outcomes. As the adoption of digital interactions in HIV care via mobile health (mHealth) tools is emerging, maintaining rapport in these asynchronous text-based communications becomes a critical yet challenging task. In this paper, we analyze 1,740 messages from an mHealth platform, categorized by experienced clinicians as either ‘rapport-building’ or ‘information-only.’ We utilize linguistic analysis to uncover key attributes of rapport-building communication. This led to a set of machine learning (ML) models and Large Language Models (LLMs) capable of classifying these communication styles. Further, we propose the application of LLMs not only to identify but also to actively rewrite ‘information only’ messages into versions that enhance rapport building without compromising information integrity. Our research demonstrates potential advancements in HIV mHealth communication by integrating linguistic analysis with language models, leading to more effective patient-provider interactions.
While audio data shows promise in addressing various health challenges, there is a lack of research on on-device audio processing for smartwatches. Privacy concerns make storing raw audio and performing post-hoc analysis undesirable for many users. Additionally, current on-device audio processing systems for smartwatches are limited in their feature extraction capabilities, restricting their potential for understanding user behavior and health. We developed a real-time system for on-device audio processing on smartwatches, which takes an average of 1.78 minutes (SD = 0.07 min) to extract 22 spectral and rhythmic features from a 1-minute audio sample, using a small window size of 25 milliseconds. Using these extracted audio features on a public dataset, we developed and incorporated models into a watch to classify foreground and background speech in real-time. Our Random Forest-based model classifies speech with a balanced accuracy of 80.3%.
Effective patient-provider communication is crucial in clinical care, directly impacting patient outcomes and quality of life. Traditional evaluation methods, such as human ratings, patient feedback, and provider self-assessments, are often limited by high costs and scalability issues. Although existing natural language processing (NLP) techniques show promise, they struggle with the nuances of clinical communication and require sensitive clinical data for training, reducing their effectiveness in real-world applications. Emerging large language models (LLMs) offer a new approach to assessing complex communication metrics, with the potential to advance the field through integration into passive sensing and just-in-time intervention systems. This study explores LLMs as evaluators of palliative care communication quality, leveraging their linguistic, in-context learning, and reasoning capabilities. Specifically, using simulated scripts crafted and labeled by healthcare professionals, we test proprietary models (e.g., GPT-4) and fine-tune open-source LLMs (e.g., LLaMA2) with a synthetic dataset generated by GPT-4 to evaluate clinical conversations, to identify key metrics such as ‘understanding’ and ‘empathy’. Our findings demonstrated LLMs’ superior performance in evaluating clinical communication, providing actionable feedback with reasoning, and demonstrating the feasibility and practical viability of developing in-house LLMs. This research highlights LLMs’ potential to enhance patient-provider interactions and lays the groundwork for downstream steps in developing LLM-empowered clinical health systems.
Quality patient-provider communication is critical to improve clinical care and patient outcomes. While progress has been made with communication skills training for clinicians, significant gaps exist in how to best monitor, measure, and evaluate the implementation of communication skills in the actual clinical setting. Advancements in ubiquitous technology and natural language processing make it possible to realize more objective, real-time assessment of clinical interactions and in turn provide more timely feedback to clinicians about their communication effectiveness. In this paper, we propose CommSense, a computational sensing framework that combines smartwatch audio and transcripts with natural language processing methods to measure selected "best-practice'' communication metrics captured by wearable devices in the context of palliative care interactions, including understanding, empathy, presence, emotion, and clarity. We conducted a pilot study involving N=40 clinician participants, to test the technical feasibility and acceptability of CommSense in a simulated clinical setting. Our findings demonstrate that CommSense effectively captures most communication metrics and is well-received by both practicing clinicians and student trainees. Our study also highlights the potential for digital technology to enhance communication skills training for healthcare providers and students, ultimately resulting in more equitable delivery of healthcare and accessible, lower cost tools for training with the potential to improve patient outcomes.
This article explores the convergence of connectionist and symbolic artificial intelligence (AI), from historical debates to contemporary advancements. Traditionally considered distinct paradigms, connectionist AI focuses on neural networks, while symbolic AI emphasizes symbolic representation and logic. Recent advancements in large language models (LLMs), exemplified by ChatGPT and GPT-4, highlight the potential of connectionist architectures in handling human language as a form of symbols. The study argues that LLM-empowered Autonomous Agents (LAAs) embody this paradigm convergence. By utilizing LLMs for text-based knowledge modeling and representation, LAAs integrate neuro-symbolic AI principles, showcasing enhanced reasoning and decision-making capabilities. Comparing LAAs with Knowledge Graphs within the neuro-symbolic AI theme highlights the unique strengths of LAAs in mimicking human-like reasoning processes, scaling effectively with large datasets, and leveraging in-context samples without explicit re-training. The research underscores promising avenues in neuro-vector-symbolic integration, instructional encoding, and implicit reasoning, aimed at further enhancing LAA capabilities. By exploring the progression of neuro-symbolic AI and proposing future research trajectories, this work advances the understanding and development of AI technologies.
Background: Rectus abdominis flap coverage of high-risk perineal wounds following extirpative pelvic procedures can result in improved perineal outcomes. However, rectus abdominis flap harvest has morbidity associated with the donor site, including hernia or bulge development. The risk–benefit profile of mesh use in this scenario is not well-defined in the literature. Methods: We performed a retrospective chart review of all patients who underwent rectus abdominis flap coverage of pelvic defects at our institution during July 2012–January 2021. Patient characteristics and postoperative outcomes were assessed. Patients were stratified into groups based on whether mesh was used and whether primary fascial closure was achieved. Donor site outcomes were analyzed between groups. Results: One hundred consecutive patients were included. When considering all patients in whom primary fascial closure was achieved, the use of mesh did not significantly decrease rates of hernia development. Mesh use in this setting was associated with significantly greater rates of infection, requiring procedural intervention (12% versus 0%, P = 0.044). When considering all patients in whom mesh was used, primary fascial closure was associated with decreased rates of hernia development, and this trended toward significance (16.1% versus 0.0%, P = 0.058). Conclusions: When closing a pedicled rectus abdominis flap donor site, if primary fascial closure is achievable, the addition of mesh to reinforce the repair does not have an added benefit. Mesh use in this setting was not shown to prevent hernia or bulge development, and was found to be associated with significantly greater rates of infection, requiring procedural intervention.
During social interactions, understanding the intricacies of the context can be vital, particularly for socially anxious individuals. While previous research has found that the presence of a social interaction can be detected from ambient audio, the nuances within social contexts, which influence how anxiety provoking interactions are, remain largely unexplored. As an alternative to traditional, burdensome methods like self-report, this study presents a novel approach that harnesses ambient audio segments to detect social threat contexts. We focus on two key dimensions: number of interaction partners (dyadic vs. group) and degree of evaluative threat (explicitly evaluative vs. not explicitly evaluative). Building on data from a Zoom-based social interaction study (N=52 college students, of whom the majority N=45 are socially anxious), we employ deep learning methods to achieve strong detection performance. Under sample-wide 5-fold Cross Validation (CV), our model distinguished dyadic from group interactions with 90% accuracy and detected evaluative threat at 83%. Using a leave-one-group-out CV, accuracies were 82% and 77%, respectively. While our data are based on virtual interactions due to pandemic constraints, our method has the potential to extend to diverse real-world settings. This research underscores the potential of passive sensing and AI to differentiate intricate social contexts, and may ultimately advance the ability of context-aware digital interventions to offer personalized mental health support.
Wearable smart devices are capable of capturing a variety of information from their users using a multitude of noninvasive sensing modalities. Using features from the raw measurements of wearable devices, sensor fusion enables us to obtain a holistic picture of the users’ context and monitor their activity state with increased accuracy. Human activity recognition using noninvasive sensors allows us to capture the natural behavior of users in their day-to-day lives. This in-the-wild activity recognition, however, poses several key challenges that must be addressed to create effective classification models. The main challenges are class imbalance, uncertainty in classifier decisions, and large feature spaces. To address them, this study further explores a probabilistic sensor fusion method called Naive Adaptive Probabilistic Sensor (NAPS) Fusion. In doing so, we establish the viability of NAPS Fusion for natural human activity recognition using noninvasive sensing modalities. NAPS Fusion handles dimensionality reduction by creating reduced feature sets and mitigates the class imbalance issue through the use of Synthetic Minority Oversampling Technique (SMOTE). Moreover, NAPS Fusion addresses uncertainty in the decisions of classifiers using a Dempster-Shafer theoretic late fusion framework. Our empirical evaluation demonstrates that NAPS Fusion has broad applications beyond its original design for cognitive state detection. It outperforms similar decision level sensor fusion methods (late fusion using averaging, LFA, and late fusion using learned weights, LFL) in the detection of exercise and sedentary activities such as walking, running, lying down, and sitting. We observe improvements of up to 56% in F1 score and up to 59% in precision with NAPS Fusion over the compared methods.
Emotion regulation (ER) diversity, defined as the variety, frequency, and evenness of ER strategies used, may predict social anxiety severity. In a sample of individuals with high (n = 113) and low (n = 42) social anxiety severity, we tested whether four trait ER diversity metrics predicted group membership. We generalized existing trait ER diversity calculations to repeated measures data to test whether state-level metrics (using 2 weeks of ecological momentary assessment [EMA] data) predicted social anxiety severity within the higher severity group. As hypothesized, higher trait ER diversity within avoidance-oriented strategies predicted greater likelihood of belonging to the higher severity group. At the state level, higher diversity across all ER strategies, and within and between avoidance- and approach-oriented strategies, predicted higher social anxiety severity (but only after analyses controlled for number of submitted EMAs). Only diversity within avoidance-oriented strategies was significantly correlated across trait and state levels. Findings suggest that high avoidance-oriented ER diversity may co-occur with higher social anxiety severity.
Wearable devices with embedded sensors can provide personalized healthcare and wellness benefits in digital phenotyping and adaptive interventions. However, the collection, storage, and transmission of biometric data (including processed features rather than raw signals) from these devices pose significant privacy concerns. This quantitative, data-driven study examines the privacy risks associated with wearable-based digital phenotyping practices, with a focus on user reidentification (ReID), which is the process of identifying participants’ IDs from deidentified digital phenotyping datasets. We propose a machine-learning-based computational pipeline to evaluate and quantify model outcomes under various configurations, such as modality inclusion, window length, and feature type and format, to investigate the factors influencing ReID risks and their predictive trade-offs. This pipeline leverages features extracted from three wearable sensors, resulting in up to 68.43% accuracy in ReID risk for a sample size of N=45 socially anxious participants based on only descriptive features of 10-second observations. Additionally, we explore the trade-offs between privacy risks and predictive benefits by adjusting various settings (e.g., the ways to process extracted features). Our findings highlight the importance of privacy in digital phenotyping and suggest potential future directions.