Self-experimentation, or using tracked data to systematically answer health and wellbeing questions via hypothesis testing, has significant potential to support personal health. However, technological support for self-experimentation has focused on expert-designed self-experiments for specific health conditions, limiting people’s ability to design their own rigorous experiments. To address this gap, we developed CASEbot (Conversation Agent for Self-Experimentation), an LLM-powered chatbot using a theory-driven approach to guide users through designing well-structured, personalized, and safe self-experiments. We conducted a within-subjects, mixed-methods study with 42 participants comparing CASEbot to a traditional worksheet-based approach. When formally comparing the experiment rigor and specificity, most participants designed better experiments using CASEbot. They appreciated CASEbot’s conversational approach, which prompted them to surface everyday constraints and proactively raised safety concerns, but some found the platform too rigid in its recommendations. We discuss opportunities for future generative AI self-experimentation systems for health to balance structured guidance with user autonomy.
Family caregivers of individuals with Alzheimer’s Disease and Related Dementia (AD/ADRD) face significant emotional and logistical challenges that place them at heightened risk for stress, anxiety, and depression. Although recent advances in generative AI—particularly large language models (LLMs)—offer new opportunities to support mental health, little is known about how caregivers perceive and engage with such technologies. To address this gap, we developed Carey, a GPT-4o–based chatbot designed to provide informational and emotional support to AD/ADRD caregivers. Using Carey as a technology probe, we conducted semi-structured interviews with 16 family caregivers following scenario-driven interactions grounded in common caregiving stressors. Through inductive coding and reflexive thematic analysis, we surface a systemic understanding of caregiver needs and expectations across six themes— on-demand information access, safe space for disclosure, emotional support, crisis management, personalization, and data privacy . For each of these themes, we also identified the nuanced tensions in the caregivers’ desires and concerns. We present a mapping of caregiver needs, AI chatbots’ strengths, gaps, and design recommendations. Our findings offer theoretical and practical insights to inform the design of proactive, trustworthy, and caregiver-centered AI systems that better support the evolving mental health needs of AD/ADRD caregivers.
Background Critical flicker frequency (CFF) is a well-validated neurophysiologic screening test for minimal hepatic encephalopathy (MHE) in cirrhosis. We evaluated the feasibility and reliability of repeated at-home, self-administered CFF testing using a novel device (Beacon). Methods We enrolled 21 participants with cirrhosis. Each received a Beacon device and smartphone preloaded with an application to measure CFF using two protocols: 1) method of limits–descending (MOL-D) and 2) Forced Choice. Participants were instructed to complete ≥5 tests per week for 2 weeks, followed by ≥1 test per week for an additional 4 weeks. Outcomes included adherence, test duration, and longitudinal CFF trends to assess potential learning effects from repeated testing. Results Participants were predominantly male (76%) with a mean age of 52.2 years; 43% had alcohol-related liver disease and 38% had prior overt HE. Adherence (completing ≥70% of recommended measurements) was observed in 19 (90.5%) for daily and 16 participants (76.2%) for weekly testing. Median CFF was 42.9 Hz (IQR: 38.4-44.4) by MOL-D and 43.7 Hz (IQR: 41.7-45.8) by Forced Choice. Forced Choice CFF was significantly lower in those with prior HE (median 40.7 Hz vs. 44.9 Hz; p = 0.02), while MOL-D CFF did not differ by HE status (42.7 Hz vs. 42.9 Hz; p = 0.44). CFF remained stable over time for both protocols (MOL-D: +0.01 Hz/day of testing, p=0.67; Forced Choice: +0.03 Hz/day, p=0.09). Forced Choice testing required more time than MOL-D to complete (median 4.67 vs. 2.48 minutes; p < 0.001). Conclusion Beacon enabled reliable, self-administered CFF testing at home with high adherence and no evidence of learning effects. Further studies incorporating other validated MHE tools are needed to determine whether at-home CFF monitoring can detect MHE early and prevent overt HE episodes and hospitalizations.
Family members caring for individuals with Alzheimer's disease and related dementias (AD/ADRD) provide the foundation of long-term care worldwide. In 2023, more than 11 million U.S. family and friends contributed 18 billion hours of unpaid care, often at the cost of their own physical and mental health. These informal caregivers – also referred as the "invisible second patients" – experience elevated rates of mental health problems. Yet research commonly reduces their complex psychosocial experiences to a single construct of caregiver burden, obscuring which specific needs are unmet or effectively supported. At the same time, digital and AI-enabled technologies are rapidly expanding, from smartphone apps and videoconferencing to sensor platforms and AI chatbots. However, the absence of shared frameworks across medicine, psychology, and technology research limits cumulative progress. This study introduces a Caregiver Mental Health and Technology Taxonomy that systematically links AD/ADRD caregiver needs with corresponding classes of technology-based interventions. Drawing from an interdisciplinary literature review and two qualitative studies with caregivers, the taxonomy identifies mismatches between caregiver priorities and existing technological support, highlights under-served domains such as relational strain and compassion fatigue, and proposes design directions for adaptive, responsive systems. The framework offers a shared vocabulary to guide clinicians, researchers, and technology designers in developing more person-centered and clinically grounded innovation in dementia care.
Caregivers seeking AI-mediated support express complex needs – information-seeking, emotional validation, and distress cues – that warrant careful evaluation of response safety and appropriateness. Existing AI evaluation frameworks, primarily focused on general risks (toxicity, hallucinations, policy violations, etc), may not adequately capture the nuanced risks of LLM-responses in caregiving-contexts. We introduce RubRIX (Rubric-based Risk Index), a theory-driven, clinician-validated framework for evaluating risks in LLM caregiving responses. Grounded in the Elements of an Ethic of Care, RubRIX operationalizes five empirically-derived risk dimensions: Inattention, Bias Stigma, Information Inaccuracy, Uncritical Affirmation, and Epistemic Arrogance. We evaluate six state-of-the-art LLMs on over 20,000 caregiver queries from Reddit and ALZConnected. Rubric-guided refinement consistently reduced risk-components by 45-98
Caregivers often turn to online communities for informational and emotional support. In these spaces, peer supporters frequently draw on personal narratives to respond to emotionally complex caregiving situations. As LLMs are increasingly designed as peer-like sources of support, they introduce a critical tension: AI can provide immediate, private, and nonjudgmental support, but it cannot authentically possess the lived experiences that make human peer support meaningful. Yet, when prompted to sound peer-like, LLMs may generate language that implies lived experience. This creates a synthetic lived experience paradox: the same experiential language that may make AI support feel warm, relatable, and peer-like can also falsely position the system as someone with lived experience. We examine this paradox in the context of family caregivers of people living with Alzheimer's Disease and Related Dementias (ADRD). Drawing on caregiver support exchanges from online communities and prompted peer-like responses from three LLMs – LLaMA, GPT-4o-mini, and MedGemma – we analyze how human peers use personal narratives and how AI incorporates similar narrative forms. Psycholinguistic analysis shows that peer responses used significantly more first-person and past-focused language than peer-like AI responses. Qualitatively, we identify seven types of personal narratives in human peer support and show that AI often captures their emotional work, but can fabricate experiential grounding. These findings reveal a narrative authenticity gap: peer-like AI can generate synthetic lived experience without the real experience that makes peer support meaningful. We argue that caregiver-support AI systems need mechanisms to distinguish supportive peer-like framing from fabricated lived experience, ensuring that models can offer warmth and validation without falsely positioning themselves as experiential peers.
Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond information-seeking: caregivers seek emotional reassurance, guidance, and help, while navigating uncertain, relationally complex care decisions. Yet most safety evaluations assess model behavior under generic prompts, leaving a critical question unexamined: does a model's safety profile change with its support role? We study this by operationalizing four expert-reviewed support roles grounded in social support theory: Inform, Coach, Relate, and Listen, and comparing them against two baseline controls: a basic prompting condition and a retrieval-augmented generation (RAG) condition. We evaluate across three language models (GPT-4o-mini, Llama-3.1-8B-Instruct, and MedGemma-1.5-4b-it) on 5,000 real-world queries from online Alzheimer's Disease and Related Dementias (ADRD) communities. We find that the LLM's support role systematically shapes both the prevalence and composition of interactional risks. Furthermore, a human evaluation study reveals a perceived quality–safety tension: more directive, information-oriented roles are rated as more helpful and trustworthy despite exhibiting elevated interactional risk profiles. We release 90,000 support role-conditioned model responses with risk annotations as an ecologically grounded resource for research on safer LLM-mediated conversational support.
Foundation models tested for clinical practice using human-designed metrics may mask fundamental differences in information processing. We investigated this using the clock drawing test (CDT), a cognitive screening tool. Three foundation models achieved 94% accuracy on conventional metrics, matching experts. However, upon decomposing the CDT into 24 questions across five cognitive domains, results diverged significantly. In cases with unanimous model agreement, they still disagreed with human raters in 22% cases. Performance varied drastically with 88% alignment with humans on rule-based executive questions but only 46% on context-dependent anticipatory thinking questions. We observed that models abstained three times more than humans, primarily owing to poor data quality. These findings show standard clinical evaluation metrics fail to capture how foundation models process information. High aggregate accuracy obscures component-level failures. We contribute a systematic evaluation of frontier models’ healthcare capabilities, demonstrate theory-driven task decomposition, and discuss design implications for better human-AI collaborative systems.
BACKGROUND:Technologies animated by artificial intelligence (AI) and machine learning proliferate in nursing's work and health care. Rapidly evolving AI technologies demand ethics and research infrastructure responsive to this dynamic landscape. PURPOSE:We founded Health Tech for the People (HT4P) to develop a critical technology ethics for more accountable, human-centered design of AI technologies. METHODS:Our HT4P framework, grounded in data feminism and design justice principles, guided projects across two priority areas: reproductive health and aging technologies. DISCUSSION:Outcomes of HT4P's first year included an ethics fellowship, five transdisciplinary symposia/workshops, community partnerships and a community-directed technology ethics seminar, and multiple projects-in-progress. CONCLUSION:Nurses have an opportunity to cultivate a radical imagination for more just and careful tech futures. This requires us to develop ethics of technology that puts values in practice by redistributing power, acknowledging the invisible and undervalued labor and resources, and repairing long-standing injustices.
Chronic liver disease can lead to neurological conditions that result in coma or death. Although early detection can allow for intervention, testing is infrequent and unstandardized. Beacon is a device for at-home patient self-measurement of cognitive function via critical flicker frequency, which is the frequency at which a flickering light appears steady to an observer. This paper presents our efforts in iterating on Beacon’s hardware and software to enable at-home use, then reports on an at-home deployment with 21 patients taking measurements over 6 weeks. We found that measurements were stable despite being taken at different times and in different environments. Finally, through interviews with 15 patients and 5 hepatologists, we report on participant experiences with Beacon, preferences around how CFF data should be presented, and the role of caregivers in helping patients manage their condition. Informed by our experiences with Beacon, we further discuss design implications for home health devices.
AI chatbots are increasingly integrated into various sectors, including healthcare. We examine their role in responding to queries related to Alzheimer’s Disease and Related Dementias (AD/ADRD). We obtained real-world queries from AD/ADRD online communities (OC)—Reddit (r/Alzheimers) and ALZConnected. First, we conducted a small-scale qualitative examination where we prompted ChatGPT, Bard, and Llama-2 with 101 OC posts to generate responses and compared them with OC responses through inductive coding and thematic analysis. We found that although AI can provide emotional and informational support like OCs, they do not engage in deeper conversations, provide references, and share personal experiences. These insights motivated us to conduct a large-scale quantitative examination of comparing AI (GPT) and OC responses (90K) to 13.5K posts, in terms of psycholinguistics, lexico-semantics, and content. AI responses tend to be more verbose, readable, and complex. AI responses exhibited greater empathy, but more formal and analytical language, lacking personal narratives and linguistic diversity. We found that various LLMs, including GPT, Llama, and Mistral, exhibit consistent patterns in responding to AD/ADRD-related queries, underscoring the robustness of our insights across LLMs. Our study sheds light on the potential of AI in digital health and underscores design considerations of AI to complement human interactions.
Alzheimer's Disease and Related Dementias (AD/ADRD) are progressive neurodegenerative conditions that impair memory, thought processes, and functioning. Family caregivers of individuals with AD/ADRD face significant mental health challenges due to long-term caregiving responsibilities. Yet, current support systems often overlook the evolving nature of their mental wellbeing needs. Our study examines caregivers' mental wellbeing concerns, focusing on the practices they adopt to manage the burden of caregiving and the technologies they use for support. Through semi-structured interviews with 25 family caregivers of individuals with AD/ADRD, we identified the key causes and effects of mental health challenges, and developed a temporal mapping of how caregivers' mental wellbeing evolves across three distinct stages of the caregiving journey. Additionally, our participants shared insights into improvements for existing mental health technologies, emphasizing the need for accessible, scalable, and personalized solutions that adapt to caregivers' changing needs over time. These findings offer a foundation for designing dynamic, stage-sensitive interventions that holistically support caregivers' mental wellbeing, benefiting both caregivers and care recipients.
BACKGROUND:Alzheimer disease (AD) is the leading type of dementia, demanding comprehensive understanding and intervention strategies. In the United States, where over 6 million people are impacted, the prevalence of AD and related dementias (AD/ADRD) presents a growing public health challenge. However, individuals living with AD/ADRD and their caregivers frequently express feelings of marginalization, describing interactions characterized by perceptions of patient infantilization and a lack of respect. OBJECTIVE:This study aimed to address 2 key research questions (RQs). For RQ1, we investigated the needs and concerns expressed by participants in online social communities focused on AD/ADRD, specifically on 2 platforms-Reddit's r/Alzheimers and ALZConnected. For RQ2, we examined the prevalence and distribution of social support corresponding to these needs and concerns, and the association between these needs and received support. METHODS:We collected 13,429 posts and comments from the r/Alzheimers subreddit spanning July 2014 to November 2023, and 90,113 posts and comments from ALZConnected between December 2020 (the community's earliest post) and November 2023. We conducted topic modeling using latent Dirichlet allocation (LDA), followed by labeling to identify the major topical themes of discussions. We used transfer learning classifiers to identify the occurrences of emotional support (ES) and informational support (IS) in the comments (or responses) in the discussions. We built regression models to examine how various topical themes are associated with the kinds of support received. RESULTS:Our analysis revealed a diverse range of topics reflecting community members' varying needs and concerns of individuals affected by AD/ADRD. These themes encapsulate the primary discussions within the online communities: memory care, nursing and caregiving, gratitude and acknowledgment, and legal and financial considerations. Our findings indicated a higher prevalence of IS compared to ES. Regression models revealed that ES primarily occurs in posts relating to nursing and caring, and IS primarily occurs in posts concerning medical conditions and diagnosis, legal and financial, and caregiving at home. CONCLUSIONS:This study reveals that online communities dedicated to AD/ADRD support engage in discussions on a wide range of topics, such as memory care, nursing, caregiving, and legal and financial challenges. The findings shed light on the key pain points and concerns faced by individuals managing AD/ADRD in their households, revealing how they leverage online platforms for guidance and support. These insights underscore the need for targeted institutional and social interventions to address the specific needs of AD/ADRD patients, caregivers, and other family members.
Collaborative care management is an evidence-based approach to integrated psychosocial care for patients with comorbid cancer and depression. Prior work highlights challenges in patient-provider collaboration in navigating parallel cancer care and psychosocial care journeys of these patients. We design and deploy SCOPE, a platform for technology-enhanced collaborative care combining a patient-facing mobile app with a provider-facing registry. We examine SCOPE through a total of 45 interviews with patients and providers conducted in SCOPE's 15 months of design and development and 24 months of SCOPE's deployment for actual care in 6 cancer clinics. We find that: (1) SCOPE supported patient engagement in its underlying collaborative care and behavioral activation interventions, (2) patient-generated data in SCOPE improved patient-provider collaboration between and within in-person sessions, (3) SCOPE supported providers in delivering care and improved care team collaboration, (4) experience with SCOPE created evolving expectations for collaboration around data, and (5) SCOPE's deployment in actual care surfaced important implementation barriers. We discuss the implications of our findings in terms of designing for engagement with behavioral health interventions, negotiating patient data sharing and provider responsiveness, supporting personalized self-tracking goals in evidence-based interventions, exploring the role of digital health navigators in technology-enhanced care, and the need for flexibility in aligning technology-supported interventions to patient needs.
Background and Objectives:To better support aging in place, we first must understand the needs of the older adult population. We conducted a systematic review to understand the needs of older adults in the home. Research Design and Methods:We queried the PubMed, CINAHL, and ProQuest databases to identify literature related to needs assessments of older adults in the home. Records were included if: (1) the population focused on older adults (aged 65 years and older); (2) a needs assessment was conducted; (3) the older adult population was aging in place and not in a long-term care facility; (4) English language publication; (5) published since 2013; and (6) pertaining solely to older adult caregivers' needs. The needs identified in each article were extracted and categorized based on emergent themes. Results:A total of 1,963 records were identified. After removing duplicate records and those not meeting the inclusion criteria, 65 articles were included in the final analysis. Six need-related theme domains were identified: health management needs; social needs; homecare and practical needs; information needs; technology needs; and healthcare system needs. Discussion and Implications:Through the systematic review, we identified a wide range of unmet needs for older adults aging in the home. The unmet needs of older adults are multifaceted and provide ideal targets for the development of novel technological solutions. In particular, recent advances in artificial intelligence (AI), especially generative AI such as large language models (LLMs), surface the potential for technology to address unmet needs across multiple domains. We discuss the potential for AI to lower barriers to technology uptake for older adults and create novel solutions to each of the need domains identified. Ultimately, AI-enabled solutions may increase independence for older adults and potentially increase the ability to age in place.
We explore some of the existing research in HCI around technology for older adults and examine the role of LLMs in enhancing it. We also discuss the digital divide and emphasize the need for inclusive technology design. At the same time, we also surface concerns regarding privacy, security, and the accuracy of information provided by LLMs, alongside the importance of user-centered design to make technology accessible and effective for the elderly. We show the transformative possibilities of LLM-supported interactions at the intersection of aging, technology, and human-computer interaction, advocating for further research and development in this area.
EEG signals contain highly sensitive information about an individual's mental state, cognitive processes, and health conditions, making privacy preservation crucial. With the rise of commercial headwear capable of capturing EEG signals, developing robust mechanisms for ensuring privacy of such data is imperative. This work aims to protect EEG data privacy in cloud-based processing systems by sending intermediate output after neural network layer splitting to the cloud. We propose a novel holistic Combined Privacy Metric (CPM) that quantifies privacy leakage between raw EEG signals and intermediate outputs. Our study focuses on EEG-based seizure detection using a 1D CNN architecture, achieving accuracy of 96.25%. We evaluate various splitting configurations to optimize the trade-off between privacy preservation and computational efficiency. We find that splitting after the second convolutional layer achieves a CPM of 0.82 with a modest client-side model size of 509kB. This approach significantly enhances EEG data privacy while enabling effective cloud-based analysis, potentially facilitating wider adoption of secure EEG technologies in healthcare and research applications.
Research at the intersection of human-computer interaction (HCI) and health is increasingly done by collaborative cross-disciplinary teams. The need for cross-disciplinary teams arises from the interdisciplinary nature of the work itself-with the need for expertise in a health discipline, experimental design, statistics, and computer science, in addition to HCI. This work can also increase innovation, transfer of knowledge across fields, and have a higher impact on communities. To succeed at a collaborative project, researchers must effectively form and maintain a team that has the right expertise, integrate research perspectives and work practices, align individual and team goals, and secure funding to support the research. However, successfully operating as a team has been challenging for HCI researchers, and can be limited due to a lack of training, shared vocabularies, lack of institutional incentives, support from funding agencies, and more; which significantly inhibits their impact. This workshop aims to draw on the wealth of individual experiences in health project team collaboration across the CHI community and beyond. By bringing together different stakeholders involved in HCI health research, together, we will identify needs experienced during interdisciplinary HCI and health collaborations. We will identify existing practices and success stories for supporting team collaboration and increasing HCI capacity in health research. We aim for participants to leave our workshop with a toolbox of methods to tackle future team challenges, a community of peers who can strive for more effective teamwork, and feeling positioned to make the health impact they wish to see through their work.