
While AI-enabled medical devices (AIeMD) are redefining healthcare, safety relies on human-AI interaction as much as algorithmic performance. This systematic review of 53 studies reveals that 58.2% of research omits Explainable AI (XAI) methods, despite transparency being a regulatory and safety necessity. Results identify a "symmetry of modality": qualitative interviews correlate with written text explanations, while Think-Aloud protocols better assess cognitively demanding tools like SHAP values. Currently, AI-specific risks like automation bias, driven by algorithmic opacity, are critically under-reported. To ensure safe adoption, practitioners must move beyond "user satisfaction" to treat transparency and XAI as safety-critical requirements for trust calibration.
Prediction of clinical deterioration in hospital settings is essential for patient safety. While there are tools, including AI tools, for predicting deterioration in acute-care facilities, there is little developed on predicting deterioration in post-acute hospitals. These environments present challenges for use of AI predictive tools, as they typically have poor data foundations and lower digital maturity. We developed an AI model that predicts clinical deterioration in a post-acute hospital with a sensitivity of 81.5%. Analysis of the variables that most contribute to prediction of deterioration reveals a strikingly physiologic pattern, consistent with known predictors of deterioration in acute care. This both validates the clinical plausibility of our findings and also provides novel evidence that acute-care predictors of deterioration may hold more broadly in the post-acute environment. We discuss the approach to developing an AI predictive model in a digitally-immature environment, including the use of workflow analysis and a modified data analytics approach.
Artificial intelligence (AI) systems, particularly large language models (LLMS) or chatbots, have been gaining increased popularity in recent years. Users will sometimes use these models for personal mental health assistance. This review aimed to expand the topic and to further understanding of AI’s effectiveness in these use cases. A total of 51 different research articles were examined, leading to themes of interaction quality, explainability, trust, ethical concerns, and anthropomorphism being discussed. In addition to the themes, a list of eight best practices was created to assist developers with designing AI systems in a way that would reduce the overall risk of harm for users attempting to use their AI for mental health cases. While AI is no substitute for a clinician, there are areas in which these systems can perform well and help foster the clinical relationship, but proper user experience design principles should be maintained throughout the process to ensure efficacy and user safety.
This work explores the potential transfer of military and professional healthcare XR training techniques to civilian settings. This is accomplished by briefly discussing key insights gleaned from military and professional healthcare applications of XR training technologies. Based on these insights, recommendations surrounding civilian lay-provider applications for this technology are provided, as well as practical considerations regarding its adoption. The lack of transfer of these XR technologies into basic medical training of the civilian lay-providers showcases the inequality in adoption of the practice. With the boom of XR exploration in healthcare and military settings, it stands that there are areas outside of professional use that may benefit from the tool. The implications of which include the potential for a greater pursuit of life-saving knowledge and more accessible training opportunities for civilian lay-providers.
The advent of artificial intelligence (AI) in healthcare operations is both promising and perilous. This paper reflects on the importance of experience, defined as our on-going, direct interactions with the world that enhance our abilities and establish useful habits, and then discusses potential implications of AI as a mediator to our experience. A case report describing an instance of diagnostic excellence is offered to illustrate the potential for AI-based passive notetaking systems to interfere with the development of mindful notetaking skills. With careful study, it is our hope that AI systems are designed to favor experience, and provide opportunities for the development of curiosity, intellect, skill, expertise, and fulfillment for current and future generations of caregivers. It is our peril that the design of AI systems will supplant experience, leaving us with a dull, automated, numb, and hollow existence.
Iconic hand gestures paired with speech can enhance word comprehension and are candidates for caregiver delivered interventions in rural settings where access to language services is limited. Whether children’s brains process gestures as communicative semantic content, and whether this differs between typically developing and neurodivergent children, is not well characterized. We recorded 20 channel EEG at 500 Hz from 13 participants across four groups (typically developing four year olds, typically developing older children, older neurodivergent children with ASD, ADHD, or both, and adults) during a passive viewing trifecta paradigm in which each spoken noun appeared with a matching or mismatching animation (Time 0) or iconic gesture (Time 1). After FASTER style preprocessing and epoching time-locked to the final target stimulus of each trial, we examined the N400 response over an eight channel centro-parietal and temporal region of interest. In Gesture word last trials, match and mismatch waveforms diverged approximately 300 ms after speech onset with the expected direction in typically developing older children (d = 0.34), neurodivergent older children (d = 0.78), and typical four year olds (d = 0.36), suggesting that gesture and speech pairings yielded the most consistent semantic differentiation across groups, evidence for semantic processing and integration; the largest single effect was observed for Animation in four year olds (d = 1.80). Animation last and gesture last trials did not show the same canonical pattern. Per word amplitudes in the N400 time window identified Goat, Umbrella, and Emoji as more stable items and Lock, Flag, and Rectangle as more variable. These preliminary findings from eight participants and 251 retained epochs support the trifecta paradigm for studying gesture and speech integration across development and motivate larger studies with normed stimuli.
Artificial intelligence (AI) systems are now used in many areas of healthcare, but stakeholders in the care like clinicians, patients and caregivers still experience mistrust, confusion, and uncertainty when AI-supported recommendations appear in their care. Traditional explainable AI (XAI) methods focus on showing how the algorithm works, which may help developers but often does not support stakeholders in real clinical situations. Human-centered AI research shows that explanations need to match user tasks, mental models, and decision needs. And the SEIPS 2.0 and SEIPS 3.0 frameworks offer a clear structure for placing explanation strategies across the patient and caregiver journey. In this report, we extend two non-algorithmic approaches called Collaborative Explainable AI (CXAI) and Cognitive Tutorials, and we use SEIPS 2.0 and SEIPS 3.0 to guide where and how these methods should be used. We also bring evidence from qualitative studies showing perspective of non-clinical stakeholders on AI in the journey, that patients prefer explanation through conversation, and that caregivers often struggle when AI outputs are unclear or incomplete. We present two short vignettes and a design guide to help human factors researchers create explanation systems that are practical and trustworthy. Our goal is to show how explainability can become part of the care system instead of being treated as a separate technical feature.
This study analyzes 28 critical incidents reported by healthcare providers caring for children with medical complexity during the COVID-19 pandemic. Four themes emerged: barriers to care, ethical dilemmas related to safety and risk, challenges of remote and modified care, and impacts on families and provider wellbeing. Findings highlight the importance of balancing infection control with continuity of care, strengthening coordination with families, and supporting resilient healthcare systems during crises.
Rapid advances in digital health technologies—including machine learning–enabled diagnostics and wearable devices—are transforming care pathways through real‑time monitoring and personalized interventions. However, static evidence requirements struggle to keep pace with products that update frequently and learn from real‑world use. This extended abstract examines the current role of the National Institute for Health and Care Excellence (NICE) Evidence Standards Framework (ESF), identifies gaps that hinder the timely adoption of high‑value innovations, and proposes a dynamic, stakeholder‑driven approach to evidence generation and validation. Recommendations focus on integrating adaptive real‑world evidence (RWE) methods, continuous post‑market learning, and transparent governance so that the NHS can safely accelerate adoption while maximizing patient benefit.
“Device design” issues have been a leading cause of Class I medical device recalls in recent years, highlighting a system-level gap despite established FDA Human Factors (HF) standards. This study analyzes recall patterns to identify recurring usability failures and limitations in current HF practices. It proposes a proactive, Closed-Looped Post-Market Surveillance approach to detect use-related issues earlier, strengthen design risk controls, and reduce patient harm. The study also emphasizes collaboration between device manufacturers and healthcare providers to enable continuous real-world feedback and improve patient safety beyond regulatory compliance.
Medical device manufacturers in the U.S. routinely design devices with consideration for user characteristics such as hand dexterity, cognitive and vision impairment, and related disease state comorbidities. We have expanded these considerations during the device development process to include language as an additional distinct user characteristic within the U.S. and to ensure representation of non-English-speaking individuals when they are part of the intended user population. This has included adapting product labeling and training materials to account for idiomatic expressions needed for the interpreted language, device-specific terminology from the original language, and culturally appropriate practices applicable to study material development and participant interviews. This paper presents lessons learned from developing non-U.S. English materials using back translation, recruiting non-English-speaking participants, and conducting research to support shared understanding among participants, moderators, and clients. Although English materials are often assumed to be sufficient for the U.S. market, our findings suggest that well-developed materials in additional languages should be considered essential when products are intended for non-English-speaking users. More broadly, this paper highlights the importance of conducting usability studies that account for native language differences, as consideration of these user characteristics is essential to support safe and effective use across diverse populations.
Near misses occur frequently in perioperative care but are inconsistently reported, limiting their contribution to system-level safety improvement. A cross-sectional survey was administered to anesthesia providers at a single academic center to assess near miss experience, reporting behavior, perceived barriers, and event classification accuracy using clinical vignettes. Forty-nine respondents (17.3% response rate) participated. Nearly all respondents (98.0%) reported exposure to a near miss, while 42.9% indicated that such events were reported. The most identified barriers included reporting process complexity (55.1%), uncertainty about reporting procedures (42.9%), and ambiguity in event classification (38.8%). Scenario-based assessment demonstrated substantial variability in distinguishing near miss from no-harm events, particularly in routine or borderline clinical scenarios. These findings define key areas for targeted intervention to improve near miss reporting.