The use of artificial intelligence (AI) to draft responses to patient portal messages has been proposed to reduce provider in-basket burden. However, little is known about its effects on patient–provider communication. In this retrospective observational study, we evaluated demographic differences in the tone of AI-generated draft replies (AI-GDRs) and care team responses to patient messages. Our study included 12,202 message triads comprising patient messages, AI-GDRs, and care team responses from three internal and family medicine practices in New York City. We found differences in tone across patient demographics. AI-GDRs had lower odds of including polite language in responses to Hispanic patients, compared to White patients. AI-GDRs also had lower odds of conveying positive affect in responses to patients who were Hispanic, preferred a non-English language, assigned female sex at birth, or lived in an area with a lower average income. Some of these tone differences were also found in care team responses. Physicians and advanced practice providers had lower odds of conveying positive affect when responding to patients who were Hispanic or preferred a non-English language. Our findings highlight that careful implementation of AI drafting is needed to ensure that this technology does not introduce or amplify inequities in patient–provider communication.
About one-third of the computed tomography (CT) scans ordered yearly to evaluate for pulmonary embolism (PE) in emergency departments (ED) in the U.S. are avoidable. Clinical guidelines recommend the use of validated PE prediction rules which reduce CT scan ordering without an increase in missed PEs, but there is low provider adoption. Clinical decision support (CDS) that incorporates these rules along with nudges (subtle, non-coercive influences on decision-making) may improve provider adoption. In our pilot trial of a PE risk CDS tool with a nudge at order entry, adoption was significantly higher (39.1
BACKGROUND:Nonadherence to antihypertensive medications is common. Mobile health (mHealth)-based behavioral economic interventions may improve adherence, but remain largely untested, especially in vulnerable populations. OBJECTIVE:The study sought to test whether an mHealth incentive lottery would lower systolic blood pressure (SBP) and improve adherence. METHODS:BETTER-BP (Behavioral Economics Trial To Enhance Regulation of Blood Pressure) was a randomized trial conducted in 3 safety-net clinics in New York City. Eligible participants were adults with hypertension prescribed at least 1 antihypertensive medication, with SBP >140 mm Hg, and poor self-reported adherence. In the intervention arm, an incentive lottery was administered via SMS messaging. All participants received passive adherence monitoring. The intervention lasted 6 months, with continued monitoring until 12 months. The primary clinical endpoint was change in SBP at 6 months. The primary process endpoint was adequate antihypertensive medication adherence (≥80% days adherent) from baseline to 6 months. RESULTS:Four-hundred participants (265 intervention:135 control) were enrolled with median age 57 years, 60.5% women, 61.5% Hispanic, and 20.3% non-Hispanic Black. Over 70% had Medicaid or no insurance. At 6 months, intervention arm participants were twice as likely to achieve adequate adherence (71% vs 34%; adjusted risk ratio: 2.04; 95% CI: 1.58-2.63), but there was no significant change in mean SBP (-6.7 mm Hg intervention vs -5.8 mm Hg control; P = 0.62). From 6 to 12 months, adherence was similar (31% intervention vs 26% control; adjusted risk ratio: 1.17; 95% CI: 0.83-1.65). CONCLUSIONS:In a diverse safety-net population, the BETTER-BP intervention doubled the rate of adequate antihypertensive medication adherence but did not reduce SBP at 6 months.
The integration of generative AI (GenAI) in patient communication presents benefits and challenges. This retrospective observational study analyzed EHR audit logs to assess how 75 healthcare professionals (HCPs) utilized AI-generated drafts for patient messages from October 2023 to August 2024 at a large health system in New York City. Overall utilization was low (19.4%), though prompt refinements improved usage (from 12% to 20%), particularly among physicians. GenAI drafts were generated for all messages, including 80% that received no response, adding to the review burden and potentially undermining efficiency. Text analysis showed HCPs preferred concise, information-rich drafts, with role-based differences-physicians favored shorter drafts, while clinical support staff preferred more empathetic responses. AI-generated drafts reduced message turnaround time by 6.76% despite a marginal increase in required steps (InBasket actions). These findings highlight the need for targeted GenAI deployment strategies, better aligned with clinician workflows and optimized draft generation for improved efficiency.
Health care has become increasingly digitized. Given that under-invested health systems and patient populations frequently have worse access to the newest innovations, there is concern that this digitalization may exacerbate preexisting health inequities. This article discusses the multiple ways that digital health may increase health inequities. Using case studies presented by digital health leaders in different roles and settings, it provides examples of how health systems can adopt and implement innovative tools to deliver care while centering health equity. The case studies highlight five guidelines that health-care systems should consider as they weigh the equity implications of adopting any digital tool: auditing benefits, institutional incentives, elevating frontline and patient perspectives, long-term community engagement, and protecting data.
Health equity is receiving increased attention in medical education. However, guidance is often lacking on how to integrate health equity into routine medical education. Journal club presents an opportunity to deepen medical educators’ and learners’ understanding of health equity principles and use it as a lens through which to critically appraise the literature. We present a health equity framework, iteratively co-created by faculty and learners, that can be applied in a journal club setting. Academic medical center in New York City, USA. Faculty, residency program directors, medical students, and residents. Authors developed the health equity journal club framework during a medical student selective course. Learner and faculty applied the framework to journal club articles; their feedback informed revisions. Framework domains included authorship, ethics, methodology, language, peer review, and references. Learner evaluations were overall positive, and 86
Background:Usability testing is valuable for assessing a new tool or system's usefulness and ease-of-use. Several established methods of usability testing exist, including think-aloud testing. Although usability testing has been shown to be crucial for successful clinical decision support (CDS) tool development, it is often difficult to conduct across multisite development projects due to its time- and labor-intensiveness, cost, and the skills required to conduct the testing. Objective:Our objective was to develop a new method of usability testing that would enable efficient acquisition and dissemination of results among multiple sites. We sought to address the existing barriers to successfully completing usability testing during CDS tool development. Methods:We combined individual think-aloud testing and focus groups into one session and performed sessions serially across 4 sites (snowball group usability testing) to assess the usability of two CDS tools designed for use by nurses in primary and urgent care settings. We recorded each session and took notes in a standardized format. Each site shared feedback from their individual sessions with the other sites in the study so that they could incorporate that feedback into their tools prior to their own testing sessions. Results:The group testing and snowballing components of our new usability testing method proved to be highly beneficial. We identified 3 main benefits of snowball group usability testing. First, by interviewing several participants in a single session rather than individuals over the course of weeks, each site was able to quickly obtain their usability feedback. Second, combining the individualized think-aloud component with a focus group component in the same session helped study teams to more easily notice similarities in feedback among participants and to discuss and act upon suggestions efficiently. Third, conducting usability testing in series across sites allowed study teams to incorporate feedback based on previous sites' sessions prior to conducting their own testing. Conclusions:Snowball group usability testing provides an efficient method of obtaining multisite feedback on newly developed tools and systems, while addressing barriers typically associated with traditional usability testing methods. This method can be applied to test a wide variety of tools, including CDS tools, prior to launch so that they can be efficiently optimized.
Shivan J. Mehta, MD, MBA, MSHP; Kevin G. Volpp, MD, PhD; Andrea B. Troxel, ScD; Joseph Teel, MD; Catherine R. Reitz, MPH; Alison Purcell, MSN, CRNP; Humphrey Shen, BA; Kiernan McNelis, BS; Christopher K. Snider, MPH; David A. Asch, MD, MBA
Background During the COVID-19 pandemic, acute respiratory infection (ARI) antibiotic prescribing in ambulatory care markedly decreased. It is unclear if antibiotic prescription rates will remain lowered. Methods We used trend analyses of antibiotics prescribed during and after the first wave of COVID-19 to determine whether ARI antibiotic prescribing rates in ambulatory care have remained suppressed compared to pre-COVID-19 levels. Retrospective data was used from patients with ARI or UTI diagnosis code(s) for their encounter from 298 primary care and 66 urgent care practices within four academic health systems in New York, Wisconsin, and Utah between January 2017 and June 2022. The primary measures included antibiotic prescriptions per 100 non-COVID ARI encounters, encounter volume, prescribing trends, and change from expected trend. Results At baseline, during and after the first wave, the overall ARI antibiotic prescribing rates were 54.7, 38.5, and 54.7 prescriptions per 100 encounters, respectively. ARI antibiotic prescription rates saw a statistically significant decline after COVID-19 onset (step change -15.2, 95% CI: -19.6 to -4.8). During the first wave, encounter volume decreased 29.4% and, after the first wave, remained decreased by 188%. After the first wave, ARI antibiotic prescription rates were no longer significantly suppressed from baseline (step change 0.01, 95% CI: -6.3 to 6.2). There was no significant difference between UTI antibiotic prescription rates at baseline versus the end of the observation period. Conclusions The decline in ARI antibiotic prescribing observed after the onset of COVID-19 was temporary, not mirrored in UTI antibiotic prescribing, and does not represent a long-term change in clinician prescribing behaviors. During a period of heightened awareness of a viral cause of ARI, a substantial and clinically meaningful decrease in clinician antibiotic prescribing was observed. Future efforts in antibiotic stewardship may benefit from continued study of factors leading to this reduction and rebound in prescribing rates.
Objective Our objective was to determine the feasibility and preliminary efficacy of a behavioral nudge on adoption of a clinical decision support (CDS) tool. Materials and Methods We conducted a pilot cluster nonrandomized controlled trial in 2 Emergency Departments (EDs) at a large academic healthcare system in the New York metropolitan area. We tested 2 versions of a CDS tool for pulmonary embolism (PE) risk assessment developed on a web-based electronic health record-agnostic platform. One version included behavioral nudges incorporated into the user interface. Results A total of 1527 patient encounters were included in the trial. The CDS tool adoption rate was 31.67%. Adoption was significantly higher for the tool that included behavioral nudges (39.11% vs 20.66%; P < .001). Discussion We demonstrated feasibility and preliminary efficacy of a PE risk prediction CDS tool developed using insights from behavioral science. The tool is well-positioned to be tested in a large randomized clinical trial.
Indiscriminate use of predictive models incorporating race can reinforce biases present in source data and lead to an exacerbation of health disparities. In some countries, such as the United States, there is therefore a push to remove race from prediction models; however, there are still many prediction models that use race as an input. Biomedical informaticists who are given the responsibility of using these predictive models in healthcare environments are likely to be faced with questions like how to deal with race covariates in these models. Thus, there is a need for a pragmatic framework to help model users think through how to include race in their chosen model so as to avoid inadvertently exacerbating disparities. In this paper, we use the case study of lung cancer screening to propose a simple framework to guide how model users can approach the use (or non-use) of race inputs in the predictive models they are tasked with leveraging in electronic health records and clinical workflows.
BACKGROUND:Overprescribing of antibiotics for acute respiratory infections (ARIs) remains a major issue in outpatient settings. Use of clinical prediction rules (CPRs) can reduce inappropriate antibiotic prescribing but they remain underutilized by physicians and advanced practice providers. A registered nurse (RN)-led model of an electronic health record-integrated CPR (iCPR) for low-acuity ARIs may be an effective alternative to address the barriers to a physician-driven model.METHODS:Following qualitative usability testing, we will conduct a stepped-wedge practice-level cluster randomized controlled trial (RCT) examining the effect of iCPR-guided RN care for low acuity patients with ARI. The primary hypothesis to be tested is: Implementation of RN-led iCPR tools will reduce antibiotic prescribing across diverse primary care settings. Specifically, this study aims to: (1) determine the impact of iCPRs on rapid strep test and chest x-ray ordering and antibiotic prescribing rates when used by RNs; (2) examine resource use patterns and cost-effectiveness of RN visits across diverse clinical settings; (3) determine the impact of iCPR-guided care on patient satisfaction; and (4) ascertain the effect of the intervention on RN and physician burnout.DISCUSSION:This study represents an innovative approach to using an iCPR model led by RNs and specifically designed to address inappropriate antibiotic prescribing. This study has the potential to provide guidance on the effectiveness of delegating care of low-acuity patients with ARIs to RNs to increase use of iCPRs and reduce antibiotic overprescribing for ARIs in outpatient settings.TRIAL REGISTRATION:ClinicalTrials.gov Identifier: NCT04255303, Registered February 5 2020, https://clinicaltrials.gov/ct2/show/NCT04255303 .
Early buzz around ChatGPT's [1] proficiency at healthcare tasks [2][3][4] has turned into rapid deployment of large language model (LLM) technology across healthcare.Multiple healthcare systems including UC San Diego Health, Stanford Health Care, and University of Wisconsin Health are already piloting the use of GPT-4 to draft responses to patient messages [5].Others, including The University of Kansas Health System, are piloting similar LLMs to support clinical documentation [6].Early conversations about LLMs in healthcare tended towards extremes, either lauding their breathtaking potential or lamenting the threat they pose to human-centric care.More recently, nuanced narratives regarding the ethical implications of LLMs have started to emerge [7,8].Still missing from the conversation, though, is a grounded discussion of the pragmatic equity implications of LLM deployment in real-world contexts.In this perspective, we explore various health equity risks generated by the practical use of LLMs in healthcare and offer risk mitigation recommendations.We define health equity using Dr. Braveman's definition: the "absence of differences in health between the most advantaged group in a given category and all others" [9].Equity risks originating in the development of LLMs are widely discussed, with substantial dialogue dedicated to intrinsic equity and general ethical considerations of how, why, and with what data and algorithmic design choices these models should be developed [10,11].Equity risks in AI model deployment are less commonly discussed.These emergent risks arise from pragmatic use by end users (patients, clinicians, and healthcare systems) in real-world settings.It is critical to understand emergent equity risks of LLMs to avoid replicating inequities that resulted from the use of other technologies in medicine [12].For patients, LLMs have a remarkable ability to generate human-like responses and summarize complex or hard-to-find information [2,4], providing automated care with potential to alleviate disparities in access, understandability, and personalization.However, if minoritized populations who have experienced barriers to accessing traditional care are disproportionately more likely to seek care from a free chatbot rather than a licensed provider this could result in a uniquely biased care experience-with higher data privacy risks [13], biased outputs, misinformation, or personally offensive content from models trained on biased data [14][15][16].This could perpetuate negative care experiences, medical disenfranchisement and distrust.For clinicians, LLM-supported tools offer efficiency gains that may help address issues of burnout and dissatisfaction.For example, DocsGPT [17] automates high-burden, low-cognition tasks such as composing prior authorization letters.However, the choice to use LLMs versus not based on patient language, health literacy, or insurance coverage could widen existing
Background The improvements in care resulting from clinical decision support (CDS) have been significantly limited by consistently low health care provider adoption. Health care provider attitudes toward CDS, specifically psychological and behavioral barriers, are not typically addressed during any stage of CDS development, although they represent an important barrier to adoption. Emerging evidence has shown the surprising power of using insights from the field of behavioral economics to address psychological and behavioral barriers. Nudges are formal applications of behavioral economics, defined as positive reinforcement and indirect suggestions that have a nonforced effect on decision-making. Objective Our goal is to employ a user-centered design process to develop a CDS tool—the pulmonary embolism (PE) risk calculator—for PE risk stratification in the emergency department that incorporates a behavior theory–informed nudge to address identified behavioral barriers to use. Methods All study activities took place at a large academic health system in the New York City metropolitan area. Our study used a user-centered and behavior theory–based approach to achieve the following two aims: (1) use mixed methods to identify health care provider barriers to the use of an active CDS tool for PE risk stratification and (2) develop a new CDS tool—the PE risk calculator—that addresses behavioral barriers to health care providers’ adoption of CDS by incorporating nudges into the user interface. These aims were guided by the revised Observational Research Behavioral Information Technology model. A total of 50 clinicians who used the original version of the tool were surveyed with a quantitative instrument that we developed based on a behavior theory framework—the Capability-Opportunity-Motivation-Behavior framework. A semistructured interview guide was developed based on the survey responses. Inductive methods were used to analyze interview session notes and audio recordings from 12 interviews. Revised versions of the tool were developed that incorporated nudges. Results Functional prototypes were developed by using Axure PRO (Axure Software Solutions) software and usability tested with end users in an iterative agile process (n=10). The tool was redesigned to address 4 identified major barriers to tool use; we included 2 nudges and a default. The 6-month pilot trial for the tool was launched on October 1, 2021. Conclusions Clinicians highlighted several important psychological and behavioral barriers to CDS use. Addressing these barriers, along with conducting traditional usability testing, facilitated the development of a tool with greater potential to transform clinical care. The tool will be tested in a prospective pilot trial. International Registered Report Identifier (IRRID) DERR1-10.2196/42653
BackgroundThrough our work, we have demonstrated how clinical decision support (CDS) tools integrated into the electronic health record (EHR) assist providers in adopting evidence-based practices. This requires confronting technical challenges that result from relying on the EHR as the foundation for tool development; for example, the individual CDS tools need to be built independently for each different EHR. ObjectiveThe objective of our research was to build and implement an EHR-agnostic platform for integrating CDS tools, which would remove the technical constraints inherent in relying on the EHR as the foundation and enable a single set of CDS tools that can work with any EHR. MethodsWe developed EvidencePoint, a novel, cloud-based, EHR-agnostic CDS platform, and we will describe the development of EvidencePoint and the deployment of its initial CDS tools, which include EHR-integrated applications for clinical use cases such as prediction of hospitalization survival for patients with COVID-19, venous thromboembolism prophylaxis, and pulmonary embolism diagnosis. ResultsThe results below highlight the adoption of the CDS tools, the International Medical Prevention Registry on Venous Thromboembolism-D-Dimer, the Wells’ criteria, and the Northwell COVID-19 Survival (NOCOS), following development, usability testing, and implementation. The International Medical Prevention Registry on Venous Thromboembolism-D-Dimer CDS was used in 5249 patients at the 2 clinical intervention sites. The intervention group tool adoption was 77.8% (4083/5249 possible uses). For the NOCOS tool, which was designed to assist with triaging patients with COVID-19 for hospital admission in the event of constrained hospital resources, the worst-case resourcing scenario never materialized and triaging was never required. As a result, the NOCOS tool was not frequently used, though the EvidencePoint platform’s flexibility and customizability enabled the tool to be developed and deployed rapidly under the emergency conditions of the pandemic. Adoption rates for the Wells’ criteria tool will be reported in a future publication. ConclusionsThe EvidencePoint system successfully demonstrated that a flexible, user-friendly platform for hosting CDS tools outside of a specific EHR is feasible. The forthcoming results of our outcomes analyses will demonstrate the adoption rate of EvidencePoint tools as well as the impact of behavioral economics “nudges” on the adoption rate. Due to the EHR-agnostic nature of EvidencePoint, the development process for additional forms of CDS will be simpler than traditional and cumbersome IT integration approaches and will benefit from the capabilities provided by the core system of EvidencePoint.