Importance:Large language models (LLMs) can assist in various health care activities, but current evaluation approaches may not adequately identify the most useful application areas. Objective:To summarize existing evaluations of LLMs in health care in terms of 5 components: (1) evaluation data type, (2) health care task, (3) natural language processing (NLP) and natural language understanding (NLU) tasks, (4) dimension of evaluation, and (5) medical specialty. Data Sources:A systematic search of PubMed and Web of Science was performed for studies published between January 1, 2022, and February 19, 2024. Study Selection:Studies evaluating 1 or more LLMs in health care. Data Extraction and Synthesis:Three independent reviewers categorized studies via keyword searches based on the data used, the health care tasks, the NLP and NLU tasks, the dimensions of evaluation, and the medical specialty. Results:Of 519 studies reviewed, published between January 1, 2022, and February 19, 2024, only 5% used real patient care data for LLM evaluation. The most common health care tasks were assessing medical knowledge such as answering medical licensing examination questions (44.5%) and making diagnoses (19.5%). Administrative tasks such as assigning billing codes (0.2%) and writing prescriptions (0.2%) were less studied. For NLP and NLU tasks, most studies focused on question answering (84.2%), while tasks such as summarization (8.9%) and conversational dialogue (3.3%) were infrequent. Almost all studies (95.4%) used accuracy as the primary dimension of evaluation; fairness, bias, and toxicity (15.8%), deployment considerations (4.6%), and calibration and uncertainty (1.2%) were infrequently measured. Finally, in terms of medical specialty area, most studies were in generic health care applications (25.6%), internal medicine (16.4%), surgery (11.4%), and ophthalmology (6.9%), with nuclear medicine (0.6%), physical medicine (0.4%), and medical genetics (0.2%) being the least represented. Conclusions and Relevance:Existing evaluations of LLMs mostly focus on accuracy of question answering for medical examinations, without consideration of real patient care data. Dimensions such as fairness, bias, and toxicity and deployment considerations received limited attention. Future evaluations should adopt standardized applications and metrics, use clinical data, and broaden focus to include a wider range of tasks and specialties.
BACKGROUND:Adequate patient awareness and understanding of cancer clinical trials is essential for trial recruitment, informed decision making, and protocol adherence. Although large language models (LLMs) have shown promise for patient education, their role in enhancing patient awareness of clinical trials remains unexplored. This study explored the performance and risks of LLMs in generating trial-specific educational content for potential participants. METHODS:Generative Pretrained Transformer 4 (GPT4) was prompted to generate short clinical trial summaries and multiple-choice question-answer pairs from informed consent forms from ClinicalTrials.gov. Zero-shot learning was used for summaries, using a direct summarization, sequential extraction, and summarization approach. One-shot learning was used for question-answer pairs development. We evaluated performance through patient surveys of summary effectiveness and crowdsourced annotation of question-answer pair accuracy, using held-out cancer trial informed consent forms not used in prompt development. RESULTS:For summaries, both prompting approaches achieved comparable results for readability and core content. Patients found summaries to be understandable and to improve clinical trial comprehension and interest in learning more about trials. The generated multiple-choice questions achieved high accuracy and agreement with crowdsourced annotators. For both summaries and multiple-choice questions, GPT4 was most likely to include inaccurate information when prompted to provide information that was not adequately described in the informed consent forms. CONCLUSIONS:LLMs such as GPT4 show promise in generating patient-friendly educational content for clinical trials with minimal trial-specific engineering. The findings serve as a proof of concept for the role of LLMs in improving patient education and engagement in clinical trials, as well as the need for ongoing human oversight.
N Engl J Med . 2025 Dec 18;393(24):2478–2482. doi: 10.1056/NEJMms2510113. In the years following Dobbs v. Jackson Women’s Health Organization , the consequences of state abortion bans are becoming increasingly visible, in expected ways, and in deeply troubling new ones. Whereas unsafe abortions in the pre-Roe era often occurred outside medical institutions, today, patients are experiencing harm inside hospitals, where physicians hesitate to act because of fear of criminal liability. The result is restricted access to abortion and disruption of established medical standards of care.
BACKGROUND:An advance directive (AD) document allows a patient to indicate their health care preferences and identify an agent to make decisions on their behalf if they lose their ability to communicate. Due to the substantially elevated risk of acute respiratory failure and death during the COVID-19 pandemic, ADs were especially relevant. The objective of this study was to describe AD completed prior to COVID-19 infection (COVID-19) among patients receiving care in a national healthcare system. METHODS:We conducted a cohort study of United States Veterans Health Administration (VA) patients with COVID-19 between March 2020 and December 2022. AD completion before COVID-19 was ascertained by progress note titles in the electronic health record. Covariates included age, sex, race/ethnicity, marital status, geographic region, health care utilization, calendar quarter of COVID-19, and VA COVID-19 (VACO) 30-day mortality index score. RESULTS:Among 422,028 COVID-19 patients (median age = 62 years; 88.5% male; 58.6% non-Hispanic White (NH-White), 23.3% non-Hispanic Black (NH-Black), 9.2% Hispanic), 67,970 (16.1%) had AD documentation which varied substantially across all covariates. AD completion increased with VACO Index quintiles ranging from 8.2 to 31.0%. In a model adjusted for covariates, relative to NH-White, NH-Black and Hispanic groups had decreased odds for AD (NH-Black odds ratio (OR)=0.77 (95% confidence interval 0.76-0.79); Hispanic OR=0.85 (0.82-0.87)). VACO index includes age, and both were strongly associated with AD completion. Women compared to men, and those who were widowed, separated/divorced and never married relative to people who were married, had increased AD completion. CONCLUSIONS:AD completion was overall low, including among patients at high risk of mortality due to COVID-19. When controlling for age, risk for mortality and other covariates, men and people who identify as Black or Hispanic were less likely to have completed an AD. Investment in interventions to facilitate AD completion are needed, particularly among historically underrepresented populations.
OBJECTIVES:This study aimed to assess clinicians' confidence in helping adolescents access abortion. STUDY DESIGN:This study used a 2024 cross-sectional, online survey of US adolescent-serving clinicians. RESULTS:Less than half of the 188 clinicians reported high confidence in their ability to help adolescents navigate seven of 11 logistical aspects of abortion access. Participants in states with post-Dobbs restrictions were less confident than those in states without such restrictions in their ability to help adolescents find abortion providers, know what documents are needed for appointments, and interpret their state's laws. CONCLUSIONS:Many clinicians lack confidence in helping adolescents navigate abortion access, particularly clinicians in states with post-Dobbs restrictions. IMPLICATIONS:Clinicians' lack of confidence in helping adolescents access abortion, regardless of state-level abortion laws, is concerning, given adolescents' reliance on clinicians for reliable abortion information. Interventions must be developed across all states to increase clinician confidence in their ability to support their patients' abortion access in the post-Dobbs shifting legal landscape.
This study examines the distribution of payments within and across specialties and the medical products associated with the largest total payments.
Background:Health care organizations, including the Veterans Health Administration (VHA), are increasingly adopting programs to address social determinants of health. As part of a comprehensive social risk screening and referral model, tailored resource guides can support efforts to address unmet social needs. However, limited guidance is available on best practices for the development of resource guides in health care settings.Observations:This article describes the development of geographically tailored resource guides for a national VHA quality improvement initiative, Assessing Circumstances and Offering Resources for Needs (ACORN), which aims to systematically screen for and address social needs among veterans. We outline the rationale for using resource guides as a social needs intervention and provide a pragmatic framework for resource guide development and maintenance. We offer guidance based on lessons learned from the development of ACORN resource guides, emphasizing a collaborative approach with VHA social workers and other frontline clinical staff, as well as with community-based organizations. Our how-to guide provides steps for identifying high-yield resources along with formatting considerations to maximize accessibility and usability among patients.Conclusions:Resource guides can serve as a valuable cross-cutting component of health care organizations' efforts to address social needs. We provide a practical approach to resource guide development that may support successful implementation within the VHA and other clinical settings.
Background: Recent studies, including those by the National Board of Medical Examiners, have highlighted the remarkable capabilities of recent large language models (LLMs) such as ChatGPT in passing the United States Medical Licensing Examination (USMLE). However, there is a gap in detailed analysis of LLM performance in specific medical content areas, thus limiting an assessment of their potential utility in medical education. Objective: This study aimed to assess and compare the accuracy of successive ChatGPT versions (GPT-3.5, GPT-4, and GPT-4 Omni) in USMLE disciplines, clinical clerkships, and the clinical skills of diagnostics and management. Methods: This study used 750 clinical vignette-based multiple-choice questions to characterize the performance of successive ChatGPT versions (ChatGPT 3.5 [GPT-3.5], ChatGPT 4 [GPT-4], and ChatGPT 4 Omni [GPT-4o]) across USMLE disciplines, clinical clerkships, and in clinical skills (diagnostics and management). Accuracy was assessed using a standardized protocol, with statistical analyses conducted to compare the models' performances. Results: GPT-4o achieved the highest accuracy across 750 multiple-choice questions at 90.4%, outperforming GPT-4 and GPT-3.5, which scored 81.1% and 60.0%, respectively. GPT-4o's highest performances were in social sciences (95.5%), behavioral and neuroscience (94.2%), and pharmacology (93.2%). In clinical skills, GPT-4o's diagnostic accuracy was 92.7% and management accuracy was 88.8%, significantly higher than its predecessors. Notably, both GPT-4o and GPT-4 significantly outperformed the medical student average accuracy of 59.3% (95% CI 58.3-60.3). Conclusions: GPT-4o's performance in USMLE disciplines, clinical clerkships, and clinical skills indicates substantial improvements over its predecessors, suggesting significant potential for the use of this technology as an educational aid for medical students. These findings underscore the need for careful consideration when integrating LLMs into medical education, emphasizing the importance of structured curricula to guide their appropriate use and the need for ongoing critical analyses to ensure their reliability and effectiveness. JMIR Med Educ 2024;10:e63430; doi: 10.2196/63430
BACKGROUND:Recent studies, including those by the National Board of Medical Examiners (NBME), have highlighted the remarkable capabilities of recent large language models (LLMs) such as ChatGPT in passing the United States Medical Licensing Examination (USMLE). However, there is a gap in detailed analysis of these models' performance in specific medical content areas, thus limiting an assessment of their potential utility for medical education. OBJECTIVE:To assess and compare the accuracy of successive ChatGPT versions (GPT-3.5, GPT-4, and GPT-4 Omni) in USMLE disciplines, clinical clerkships, and the clinical skills of diagnostics and management. METHODS:This study used 750 clinical vignette-based multiple-choice questions (MCQs) to characterize the performance of successive ChatGPT versions [ChatGPT 3.5 (GPT-3.5), ChatGPT 4 (GPT-4), and ChatGPT 4 Omni (GPT-4o)] across USMLE disciplines, clinical clerkships, and in clinical skills (diagnostics and management). Accuracy was assessed using a standardized protocol, with statistical analyses conducted to compare the models' performances. RESULTS:GPT-4o achieved the highest accuracy across 750 MCQs at 90.4%, outperforming GPT-4 and GPT-3.5, which scored 81.1% and 60.0% respectively. GPT-4o's highest performances were in social sciences (95.5%), behavioral and neuroscience (94.2%), and pharmacology (93.2%). In clinical skills, GPT-4o's diagnostic accuracy was 92.7% and management accuracy 88.8%, significantly higher than its predecessors. Notably, both GPT-4o and GPT-4 significantly outperformed the medical student average accuracy of 59.3% (95% CI: 58.3-60.3). CONCLUSIONS:ChatGPT 4 Omni's performance in USMLE preclinical content areas as well as clinical skills indicates substantial improvements over its predecessors, suggesting significant potential for the use of this technology as an educational aid for medical students. These findings underscore the necessity of careful consideration of LLMs' integration into medical education, emphasizing the importance of structured curricula to guide their appropriate use and the need for ongoing critical analyses to ensure their reliability and effectiveness. CLINICALTRIAL:
Artificial Intelligence (AI) holds the promise of transforming healthcare by improving patient outcomes, increasing accessibility and efficiency, and decreasing the cost of care but realizing this vision of a healthier world for everyone everywhere requires partnerships and trust between healthcare systems, clinicians, payers, technology companies, pharmaceutical companies, and governments to help bring innovations in machine learning and artificial intelligence to patients. Google is a technology company that is partnering with healthcare systems, clinicians, and researchers to develop technology solutions to directly improve the lives of patients. This chapter reviews the use of AI in healthcare from Google's perspective. Landmark studies of AI application in healthcare are shared and the application of Google's novel system of organizing information to unify data in electronic health records (EHRs) and bring an integrated view of patient records to clinicians is described. Google's consumer-focused innovation in dermatology to help guide search journeys for personalized information about skin conditions is also summarized. Finally how to embed ethics and a concern for all patients into the development of AI is reviewed from Google's perspective.
Importance: Large Language Models (LLMs) can assist in a wide range of healthcare-related activities. Current approaches to evaluating LLMs make it difficult to identify the most impactful LLM application areas. Objective: To summarize the current evaluation of LLMs in healthcare in terms of 5 components: evaluation data type, healthcare task, Natural Language Processing (NLP)/Natural Language Understanding (NLU) task, dimension of evaluation, and medical specialty. Data Sources: A systematic search of PubMed and Web of Science was performed for studies published between 01-01-2022 and 02-19-2024. Study Selection: Studies evaluating one or more LLMs in healthcare. Data Extraction and Synthesis: Three independent reviewers categorized 519 studies in terms of data used in the evaluation, the healthcare tasks (the what) and the NLP/NLU tasks (the how) examined, the dimension(s) of evaluation, and the medical specialty studied. Results: Only 5% of reviewed studies utilized real patient care data for LLM evaluation. The most popular healthcare tasks were assessing medical knowledge (e.g. answering medical licensing exam questions, 44.5%), followed by making diagnoses (19.5%), and educating patients (17.7%). Administrative tasks such as assigning provider billing codes (0.2%), writing prescriptions (0.2%), generating clinical referrals (0.6%) and clinical notetaking (0.8%) were less studied. For NLP/NLU tasks, the vast majority of studies examined question answering (84.2%). Other tasks such as summarization (8.9%), conversational dialogue (3.3%), and translation (3.1%) were infrequent. Almost all studies (95.4%) used accuracy as the primary dimension of evaluation; fairness, bias and toxicity (15.8%), robustness (14.8%), deployment considerations (4.6%), and calibration and uncertainty (1.2%) were infrequently measured. Finally, in terms of medical specialty area, most studies were in internal medicine (42%), surgery (11.4%) and ophthalmology (6.9%), with nuclear medicine (0.6%), physical medicine (0.4%) and medical genetics (0.2%) being the least represented. Conclusions and Relevance: Existing evaluations of LLMs mostly focused on accuracy of question answering for medical exams, without consideration of real patient care data. Dimensions like fairness, bias and toxicity, robustness, and deployment considerations received limited attention. To draw meaningful conclusions and improve LLM adoption, future studies need to establish a standardized set of LLM applications and evaluation dimensions, perform evaluations using data from routine care, and broaden testing to include administrative tasks as well as multiple medical specialties. Keywords: Large Language Models, Generative Artificial Intelligence, Healthcare, Dimensions of Evaluation, Evaluation Metrics. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
The Department of Veterans Affairs (VA) healthcare system routinely screens Veterans for food insecurity, housing instability, and intimate partner violence, but does not systematically screen for other health-related social needs (HRSNs). To (1) develop a process for systematically identifying and addressing Veterans’ HRSNs, (2) determine reported prevalence of HRSNs, and (3) assess the acceptability of HRSN screening among Veterans. “Assessing Circumstances and Offering Resources for Needs” (ACORN) is a Veteran-tailored HRSN screening and referral quality improvement initiative. Veterans were screened via electronic tablet for nine HRSNs (food, housing, utilities, transportation, legal needs, social isolation, interpersonal violence, employment, and education) and provided geographically tailored resource guides for identified needs. Two-week follow-up interviews with a purposive sample of Veterans explored screening experiences. Convenience sample of Veterans presenting for primary care at a VA urban women’s health clinic and suburban community-based outpatient clinic (October 2019–May 2020). Primary outcomes included prevalence of HRSNs, Veteran-reported acceptability of screening, and use of resources guides. Data were analyzed using descriptive statistics, chi-square tests, and rapid qualitative analysis. Of 268 Veterans screened, 50
With abortion remaining legal in over half of the country and a proliferation of websites offering information on how to access abortion medications, for those who know where to look, there are sound options for safely ending an unwanted early-stage pregnancy. But not all patients have equal access to reliable information. This Article addresses the urgent downstream harms caused by the lack of access to abortion information, and argues that in view of these consequences, regardless of abortion's legal status, clinicians have a duty to provide their patients with abortion information. We begin by documenting clinicians' hesitation to share abortion information, drawing on our interviews with 25 doctors practicing medicine in a state where abortion is criminalized. Next, we explain why clinicians are duty-bound to provide all-options counseling. We then consider whether such duties shift where abortion is criminalized. After identifying the limited legal risks associated with supplying abortion information, and showing how, by requiring all-options counseling, professional societies might reduce risks to patients and clinicians, we conclude that, regardless of the legal status of abortion, clinicians have a professional responsibility to share basic abortion information - including treatment options and how to access those options.
This American College of Physicians position paper aims to inform ethical decision making for the integration of precision medicine and genetic testing into clinical care. Although the positions are primarily intended for practicing physicians, they may apply to other health care professionals and can also inform how health care systems, professional schools, and residency programs integrate genomics into educational and clinical settings. Addressing the challenges of precision medicine and genetic testing will guide ethical and responsible implementation to improve health outcomes.