Importance:High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While retrospective studies suggest that large language models (LLMs) can generate clinical summaries of comparable quality to those by physicians, prospective data on their safety, utility, and association with clinician well-being in clinical environments are lacking. Objective:To evaluate the safety, use, and association with clinician burden of MedAgentBrief, an LLM-based agentic workflow for generating hospital course summaries, during prospective clinical deployment. Design, Setting, and Participants:This single-arm prospective pilot quality improvement study encompassed hospital discharges at 1 academic inpatient medicine unit from August 1 to October 11, 2025, with baseline comparisons drawn from April 9 to July 31, 2025. Intervention:A custom agentic LLM workflow using Gemini 2.5 Pro generated draft hospital course summaries nightly using patient history and physical and daily progress notes. Drafts were securely emailed to physicians daily for review and optional use. Main Outcomes and Measures:The primary outcome was physician-reported potential for and severity of harm from unedited summaries (Agency for Healthcare Research and Quality Common Format Harm Scale). Secondary outcomes included use rate, error types (omissions, inaccuracies, and hallucinations), time spent in discharge summaries (electronic health record logs), and changes in cognitive burden (NASA Task Load Index; score range, 0-100, with higher scores indicating greater cognitive burden) and burnout (Stanford Professional Fulfillment Index Work Exhaustion Scale; score range, 0-4, with higher scores indicating greater burnout). Results:Among 384 hospital discharges, the system generated 1274 summaries. Physicians used artificial intelligence (AI) content in 219 cases (57.0%). Feedback on 100 summaries (88 of 219 used summaries [40.2%] and 12 of 165 unused summaries [7.3%]) noted omissions (25 summaries [25.0%]) and inaccuracies (20 summaries [20.0%]) but rare hallucinations (2 summaries [2.0%]). Physicians rated 88 unedited summaries (88.0%) as having no harm potential and 1 (1.0%) as likely to cause moderate harm; no severe harm was reported. Mean physician burnout scores decreased significantly from before to after the intervention (1.75; 95% CI, 1.16-2.34 vs 1.20; 95% CI, 0.71-1.69; P = .03). Time savings were heterogeneous, with 5 of 7 physicians with matched baseline data (71.4%) seeing reductions in median documentation time; changes from baseline to pilot were up to 2.9 minutes, which was a nonsignificant difference (10.7 minutes; 95% CI, 7.4-13.3 minutes vs 7.8 minutes; 95% CI, 5.1-11.7 minutes; P = .13). Conclusions and Relevance:In this study, an LLM-based agentic workflow produced hospital course summaries that were frequently used with minimal risk of harm identified. The intervention was associated with a reduction in physician burnout, supporting the viability of AI summarization to mitigate documentation burden.
Generative models trained on electronic health records are viewed as ‘zero-shot predictors’ for clinical outcomes — but this interpretation is misleading.
Objectives: The landscape of AI in healthcare is characterized by a proliferation of guidelines. While many articulate similar principles, they do not always overlap and often lack actionable guidance across the AI lifecycle. This article describes the development of the CHAI Responsible AI Guide and examines how large-scale, multi-stakeholder consensus methods can translate ethical principles into coordinated, adoptable practices across institutional contexts. Methods: CHAI conducted a yearlong consensus process that combined deliberation across five workgroups, a community-wide survey, multistakeholder review, and independent expert review. The approach drew on modified Delphi methods and is reported using the CREDES framework, with adaptations for participatory design and editorial responsibility analogous to the National Academies. Results: The process applied five principle-based themes to a six-stage lifecycle framework, generating considerations with representative use cases illustrating how priorities can shift across contexts and stages. Together, these outputs provide a shared vocabulary and practical framework for applying responsible AI principles across development, evaluation, deployment, and monitoring. Discussion: Unlike prior consensus efforts focused on discrete phases of AI evaluation, this initiative addressed the challenges of coordination, translation, and accountability across the full AI lifecycle. The adaptive consensus design balanced depth of deliberation with broad participation, supporting iterative refinement while incorporating perspectives from a large and diverse stakeholder community. Conclusion: The CHAI Responsible AI Guide demonstrates how structured, multi-stakeholder collaboration can support lifecycle-based governance for health AI. Its development offers insights for future consensus efforts in complex, rapidly evolving domains where responsibility is distributed across institutions and roles.
Postdeployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit - and to support governance decisions about which systems to update, modify, or decommission. Motivated by these needs, the authors developed a framework for monitoring deployed AI systems organized around three complementary principles: system integrity, performance, and impact. System integrity monitoring focuses on maximizing system uptime, detecting runtime errors, and identifying when changes to the surrounding information technology ecosystem have unintended effects. Performance monitoring focuses on maintaining accurate and equitable system behavior in the face of changing health care practices (and thus input data) over time. Impact monitoring assesses whether a deployed system continues to have value in the form of benefit to clinicians, staff, and patients. Drawing on examples of deployed AI systems at their academic medical center, the authors provide practical guidance for creating monitoring plans based on these principles that specify which metrics to measure and at what cadence, who is responsible for acting when metrics change, and what concrete follow-up actions should be taken - for both traditional and generative AI. They also discuss challenges in implementing this framework, including the effort of monitoring for health systems with limited resources, and the difficulty of incorporating data-driven monitoring practices into complex organizations where conflicting priorities and definitions of success often coexist. This framework offers a starting point for health systems seeking to ensure that AI deployments remain safe and effective over time.
Small language models can serve as a more scalable, practical and efficient alternative for medical applications compared to large language models. Here we discuss the strengths and limitations, challenges and future directions of small language models in healthcare implementation, highlighting their potential to enable broader real-world adoption across diverse clinical settings.
While data standards have been well adopted and highly impactful for observational health informatics, the emerging application of artificial intelligence (AI) to electronic health record (EHR) data - known broadly as health AI - still lacks broadly adopted data standards. This gap limits reproducibility and the development of open-source ecosystems for health AI, as well as limiting the ability to collaboratively develop and characterize emerging capabilities such as the use of foundation models. The Medical Event Data Standard (MEDS) is an open- source data standard and ecosystem designed and promulgated by us (and others) to address this gap. MEDS was first presented at a workshop in May 2024 and has since been described in online materials and preprints. Here, we describe MEDS in detail and review its use in the community, highlighting its strengths and weaknesses, and the extent of its adoption to date. We designed MEDS to emphasize simplicity, algorithm transportability, and support for workflows used in training foundation models, and to offer complementary strengths compared with existing health data standards. As of March 2026, we find usage of MEDS across 21 institutions, in at least 27 academic papers and preprints, and to support work with 17 datasets and 12 AI algorithms. Tools in the MEDS ecosystem have reported computational speedups that range from 1.9 to around 40,000 times faster than prior tools or common individual workflows. In novel case study comparisons, we find that codebases leveraging MEDS show reductions in necessary lines of functional Python code by up to 70%. Further, the MEDS standard has supported the development of key frontier models for EHR data. (Supported by Canadian Institutes of Health Research and others.).
As Academic Medicine marks its centennial, academic medical centers (AMCs) face a defining inflection point driven by the convergence of generative artificial intelligence (AI), machine learning, computational biology, digital health platforms, autonomous systems, and increasingly continuous streams of clinical and behavioral data. Together, these technologies are transforming how health professionals are educated, how biomedical knowledge is generated, how care is delivered, and how health outcomes are measured. In this Commentary, the authors argue that these changes require more than technological adoption; they demand a redefinition of the mission and responsibilities of academic medicine. AMCs must move beyond their traditional roles in education, research, and clinical care to become leaders in responsible innovation, ethical governance, workforce transformation, public trust, community engagement, health equity, and global stewardship. The authors examine implications for individualized education, AI-enabled clinical care, accelerated scientific discovery, and the democratization of health knowledge across diverse settings. The authors contend that AMCs must deliberately redesign curricula, research ecosystems, care models, and institutional incentives to ensure that technological advances strengthen rather than diminish human judgment, compassion, and equity. The next century of academic medicine will be defined not by whether AMCs adopt emerging technologies, but by whether they shape their development and use in ways that preserve the human relationships and societal responsibilities at the heart of medicine.
As artificial intelligence (AI) tools get increasingly deployed for mental healthcare, public trust in these systems remains uncertain. It is unclear how clients perceive AI involvement in counseling interactions, particularly in moments of crisis that require empathy and connection. To address this gap, we analyzed 75,777 crisis counseling conversations from a human-staffed WhatsApp helpline in India to characterize how often clients suspected they were speaking to AI, what triggered those doubts, and how counselors responded. Though no conversations actually involved AI assistance, the proportion of conversations where clients suspected AI use increased from 0.8
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignment with human raters -- a property that itself depends on costly human annotations. In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations. Metric Match selects a subset of samples for human annotation such that the subset matches the population reliability metric with respect to acquired synthetic labels. We empirically show that Metric Match achieves a win-rate of 0.838 against random subset selection across four different correlation metrics and 15 datasets, with an 18.7% decrease in average estimation error and reduces annotation needs by 32.5%. We provide a cost model and highlight a medical case study where our method saves $1,041.67 compared to random selection for expert annotation. Further, we shift our task from reliability estimation to reliability classification of whether a given judge is above a deployment threshold, outperforming random selection with Metric Match. All project code is publicly available, and we additionally provide an installable package for ease of use.
Clinical evaluation of large language models (LLMs) currently relies on static datasets and isolated scenarios that fail to capture the cascading effects of healthcare decisions. We propose the Clinical Environment Simulator (CES), a framework that evaluates clinical LLMs within digital hospital environments where every decision dynamically alters future states. The CES would use a parallel simulation architecture: a 'hospital engine' that tracks bed availability, staff workloads and equipment status in real time, and a 'patient engine' that simulates disease progression and treatment responses based on LLM interventions. Unlike current benchmarks, the CES framework requires clinical LLMs to execute decisions through realistic electronic health record interfaces, while managing trade-offs between individual patient optimization and system-wide efficiency. The CES enables three critical evaluations absent from current benchmarks: temporal reasoning under evolving constraints, where delayed diagnostics can lead to patient deterioration; resource-aware decision-making, where aggressive workups for one patient may exhaust capacity needed by others; and operational resilience, through adversarial testing with simultaneous emergencies and system failures. By scoring LLM performance on both clinical outcomes and operational metrics, the CES represents a shift toward evaluating clinical LLMs as a dynamic and integrated component of healthcare delivery systems.
Tumor boards are multidisciplinary conferences dedicated to producing actionable patient care recommendations with live review of primary radiology and pathology data. Succinct patient case summaries are needed to drive efficient and accurate case discussions. We developed a manual AI-based workflow to generate patient summaries to display live at the Stanford Thoracic Tumor board. To improve on this manually intensive process, we developed several automated AI chart summarization methods and evaluated them against physician gold standard summaries and fact-based scoring rubrics. We report these comparative evaluations as well as our deployment of the final state automated AI chart summarization tool along with post-deployment monitoring. We also validate the use of an LLM as a judge evaluation strategy for fact-based scoring. This work is an example of integrating AI-based workflows into routine clinical practice.
Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance decisions about which systems to update, modify, or decommission. Motivated by these needs, we developed a framework for monitoring deployed AI systems that is organized around three complementary principles: system integrity, performance, and impact. System integrity monitoring focuses on maximizing system uptime, detecting runtime errors, and identifying when changes to the surrounding IT ecosystem have unintended effects. Performance monitoring focuses on maintaining accurate and equitable system behavior in the face of changing health care practices (and thus input data) over time. Impact monitoring assesses whether a deployed system continues to have value in the form of benefit to clinicians, staff, and patients. Drawing on examples of deployed AI systems at our academic medical center, we provide practical guidance for creating monitoring plans based on these principles that specify which metrics to measure, when those metrics should be reviewed, who is responsible for acting when metrics change, and what concrete follow-up actions should be taken-for both traditional and generative AI. We also discuss challenges in implementing this framework, including the effort and cost of monitoring for health systems with limited resources as well as the difficulty of incorporating data-driven monitoring practices into complex organizations where conflicting priorities and definitions of success often coexist. This framework offers a practical template and starting point for health systems seeking to ensure that AI deployments remain safe and effective over time.
Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assurance these systems require. We address this challenge with two complementary contributions. First, we introduce MedFactEval, a framework for scalable, fact-grounded evaluation where clinicians define high-salience key facts and an "LLM Jury"-a multi-LLM majority vote-assesses their inclusion in generated summaries. Second, we present MedAgentBrief, a model-agnostic, multi-step workflow designed to generate high-quality, factual discharge summaries. To validate our evaluation framework, we established a gold-standard reference using a seven-physician majority vote on clinician-defined key facts from inpatient cases. The MedFactEval LLM Jury achieved almost perfect agreement with this panel (Cohen's κ = 81%), a performance statistically non-inferior to that of a single human expert (κ = 67%, P < 0.001). Our work provides both a robust evaluation framework (MedFactEval) and a high-performing generation workflow (MedAgentBrief), offering a comprehensive approach to advance the responsible deployment of generative AI in clinical workflows.
Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce unsupported claims. Existing error-detection methods produce heuristic or uncalibrated scores, providing no formal control over missed errors and no principled way to trade off safety against clinician review burden. We introduce Conformal Assessment for Risk Evaluation (CARE), a post-hoc, model-agnostic safety layer that uses conformal risk control to overlay calibrated omission and hallucination flags onto summaries from any LLM without retraining. CARE provides finite-sample, distribution-free guarantees through two controllers: a hallucination controller that bounds the probability of a document containing any unflagged hallucinated sentence, and an omission controller that bounds the expected fraction of important omissions not surfaced for review. Unlike hallucination detection, omissions depend jointly on whether a source sentence is important and whether it is covered by the summary. We show that calibrating only one dimension can violate the target risk bound, while marginal decompositions remain valid but overly conservative. By jointly calibrating over the full (τ,γ) threshold space, CARE preserves formal guarantees while surfacing up to 5× fewer sentences than alternative calibrated baselines. Across five medical summarization tasks, CARE satisfies the target risk bound at α= 0.15 with 95
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts—Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous—the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1
Importance High-quality discharge summaries are essential for safe care transitions but contribute substantially to clinician documentation burden and burnout. While retrospective studies suggest large language models (LLMs) can generate clinical summaries of comparable quality to physicians, prospective data on their safety, utility, and impact on clinician well-being in real-world environments are lacking. Objective To evaluate the safety, utilization, and impact on clinician burden of MedAgentBrief, an LLM-based agentic workflow for generating hospital course summaries, during prospective clinical deployment. Design, Setting, and Participants Single-arm prospective pilot study of 11 attending hospitalist physicians at one inpatient unit from August 1 to October 11, 2025, with baseline comparisons drawn from April 9 to July 31, 2025. Intervention MedAgentBrief, a custom agentic AI workflow utilizing Gemini 2.5 Pro, generated draft hospital course summaries nightly using the patient’s history and physical and daily progress notes. Drafts were securely emailed to physicians daily for review and optional use. Main Outcomes and Measures The primary outcome was physician-reported potential for and severity of harm from unedited summaries (AHRQ Common Format Harm Scale). Secondary outcomes included utilization rate, error types (omissions, inaccuracies, hallucinations), time spent in discharge summaries (EHR logs), and changes in cognitive burden (NASA Task Load Index [NASA-TLX]) and burnout (Stanford Professional Fulfillment Index [PFI] Work Exhaustion Scale). Results The system generated 1274 summaries. Of 384 discharges, physicians utilized AI content in 219 (57%) cases. Feedback on 100 summaries (40.2%) noted omissions (25%) and inaccuracies (20%) but rare hallucinations (2%). Physicians rated 88% of unedited summaries as having no harm potential and 1% as likely to cause moderate harm; no severe harm was reported. Physician burnout scores decreased significantly (1.75 vs 1.20; P = .03). Time savings were heterogeneous: 71% of physicians saw reductions in median documentation time (up to 2.9 minutes). Conclusions and Relevance An LLM-based agentic workflow produced hospital course summaries that were frequently utilized with mild to minimal risk of harm identified. The intervention was associated with a significant reduction in physician burnout, supporting the viability of AI summarization to mitigate documentation burden. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Dr. Chen has received research funding support in part by: -NIH/National Institute of Allergy and Infectious Diseases (1R01AI17812101) -NIH-NCATS-Clinical & Translational Science Award (UM1TR004921) -Stanford Bio-X Interdisciplinary Initiatives Seed Grants Program (IIP) \[R12\] \[JHC\] -NIH/Center for Undiagnosed Diseases at Stanford (U01 NS134358) -Stanford Institute for Human-Centered Artificial Intelligence (HAI) -Stanford RAISE Health Seed Grant (2024) -Josiah Macy Jr. Foundation (AI in Medical Education) ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This project was conducted as a quality improvement initiative; the Stanford University Institutional Review Board determined that it did not meet the criteria for human subjects research and was therefore exempt from review. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
The emergence of large language models (LLMs) offers transformative potential for medical research. Current approaches often focus on LLMs as a replacement for researchers or as a supporting tool. In this Viewpoint, we discuss the concept of co-intelligence, leveraging the strengths of both humans and LLMs for complementary and synergistic collaboration to accelerate scientific advancement while mitigating limitations. We discuss the theoretical underpinnings of co-intelligence and advocate for its role in medical research. Through this Viewpoint, we also aim to highlight key considerations, discuss potential unintended consequences, and provide actionable insights for researchers and clinicians seeking to maximise the use of LLMs while maintaining the rigour and integrity of their work.