Background Context. Low back pain (LBP) is common and a major cause of disability globally. Generative artificial intelligence (GenAI) such as ChatGPT may have the potential to improve LBP, but it is unknown how GenAI is being used in clinical care settings. Purpose. This scoping review aimed to map and synthesise the existing literature on the use of GenAI in the management of LBP across all settings. Study Design. This review followed the scoping review methodology of the Joanna Briggs Institute and Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines. Methods. We searched the Cochrane Library, Medline, Scopus, Embase, Web of Science, CINAHL, and preprint platforms for studies published between January 2012 and July 2025, on the use of GenAI in the management of LBP across all healthcare settings. The included studies were categorised according to the primary purpose of GenAI implementation. Results. A total of 31 studies were included. Seventeen studies evaluated GenAI as a potential support tool for diagnosis and treatment, including answering specific questions about LBP (n = 8), generating advice from clinical cases (n = 4), and generating advice from diagnostic imaging (n = 5). Eight studies examined the use of GenAI to provide educational advice for patients. Other uses included research support (n = 5) and clinical documentation support (n = 1). GenAI advice for the management of LBP demonstrated higher agreement when using simpler tasks, more detailed prompts, or newer generation models. Reported errors in GenAI advice included inconsistency with guidelines or clinicians’ advice, hallucination, low accuracy for complex questions, sensitivity to prompt design, and challenges with interpretability. Conclusion. Thirty-one studies reported the use of GenAI in LBP management, mainly supporting diagnosis and treatment planning, patient education, research, and clinical documentation. Despite authors’ claims about promising uses of GenAI, it typically requires further refinement or review by human end-users. Critically, no studies addressed end-user experience, or clinical effectiveness, underscoring the need for robust clinical evaluation beyond the lab-based environment.
OBJECTIVE:To evaluate and compare the performance of large language models (LLMs) in identifying contributing factors (CFs) underlying patient safety incident investigations. MATERIALS AND METHODS:Four open-source, lightweight LLMs, including BERT, LLaMA2, GPT2, and Phi-2 were applied to classify CFs across 6 sociotechnical system-levels encompassing 12 categories (eg, person, task, and organizational factors). Reports of real-world patient safety investigations from public health systems were extracted and labelled by domain experts (n_report/CFs = 300/1338). Data were split into training (n = 852), validation (n = 98), and test sets (n = 388). Performance was evaluated using specificity, precision, recall, and F1 scores. RESULTS:The fine-tuned encoder-based BERT model achieved the highest performance, with a micro-averaged F1 score of 63.6%, outperforming all decoder-based models. Among the decoder models, Phi-2 demonstrated the strongest performance (F1 = 54.9%), exceeding both LLaMA2 and GPT2. BERT performed consistently across 6 system-levels but often misclassified "organization" as "person". DISCUSSION:LLMs hold promise for automating the extraction of CFs from complex safety narratives, particularly for frequently reported system-levels such as "person" and "tasks". Such automation may substantially reduce the manual effort required to analyse reports of patient safety investigations while supporting more consistent analysis across large incident datasets. CONCLUSION:Applying LLMs to analyse the underlying causes of patient safety incidents depends on developing high-quality, domain-specific datasets that enhance the representation of patient safety knowledge and improve model understanding of incident causation. Improving data coverage for rare system-levels is essential to address the current limitations of LLMs in capturing nuanced patient safety concepts and domain-specific reasoning.
While the use of AI technologies in healthcare is increasing, there is legal uncertainty about whether, and to what extent, patients should be informed about clinicians' use of these cuttingedge technologies under current medical negligence law. Transparency around AI use has been advocated in numerous policy documents and academic papers on ethical AI, but there is a disagreement in medical negligence law literature on how much transparency-or informationshould be provided around AI use in healthcare to patients. Some legal commentators argue that patients need to know about AI use in all cases, while others suggest that clinicians do not traditionally tell patients about all the tools they use in the provision of healthcare, and thus AI use needs to be disclosed in exceptional cases only. This paper presents results of focus groups with patients that examined their information needs about AI use in their healthcare. After finding that focus group participants generally want to be informed about AI, the paper then analyses these findings in light of current medical negligence law and legal commentary. It concludes that current medical negligence law is unlikely to require disclosure of AI in all AI use cases and that the revision of the material risk standard under medical negligence law might not be required. At the same time, healthcare professional standards on informed consent might set stronger AI disclosure requirements to help increase trust and acceptance of new AI technologies.
Ramaswamy et al. reported in Nature Medicine that ChatGPT Health under-triages 51.6% of emergencies, concluding that consumer-facing AI triage poses safety risks. However, their evaluation used an exam-style protocol – forced A/B/C/D output, knowledge suppression, and suppression of clarifying questions – that differs fundamentally from how consumers use health chatbots. We tested five frontier LLMs (GPT-5.2, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 3 Flash, Gemini 3.1 Pro) on a 17-scenario partial replication bank under constrained (exam-style, 1,275 trials) and naturalistic (patient-style messages, 850 trials) conditions, with targeted ablations and prompt-faithful checks using the authors' released prompts. Naturalistic interaction improved triage accuracy by 6.4 percentage points (p = 0.015). Diabetic ketoacidosis was correctly triaged in 100% of trials across all models and conditions. Asthma triage improved from 48% to 80%. The forced A/B/C/D format was the dominant failure mechanism: three models scored 0–24% with forced choice but 100% with free text (all p < 10^-8), consistently recommending emergency care in their own words while the forced-choice format registered under-triage. Prompt-faithful checks on the authors' exact released prompts confirmed the scaffold produces model-dependent, case-dependent results. The headline under-triage rate is highly contingent on evaluation format and should not be interpreted as a stable estimate of deployed triage behavior. Valid evaluation of consumer health AI requires testing under conditions that reflect actual use.
Background: Artificial intelligence (AI)-driven clinical decision support (CDS) tools offer promising solutions for health care delivery by optimizing resource allocation, detecting deterioration, and enabling early interventions. However, adoption remains limited due to insufficient validation and a lack of transparency and trust. Explainable AI (XAI) seeks to improve user understanding of AI outputs; however, how clinicians interpret and integrate these explanations into their decision-making remains underexplored. Furthermore, discrepancies in explanations, known as the "disagreement problem," can undermine trust and, at worst, lead to poor clinical decisions. Objective: This study examines clinicians' perspectives on the role and value of explainability in AI-driven CDS tools within Australian critical care settings and the impact of discrepancies in AI-generated explanations on clinical decision-making. Methods: Qualitative data were collected using semistructured interviews with 14 clinical experts, incorporating scenario-based exercises, and were analyzed using inductive thematic analysis. Results: Clinicians valued explainability, particularly in complex or unfamiliar situations, when explanations were clear, plausible, and actionable. Trust and perceived usefulness extended beyond explanation quality, encompassing factors such as system accuracy, alignment with clinicians' reasoning, workflow integration, and perceived reliability. Discrepancies in explanations generated by different XAI methods were not a major concern, provided that the AI-generated predictive alerts were accurate. Conclusions: This study provides design recommendations for developing trustworthy, user-centric CDS tools that incorporate XAI. Findings highlight that explainability is critical for establishing initial trust in AI-driven tools by supporting perceived usefulness, but its importance diminishes over time and with user expertise and familiarity, as learned usefulness takes precedence. Recommendations highlight the importance of aligning the design and implementation of AI tools with clinicians' needs to enhance trust, mitigate risks, and promote successful adoption for improved patient outcomes.
In the resource-constrained settings, shortages of health workers and underdeveloped infrastructure hinder the delivery of equitable, high quality care. This perspective outlines strategic principles to support the design of AI tools that are genuinely beneficial in low-resource contexts: adopting problem-driven approaches, understanding the socio-technical context, selecting appropriate clinical tasks, and ensuring point-of-care accessibility, clinical comprehensibility, and actionable recommendations, ultimately improving clinicians’ decision-making and promoting health equity in low- and middle-income countries.
Background Globally, up to 17% of hospitalised people suffer a patient safety incident. Learning from adverse events through patient safety investigation is critical to prevention; however, their utility is still questioned. Two key investigation outputs include identifying contributing factors (CFs) and proposing recommendations to prevent future occurrences. Criticisms of current methods include incomplete analysis of CFs and weak incident prevention strategies. A proposed solution is systems thinking analysis, which recognises healthcare complexity. However, it is not clear whether such methods are being applied in practice.Objective This study aimed to assess current use of systems thinking-based strategies by examining a set of Australian patient safety incident investigations.Methods Investigations (n=300) from 56 different Australian health services were deductively analysed. Identified CFs were classified by healthcare system level using a framework combining Systems Engineering Initiative for Patient Safety (SEIPS) principles and AcciMap's hierarchical structure. Recommendation sustainability and effectiveness were classified as weak, medium or strong using US Department of Veteran Affairs' criteria.Results 51% of incidents were issues with clinical processes and procedures. The investigations identified CFs that disproportionally focused on the people involved in those processes (n=677, 47%) rather than other system levels and as a consequence, most recommendations were of medium (n=665, 51%) and weak (n=560, 43%) strength. Notably, 10% of investigations lacked any CFs or recommendations.Conclusion The focus on individual actions highlighted that simple linear thinking persists in patient safety incident investigations. This study proposes five key areas of effective incident analysis and investigation: a sociotechnical focus; improved data collection techniques; investigative independence; the professionalisation of investigators; and the aggregation of data. Learning from incidents is key to maximising their preventative effectiveness, especially in an increasingly complex healthcare system.
Adverse events affect one in eight hospitalised patients, leading to significant harm and healthcare burden. Many incidents, including documentation errors and pressure injuries, are preventable but require timely identification and continuous learning to keep pace with rapid care delivery. Manual review cannot support real-time safety surveillance given the complexity and growing volume of incident reports; automated incident classification is therefore critical for timely cluster detection and identification of emerging risks. We present the first application of large language models (LLMs) to classify a comprehensive set of 22 incident types aligned with the World Health Organisation’s International Classification for Patient Safety. To address data privacy requirements, we fine-tuned four open-source, lightweight LLMs (BERT, LLaMA2, GPT-2, and Phi-2) using expert-annotated incident reports from an Australian state-level reporting system (n = 5773) and evaluated model generalisability using an independent dataset from another state ( n = 569). BERT, LLaMA2, and Phi-2 demonstrated strong feasibility for classifying the 22 incident types and showed consistent performance across two independent datasets. Performance exceeded 80% for common, well-defined types, including falls, nutrition, oxygen gas/vapour , and aggression . Incident types with moderate performance (∼70%), including accident and deteriorating patient , also generalised reliably, suggesting that LLMs effectively recognise clearly articulated concepts. However, performance declined (<60%) for incident types involving complex or overlapping clinical processes, such as clinical handover, organisational management, and complaints , with persistent weaknesses observed across datasets. Error analysis showed that misclassification frequently centred on clinical management , a broad category covering multiple incident domains (e.g. blood products, infections, pressure injuries, medication ). LLMs show promise to assist with the classification of well-defined incident types but face substantial challenges with heterogeneous or conceptually overlapping types. These findings highlight the need for domain-informed, retrieval-augmented approaches and publicly available benchmark datasets to improve the scalability, robustness, and generalisability of open-source LLMs across diverse reporting cultures, languages, and jurisdictions.
This scoping review synthesises current evidence on artificial intelligence (AI) governance in healthcare organisations, outlining key components of AI governance frameworks. Following PRISMA-ScR guidelines, we searched MEDLINE, Embase, and Scopus (April 2024, updated March 2025) for AI governance frameworks in acute care. Seventy-seven frameworks were identified and examined for four components: (1) Guiding principles (ethics or governance-related); (2) Assessment methods; (3) AI lifecycle stages; and (4) Oversight mechanisms. Most frameworks were not applicable to real-world healthcare settings and missed key principles or components, such as an oversight mechanism. Only 10 frameworks (13.0%) included all four framework components, with oversight mechanisms (e.g. AI-specific governance committee) being the least common (n = 15, 19.5%). There is a need to move beyond principles to implementing AI governance frameworks in healthcare organisations and evaluating their real-world impact.
Problem description Measuring quality of care, represented by quality indicators (QIs), in long term care (LTC) residents is a key imperative for care systems globally. To measure at scale, algorithms can be developed to automatically extract QIs from residents' records. However, much information documented in residents' care records is free-text, making the automated extraction of QIs challenging. OBJECTIVES:To analyse the proportion of QIs that are extracted from structured data fields in LTC residents' care records and therefore are amenable to automated extraction. METHODS:A total of 236 QIs were developed for 16 conditions, e.g., Cognitive impairment, End of life care. Trained LTC nurses then undertook a manual care record review to assess the percentage of CareTrack Aged QIs that were collected and stored in structured data fields in the care record. RESULTS:The underlying QI data field type was extracted from six electronic care record systems across 8,333 QI assessments (encounters of care) for 118 residents in 42 facilities. To determine QI eligibility, 17 % (n = 1425/8333) of QIs used structured data fields, with 39 % hybrid (structured and free-text) and 44 % free-text. QI eligibility assessment using structured data fields varied across the care record systems (Range 11-26 %). To determine QI adherence, 5 % (n = 251/4930) of QIs used structured data fields, with 60 % hybrid and 35 % free-text. QI adherence assessment using structured data fields varied across the care record systems (Range 1-12 %). Combining QI eligibility and adherence, 13 % (n = 1676/13263) of indicator question information used structured data fields. CONCLUSION:This study challenges the assumption that QIs that are representative of care delivered in LTCs can be largely algorithmically electronically extracted from structured data fields. QI assessments for evidence-based care were extracted predominantly from free-text and hybrid data fields in LTC residents' records.
Objective To evaluate the transferability of BERT (Bidirectional Encoder Representations from Transformers) to patient safety, we use it to classify incident reports characterised by limited data and encompassing multiple imbalanced classes.Methods BERT was applied to classify 10 incident types and 4 severity levels by (1) fine-tuning and (2) extracting word embeddings for feature representation. Training datasets were collected from a state-wide incident reporting system in Australia (n_type/severity=2860/1160). Transferability was evaluated using three datasets: a balanced dataset (type/severity: n_benchmark=286/116); a real-world imbalanced dataset (n_original=444/4837, rare types/severity<=1%); and an independent hospital-level reporting system (n_independent=6000/5950, imbalanced). Model performance was evaluated by F-score, precision and recall, then compared with convolutional neural networks (CNNs) using BERT embeddings and local embeddings from incident reports.Results Fine-tuned BERT outperformed small CNNs trained with BERT embedding and static word embeddings developed from scratch. The default parameters of BERT were found to be the most optimal configuration. For incident type, fine-tuned BERT achieved high F-scores above 89% across all test datasets (CNNs=81%). It effectively generalised to real-world settings, including rare incident types (eg, clinical handover with 11.1% and 30.3% improvement). For ambiguous medium and low severity levels, the F-score improvements ranged from 3.6% to 19.7% across all test datasets.Discussion Fine-tuned BERT led to improved performance, particularly in identifying rare classes and generalising effectively to unseen data, compared with small CNNs.Conclusion Fine-tuned BERT may be useful for classification tasks in patient safety where data privacy, scarcity and imbalance are common challenges.
The U.S. Food and Drug Administration (FDA) plays an important role in ensuring safety and effectiveness of AI/ML-enabled devices through its regulatory processes. In recent years, there has been an increase in the number of these devices cleared by FDA. This study analyzes 104 FDA-approved ML-enabled medical devices from May 2021 to April 2023, extending previous research to provide a contemporary perspective on this evolving landscape. We examined clinical task, device task, device input and output, ML method and level of autonomy. Most approvals (n = 103) were via the 510(k) premarket notification pathway, indicating substantial equivalence to existing devices. Devices predominantly supported diagnostic tasks (n = 81). The majority of devices used imaging data (n = 99), with CT and MRI being the most common modalities. Device autonomy levels were distributed as follows: 52% assistive (requiring users to confirm or approve AI provided information or decision), 27% autonomous information, and 21% autonomous decision. The prevalence of assistive devices indicates a cautious approach to integrating ML into clinical decision-making, favoring support rather than replacement of human judgment.
BackgroundArtificial intelligence (AI) has the potential to improve health care delivery through enhanced diagnostics, streamlined operations, and predictive analytics. However, health care organizations face substantial challenges in implementing AI safely and responsibly. This is due to regulatory complexity, ethical considerations, and a lack of practical governance frameworks. While many theoretical frameworks exist, few have been tested or adapted for real-world application in health care settings. ObjectiveThis study aims to develop and validate a practical AI governance framework to support the safe and responsible use of AI in health care organizations. The specific objectives are to identify governance requirements for AI in health care, examine existing AI governance processes and best practices, codevelop an AI governance framework to meet the needs of health care organizations, and test and refine the framework through real-world application. MethodsA multimethod research design will be used, comprising four key stages: (1) a scoping review and document analysis to identify governance needs and current processes, (2) in-depth interviews with health care stakeholders as well as national and international AI governance experts, (3) development of a draft AI governance framework through a synthesis of findings, and (4) validation and refinement of the framework through stakeholder workshops and application to case studies of AI tools. Data will be analyzed using qualitative methods informed by grounded theory. ResultsThe project received funding in October 2023. Ethics approval was obtained from the Alfred Health Human Research Ethics Committee (project 171/24) and the Macquarie University Human Research Ethics Committee (project 16508). Data collection commenced in April 2024, with the scoping review and document analysis being finalized. As of March 2025, a total of 43 interviews have been completed. The final AI governance framework is expected to be completed and ready for dissemination by June 2025. ConclusionsThis study will deliver a comprehensive AI governance framework co-designed with health care stakeholders to address real-world challenges in AI oversight. The framework will offer practical guidance to support health care organizations in adopting AI technologies safely, ethically, and in alignment with regulatory requirements. Outcomes from this study will inform local and international discussions on AI governance and promote the responsible integration of AI in health systems. International Registered Report Identifier (IRRID)DERR1-10.2196/75702
Background Artificial intelligence (AI) has the potential to improve health care delivery through enhanced diagnostics, streamlined operations, and predictive analytics. However, health care organizations face substantial challenges in implementing AI safely and responsibly. This is due to regulatory complexity, ethical considerations, and a lack of practical governance frameworks. While many theoretical frameworks exist, few have been tested or adapted for real-world application in health care settings. Objective This study aims to develop and validate a practical AI governance framework to support the safe and responsible use of AI in health care organizations. The specific objectives are to identify governance requirements for AI in health care, examine existing AI governance processes and best practices, codevelop an AI governance framework to meet the needs of health care organizations, and test and refine the framework through real-world application. Methods A multimethod research design will be used, comprising four key stages: (1) a scoping review and document analysis to identify governance needs and current processes, (2) in-depth interviews with health care stakeholders as well as national and international AI governance experts, (3) development of a draft AI governance framework through a synthesis of findings, and (4) validation and refinement of the framework through stakeholder workshops and application to case studies of AI tools. Data will be analyzed using qualitative methods informed by grounded theory. Results The project received funding in October 2023. Ethics approval was obtained from the Alfred Health Human Research Ethics Committee (project 171/24) and the Macquarie University Human Research Ethics Committee (project 16508). Data collection commenced in April 2024, with the scoping review and document analysis being finalized. As of March 2025, a total of 43 interviews have been completed. The final AI governance framework is expected to be completed and ready for dissemination by June 2025. Conclusions This study will deliver a comprehensive AI governance framework co-designed with health care stakeholders to address real-world challenges in AI oversight. The framework will offer practical guidance to support health care organizations in adopting AI technologies safely, ethically, and in alignment with regulatory requirements. Outcomes from this study will inform local and international discussions on AI governance and promote the responsible integration of AI in health systems. International Registered Report Identifier (IRRID) DERR1-10.2196/75702
BACKGROUND:Artificial intelligence (AI) tools could assist emergency doctors interpreting chest X-rays to inform urgent care. However, the impact of AI assistance on clinical decision-making, a precursor to enhanced care and patient outcomes, remains understudied. This study evaluates the effect of AI assistance on clinical decisions of emergency doctors interpreting chest X-rays. METHOD:Junior and senior residents, emergency registrars and consultants working in Australian emergency departments were eligible. Doctors completed 18 clinical vignettes involving chest X-ray interpretation, representative of typical patient presentations. Vignettes were randomly selected from a bank of 49 based on the emergency medicine curriculum and contained a chest X-ray, presenting complaint, relevant symptoms and observations. Of the 18 vignettes, each doctor was randomly assigned to have half assisted by a commercial AI tool capable of detecting 124 different chest X-ray findings. Four vignettes contained X-rays known to produce incorrect AI findings. Primary outcomes were correct diagnosis and patient management. X-ray interpretation time, confidence of diagnosis, perceptions about the AI tool and the differential impact of AI assistance by seniority were also examined. RESULTS:200 doctors participated. AI assistance increased correct diagnosis by 5.9% (95% CI 2.7 to 9.2%) compared with unassisted vignettes, with the largest increase among senior residents (11.8%; 95% CI 5.2% to 18.3%). Patient management increased by 3.2% (95% CI 0.1% to 6.4%). Confidence in diagnosis increased by 5% (95% CI 3.4% to 6.6%; p<0.001) and interpretation time increased by 4.9 s (p=0.08). Incorrect AI findings decreased correct diagnosis by 1% for false-positive (p=0.9) and 9% for false-negative findings (p=0.1). Participants found the AI tool helpful for interpreting chest X-rays, highlighting missed findings, but were neutral on its accuracy. CONCLUSION:Improvements in diagnosis and patient management without meaningful increases in interpretation time suggest AI assistance could benefit clinical decisions involving chest X-ray interpretation. Further studies are required to ascertain if such improvements translate to improved patient care.
This study explores the governance challenges and opportunities associated with the implementation and use of AI in healthcare, offering practical insights informed by extensive stakeholder interviews. By analyzing perspectives from academia, government, clinicians, healthcare associations, and consumer groups, the research highlights critical themes, including data governance, ethical considerations, accountability, risk management, and equity. Findings highlight the need for robust frameworks to address fragmented data systems, privacy concerns, and systemic biases while recommending tiered risk management approaches to AI implementation. Enhanced AI literacy, improved infrastructure, and centralized governance structures are proposed to streamline decision-making and ensure safe, equitable, and sustainable AI adoption in healthcare organizations.
The rapid growth of clinical explainable AI (XAI) models raised concerns over unclear purposes and false hope regarding explanations. Currently, no standardised metrics exist for XAI evaluation. We developed a clinician-informed, 14-item checklist including clinical, machine and decision attributes. This is the first step toward XAI standardisation and transparent reporting XAI methods to enhance trust, reduce risks, foster AI adoption, and improve decisions to determine the true clinical potential of applied XAI.
Despite the excitement surrounding Artificial Intelligence (AI) in health care, one of the key concerns is the lack of transparency, which is essential for ensuring quality, safety, and trust around AI technologies. This article examines the notion of AI transparency and its traditional role in health care. It examines its ambiguous meaning and identifies the emerging consensus in recent policy documents to distinguish it from related concepts such as AI explainability. The article explores the rationales underlying the AI transparency principle for different stakeholders and investigates potential challenges for enhancing transparency to users of AI-based medical devices. It concludes that, although the need for transparency around AI-based medical devices is widely recognised, the main hurdle is a lack of clarity about the level of information to be provided across different contexts and stakeholders, and how best to provide this information to achieve transparency goals.
Elske Ammenwerth合作论文数Health Informatics and the Institute for Health Information Systems at10