
BACKGROUND:Artificial intelligence (AI) is increasingly being used in healthcare settings, yet evidence of its real-world value remains inconsistent. Current evaluation paradigms often emphasize methodological rigor and technical validity over measurable improvements in patient outcomes or system performance. OBJECTIVE:To examine limitations in prevailing approaches to health AI evaluation and propose a framework prioritizing outcomes-based, systems-level assessment aligned with healthcare delivery goals. METHODS:This perspective analyzes current evaluation practices through conceptual and ethical lenses, contrasting a deontological focus on methodological standards with a consequentialist framework emphasizing real-world impact. RESULTS:A persistent gap exists between how AI systems are evaluated and how their value is realized. Technical metrics are necessary but insufficient; meaningful evaluation requires measuring clinical and operational outcomes. Strategies include standardized outcome frameworks, evaluation infrastructure, multistakeholder governance, and aligned incentives. CONCLUSIONS:Advancing health AI requires shifting from process-focused evaluation toward outcome-based assessment embedded within healthcare systems.
OBJECTIVES:Artificial intelligence (AI) is increasingly being deployed in health-care settings, yet health systems remain at an early stage in developing approaches to scale effective applications. Much of current activity is characterized by early-stage use cases and pilot projects. In this study, we explored how different health systems are approaching the development, adoption, implementation, and scaling of AI in healthcare to identify transferable lessons for system-level change. MATERIALS AND METHODS:We conducted 4 qualitative case studies of health systems and their approaches to scaling AI in healthcare: Catalonia, Norway, Singapore, and Queensland. Data were generated through documents, qualitative interviews and focus groups with strategic decision-makers, policymakers, and lead clinicians. The study explored participants' perspectives on evolving strategies and needs, rationales for change, characteristics of approaches and technologies, expected and experienced benefits, lessons learned, perceived challenges, and elements considered transferable across contexts. Analysis was conducted in 2 stages: an initial within-case thematic analysis, followed by a cross-case analysis informed by the Technology, People, Organization, and Macroenvironment framework, while also allowing inductive themes to emerge. RESULTS:The dataset consisted of 60 documents, 34 interviews, and 5 focus groups. We consulted a total of 50 different participants from case study sites across data collection activities. Our findings indicate that AI scaling trajectories in health-care systems are strongly shaped by existing digital strategies, historical and cultural contexts of digitalization, funding arrangements, and legacy technological infrastructures. Approaches to accelerating safe AI adoption at scale were characterized by efforts to orchestrate and coordinate change across multiple stakeholder groups with differing and evolving spheres of influence. Within this context, scaling was shaped by opportunities for experimentation, incentive structures, and the distribution of decision-making authority. Sustained use and wider spread of AI applications can be supported through ongoing post-deployment monitoring and the central collation and dissemination of emerging evidence. Crucially, these activities need to be embedded within a broader ecosystem that foregrounds learning, partnership, governance, and procurement as integral components of scaling. CONCLUSION:Our findings suggest that strategic decision-makers need to move beyond the conventional dichotomy of "top-down" and "bottom-up" approaches to scaling AI and instead conceptualize scaling as a process of multilevel orchestration.
OBJECTIVE:To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). MATERIALS AND METHODS:We defined 3 EHR-based tasks that are replicable across health systems and vary in reasoning complexity: (1) extracting imaging procedures (modality, date, and anatomic site), (2) generating timelines of therapeutic antibiotic use, and (3) identifying the key diagnoses for a hospitalization. Using real inpatient clinical notes from a US academic health system, we evaluated 3 large language models (GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1) with varying amounts of provided context, comparing targeted retrieval to using the most recent clinical notes. RESULTS:For Imaging Procedures, RAG strongly outperformed recent-note inputs and exceeded long-context performance (by 0.17-9.83 F1 across all models) using fewer than 8K tokens. Similar benefits were observed for Antibiotic Timelines, where <8K of retrieved tokens matched long-context recent-notes performance (between -3.26 and +3.24 Jaccard). Error analysis revealed that missing information in the clinical notes-often due to inter-hospital transfers-limited performance to some extent. However, performance on the Diagnosis Generation task remains largely static across methods and models. DISCUSSION:RAG demonstrated strong token efficiency across tasks, with the clearest and most consistent gains observed for imaging extraction and antibiotic timeline reconstruction. Diagnosis generation proved the most challenging task, suggesting ceiling effects imposed by documentation variability and evaluation constraints. CONCLUSION:Our results suggest that RAG remains a competitive and efficient approach for clinical tasks over large amounts of EHR, even as newer models become capable of handling increasingly longer amounts of text.
OBJECTIVE:To systematically examine different clinical data modalities in large language models (LLMs) and multimodal large language models (MLLMs), and to quantify the contribution of data modalities in early inpatient risk prediction and decision support tasks. MATERIALS AND METHODS:We conducted a systematic analysis using MIMIC-IV, MIMIC-IV-Note, and MIMIC-CXR-JPG datasets to create a unified cohort of 22 254 hospital admissions containing structured electronic health records (EHRs), radiology reports (clinical notes), and chest X-ray images. We evaluated general-purpose and medical-adapted LLM/VLMs across uni-, bi-, and tri-modal configurations on 2 risk prediction tasks (in-hospital mortality and length of stay [LOS] prediction) and 2 clinical decision support (CDS) tasks (discharge diagnosis phenotyping and medication-use prediction). RESULTS:For risk prediction tasks, structured EHR data alone achieved the best or comparable performance (best mortality AUROC: 0.849; LOS AUROC: 0.868), with limited incremental benefit observed from adding radiology reports or medical images. For CDS tasks, multimodal integration yielded substantial improvements: the best tri-modal configuration achieved F1-scores of 0.589 (diagnosis) and 0.405 (medication), representing 21.4% and 18.4% improvement over the best unimodal approach. Radiology reports consistently outperformed raw single-view chest radiographs as a supplementary modality. MLLMs demonstrated better zero- and few-shot performance than unimodal LLMs. Multi-view imaging consistently improved performance over single view across all tasks. CONCLUSION:The benefits of multimodal data integration are task dependent. Healthcare LLMs should examine clinical data modalities according to specific tasks for efficient integration. These findings provide practical guidance for designing efficient clinical decision support systems.
OBJECTIVE:Clinical AI systems and models increasingly need traceable documentation that makes intended use, performance evidence, transparency, and lifecycle status explicit, yet existing model cards, datasheets, reporting guidelines, and other documentation approaches do not consistently bring together structured clinical fields, vocabulary-linked cohort descriptions, structured performance reporting, and governed documentation history. We introduce Structured, Meaningful, Auditable, Responsible, and Transparent (SMART), a framework that adapts the model card concept into structured, lifecycle-aware documentation with OMOP integration, role-based lifecycle governance, and a blockchain-backed audit trail for verifying documentation integrity and lifecycle history. MATERIALS AND METHODS:Following a design-science-informed approach, we analyzed documentation requirements, reviewed existing approaches, implemented SMART through open-source packages and smart contracts, and evaluated it using example SMART model cards, an illustrative COPD exacerbation risk-prediction documentation case study, and blockchain-governed lifecycle experiments. RESULTS:The design process resulted in the SMART framework in which documentation provisions are reflected across Schema, Lifecycle, and Chain (ie, Blockchain) components. Evaluation showed that, in multi-party settings, documentation history remains independently verifiable because lifecycle actions are role-attributed and linked to content hashes; median testnet gas use ranged from 352 722 to 852 399 gas across the evaluated lifecycle operations. DISCUSSION:SMART positions clinical AI documentation as a shared infrastructure for preliminary model assessment, making role attribution, lifecycle state, and documentation changes more transparent and traceable over time. CONCLUSION:By providing an implementable framework, SMART illustrates how clinical AI documentation can be structured, clinically contextualized, auditable, responsible, transparent, and lifecycle-aware.
OBJECTIVES:Accurate citation of relevant publications is essential for scientific integrity in biomedical research. Large language models (LLMs) excel at text generation but often hallucinate fabricated or inaccurate citations. Retrieval-augmented generation (RAG) can mitigate these errors, yet current approaches lack semantic precision in evidence retrieval. This study aims to develop a domain-specific RAG system for reliable, context-specific biomedical citation recommendations. MATERIALS AND METHODS:We introduce CiteSure, a sentence-level citation recommendation tool designed to deliver reliable, evidence-based, and context-specific references using LLMs. CiteSure utilizes a 2-stage retrieval-augmented generation (RAG) framework, combining a domain-specific dense retriever (BioLLM2Vec) and reranker (BioRankLLaMA), adapted from LLaMA3-8B-Instruct using biomedical-specific training data. CiteSure leverages the complementary strengths of retrieval and generative LLM models, ensuring factual precision and contextual alignment. We evaluated CiteSure on a curated Alzheimer's disease dataset, comparing it to standalone LLMs and traditional retrieval-based methods. RESULTS:CiteSure achieved 100% factual accuracy and the highest relevance score of 77.50%, outperforming all baselines. BioLLM2Vec retrieved relevant articles with over 80% accuracy in the top 100 candidates. BioRankLLaMA consistently outperformed baseline rerankers across MAP, MRR, and Precision@5 metrics, confirming the benefit of domain-specific adaptation and contrastive fine-tuning. DISCUSSION AND CONCLUSION:Our results demonstrate that CiteSure, built on a 2-stage retrieval-augmented generation framework, effectively integrates domain-specific retrieval with LLM-based generation to achieve substantial improvements over baseline approaches. Our work underscores the importance of domain-specific adaptation in biomedical citation recommendation and provides publicly available datasets, models, and code for support future research.
OBJECTIVE:Information extraction (IE) from clinical texts has advanced rapidly with recent advances in natural language processing, particularly the advent of large language models (LLMs). However, inconsistent and incomplete reporting of methodologies limits reproducibility, comparability, and clinical translation. We aimed to develop a consensus-based reporting guideline tailored to clinical IE studies. MATERIAL AND METHODS:We developed the Clinical Information Extraction Reporting Guideline (CINEX) through a multi-phase process. The initiative was prospectively registered on the EQUATOR Network as a reporting guideline under development, and a detailed Delphi study protocol was published in advance. A scoping review informed an initial set of items, which was refined through a 3-round electronic Delphi study with 20 international experts, followed by a final consensus meeting. Items were iteratively refined based on predefined inclusion criteria and expert feedback. RESULTS:The final CINEX guideline comprises 29 checklist items grouped into 5 domains: information model, architecture, data, annotation, and outcomes. The 3 eDelphi rounds included 20, 15, and 12 experts, respectively. Two items were added after round one. Consensus for inclusion was reached for a 21 and additional 7 items after rounds 2 and 3, respectively. DISCUSSION:CINEX provides a structured framework to improve transparency, reproducibility, and interpretability in clinical IE research. By standardizing reporting of key methodological components, such as data provenance, annotation processes, and evaluation strategies, it facilitates meaningful comparison across studies and supports safer clinical implementation. CONCLUSION:CINEX complements existing AI reporting standards by addressing domain-specific challenges in clinical IE.
OBJECTIVES:The potential of "big data" in health research remains largely untapped, particularly concerning real-world data sources such as administrative health data and electronic medical records. While healthcare insurance claims data have been essential for assessing medication safety and effectiveness within indicated patient populations, exploring broader drug-outcome associations could uncover significant insights. MATERIALS AND METHODS:The REWARD (REal-World Evidence and Research of Drug performance) framework was established to perform large-scale analytics on medication benefits beyond their original indications. Employing an "all-by-all" approach, it investigates all medication-outcome pairs using standardized vocabularies from the OMOP common data model and implements causal inference methods, including self-controlled cohort and active comparator new-user designs. Negative control calibration and large-scale propensity scores are used to control for systematic bias in the study designs. RESULTS:Our framework enables the identification of new benefits associated with thousands of medications across millions of patients. In particular, REWARD facilitates insights into disease mechanisms to guide the development of novel interventions, identifies opportunities for drug repurposing, and informs potential additional indications for drugs in clinical development. REWARD has led to various publications with the goal of informing new drug development. DISCUSSION:REWARD is distinguished by its open-source implementation, use of standardized OMOP CDM vocabularies, and integration of best-practice pharmacoepidemiologic methods. Results are best interpreted as hypothesis-generating signals, with the active comparator new-user design providing higher causal rigor for prioritized drug-outcome pairs. CONCLUSION:The REWARD framework demonstrates how real-world evidence can be harnessed to address unmet medical needs, particularly for diseases lacking effective approved treatments. By making the REWARD analytic package open-source and accessible, we promote an open scientific approach, while maintaining best practices in pharmacoepidemiology.
OBJECTIVE:This study aimed to identify trajectories of multimorbidity following acute myocardial infarction (AMI), using explainable temporal machine-learning methods, and assess their clinical, prognostic, and biological significance. MATERIALS AND METHODS:Dynamic Time Warping k-means clustering was applied to post-AMI diagnostic sequences from 12 701 UK-Biobank participants. Latent Dirichlet Allocation characterised cluster themes. Multiclass classifiers (CatBoost, XGBoost, random forest, logistic regression) trained on pre-AMI diagnoses and demographics predicted trajectory membership, with SHAP interpretability. SMART scores and Cox models evaluated 5-year mortality; Phenotype-Wide Association (PheWAS) and Reactome pathway enrichment were used to identify associated biological mechanisms. RESULTS:Three trajectories of multimorbidity were identified: acute cardiorenal-respiratory with metabolic disease (ACUTE-CARD; 63.4%), cardiometabolic disease with arrhythmic-ischemic burden (CARDIOMIX; 13.5%), and smoking-related multisystem multimorbidity (SMO-CARD; 23.1%). XGBoost achieved the highest discrimination (AUC-ROC 0.906; 95% CI, 0.895-0.916), with CatBoost showing comparable performance (AUC-ROC 0.900; 95% CI, 0.889-0.910). SMO-CARD had the highest 5-year mortality (43.9%). The established SMART cardiovascular risk score remained the dominant predictor of mortality while the identified trajectories provided complementary prognostic signals; these associations were weaker after full adjustment for potential confounders with each profile displaying distinct genetic and pathway signatures. DISCUSSION:The SMART score outperformed in capturing mortality risk, whereas the trajectories complement it by revealing nuanced clustering of risk factors across organ systems and in identifying trajectory-specific intervention priorities. CONCLUSION:Explainable temporal modeling of EHR data reveals clinically interpretable, biologically grounded multimorbidity trajectories after AMI that complement established risk scores and provide a reproducible approach to mechanistic phenotyping and precision care.
OBJECTIVE:To develop and systematically compare a human-led and LLM-assisted hybrid deductive-inductive workflow for qualitative analyses. MATERIALS AND METHODS:We analyzed 122 transcripts (n = 61 research clinical consultations; n = 61 reflexive interviews) from a video ethnography study of patients with heart failure. Human-led thematic analysis used Dedoose software, and LLM-based analysis was conducted using ChatGPT Edu (GPT-5.2; OpenAI) with an eleven-prompt protocol. Both applied a hybrid deductive-inductive approach. The research team compared outputs across 63 human-LLM theme pairs using human consensus and LLM-based evaluation, integrated themes into a final framework, and manually verified quotation fidelity against the original transcripts. RESULTS:Human-led and LLM-generated analyses produced complementary cross-cutting themes (7 human; 9 LLM), all judged valid and integrated into 14 final themes across three domains. Thematic overlap was moderate to substantial (Hit Rate 1.00; Jaccard 0.44-0.51). Robustness testing across three runs revealed recurrence of five core concepts alongside variability in theme labels and counts. Quotation fidelity showed 68% verbatim, 20% paraphrased, 6% partial and 3% full hallucinations, and 3% truncated excerpts; verbatim quotations did not always clearly support their assigned themes. DISCUSSION:LLM-assisted analysis is feasible for large-scale qualitative health research within a HIPAA-compliant environment using an adaptable eleven-prompt protocol. Human oversight remained essential for contextual interpretation, quotation verification, and assessment of theme-quotation support. CONCLUSIONS:LLMs are best positioned as analytic partners rather than autonomous coders. Transparent workflows with human-in-the-loop validation are essential for responsible AI integration in health and biomedical informatics.
BACKGROUND:Large-scale propensity score (LSPS) models are increasingly used to control confounding in observational studies, but their reliability in small-sample settings is unclear. Small samples can limit a study's ability to estimate propensity scores accurately and achieve adequate covariate balance, raising concerns about insufficient confounding adjustment. Resulting in small datasets being excluded, despite potentially containing valid information. METHODS:We evaluated the balancing and bias reduction performance of LSPS models using six target-comparator pairs. To emulate a distributed data network analysis setting, we partitioned two large real-world data sources into smaller subsets. Within each subset, LSPS models were fit locally and evaluated on covariate balance, treatment effect bias and precision after empirical calibration, using real negative controls and synthetic positive controls. Performance was compared to a global LSPS model trained on the full dataset. RESULTS:Across most target-comparator pairs, LSPS models substantially reduced bias and, when applying empirical calibration, improved precision. PS matching was more resilient to small sample sizes than PS stratification, achieving adequate balance even at n = 500, while stratification sometimes failed at n = 4000. A slight bias increase was observed at the smallest sample sizes, though not universally. Standard balance diagnostics consistently failed below n = 20 000, while a recently proposed diagnostic accounting for chance imbalance did not. CONCLUSIONS:LSPS models generally provide reliable bias reduction in small-sample settings, supporting their use in federated analyses. However, standard balance diagnostics may be misleading in small samples, and alternatives should be considered, such as significance checking. When LSPS fails to reduce bias adequately, additional adjustment strategies are required.
OBJECTIVE:Long COVID (LC) remains poorly understood, and there is a critical need for advanced computational tools to better identify and characterize patients. In this study, we use summarized symptom reports by RECOVER-Adult cohort participants linked to EHR data to characterize patients and train a computable phenotype algorithm of LC. MATERIALS AND METHODS:The study included adult participants with linked FHIR-sourced EHR data. We characterized EHR diagnoses, procedures, medications, lab tests, and vital sign features associated with LC. A computable phenotyping algorithm was trained and validated against patient-reported symptoms. MAIN OUTCOME AND MEASURES:We assessed model discrimination and calibration in a held-out test set. We describe important model features and evaluate model discrimination and calibration. RESULTS:The study included 1,501 RECOVER-Adult cohort participants with linked EHR data. 376 (25%) met criteria for highly symptomatic LC based on the RECOVER Long COVID Research Index (LCRI). EHR features associated with LC included clinician diagnosis of shortness of breath, malaise and fatigue, and cardiac dysrhythmias; documented treatment with albuterol, gabapentin, or duloxetine; or elevated heart rate. The algorithm identifying patients with highly symptomatic LC had an AUROC of 0.80 (95% confidence interval (CI) 0.74-0.85), and AUPRC of 0.58 (95% CI, 0.47-0.69). CONCLUSION AND RELEVANCE:These findings demonstrate that, using EHR data, a machine-learning model can accurately select patients with sets of self-reported LC symptoms. The model could help identify patients within a health system with the highest probability of the condition and facilitate screening, recruitment for clinical trials, and etiologic studies.
OBJECTIVE:The analysis of care trajectories derived from electronic health records and claims data has become increasingly common in biomedical informatics. This has enabled large-scale studies of care processes, yet widely used binary code representations result in high-dimensional, sparse data that fail to capture semantic relationships between medical concepts. Learning dense vector representations (embeddings) has emerged as a promising approach to address these limitations. We aimed to construct and share joint embeddings for the International Classification of Diseases (ICD-10) and the Anatomical Therapeutic Chemical (ATC) classification system, providing reusable semantic representations of diagnoses and treatments from real-world claims data. MATERIALS AND METHODS:Using claims records from 1.5 million patients, we defined code co-occurrences within temporal windows and constructed a Positive Pointwise Mutual Information (PPMI) matrix spanning ICD-10 and ATC codes. Singular Value Decomposition (SVD) was applied to derive a low-dimensional embedding space. Evaluation combined UMAP visualization, nearest-neighbor retrieval, and a code-level classification task based on ICD chapters and ATC classes. RESULTS:The embeddings reflected the hierarchical organization of ICD-10 and ATC and revealed associations across coding systems, including clinically relevant diagnosis-treatment relationships. The classification task achieved mean AUCs of 0.93 for ICD-10 and 0.90 for ATC, indicating strong grouping of semantically related codes. DISCUSSION:The embeddings provide a reusable, code-level semantic representation that can support code retrieval, reduce manual code grouping, and be aggregated into patient-level features without training a task-specific model. CONCLUSION:We release the first openly available joint ICD-10-ATC embedding space derived from real-world claims data, providing a reusable resource for biomedical informatics research.
OBJECTIVES:To describe considerations for integration of human and artificial intelligence for creating a postmarketing surveillance system capable of timely and reliably identifying causal effects of medications on safety endpoints. MATERIALS AND METHODS:The FDA has prioritized more extensive Electronic Health Records (EHR) integration along with generative artificial intelligence and machine learning (Gen AI/ML) into the national active surveillance program for medical products-the Sentinel Initiative. Based on our experience of leading these efforts, we provide perspectives on the opportunities and challenges of Gen AI/ML integration into Sentinel. RESULTS:Using specific examples, we outline the role of Gen AI and ML in a causal inference framework for scalable information extraction, assessment of fitness-for-purpose of data sources, diagnosing residual confounding, and enhancing confounding adjustment. Critically, we outline steps and checkpoints along the way where human involvement remains indispensable. DISCUSSION AND CONCLUSION:In public health applications where stakes are high, use of Gen AI/ML needs to be carefully considered with appropriate guardrails ensuring human expert involvement.
OBJECTIVE:The successful integration of Machine Learning (ML) models into clinical practice remains limited, as they often lack the standardized, quantifiable risk measures essential for clinical workflows. This study, therefore, aims to demonstrate the Unified Auto Clinical Scores (Uni-ACS) method as a means to translate ML predictions into interpretable biostatistical formats, a translation critical for clinical adoption and alignment with evidence-based practice guidelines. MATERIALS AND METHODS:We employed the Uni-ACS post hoc methodology to convert ML model outputs into clinical scores and odds ratios. Validation was performed on a retrospective cohort of Chronic Obstructive Pulmonary Disease patients admitted between 2016 and 2018, predicting adverse outcomes from 2019 to 2020. The performance of Uni-ACS was benchmarked against the original ML models and logistic regression. RESULTS:The successful application of Uni-ACS enabled the conversion of complex ML model outputs into readily interpretable clinical scores and metrics, like odds ratios. Clinical scores derived from the ML models' SHAP values using Uni-ACS maintained a strong, interpretable predictive performance (AUROC 0.69-0.80), which was comparable to the original ML models (AUROC 0.71-0.83) and logistic regression (AUROC 0.72-0.81). DISCUSSION:Uni-ACS addresses the fundamental clinical requirement for standardized, meaningful risk stratification. It efficiently translates complex ML predictions into validated clinical scores with minimal computational burden. CONCLUSION:In conclusion, this approach facilitates the widespread adoption of ML-driven risk assessment across diverse healthcare settings while ensuring compatibility with existing clinical guidelines and regulatory frameworks.
OBJECTIVES:We examine how hospital characteristics relate to clinical and operational artificial intelligence (AI) adoption and implementation stages and characterize AI deserts and spatial clustering patterns to highlight place-based AI access gaps among United States (US) hospitals. MATERIALS AND METHODS:We used the 2024 American Hospital Association Annual Survey data from 2720 hospitals with at least one AI response. We applied logistic regression models to examine the associations between hospital characteristics and AI adoption, local indicators of spatial association to identify local clusters, and distance analyses to locate AI desert hospitals (ie, geographically isolated nonadopters located >50 miles from the nearest AI adopters). RESULTS:We found that system membership, larger bed size, and nurse staffing intensity are positively associated with clinical and operational AI adoption and implementation stages, whereas rural location, for-profit ownership, and physician intensity are negatively associated with them. Teaching status is more strongly associated with clinical AI, while system membership is more positively associated with operational AI. Although 68.0% of hospitals adopted ≥1 clinical AI functionalities and 60.7% adopted ≥1 operational AI functionalities, 12.1% of clinical and 13.2% of operational non-adopters are classified as AI desert hospitals. DISCUSSION:AI deserts reveal regional implementation gaps; sustained diffusion requires building shared regional capacity rather than relying only on hospital-level incentives. Policy implications to emerging divides may include tracking implementation stage, providing domain-specific support, and funding regional partnerships. CONCLUSION:US hospital AI diffusion is uneven and spatially structured. AI deserts mark regional gaps in access to AI-adopting hospitals.
OBJECTIVE:Human Factors (HF) methods have potential to ensure the design of health information technology (HIT) is safe and usable, but limited research has examined what HF methods have been applied or could be applied to HIT. This study aimed to identify key HF and safety analysis methods used and/or recommended by HF experts working in healthcare or non-healthcare industries that may be suitable for supporting the design, redesign and configuration of HIT. MATERIALS AND METHODS:We ran semi-structured interviews with 21 HF experts working across health and non-health industries to identify key HF and safety analysis methods used and/or recommended for supporting the design, redesign and configuration of HITs. RESULTS:Forty-seven HF methods across 11 method categories were reported by HF experts. Non-health participants reported a wider range of Human Error Identification and Risk Assessment Methods and were the only cohort to report workload assessment methods. Both health and non-health participants described their approaches for selecting methods and recommended that methods be applied flexibly. DISCUSSION:This study highlights the need to uplift HF capability by teaching methods to system designers, particularly where access to HF experts is limited and undertaking efforts to learn from HF practice in safety-critical industries outside of healthcare. CONCLUSION:Further work is required to integrate HF methods and approaches, especially those focused on safety, into the work of HIT designers.
Objective Secondary headaches require urgent recognition due to potentially devastating consequences if untreated. Despite established clinical "red flag" criteria, identifying patients needing immediate evaluation remains challenging in primary care. This study developed and evaluated a large language model (LLM)-based multi-agent clinical decision support system for interpretable secondary headache diagnosis.Materials and Methods We first established 7 clinically relevant secondary headache red flag domains through manual review and synthesis of clinical guidelines. Based on these domains, we designed an LLM-based system using an orchestrator-specialist multi-agent architecture that decomposes diagnostic reasoning into 7 guideline-aligned agents corresponding to key red flag features. Each agent generates structured, evidence-grounded reasoning, coordinated by a central orchestrator. The system was evaluated on 90 expert-validated secondary headache cases and compared with a single-LLM baseline under 2 prompting strategies: question-based prompting (QPrompt) and guideline-based prompting (GPrompt). Five open-source LLMs (Qwen-8b, Qwen-14b, Qwen-30b, GPT-OSS-20b, and Llama-3.1-8b) were tested.Results The orchestrated multi-agent system with GPrompt achieved the highest red flag classification performance across models, measured by F1 score. Performance gains were consistent and more pronounced in smaller LLMs, suggesting that structured reasoning improves efficiency and accuracy beyond prompt engineering alone. The framework also produced transparent and guideline-aligned intermediate reasoning.Discussion Decomposing clinical reasoning into specialized agents enhances interpretability and diagnostic reliability compared with monolithic LLM approaches. Multi-agent orchestration provides a clinically aligned framework for explainable decision support.Conclusion An orchestrator-specialist multi-agent LLM framework improves secondary headache diagnosis accuracy and transparency, supporting the development of explainable AI systems for time-constrained clinical decision-making in primary care.