Medical informatics training consists of two types of fellowship programs: clinical informatics (CI) and imaging informatics (II). Historically, there has been little collaborative learning between these fellowships despite significant overlap in content. This project created a collaborative workshop series for CI and II fellows, which taught mutually-relevant topics and provided an opportunity for cross-disciplinary collaboration in designing workflows for enterprise imaging scenarios. Content was informed by the American Board of Preventive Medicine CI Test Content Outline and an II fellowship curriculum previously published by II fellowship directors. Outputs of the exercise included a Venn diagram displaying similarities and differences between the two specialties, as well as the enterprise imaging workflow designs of each small group. The workshop series was well-received by participating fellows in a follow-up survey, with most (77%) indicating that the content was effective or very effective at showing the need for collaboration between CI and II disciplines and that the knowledge gained would likely be implemented in their future work. Potential areas for improvement include: clarifying expectations at the outset, translating ideas generated in the curriculum into real-world projects, and optimizing learner experience in a partially virtual platform. A collaborative workshop series between imaging and clinical informatics fellows is the first of its kind in informatics, both in content and teaching style. Continued efforts to increase the quality and reach of such integrated learning are needed to prepare future clinical and imaging informaticists to meet the demands of the highly interdependent modern healthcare system.
Multimodal large language models have demonstrated comparable performance to that of radiology trainees on multiple-choice board-style exams. However, to develop clinically useful multimodal LLM tools, high-quality benchmarks curated by domain experts are essential. To curate released and holdout datasets of 100 chest radiographic studies each and propose an artificial intelligence (AI)-assisted expert labeling procedure to allow radiologists to label studies more efficiently. A total of 13,735 deidentified chest radiographs and their corresponding reports from the MIDRC were used. GPT-4o extracted abnormal findings from the reports, which were then mapped to 12 benchmark labels with a locally hosted LLM (Phi-4-Reasoning). From these studies, 1,000 were sampled on the basis of the AI-suggested benchmark labels for expert review; the sampling algorithm ensured that the selected studies were clinically relevant and captured a range of difficulty levels. Seventeen chest radiologists participated, and they marked "Agree all", "Agree mostly" or "Disagree" to indicate their assessment of the correctness of the LLM suggested labels. Each chest radiograph was evaluated by three experts. Of these, at least two radiologists selected "Agree All" for 381 radiographs. From this set, 200 were selected, prioritizing those with less common or multiple finding labels, and divided into 100 released radiographs and 100 reserved as the holdout dataset. The holdout dataset is used exclusively by RSNA to independently evaluate different models. A benchmark of 200 chest radiographic studies with 12 benchmark labels was created and made publicly available https://imaging.rsna.org, with each chest radiograph verified by three radiologists. In addition, an AI-assisted labeling procedure was developed to help radiologists label at scale, minimize unnecessary omissions, and support a semicollaborative environment.
Objective To describe the technical workflow enabling scalable automated artificial intelligence (AI) monitoring in the first national imaging AI registry, Assess-AI. Materials and Methods Large language model (LLM) prompts are developed to extract clinically relevant findings from radiology reports through collaboration between data scientists and subspecialty radiologists. Prompts are optimized using tuning cohorts of use case-specific radiology reports and LLMs available through AWS Bedrock. Such cohorts are used to evaluate prompt accuracy and consistency across 10 repeated runs. Report-AI result pairs submitted to the Assess-AI registry for actively monitored use cases are additionally used to further optimize corresponding prompts. Results Prompts were developed for nine use cases. In Stage 2 development cohorts, final-prompt agreement with hybrid report-derived reference standard labels was 0.985 for ICH and 0.997 for PE. Because these cohorts informed prompt refinement and label construction, they were not independent validation sets. Discussion The workflow demonstrates feasible report-finding extraction at scale; independent accuracy and clinical utility remain unestablished. Conclusion LLM-based extraction within Assess-AI enables scalable, report-anchored AI performance monitoring in radiology.
To enhance automated de-identification of radiology reports by scaling transformer-based models through extensive training datasets and benchmarking commercial cloud vendor systems using a privacy-preserving evaluation framework for protected health information (PHI) detection. In this retrospective study, we built upon our state-of-the-art, transformer-based, PHI de-identification pipeline by further training on two large annotated radiology corpora from Stanford University, encompassing chest X-ray, chest CT, abdomen/pelvis CT, and brain MR reports and introducing an additional PHI category (AGE ≥ 90) into the architecture. Model performance was evaluated on test sets from Stanford and the University of Pennsylvania (Penn) for token-level PHI detection. We further assessed (1) the stability of synthetic PHI generation using a “hide-in-plain-sight” method and (2) agreement with commercial systems under a privacy-preserving synthetic-label evaluation framework. Precision, recall, and F1 scores were computed across all PHI categories. Our model achieved overall F1 scores of 0.973 on the Penn dataset and 0.999 on the Stanford dataset, achieving increases of up to approximately 0.7 F1 on the Stanford test set in the HOSPITAL, PATIENT, and VENDOR categories while maintaining strong overall performance on the independently annotated Penn dataset. Synthetic PHI evaluation showed consistent detectability (overall F1: 0.959 [0.958–0.960]) across 50 independently de-identified Penn datasets. Under a privacy-preserving evaluation framework using synthetic PHI surrogate labels as reference markers, our model demonstrated higher agreement with the synthetic labels than the evaluated cloud vendor systems on synthetic Penn reports (overall F1: 0.960 vs. 0.632–0.754). Because these reference labels were generated by our own de-identification pipeline, these findings reflect relative agreement rather than independent comparative detection accuracy. Large-scale, multimodal training supported strong performance across two institutions while synthetic PHI generation enabled privacy-preserving evaluation. A transformer-based de-identification model that performs highly in PHI detection on diverse radiology datasets, supporting privacy-preserving evaluation of radiology report de-identification systems.
IntroductionOnly two-thirds of patients with ovarian cancer ever see a gynecologic oncologist. Our objective was to examine the feasibility of an electronic health record-based nudge to clinicians for referral to gynecologic oncology at suspected ovarian cancer by imaging.MethodsWe developed a nudge, a short behavioral economics informed best practice advisory with a pended referral order for gynecologic oncology, for primary care, emergency medicine, and obstetrician/gynecology clinicians for when a patient had a O-RADS 4 or 5 lesion on imaging and had not already seen gynecologic oncology. In 2024, clinicians were sent the nudge within 2 business days of a patient's abnormal imaging through the electronic health record. Our primary outcome was referral rate to gynecologic oncology compared to a historic cohort of patients with O-RADS 4 or 5 lesions from 2020-2023.ResultsIn this prospective cohort study, we sent 20 clinician nudges for gynecologic oncology referral; six clinicians (30%) responded that the nudge changed their referral behavior. The 90-day referral rate was 75% compared to historic baseline of 61%. In the pilot, 92% patients undergoing surgery for complex adnexal mases had surgery with gynecologic oncology compared to historic baseline of 82%. One in four patients in the pilot were diagnosed with cancer, all early-stage disease.ConclusionsA clinician nudge for gynecologic oncology referral at suspected ovarian cancer diagnosis was acceptable and associated with 75% referral rate. A clinician nudge standardizes gynecologic oncology referral and may improve early detection of ovarian cancer. A randomized controlled trial of the clinician nudge is warranted.
As artificial intelligence (AI) tools are increasingly integrated into imaging workflows, understanding how AI results are generated and presented to the end user can equip radiologists to optimize interactions with AI results in practice. Although AI solutions are marketed as high-performing options that promise efficiency and diagnostic gains, issues arising at the radiologist-AI interface can lead to diminished returns due to unintentional cognitive burdens or misalignments with clinical workflows. A foundational understanding of how images are processed by AI solutions, presented in imaging workflows, and documented can allow radiologists to troubleshoot shortcomings in practice after clinical deployment. This article is the second in a primer series providing a foundation in AI literacy for radiologists. Building on the first article in this primer series, this article addresses questions raised by end users while using AI in practice. It covers how to consider the intended use of AI solutions, the process through which AI results are generated, the reasons why results may not be available at the time of study interpretation, and how to align the presentation of AI results with end user workflows. The discussion also explores emerging topics and challenges, including considerations for storing AI results and the related medicolegal considerations.
Climate change is driving increased cardiovascular risk and demand for medical imaging. While cardiovascular magnetic resonance (CMR) plays a critical role in diagnosing and managing cardiovascular disease, it is also among the most environmentally intensive imaging modalities. This review outlines the environmental impact of CMR and presents strategies to reduce emissions, conserve resources, and improve sustainability. From operational efficiencies to artificial intelligence innovation and systems-level reform, CMR professionals and industry partners all have a role to play. Implementing sustainable practices will be essential to support both patient care and planetary health as global CMR demand continues to rise.
Clinical histories are essential for accurate radiologic interpretation but are often lengthy and time-consuming to review amid increasing radiologist workloads. Large language models (LLMs) offer a potential solution by generating concise, clinically focused summaries from existing documentation; however, the optimal presentation format for radiologists remains unclear. This retrospective reader study evaluated the readability, efficiency, and radiologist preference for a two-stage LLM-based summarization pipeline designed to identify the optimal summary format for radiologic use. Ninety imaging studies across multiple modalities and care settings were processed to generate structured timeline summaries (stage 1) and brief narrative summaries (stage 2). Text length was reduced by approximately 86.8
OBJECTIVE:To describe the design, infrastructure, and functionality of the ACR's Assess-AI registry, a national quality registry created to monitor the real-world performance of clinical imaging artificial intelligence (AI) models. METHODS:Assess-AI is a registry within the National Radiology Data Registry that enables participating facilities to submit de-identified AI output, radiology report text, and DICOM study metadata through the ACR Connect platform. Data are normalized and compared with surrogate labels extracted from radiology reports using large language model-based prompting pipelines. Concordance between AI outputs and extracted surrogate labels is computed centrally, with results delivered through interactive dashboards. Facilities may locally re-identify studies via the Forensics App for quality review. RESULTS:Assess-AI currently supports multiple imaging AI use cases including intracranial hemorrhage, pulmonary embolism, pneumothorax, large-vessel occlusion, bone age, and cervical spine fracture. Participating facilities can visualize data completeness, monitor longitudinal concordance trends, compare performance with registry benchmarks, and explore discordance by demographic or technical factors. Local review workflows allow detailed evaluation of discordant cases, supporting root-cause analysis and AI governance. DISCUSSION:Assess-AI provides a scalable, privacy-preserving framework for postdeployment performance monitoring of imaging AI models. By combining standardized data ingestion, large language model-based surrogate labeling, interactive analytics, and local adjudication, the registry addresses critical gaps in evaluating real-world AI performance. Ongoing expansion will incorporate additional modalities, model types, and risk-adjusted benchmarking to further enhance clinical utility.
To perform a systematic review on the impact of deep learning (DL)-based triage for reducing diagnostic delays and improving patient outcomes in peer-reviewed and pre-print publications. A search was conducted of primary research studies focused on DL-based worklist optimization for diagnostic imaging triage published on multiple databases from January 2018 until July 2024. Extracted data included study design, dataset characteristics, workflow metrics including report turnaround time and time-to-treatment, and patient outcome differences. Further analysis between clinical settings and integration modality was investigated using nonparametric statistics. Risk of bias was assessed with the risk of bias in non-randomized studies-of interventions (ROBINS-I) checklist. A total of 38 studies from 20 publications, involving 138,423 images, were analyzed. Workflow interventions concerned pulmonary embolism (n = 8), stroke (n = 3), intracranial hemorrhage (n = 12), and chest conditions (n = 15). Patients in the post DL-triage group had shorter median report turnaround times: a mean difference of 12.3 min (IQR: −25.7, −7.6) for pulmonary embolism, 20.5 min (IQR: −32.1, −9.3) for stroke, 4.3 min (IQR: −8.6, 1.3) for intracranial hemorrhage and 29.7 min (IQR: −2947.7, −18.3) for chest diseases. Sub-group analysis revealed that reductions varied per clinical environment and relative prevalence rates but were the highest when algorithms actively stratified and reordered the radiological worklist, with reductions of −43.7
Large language models (LLMs) are reshaping radiology through their advanced capabilities in tasks such as medical report generation and clinical decision support. However, their effectiveness is heavily influenced by prompt engineering-the design of input prompts that guide the model's responses. This review aims to illustrate how different prompt engineering techniques, including zero-shot, one-shot, few-shot, chain of thought, and tree of thought, affect LLM performance in a radiology context. In addition, we explore the impact of prompt complexity and temperature settings on the relevance and accuracy of model outputs. This article highlights the importance of precise and iterative prompt design to enhance LLM reliability in radiology, emphasizing the need for methodological rigor and transparency to drive progress and ensure ethical use in health care.
Radiology education is challenged by increasing clinical workloads, limiting trainee supervision time and hindering real-time feedback. Large language models (LLMs) can enhance radiology education by providing real-time guidance, feedback, and educational resources while supporting efficient clinical workflows. We present an interpretation-centric framework for integrating LLMs into radiology education subdivided into distinct phases spanning predictation preparation, active dictation support, and postdictation analysis. In the predictation phase, LLMs can analyze clinical data and provide context-aware summaries of each case, suggest relevant educational resources, and triage cases based on their educational value. In the active dictation phase, LLMs can provide real-time educational support through processes such as differential diagnosis support, completeness guidance, classification schema assistance, structured follow-up guidance, and embedded educational resources. In the postdictation phase, LLMs can be used to analyze discrepancies between trainee and attending reports, identify areas for improvement, provide targeted educational recommendations, track trainee performance over time, and analyze the radiologic entities that trainees encounter. This framework offers a comprehensive approach to integrating LLMs into radiology education, with the potential to enhance trainee learning while preserving clinical efficiency.
Medical imaging is undergoing a transformation driven by the advent of new, highly effective, machine learning techniques paired with increases in computational capabilities (Cheng et al. 2021; Gilson et al. 2023; Almeida et al. 2024; Krishna et al. 2024). These advanced algorithms have the potential to improve disease detection, diagnosis, prognosis, and treatment outcomes. However, the complexity of machine learning models, the large amounts of curated and annotated data required by some methods, and the potential for bias and error make it challenging for individuals to safely and effectively leverage these methods (Lin et al. 2024; Guo et al. 2024; Xu et al. 2024; Linguraru et al. 2024; Wood et al. 2019). To address these challenges, the American Association of Physicists in Medicine (AAPM), American College of Radiology (ACR), Radiological Society of North America (RSNA), and Society for Imaging Informatics in Medicine (SIIM) have worked together to develop a syllabus detailing a recommended set of competencies for medical imaging professionals interacting with these systems. This guide is aimed at four different personas: users of AI systems, purchasers of AI systems, individuals who provide clinical expertise during the development of AI systems ("clinical collaborators"), and developers of AI systems.1 This is a syllabus, not a curriculum, and is intentional in this scope. Recognizing that individuals may benefit from different presentations of the same material, this work enumerates a series of relevant competencies but does not prescribe, nor offer, a method of instruction (Schuur, Rezazade Mehrizi, and Ranschaert 2021; Garin et al. 2023). By addressing the task-specific demands of each role, this guide will enable medical imaging professionals to utilize machine learning systems more safely and effectively, ultimately improving patient care and outcomes.
Diagnostic radiologists are uniquely positioned to make a difference in the reduction of diagnostic errors in medicine, which was identified as a national priority in a special report of the National Academies of Medicine in 2015. Interest in diagnostic error reduction within the specialty of diagnostic radiology has been accelerated in recent years by the adoption of rapidly evolving technological and process-based solutions for previously identified vulnerabilities (i.e., potential failure modes) in the radiology workflow that are known to increase the risk of errors. Here we describe a range of such potential failure modes contributing to diagnostic error in the practice of diagnostic radiology and summarize evolving efforts to mitigate them. These include fostering the adoption of peer learning and other feedback and education programs, a range of measures under development for ongoing performance evaluation of radiologists and their work-processes, and the deployment of new technologies, most notably artificial intelligence tools, to augment human performance and thereby improve diagnostic accuracy in Radiology.
Artificial intelligence (AI) has transformed medical imaging across disciplines. In cardiothoracic imaging, AI has the potential to impact all modalities, including chest radiography, echocardiography, computed tomography, magnetic resonance imaging, and nuclear cardiology. Use cases include automated structural and functional quantification of anatomic structures, lesion detection and characterization, and novel disease screening opportunities. Here, we discuss techniques and applications specific to cardiovascular diseases, including multimodal AI, coronary artery disease and coronary calcium detection, cardiac structural and functional evaluation, analysis of acute aortic syndromes, and evolving use cases such as pericardial disease evaluation and pulmonary hypertension assessment. We also discuss population health-related aspects of AI, model development challenges, and ethical considerations.