Background: Due to the ongoing rapid advancement of artificial intelligence (AI), including large language models (LLMs), radiologists will soon face the challenge of the responsible clinical integration of these models. Objectives: The aim of this work is to provide an overview of current developments regarding LLMs, potential applications in radiology, and their (future) relevance and limitations.Materials and methodsThis review analyzes publications on LLMs for specific applications in medicine and radiology. Additionally, literature related to the challenges of clinical LLM use was reviewed and summarized. Results: In addition to a general overview of current literature on radiological applications of LLMs, several particularly noteworthy studies on the subject are recommended. Conclusions: In order to facilitate the forthcoming clinical integration of LLMs, radiologists need to engage with the topic, understand various application areas, and be aware of potential limitations in order to address challenges related to patient safety, ethics, and data protection.
Importance Given the widespread use of medical imaging, evaluating the effectiveness of interventions to improve appropriateness is crucial for optimizing health care resources and patient outcomes. Objective To assess the effects of implementing a clinical decision support system (CDSS), the European Society of Radiology iGuide, on the appropriateness of the medical imaging ordering behavior of physicians. Design and Setting A cluster randomized clinical trial with 26 departments at 3 German university hospitals acting as clusters, incorporating a before and after discontinued design. All imaging requests originating from physicians in the participating departments over a 2.5-year period were included (between December 2021 and June 2024). Interventions All departments started without a CDSS and required structured clinical indication data entry and tracking of requested imaging. After randomization, 13 clusters (departments at hospitals) received the CDSS intervention (intervention clusters) and 13 clusters did not (control clusters). The CDSS intervention provided ordering physicians with information as to whether their imaging requests were appropriate, appropriate under certain conditions, or inappropriate; in addition, alternative diagnostic tests, including the corresponding appropriateness score, were suggested by the CDSS, after which physicians could choose to modify their imaging requests. Main Outcomes and Measures The primary outcome measure was the proportion of inappropriate imaging requests made per department. A difference-in-differences analysis was used to investigate changes in the proportion of inappropriate imaging requests between departments with vs those without the CDSS. Results A total of 65 764 imaging requests were scored using the CDSS; 50.1% of imaging requests were for female patients and the mean patient age was 64 years (SD, 17.1 years). Prior to implementation of the CDSS, there were 21 625 imaging requests from the control clusters, 1367 (6.3%) of which were categorized as inappropriate; and there were 13 338 imaging requests from the intervention clusters, 1007 (7.6%) of which were categorized as inappropriate. After implementation of the CDSS, there were 10 055 imaging requests from the control clusters, 518 (5.2%) of which were categorized as inappropriate; and there were 7206 imaging requests from the intervention clusters, 461 (6.4%) of which were categorized as inappropriate. The intervention clusters showed a similar reduction (mean difference, −0.5% [99% CI, −2.4% to 0.4%]) in inappropriate imaging requests compared with the control clusters (mean difference, −1.8% [99% CI, −4.3% to −0.4%]) and there was a difference-in-differences value of 1.3 percentage points (99% CI, −2.0 to 1.8 percentage points; P = .69), which was not statistically significant. Conclusions and Relevance The CDSS did not reduce the number of inappropriate imaging requests ordered by physicians in academic hospital settings. Trial Registration ClinicalTrials.gov Identifier: NCT05490290
Importance Given the widespread use of medical imaging, evaluating the effectiveness of interventions to improve appropriateness is crucial for optimizing health care resources and patient outcomes. Objective To assess the effects of implementing a clinical decision support system (CDSS), the European Society of Radiology iGuide, on the appropriateness of the medical imaging ordering behavior of physicians. Design and Setting A cluster randomized clinical trial with 26 departments at 3 German university hospitals acting as clusters, incorporating a before and after discontinued design. All imaging requests originating from physicians in the participating departments over a 2.5-year period were included (between December 2021 and June 2024). Interventions All departments started without a CDSS and required structured clinical indication data entry and tracking of requested imaging. After randomization, 13 clusters (departments at hospitals) received the CDSS intervention (intervention clusters) and 13 clusters did not (control clusters). The CDSS intervention provided ordering physicians with information as to whether their imaging requests were appropriate, appropriate under certain conditions, or inappropriate; in addition, alternative diagnostic tests, including the corresponding appropriateness score, were suggested by the CDSS, after which physicians could choose to modify their imaging requests. Main Outcomes and Measures The primary outcome measure was the proportion of inappropriate imaging requests made per department. A difference-in-differences analysis was used to investigate changes in the proportion of inappropriate imaging requests between departments with vs those without the CDSS. Results A total of 65 764 imaging requests were scored using the CDSS; 50.1% of imaging requests were for female patients and the mean patient age was 64 years (SD, 17.1 years). Prior to implementation of the CDSS, there were 21 625 imaging requests from the control clusters, 1367 (6.3%) of which were categorized as inappropriate; and there were 13 338 imaging requests from the intervention clusters, 1007 (7.6%) of which were categorized as inappropriate. After implementation of the CDSS, there were 10 055 imaging requests from the control clusters, 518 (5.2%) of which were categorized as inappropriate; and there were 7206 imaging requests from the intervention clusters, 461 (6.4%) of which were categorized as inappropriate. The intervention clusters showed a similar reduction (mean difference, -0.5% [99% CI, -2.4% to 0.4%]) in inappropriate imaging requests compared with the control clusters (mean difference, -1.8% [99% CI, -4.3% to -0.4%]) and there was a difference-in-differences value of 1.3 percentage points (99% CI, -2.0 to 1.8 percentage points; P = .69), which was not statistically significant. Conclusions and Relevance The CDSS did not reduce the number of inappropriate imaging requests ordered by physicians in academic hospital settings. Trial RegistrationClinicalTrials.gov Identifier: NCT05490290
Lung cancer is the leading cause of cancer-related mortality. While early detection improves survival, distinguishing malignant from benign pulmonary nodules remains challenging. Artificial intelligence (AI) has been proposed to enhance diagnostic accuracy, but its clinical reliability is still under investigation. Here, we aimed to evaluate the diagnostic performance of AI models in classifying pulmonary nodules. This single-center retrospective study analyzed pulmonary nodules (4–30 mm) detected on CT scans, using three AI software models. Sensitivity, specificity, false-positive and false-negative rates were calculated. The diagnostic accuracy was assessed using the area under the receiver operating characteristic (ROC) curve (AUC), with histopathology serving as the gold standard. Subgroup analyses were based on nodule size and histopathological classification. The impact of imaging parameters was evaluated using regression analysis. A total of 158 nodules (n = 30 benign, n = 128 malignant) were analyzed. One AI model classified most nodules as intermediate risk, preventing further accuracy assessment. The other models demonstrated moderate sensitivity (53.1–70.3
Aufgrund des andauernden, rapiden Fortschritts künstlicher Intelligenz (KI) inklusive Large Language Models (LLMs) werden Radiolog*innen in absehbarer Zeit vor die Herausforderung der verantwortungsvollen klinischen Integration dieser Modelle gestellt. Ziel der Arbeit ist es, einen Überblick über aktuelle Entwicklungen zum Thema LLMs, mögliche Einsatzgebiete in der Radiologie sowie ihre (zukünftige) Relevanz und Limitationen zu liefern. In dieser Übersichtsarbeit wurden Publikationen zu LLMs für spezifische Anwendungen in der Medizin und Radiologie analysiert. Zusätzlich wurde Literatur zu den Herausforderungen im Zusammenhang mit einer klinischen LLM-Nutzung gesichtet und zusammengefasst. Neben einem generellen Überblick über aktuelle Literatur zu radiologischen Anwendungsbeispielen von LLMs werden verschiedene besonders spannende Arbeiten zum Thema empfohlen. Um die anstehende klinische Integration von LLMs zu ermöglichen, müssen sich Radiolog*innen mit der Thematik auseinandersetzen, verschiedene Anwendungsgebiete und möglicher Limitationen kennen, um Herausforderungen im Hinblick auf Patientensicherheit, Ethik und Datenschutz bewältigen zu können.
Advanced imaging techniques play a pivotal role in oncology. A large variety of computed tomography (CT) scanners, scan protocols, and acquisition techniques have led to a wide range in image quality and radiation exposure. This study aims at implementing verifiable oncological imaging by quality assurance and optimization (i-Violin) through harmonizing image quality and radiation dose across Europe. The 2‑year multicenter implementation study outlined here will focus on CT imaging of lung, stomach, and colorectal cancer and include imaging for four radiological indications: diagnosis, radiation therapy planning, staging, and follow-up. Therefore, 480 anonymized CT data sets of patients will be collected by the associated university hospitals and uploaded to a repository. Radiologists will determine key abdominopelvic structures for image quality assessment by consensus and subsequently adapt a previously developed lung CT tool for the objective evaluation of image quality. The quality metrics will be evaluated for their correlation with perceived image quality and the standardized optimization strategy will be disseminated across Europe. The results of the outlined study will be used to obtain European reference data, to build teaching programs for the developed tools, and to create a culture of optimization in oncological CT imaging. The study protocol and rationale for i‑Violin, a European approach for standardization and harmonization of image quality and optimization of CT procedures in oncological imaging, is presented. Future results will be disseminated across all EU member states, and i‑Violin is thus expected to have a sustained impact on CT imaging for cancer patients across Europe.
Objectives Artificial intelligence (AI) has tremendous potential to help radiologists in daily clinical routine. However, a seamless, standardized, and time-efficient way of integrating AI into the radiology workflow is often lacking. This constrains the full potential of this technology. To address this, we developed a new reporting pipeline that enables automated pre-population of structured reports with results provided by AI tools. Methods Findings from a commercially available AI tool for chest X-ray pathology detection were sent to an IHE-MRRT-compliant structured reporting (SR) platform as DICOM SR elements and used to automatically pre-populate a chest X-ray SR template. Pre-populated AI results could be validated, altered, or deleted by radiologists accessing the SR template. We assessed the performance of this newly developed AI to SR pipeline by comparing reporting times and subjective report quality to reports created as free-text and conventional structured reports. Results Chest X-ray reports with the new pipeline could be created in significantly less time than free-text reports and conventional structured reports (mean reporting times: 66.8 s vs. 85.6 s and 85.8 s, respectively; both p < 0.001). Reports created with the pipeline were rated significantly higher quality on a 5-point Likert scale than free-text reports ( p < 0.001). Conclusion The AI to SR pipeline offers a standardized, time-efficient way to integrate AI-generated findings into the reporting workflow as parts of structured reports and has the potential to improve clinical AI integration and further increase synergy between AI and SR in the future. Critical relevance statement With the AI-to-structured reporting pipeline, chest X-ray reports can be created in a standardized, time-efficient, and high-quality manner. The pipeline has the potential to improve AI integration into daily clinical routine, which may facilitate utilization of the benefits of AI to the fullest. Key points • A pipeline was developed for automated transfer of AI results into structured reports. • Pipeline chest X-ray reporting is faster than free-text or conventional structured reports. • Report quality was also rated higher for reports created with the pipeline. • The pipeline offers efficient, standardized AI integration into the clinical workflow. Graphical Abstract
Purpose Structured reporting (SR) not only offers advantages regarding report quality but, as an IT-based method, also the opportunity to aggregate and analyze large, highly structured datasets (data mining). In this study, a data mining algorithm was used to calculate epidemiological data and in-hospital prevalence statistics of pulmonary embolism (PE) by analyzing structured CT reports. Methods All structured reports for PE CT scans from the last 5 years (n = 2790) were extracted from the SR database and analyzed. The prevalence of PE was calculated for the entire cohort and stratified by referral type and clinical referrer. Distributions of the manifestation of PEs (central, lobar, segmental, subsegmental, as well as left-sided, right-sided, bilateral) were calculated, and the occurrence of right heart strain was correlated with the manifestation. Results The prevalence of PE in the entire cohort was 24% (n = 678). The median age of PE patients was 71 years (IQR 58-80), and the sex distribution was 1.2/1 (M/F). Outpatients showed a lower prevalence of 23% compared to patients from regular wards (27%) and intensive care units (30%). Surgically referred patients had a higher prevalence than patients from internal medicine (34% vs. 22%). Patients with central and bilateral PEs had a significantly higher occurrence of right heart strain compared to patients with peripheral and unilateral embolisms. Conclusion Data mining of structured reports is a simple method for obtaining prevalence statistics, epidemiological data, and the distribution of disease characteristics, as demonstrated by the PE use case. The generated data can be helpful for multiple purposes, such as for internal clinical quality assurance and scientific analyses. To benefit from this, consistent use of SR is required and is therefore recommended.
BACKGROUND:Structured reporting (SR) is recommended in radiology, due to its advantages over free-text reporting (FTR). However, SR use is hindered by insufficient integration of speech recognition, which is well accepted among radiologists and commonly used for unstructured FTR. SR templates must be laboriously completed using a mouse and keyboard, which may explain why SR use remains limited in clinical routine, despite its advantages. Artificial intelligence and related fields, like natural language processing (NLP), offer enormous possibilities to facilitate the imaging workflow. Here, we aimed to use the potential of NLP to combine the advantages of SR and speech recognition.RESULTS:We developed a reporting tool that uses NLP to automatically convert dictated free text into a structured report. The tool comprises a task-oriented dialogue system, which assists the radiologist by sending visual feedback if relevant findings are missed. The system was developed on top of several NLP components and speech recognition. It extracts structured content from dictated free text and uses it to complete an SR template in RadLex terms, which is displayed in its user interface. The tool was evaluated for reporting of urolithiasis CTs, as a use case. It was tested using fictitious text samples about urolithiasis, and 50 original reports of CTs from patients with urolithiasis. The NLP recognition worked well for both, with an F1 score of 0.98 (precision: 0.99; recall: 0.96) for the test with fictitious samples and an F1 score of 0.90 (precision: 0.96; recall: 0.83) for the test with original reports.CONCLUSION:Due to its unique ability to integrate speech into SR, this novel tool could represent a major contribution to the future of reporting.
Einleitung Die ERCP als intraluminale Diagnostik und ggf. Therapie einer Abflussstörung stellt einen Grundpfeiler der Bildgebung von perihilären Cholangiokarzinomen (pCCA) dar. Für die operative Planung ist jedoch eine Mehrphasen-Computertomographie zur Beurteilung der Hilusgefäße und des zukünftigen Restlebergewebes notwendig. Die 3D-Darstellung der Gallenwege, basierend auf einer CT, ist herausfordernd, wobei die photon counting-CT (pc-CT) eine Möglichkeit darstellt, diese zu optimieren. Eine Stentversorgung verhindert jedoch die adäquate 3D-Rekonstruktion der Gallenwege. Ziel dieser Untersuchung war eine monozentrische Analyse der Versorgungsrealität der pCCA in Hinblick auf die Optimierung der präoperativen Bildgebung.
Interdisziplinäre Fallbesprechungen, insbesondere Tumorboards, stellen einen großen Anteil der täglichen Arbeit des klinischen Radiologen dar. Die Radiologie nimmt im Tumorboard eine Schlüsselrolle ein, da bildgebend erhobene Befunde direkten Einfluss auf Therapieentscheidungen haben. Dieser Artikel soll die Anforderungen an den Radiologen bei der Vorbereitung und Durchführung von Tumorboards erörtern. Weiter werden Rahmenbedingungen und Durchführungsformen von Tumorboards beleuchtet. IT-Tools zur Prozessautomatisierung und verschiedene Systeme zur Verlaufsbeurteilung von Tumorerkrankungen werden vorgestellt. Eine ausführliche Vorbereitung des Tumorboards und eine klare Kommunikation von Befunden ist unerlässlich. Durch die radiologische Expertise im Tumorboard kommt es oft zu Änderungen oder Anpassungen von initial geplanten Therapien. Neben klassischen Präsenzveranstaltungen haben sich Hybridlösungen bei der Durchführung von Tumorboards etabliert, bei denen das Kernteam vor Ort ist und weitere Teilnehmer (externe Zuweiser, interne Teilnehmer außerhalb des Kernteams) per Videokonferenz zugeschaltet sind. Zur Verlaufsbeurteilung von Tumorerkrankungen sind verschiedene Systeme etabliert. Aufgrund der breiten Anwendbarkeit wird vor allem RECIST 1.1. genutzt. IT-Tools ermöglichen es, zuvor markierte Tumorherde im zeitlichen Verlauf in einer Matrixansicht darzustellen (Läsionsnachverfolgung). Durch den Einsatz künstlicher Intelligenz (KI) können Herde zudem automatisch erkannt und volumetriert werden. Die Vorbereitung und Durchführung von Tumorboards sind für den Radiologen zeitaufwändig. IT-Tools können dabei die Abläufe automatisieren und somit vereinfachen. Hybridlösungen aus Präsenzveranstaltungen und Videokonferenzen vereinfachen es externen Zuweisern, ihre Patienten im Tumorboard vorzustellen.
Preparing and conducting tumor conferences is time-consuming for radiologists. IT tools can automate and thus facilitate the processes. Hybrid solutions combining face-to-face meetings and video conferences make it easier for external referring physicians to present their patients in tumor conferences.
Zielsetzung Die Mehrzahl der schriftlichen medizinischen Examina basieren auf Multiple-Choice-Fragen und/oder Freitextantworten. Die Aus- und Bewertung derartiger Freitextantworten erfolgt händisch und ist dementsprechend 1. zeitintensiv und 2. fehleranfällig. Ziel dieser Studie ist es daher zu untersuchen, ob es möglich ist, Freitextantworten mittels Natural Language Processing (NLP) automatisiert zu analysieren und anschließend eine Benotung vorzuschlagen, um den Auswerteprozess zu unterstützen.
Purpose Kidney volume is important in the management of renal diseases. Unfortunately, the currently available, semi-automated kidney volume determination is time-consuming and prone to errors. Recent advances in its automation are promising but mostly require contrast-enhanced computed tomography (CT) scans. This study aimed at establishing an automated estimation of kidney volume in non-contrast, low-dose CT scans of patients with suspected urolithiasis. Methods The kidney segmentation process was automated with 2D Convolutional Neural Network (CNN) models trained on manually segmented 2D transverse images extracted from low-dose, unenhanced CT scans of 210 patients. The models’ segmentation accuracy was assessed using Dice Similarity Coefficient (DSC), for the overlap with manually-generated masks on a set of images not used in the training. Next, the models were applied to 22 previously unseen cases to segment kidney regions. The volume of each kidney was calculated from the product of voxel number and their volume in each segmented mask. Kidney volume results were then validated against results semi-automatically obtained by radiologists. Results The CNN-enabled kidney volume estimation took a mean of 32 s for both kidneys in a CT scan with an average of 1026 slices. The DSC was 0.91 and 0.86 and for left and right kidneys, respectively. Inter-rater variability had consistencies of ICC = 0.89 (right), 0.92 (left), and absolute agreements of ICC = 0.89 (right), 0.93 (left) between the CNN-enabled and semi-automated volume estimations. Conclusion In our work, we demonstrated that CNN-enabled kidney volume estimation is feasible and highly reproducible in low-dose, non-enhanced CT scans. Automatic segmentation can thereby quantitatively enhance radiological reports.
Background The radiological report is the cornerstone of communication between radiologists and referring physicians and patients, respectively. The report is comprised of image interpretation on the one hand and communication of this interpretation on the other hand. Objectives and methods To outline different types of radiological reports (regarding content as well as structure) and their communication. To this end, current guidelines are summarized and clinical examples are presented. Results The radiological report is typically a written piece of free text prose and highly individualized regarding its quality, precision, and structure. In order to improve the understanding of the written report, additional material (e.g., annotations, images, tables) can be supplemented (multimedia-enhanced reporting). In terms of standardization, national and international radiological associations promote structured reporting in radiology. However, this is not without issues. Conclusion Effective communication should improve patient care and it should be clear and provided in a timely manner. As communication in clinical reality is often hampered by various factors, internal standard operating procedures (SOPs) should be developed to improve communication workflows.to improve communication procedures.
Background: The hype around artificial intelligence (AI) in radiology continues and the number of approved AI tools is growing steadily. Despite the great potential, integration into clinical routine in radiology remains limited. In addition, the large number of individual applications poses a challenge for clinical routine, as individual applications have to be selected for different questions and organ systems, which increases the complexity and time required.Objectives: This review will discuss the current status of validation and implementation of AI tools in clinical routine, and identify possible approaches for an improved assessment of the generalizability of results of AI tools.Materials and methods: A literature search in various literature and product databases as well as publications, position papers, and reports from various stakeholders was conducted for this review.Results: Scientific evidence and independent validation studies are available for only a few commercial AI tools and the generalizability of the results often remains questionable.Conclusions: One challenge is the multitude of offerings for individual, specific application areas by a large number of manufacturers, making integration into the existing site-specific IT infrastructure more difficult. Furthermore, remuneration for the use of AI tools in clinical routine by health insurance companies in Germany is lacking. But in order for reimbursement to be granted, the clinical utility of new applications must first be proven. Such proof, however, is lacking for most applications.