Automated quality assurance is essential for low-dose computed tomography (LDCT) lung screening, yet manual checks strain clinical workflows. We present a fully automated artificial intelligence tool that quantifies scan coverage and image noise in LDCT without user input. Lungs and the aorta are segmented to measure cranial/caudal over- and underscanning, and noise is computed as the standard deviation of Hounsfield units (HUs) within descending aortic blood, normalized to a 1 mm3 voxel. Performance was verified in a reader study of 98 LDCT scans from the National Lung Screening Trial (NLST), and then applied to 38,834 NLST scans reconstructed with a standard kernel. In the reader study, lung masks were rated ≥“Nearly Perfect” in 90.8% and aorta-blood masks in 96.9% of cases. Across 38,834 scans, mean overscanning distances were 31.21 mm caudally and 14.54 mm cranially; underscanning occurred in 4.36% (caudal) and 0.89% (cranial). The tool enables objective, large-scale monitoring of LDCT quality—reducing routine manual workload through exception-based human oversight, flagging protocol deviations, and supporting cross-center benchmarking—and may facilitate dose optimization by reducing systematic over- and underscanning.
Automated analysis of bedside chest radiographs remains challenging due to limited large-scale datasets with expert annotations and standardized severity grading. We provide TAIX-Ray, a comprehensive dataset of 215,381 bedside chest radiographs collected from 47,724 intensive care unit patients (30,306 male, 17,418 female, median age 68 years) collected over 14 years (01/2010-12/2023) at the University Hospital Aachen, Germany. During routine clinical reporting, 134 trained radiologists provided structured, itemized reports using a standardized template. They systematically assessed eight pathological findings: heart size (cardiomegaly), pulmonary congestion, pleural effusion (left/right), pulmonary opacities (left/right), and atelectasis (left/right) using a five-point ordinal severity scale (absent, questionable, mild, moderate, severe). The dataset includes (i) bedside chest radiographs (anteroposterior projections), (ii) structured, itemized reports, (iii) patient demographics (age and sex), and (iv) the temporal metadata. To facilitate immediate research adoption, we provide a baseline transformer model, implementation code, and predefined data splits, ensuring reproducible benchmarking. This resource enables the development of clinical AI models for automated pathology detection and severity assessment in critical care settings.
To determine whether using discrete semantic entropy (DSE) to reject questions likely to generate hallucinations can improve the accuracy of black-box vision-language models (VLMs) in radiologic image-based visual question answering (VQA). This retrospective study evaluated DSE using two publicly available, de-identified datasets: the VQA-Med 2019 benchmark (500 images with clinical questions and short-text answers) and a diagnostic radiology dataset (206 cases: 60 computed tomography scans, 60 magnetic resonance images, 60 radiographs, 26 angiograms) with corresponding ground-truth diagnoses. GPT-4o and GPT-4.1 (Generative Pretrained Transformer) answered each question 15 times using a temperature of 1.0. Baseline accuracy was determined using low-temperature answers (0.1). Meaning-equivalent responses were grouped using bidirectional entailment checks, and DSE was computed from the relative frequencies of the resulting semantic clusters. Accuracy was recalculated after excluding questions with DSE > 0.6 or > 0.3. p values and 95
Visual large language models (VLLMs) are discussed as potential tools for assisting radiologists in image interpretation, yet their clinical value remains unclear. This study provides a systematic and comprehensive comparison of general-purpose and biomedical VLLMs in radiology. We evaluated 180 representative clinical images with validated reference diagnoses (radiography, CT, MRI; 60 each) using seven VLLMs (ChatGPT-4o, Gemini 2.0, Claude Sonnet 3.7, Perplexity AI, Google Vision AI, LLaVA-1.6, LLaVA-Med-v1.5). Each model interpreted the image without and with clinical context. Mixed-effects logistic regression models assessed the influence of model, modality, and context on diagnostic performance and hallucinations (fabricated findings or misidentifications). Diagnostic accuracy varied significantly across all dimensions (p ≤ 0.001), ranging from 8.1% to 29.2% across models, with Gemini 2.0 performing best and LLaVA performing weakest. CT achieved the best overall accuracy (20.7%), followed by radiography (17.3%) and MRI (13.9%). Clinical context improved accuracy from 10.6% to 24.0% (p < 0.001) but shifted the model to rely more on textual information. Hallucinations were frequent (74.4% overall) and model-dependent (51.7–82.8% across models; p ≤ 0.004). Current VLLMs remain diagnostically unreliable, heavily context-biased, and prone to generating false findings, which limits their clinical suitability. Domain-specific training and rigorous validation are required before clinical integration can be considered.
Head and neck cancer (HNC) patients face an increased risk of malnutrition due to lifestyle, tumor localization, and treatment effects. While skeletal muscle area (SMA) and radiation attenuation (SM-RA) at the third lumbar vertebra (L3) are established prognostic markers, L3 is not routinely available in head and neck imaging. The prognostic value of SM-RA at the third cervical vertebra (C3) remains unclear. This study assesses whether SMA and SM-RA at C3 predict locoregional control (LRC) and overall survival (OS) in HNC. We analyzed 904 HNC cases with head and neck CT scans. A deep learning pipeline identified C3, and SMA/SM-RA were quantified via automated segmentation with manual verification. Cox proportional hazards models assessed associations with LRC and OS, adjusting for clinical factors. Median SMA and SM-RA were 36.64 cm² (IQR: 30.12–42.44) and 50.77 HU (IQR: 43.04–57.39). In multivariate analysis, lower SMA (HR 1.62, 95% CI: 1.02–2.58, p = 0.04), lower SM-RA (HR 1.89, 95% CI: 1.30–2.79, p < 0.001), and advanced T stage (HR 1.50, 95% CI: 1.06–2.12, p = 0.02) were prognostic for LRC. OS predictors included advanced T stage (HR 2.17, 95% CI: 1.64–2.87, p < 0.001), age ≥70 years (HR 1.40, 95% CI: 1.00–1.96, p = 0.05), male sex (HR 1.64, 95% CI: 1.02–2.63, p = 0.04), and lower SM-RA (HR 2.15, 95% CI: 1.56–2.96, p < 0.001). Deep learning-assisted SM-RA assessment at C3 outperforms SMA for LRC and OS in HNC, supporting its use as a routine biomarker and L3 alternative.
Developing a deep-learning model for automated multi-tissue, multi-condition knee MRI analysis and assessing its clinical potential. This retrospective dual-center study included 3121 MRI studies from 3018 adults, who underwent routine knee MRI examinations at a radiologic practice (2012–2019). Twenty-three conditions across cartilage, menisci, bone marrow, ligaments, and other soft tissues were manually labeled. A 3D slice transformer network was trained for binary classification and evaluated in terms of the area under the receiver operating characteristic curve (AUC), sensitivity, and specificity using a five-fold cross-validation and an external test set of 448 MRI studies (429 adults) from a university hospital (2022–2023). To assess differences in diagnostic performance, two inexperienced and two experienced radiology residents read 50 external test studies with and without model assistance. Paired t-tests were used for statistical analysis. Averaged over cross-validation tests, the model’s AUC was at least 0.85 for 8 conditions and at least 0.75 for 18 conditions. Generalization on the external test set was robust, with a mean absolute AUC difference of 0.05 ± 0.03 per condition. Model assistance improved accuracy and sensitivity for inexperienced residents, increased inter-reader agreement for both groups, and increased sensitivity and shortened reading times by 10
Purpose To evaluate diagnostic accuracy of various large language models (LLMs) when answering radiology-specific questions with and without access to additional online, up-to-date information via retrieval-augmented generation (RAG). Materials and Methods The authors developed radiology RAG (RadioRAG), an end-to-end framework that retrieves data from authoritative radiologic online sources in real-time. RAG incorporates information retrieval from external sources to supplement the initial prompt, grounding the model's response in relevant information. Using 80 questions from the RSNA Case Collection across radiologic subspecialties and 24 additional expert-curated questions with reference standard answers, LLMs (GPT-3.5-turbo [OpenAI], GPT-4, Mistral 7B, Mixtral 8×7B [Mistral], and Llama3-8B and -70B [Meta]) were prompted with and without RadioRAG in a zero-shot inference scenario (temperature ≤ 0.1, top-p = 1). RadioRAG retrieved context-specific information from www.radiopaedia.org. Accuracy of LLMs with and without RadioRAG in answering questions from each dataset was assessed. Statistical analyses were performed using bootstrapping while preserving pairing. Additional assessments included comparison of model with human performance and comparison of time required for conventional versus RadioRAG-powered question answering. Results RadioRAG improved accuracy for some LLMs, including GPT-3.5-turbo (74% [59 of 80] vs 66% [53 of 80], false discovery rate [FDR] = 0.03) and Mixtral 8×7B (76% [61 of 80] vs 65% [52 of 80], FDR = 0.02) on the RSNA radiology question answering (RSNA-RadioQA) dataset, with similar trends in the ExtendedQA dataset. Accuracy exceeded that of a human expert (63% [50 of 80], FDR ≤ 0.007) for these LLMs, although not for Mistral 7B-instruct-v0.2, Llama3-8B, and Llama3-70B (FDR ≥ 0.21). RadioRAG reduced hallucinations for all LLMs (rate, 6%-25%). RadioRAG increased estimated response time fourfold. Conclusion RadioRAG shows potential to improve LLM accuracy and factuality in radiology QA by integrating real-time, domain-specific data. Keywords: Retrieval-augmented Generation, Informatics, Computer-aided Diagnosis, Large Language Models Supplemental material is available for this article. © RSNA, 2025.
Objective: To assess response-to-loading in a human cadaveric knee joint model under different loading conditions before and after meniscectomy Design: In this prospective study, stress magnetic resonance imaging was performed using an MR-compatible loading device and quantitative T2 mapping in unloaded (UL), 0° neutrally loaded (LN) and 10° varus loaded (LV) condition before and after meniscectomy. Mean T2 values and four radiomic texture parameters were assessed within the cartilage of medial femur (MF) and medial tibia (MT) for all conditions. Results: Medial joint space width decreased from UL to LN to LV and after meniscectomy (all p<0.05). T2 values did not show any significant dependency on pressure or meniscectomy (all p>0.05). The radiomic parameter variance could assess loading induced textural T2 changes in the MF (UL vs. LN: p=0.042; UL vs. LV: p<0.001; LN vs. LV: p=0.022), and, in part, in the MT (LN vs. LV: p<0.013). Meniscectomy did not significantly alter the T2 mean values or radiomic parameters, respectively. Conclusions: T2-based radiomic features were more sensitive to assess cartilage response to loading than T2-mapping alone.
Background: Sarcopenia assessed by skeletal muscle area (SMA) at the third lumbar vertebra (L3) is an established prognostic marker in many malignancies, including head and neck cancer (HNC). However, in HNC, L3 is rarely assessed. The prognostic value of myosteatosis, measured by skeletal muscle radiation attenuation (SMRA) remains largely unexplored. This study evaluated both muscle metrics at the third cervical vertebra (C3) for locoregional control (LRC) and overall survival (OS) in HNC. Methods: SMA and SMRA at C3 were quantified in CT scans of 904 HNC cases by a deep learning-based segmentation pipeline with manual verification. Cox proportional hazards models assessed associations with LRC and OS. Results: Median SMA was 36.64 cm2 (IQR: 30.12–42.44). Median SMRA was 50.77 HU (IQR: 43.04–57.39). In multivariable analysis, lower SMA (HR 1.85, 95% CI: 1.19–2.88, p ≤ 0.001) and lower SMRA (HR 1.76, 95% CI: 1.22–2.54, p < 0.001) were associated with lower LRC. For OS, lower SMA (HR 1.53, 95% CI:1.06–2.20, p = 0.02) and lower SMRA (HR 2.13, 95% CI: 1.58–2.88, p < 0.001) were associated with a worse outcome in multivariable analysis. Conclusions: Both SMRA and SMA assessed at C3 correlate with worse LRC and OS in HNC.
Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) are essential clinical cross-sectional imaging techniques for diagnosing complex conditions. However, large 3D datasets with annotations for deep learning are scarce. While methods like DINOv2 are encouraging for 2D image analysis, these methods have not been applied to 3D medical images. Furthermore, deep learning models often lack explainability due to their "black-box" nature. This study aims to extend 2D self-supervised models, specifically DINOv2, to 3D medical imaging while evaluating their potential for explainable outcomes. We introduce the Medical Slice Transformer (MST) framework to adapt 2D self-supervised models for 3D medical image analysis. MST combines a Transformer architecture with a 2D feature extractor, i.e., DINOv2. We evaluate its diagnostic performance against a 3D convolutional neural network (3D ResNet) across three clinical datasets: breast MRI (651 patients), chest CT (722 patients), and knee MRI (1199 patients). Both methods were tested for diagnosing breast cancer, predicting lung nodule dignity, and detecting meniscus tears. Diagnostic performance was assessed by calculating the Area Under the Receiver Operating Characteristic Curve (AUC). Explainability was evaluated through a radiologist's qualitative comparison of saliency maps based on slice and lesion correctness. P-values were calculated using Delong's test. MST achieved higher AUC values compared to ResNet across all three datasets: breast (0.94 ± 0.01 vs. 0.91 ± 0.02, P = 0.02), chest (0.95 ± 0.01 vs. 0.92 ± 0.02, P = 0.13), and knee (0.85 ± 0.04 vs. 0.69 ± 0.05, P = 0.001). Saliency maps were consistently more precise and anatomically correct for MST than for ResNet. Self-supervised 2D models like DINOv2 can be effectively adapted for 3D medical imaging using MST, offering enhanced diagnostic accuracy and explainability compared to convolutional neural networks.
Structured reporting (SR) and artificial intelligence (AI) may transform how radiologists interact with imaging studies. This prospective study (July to December 2024) evaluated the impact of three reporting modes: free-text (FT), structured reporting (SR), and AI-assisted structured reporting (AI-SR), on image analysis behavior, diagnostic accuracy, efficiency, and user experience. Four novice and four non-novice readers (radiologists and medical students) each analyzed 35 bedside chest radiographs per session using a customized viewer and an eye-tracking system. Outcomes included diagnostic accuracy (compared with expert consensus using Cohen's κ), reporting time per radiograph, eye-tracking metrics, and questionnaire-based user experience. Statistical analysis used generalized linear mixed models with Bonferroni post-hoc tests with a significance level of (P ≤ .01). Diagnostic accuracy was similar in FT (κ= 0.58) and SR (κ= 0.60) but higher in AI-SR (κ= 0.71, P < .001). Reporting times decreased from 88 ± 38 s (FT) to 37 ± 18 s (SR) and 25 ± 9 s (AI-SR) (P < .001). Saccade counts for the radiograph field (205 ± 135 (FT), 123 ± 88 (SR), 97 ± 58 (AI-SR)) and total fixation duration for the report field (11 ± 5 s (FT), 5 ± 3 s (SR), 4 ± 1 s (AI-SR)) were lower with SR and AI-SR (P < .001 each). Novice readers shifted gaze towards the radiograph in SR, while non-novice readers maintained their focus on the radiograph. AI-SR was the preferred mode. In conclusion, SR improves efficiency by guiding visual attention toward the image, and AI-prefilled SR further enhances diagnostic accuracy and user satisfaction.
AI requires extensive datasets, while medical data is subject to high data protection. Anonymization is essential, but poses a challenge for some regions, such as the head, as identifying structures overlap with regions of clinical interest. Synthetic data offers a potential solution, but studies often lack rigorous evaluation of realism and utility. Therefore, we investigate to what extent synthetic data can replace real data in segmentation tasks. We employed head and neck cancer CT scans and brain glioma MRI scans from two large datasets. Synthetic data were generated using generative adversarial networks and diffusion models. We evaluated the quality of the synthetic data using MAE, MS-SSIM, Radiomics and a Visual Turing Test (VTT) performed by 5 radiologists and their usefulness in segmentation tasks using DSC. Radiomics indicates high fidelity of synthetic MRIs, but fall short in producing highly realistic CT tissue, with correlation coefficient of 0.8784 and 0.5461 for MRI and CT tumors, respectively. DSC results indicate limited utility of synthetic data: tumor segmentation achieved DSC=0.064 on CT and 0.834 on MRI, while bone segmentation a mean DSC=0.841. Relation between DSC and correlation is observed, but is limited by the complexity of the task. VTT results show synthetic CTs' utility, but with limited educational applications. Synthetic data can be used independently for the segmentation task, although limited by the complexity of the structures to segment. Advancing generative models to better tolerate heterogeneous inputs and learn subtle details is essential for enhancing their realism and expanding their application potential.
Background To validate pulmonary computed tomography (CT) perfusion in a porcine model by invasive monitoring of cardiac output (CO) using thermodilution method. Methods Animals were studied at a single center, using a Swan-Ganz catheter for invasive CO monitoring as a reference. Fifteen pigs were included. Contrast-enhanced CT perfusion of the descending aorta and right and left pulmonary artery was performed. For variation purposes, a balloon catheter was inserted to block the contralateral pulmonary vascular bed; additionally, two increased CO settings were created by intravenous administration of catecholamines. Finally, stepwise capillary occlusion was performed by intrapulmonary arterial injection of 75-μm microspheres in four stages. A semiautomatic selection of AFs and a recirculation-aware tracer-kinetics model to extract the first-pass of AFs, estimating blood flow with the Stewart-Hamilton method, was implemented. Linear mixed models (LMM) were developed to calibrate blood flow calculations accounting with individual- and cohort-level effects. Results Nine of 15 pigs had complete datasets. Strong correlations were observed between calibrated pulmonary (0.73, 95% confidence interval [CI] 0.6–0.82) and aortic blood flow measurements (0.82, 95% CI, 0.73–0.88) and the reference as well as agreements (± 2.24 L/min and ± 1.86 L/min, respectively) comparable to the state of the art, on a relatively wide range of right ventricle-CO measurements. Conclusions CT perfusion validly measures CO using LMMs at both individual and cohort levels, as demonstrated by referencing the invasive CO. Relevance statement Possible clinical applications of CT perfusion for measuring CO could be in acute pulmonary thromboembolism or to assess right ventricular function to show impairment or mismatch to the left ventricle. Key points • CT perfusion measures flow in vessels. • CT perfusion measures cumulative cardiac output in the aorta and pulmonary vessels. • CT perfusion validly measures CO using LMMs at both individual and cohort levels, as demonstrated by using the invasive CO as a reference standard. Graphical Abstract
ChatGPT-4 Vision (GPT-4V) is a state-of-the-art multimodal large language model (LLM) that may be queried using images. We aimed to evaluate the tool’s diagnostic performance when autonomously assessing clinical imaging studies. A total of 206 imaging studies (i.e., radiography (n = 60), CT (n = 60), MRI (n = 60), and angiography (n = 26)) with unequivocal findings and established reference diagnoses from the radiologic practice of a large university hospital were accessed. Readings were performed uncontextualized, with only the image provided, and contextualized, with additional clinical and demographic information. Responses were assessed along multiple diagnostic dimensions and analyzed using appropriate statistical tests. With its pronounced propensity to favor context over image information, the tool’s diagnostic accuracy improved from 8.3
Background Limited statistical knowledge can slow critical engagement with and adoption of artificial intelligence (AI) tools for radiologists. Large language models (LLMs) such as OpenAI's GPT-4, and notably its Advanced Data Analysis (ADA) extension, may improve the adoption of AI in radiology. Purpose To validate GPT-4 ADA outputs when autonomously conducting analyses of varying complexity on a multisource clinical dataset. Materials and Methods In this retrospective study, unique itemized radiologic reports of bedside chest radiographs, associated demographic data, and laboratory markers of inflammation from patients in intensive care from January 2009 to December 2019 were evaluated. GPT-4 ADA, accessed between December 2023 and January 2024, was tasked with autonomously analyzing this dataset by plotting radiography usage rates, providing descriptive statistics measures, quantifying factors of pulmonary opacities, and setting up machine learning (ML) models to predict their presence. Three scientists with 6-10 years of ML experience validated the outputs by verifying the methodology, assessing coding quality, re-executing the provided code, and comparing ML models head-to-head with their human-developed counterparts (based on the area under the receiver operating characteristic curve [AUC], accuracy, sensitivity, and specificity). Statistical significance was evaluated using bootstrapping. Results A total of 43 788 radiograph reports, with their laboratory values, from University Hospital RWTH Aachen were evaluated from 43 788 patients (mean age, 66 years ± 15 [SD]; 26 804 male). While GPT-4 ADA provided largely appropriate visualizations, descriptive statistical measures, quantitative statistical associations based on logistic regression, and gradient boosting machines for the predictive task (AUC, 0.75), some statistical errors and inaccuracies were encountered. ML strategies were valid and based on consistent coding routines, resulting in valid outputs on par with human specialist-developed reference models (AUC, 0.80 [95% CI: 0.80, 0.81] vs 0.80 [95% CI: 0.80, 0.81]; P = .51) (accuracy, 79% [6910 of 8758 patients] vs 78% [6875 of 8758 patients], respectively; P = .27). Conclusion LLMs may facilitate data analysis in radiology, from basic statistics to advanced ML-based predictive modeling. © RSNA, 2024 Supplemental material is available for this article.
Objectives Large language models (LLMs) have shown potential in radiology, but their ability to aid radiologists in interpreting imaging studies remains unexplored. We investigated the effects of a state-of-the-art LLM (GPT-4) on the radiologists' diagnostic workflow. Materials and methods In this retrospective study, six radiologists of different experience levels read 40 selected radiographic [n = 10], CT [n = 10], MRI [n = 10], and angiographic [n = 10] studies unassisted (session one) and assisted by GPT-4 (session two). Each imaging study was presented with demographic data, the chief complaint, and associated symptoms, and diagnoses were registered using an online survey tool. The impact of Artificial Intelligence (AI) on diagnostic accuracy, confidence, user experience, input prompts, and generated responses was assessed. False information was registered. Linear mixed-effect models were used to quantify the factors (fixed: experience, modality, AI assistance; random: radiologist) influencing diagnostic accuracy and confidence. Results When assessing if the correct diagnosis was among the top-3 differential diagnoses, diagnostic accuracy improved slightly from 181/240 (75.4%, unassisted) to 188/240 (78.3%, AI-assisted). Similar improvements were found when only the top differential diagnosis was considered. AI assistance was used in 77.5% of the readings. Three hundred nine prompts were generated, primarily involving differential diagnoses (59.1%) and imaging features of specific conditions (27.5%). Diagnostic confidence was significantly higher when readings were AI-assisted (p > 0.001). Twenty-three responses (7.4%) were classified as hallucinations, while two (0.6%) were misinterpretations. Conclusion Integrating GPT-4 in the diagnostic process improved diagnostic accuracy slightly and diagnostic confidence significantly. Potentially harmful hallucinations and misinterpretations call for caution and highlight the need for further safeguarding measures. Clinical relevance statement Using GPT-4 as a virtual assistant when reading images made six radiologists of different experience levels feel more confident and provide more accurate diagnoses; yet, GPT-4 gave factually incorrect and potentially harmful information in 7.4% of its responses.
Quantitative MRI techniques such as T2 and T1ρ mapping are beneficial in evaluating knee joint pathologies; however, long acquisition times limit their clinical adoption. MIXTURE (Multi-Interleaved X-prepared Turbo Spin-Echo with IntUitive RElaxometry) provides a versatile turbo spin-echo (TSE) platform for simultaneous morphologic and quantitative joint imaging. Two MIXTURE sequences were designed along clinical requirements: “MIX1”, combining proton density (PD)-weighted fat-saturated (FS) images and T2 mapping (acquisition time: 4:59 min), and “MIX2”, combining T1-weighted images and T1ρ mapping (6:38 min). MIXTURE sequences and their reference 2D and 3D TSE counterparts were acquired from ten human cadaveric knee joints at 3.0 T. Contrast, contrast-to-noise ratios, and coefficients of variation were comparatively evaluated using parametric tests. Clinical radiologists (n = 3) assessed diagnostic quality as a function of sequence and anatomic structure using five-point Likert scales and ordinal regression, with a significance level of α = 0.01. MIX1 and MIX2 had at least equal diagnostic quality compared to reference sequences of the same image weighting. Contrast, contrast-to-noise ratios, and coefficients of variation were largely similar for the PD-weighted FS and T1-weighted images. In clinically feasible scan times, MIXTURE sequences yield morphologic, TSE-based images of diagnostic quality and quantitative parameter maps with additional insights on soft tissue composition and ultrastructure.