OBJECTIVES:To evaluate whether an artificial intelligence-based denoising (AID) algorithm can effectively attenuate body mass index (BMI)-related noise degradation in cardiac photon-counting detector CT (PCD-CT) while maintaining strict quantitative diagnostic equivalence and clinical interchangeability compared with quantum iterative reconstruction (QIR). MATERIALS AND METHODS:This retrospective study included 100 patients undergoing non-contrast coronary artery calcium (CAC) scoring and contrast-enhanced coronary CT angiography (CCTA) on a dual-source PCD-CT. Images were reconstructed using standard QIR and a PCD-CT-weighted AID algorithm. Subjective image quality was assessed by four blinded raters. Objective metrics included noise magnitude (noise power spectrum) and spatial resolution (edge rise distance). Frequentist generalized estimating equations (GEE) evaluated reconstruction effects and BMI interactions. Clinical interchangeability (Agatston scoring, coronary age, CAD-RADS, and plaque characterization) was assessed using Bayesian regions of practical equivalence (ROPE). RESULTS:AID significantly improved subjective image quality and reduced overall noise magnitude by approximately 64 % in CAC and 78 % in CCTA (p<0.001). GEE models revealed that AID robustly attenuated BMI-driven noise escalation by 44 % (CAC) and 78 % (CCTA, p<0.001 for interactions in both). Noise texture was stable in CCTA datasets, but revealed smoothening in CAC datasets. However, spatial resolution was preserved and remained stable across varying patient habitus. ROPE analysis proved strict diagnostic equivalence between QIR and AID, yielding near-certain probabilities (>0.999) of practical equivalence for continuous Agatston scores, categorical CAD-RADS severity classification, and specific coronary plaque composition. CONCLUSIONS:The PCD-CT-weighted AID algorithm improves subjective and objective image quality in both CAC and CCTA, while maintaining diagnostic equivalence and workflow efficiency compared to QIR.
ObjectivesIschemic heart disease is a major global health burden requiring timely diagnosis. Although coronary artery calcium (CAC) scoring and coronary computed tomography angiography (CCTA) are valuable tools, image noise can distort assessment. Deep learning-based denoising (DLD) algorithms may enhance quality, yet their impact on cardiac CT workflows remains unclear. This study evaluated DLD effects on CAC and CCTA image quality, clinical interchangeability, and workflow efficiency compared with iterative reconstruction (IR).Materials and methodsA retrospective analysis of 100 patients with CAC and CCTA scans from the same CT scanner was performed. IR and DLD reconstructions yielded 400 datasets, rated by two radiologists using a semiquantitative scoring system. Objective metrics (CT number stability, noise, contrast-to-noise ratio) were measured. Clinical interchangeability and Workflow efficiency were evaluated.ResultsDLD showed significantly higher overall image quality than IR (p < 0.001), while preserving CT attenuation and improving noise and CNR. Although cardiac age classifications did not differ between reconstruction methods, Agatston scores were significantly higher in IR before manual correction (p < 0.001. After manual correction, Agatston scores were not significantly different between IR and DLD (p ≥ 0.158), supporting clinical comparability of the derived clinical metrics between the two approaches. DLD also reduced manual correction time (p < 0.001).ConclusionsThe investigated DLD algorithm improves image quality and radiological workflows thus potentially enhancing overall patient care. Key limitations include the retrospective single-center design and inherent subjectivity in image-quality evaluation; therefore, findings should be confirmed in prospective studies with more objective measures.
To compare the diagnostic performance of portal-venous (PV) blended CT images alone vs PV blended images supplemented with spectral reconstructions for detecting acute bowel ischemia, stratified by scanner platform (dual-energy CT [DECT] vs photon-counting CT [PCCT]). This retrospective single-center, multireader, crossover diagnostic accuracy study compared PV blended images alone (image set A) with PV blended images supplemented by spectral reconstructions (image set B: 40-keV virtual monoenergetic images, iodine maps, and virtual non-contrast images) acquired on DECT and PCCT. The reference standard was a prespecified composite derived from surgical and index-hospitalization documentation. Diagnostic performance was analyzed using mixed-effects logistic models for accuracy, sensitivity, and specificity, cumulative link mixed models for 7-point suspicion scores, and ROC analyses using a multireader multicase framework. A total of 378 patients (mean age, 68.1 years ± 7.4, 206 men) were evaluated. 150/378 (39.7
To evaluate whether deep learning-based segmentation can improve visualization and diagnostic interpretation of MR cholangiopancreatography (MRCP) maximum intensity projections (MIP) by suppressing overlapping high-intensity anatomy while preserving the pancreatobiliary system. A total of 322 3D MRCP datasets from 162 patients were included. The training set comprised 265 cases from 127 patients, allowing multiple acquisitions per patient, whereas the evaluation set consisted of 35 cases with a single acquisition per patient. A deep learning-based segmentation model was trained using manual annotations with three distinct labels: background, primary structures, and secondary structures. Conservative safety margins were applied to reduce segmentation omissions. Two board-certified radiologists independently rated processed and original MIP images using 4-point Likert scales (1 = poor, 4 = excellent). Two-sided Wilcoxon signed-rank tests with Benjamini–Hochberg correction were used, and inter-reader agreement was assessed using linearly weighted Cohen’s κ . Segmentation suppressed obscuring structures and improved biliary visualization. After correction for multiple comparisons, processed images showed significantly improved structure overlap scores compared with original images (3.5 ± 0.9 vs. 2.5 ± 0.6) and higher diagnostic confidence (3.4 ± 0.7 vs. 3.2 ± 0.7). Deep learning-based segmentation improves MRCP MIP visualization by reducing anatomical overlap while preserving clinically relevant pancreatobiliary anatomy.
Multidisciplinary tumor boards (MDTs) are critical for the personalized management of soft tissue sarcomas (STS), but they are limited by time, costs, and resource demands. With recent advances in large language models (LLMs) like ChatGPT, there is growing interest in evaluating their potential role in augmenting MDT workflows. This study aimed to assess the clinical performance of ChatGPT-4o in real-world STS cases using predefined evaluation criteria, comparing its treatment suggestions with expert MDT decisions. This retrospective study included 152 patients presented to the multidisciplinary sarcoma tumor board. ChatGPT-4o was prompted to generate guideline-based treatment recommendations based on anonymized tumor board registration letters. Outputs were scored by blinded expert reviewers using a five-domain framework: diagnostic modalities, therapeutic modalities, treatment sequencing/timing, chemotherapy regimen, and clinical contextualization. Descriptive statistics and non-parametric ANOVA with post hoc tests assessed performance, including subgroup analysis by sarcoma subtype. ChatGPT-4o scores were significantly lower than the maximum achievable value of 1.0 across all five criteria (all p < 0.0001). Among individual domains, clinical contextualization significantly outperformed all other criteria in pairwise comparisons (all p < 0.05). No significant performance differences were observed across sarcoma subtypes (H = 19.74, p = 0.138). ChatGPT-4o demonstrated substantial expert-rated performance in generating tumor board recommendations for soft tissue sarcoma cases, particularly excelling in personalized contextualization. Discrepancies in treatment sequencing and chemotherapy selection highlight the need for expert oversight. These findings support the feasibility of LLM integration into oncology workflows, warranting further refinement toward safe, supportive clinical use.
Diffusion-weighted imaging (DWI) of the head and neck is essential for various clinical applications but is often hampered by artifacts and reduced image quality. Deep learning (DL) reconstruction has the potential to enhance the quality of head and neck DWI. This study aims to evaluate the performance of an accelerated, DL-reconstructed DWI (DWIDL) in terms of image quality and diagnostic confidence. This retrospective study included patients who underwent clinically indicated head and neck DWI at 1.5 T and 3 T between August 2023 and January 2024 at a tertiary care center. Imaging was performed at low b‑values (0 or 50 sec/mm2) and high b‑values (800 sec/mm2), and apparent diffusion coefficient (ADC) maps were computed. After acquiring standard single-shot echoplanar imaging DWI sequences, the raw MR datasets underwent simulated acceleration by reducing the number of signal averages. These accelerated exams were then reconstructed using a novel DL-based algorithm that combined DL-based k‑space to image reconstruction with DL-based super-resolution processing (DWIDL). Three readers analyzed the images using a visual Likert score to evaluate image sharpness, artifacts, noise, overall image quality, and diagnostic confidence. Comparisons were made using the Wilcoxon signed-rank test. A quantitative analysis of signal-to-noise ratio (SNR), contrast-to-noise ratio (CNR) and apparent diffusion coefficient values (ADC) was also performed. The study included 30 patients (mean age, 55 ± 19 years; range, 24–84; 18 men) with various pathologies. Scan times were reduced by 67
Noise and artifacts often compromise image quality and diagnostic accuracy in liver Volume Perfusion CT (VPCT), which can impact clinical decision-making in the diagnosis of hepatocellular carcinoma (HCC). AI-based denoising can potentially improve image quality in computed tomography imaging. Therefore, we aimed to evaluate the effects of AI denoising (AID) on image quality and Milan classification in VPCT compared to standard methods. VPCT examinations from 100 patients acquired between 2017 and 2021 were retrospectively included in the analysis. Perfusion maps were reconstructed using original data (Origin), vendor-specific median filtration (Vendor), and AID. Two radiologists independently scored subjective image quality (image quality, diagnostic confidence, contrast, and sharpness). Objective image quality parameters (CT numbers, noise, contrast-to-noise ratios) and diameters of all HCC lesions were measured across all datasets. Additionally, the Milan classification was determined for each patient. AID enhanced subjective image quality (0.48 ± 0.29) compared to Origin (-0.31 ± 0.29, p-value < 0.001) and Vendor (-0.17 ± 0.24, p-value < 0.001). CNR was higher in AID (27.40 ± 2.98) compared to Origin (16.98 ± 1.54, p < 0.001) and Vendor (19.43 ± 1.79, p < 0.001). No significant differences were found regarding lesion diameter between Origin, Vendor, and AID (p > 0.999). The number and localization of HCC lesions were equal between Origin, Vendor, and AID. The different reconstruction methods did not affect Milan classification. AID enhances the image quality of HCC liver VPCT without compromising diagnostic capabilities and may support the evaluation of VPCT in the diagnosis of HCC.
Abstract Objectives Gastrointestinal stromal tumors (GISTs) are molecularly heterogeneous neoplasms whose management depends on individualized, multidisciplinary decision-making. While multidisciplinary tumor boards (MTBs) represent the standard of care, access remains limited in many clinical settings. This study evaluates the performance of two large language models in generating GIST MTB recommendations and assesses their agreement with expert MTB decisions using predefined clinical evaluation criteria. Materials and methods This retrospective single-center study included 99 GIST cases discussed at an institutional MTB. A structured prompt was developed to extract clinical variables and generate treatment recommendations. ChatGPT-5 and Qwen3 were independently evaluated across five predefined domains: diagnostic recommendations, therapeutic modalities, treatment sequence and timing, systemic therapy regimen selection, and clinical contextualization. Two expert reviewers scored all outputs in a blinded fashion. Normalized scores, inter-model comparisons, perfect-case rates, and inter-rater agreement were analyzed. Results Both models demonstrated high concordance with expert MTB recommendations, with mean total normalized scores of 0.901 for ChatGPT-5 and 0.875 for Qwen3, without a significant difference between models ( p > 0.05). Perfect agreement was observed in 52.5% of ChatGPT-5 cases and 48.5% of Qwen3 cases ( p > 0.05). Diagnostic recommendations scored significantly lower than all other domains in both models (all adjusted p < 0.05). Overall inter-rater agreement was almost perfect (weighted Cohen’s kappa=0.978). Conclusions Both models demonstrated high agreement with expert GIST MTB recommendations, with no significant performance difference between them. Diagnostic reasoning represented the weakest domain, reflecting the challenge of reconstructing context-dependent workup decisions from tumor board documentation. These findings support a potential assistive role for LLMs in GIST MTB workflows, while underscoring the continued necessity of expert oversight.
This study evaluates the impact of deep learning-enhanced T1-weighted VIBE sequences (DL-VIBE) on image quality and procedural parameters during MR-guided thermoablation of liver malignancies, compared to standard VIBE (SD-VIBE). Between September 2021 and February 2023, 34 patients (mean age: 65.4 years; 13 women) underwent MR-guided microwave ablation on a 1.5 T scanner. Intraprocedural SD-VIBE sequences were retrospectively processed with a deep learning algorithm (DL-VIBE) to reduce noise and enhance sharpness. Two interventional radiologists independently assessed image quality, noise, artifacts, sharpness, diagnostic confidence, and procedural parameters using a 5-point Likert scale. Interrater agreement was analyzed, and noise maps were created to assess signal-to-noise ratio improvements. DL-VIBE significantly improved image quality, reduced artifacts and noise, and enhanced sharpness of liver contours and portal vein branches compared to SD-VIBE (p < 0.01). Procedural metrics, including needle tip detectability, confidence in needle positioning, and ablation zone assessment, were significantly better with DL-VIBE (p < 0.01). Interrater agreement was high (Cohen κ = 0.86). Reconstruction times for DL-VIBE were 3 s for k-space reconstruction and 1 s for superresolution processing. Simulated acquisition modifications reduced breath-hold duration by approximately 2 s. DL-VIBE enhances image quality during MR-guided thermal ablation while improving efficiency through reduced processing and acquisition times.
Background/Objectives: Assessment of a novel deep-learning (DL)-based T1w volumetric interpolated breath-hold (VIBEDL) sequence in breast MRI in comparison with standard VIBE (VIBEStd) for image quality evaluation. Methods: Prospective study of 52 breast cancer patients examined at 1.5T breast MRI with T1w VIBEStd and T1 VIBEDL sequence. T1w VIBEDL was integrated as an additional early non-contrast and a delayed post-contrast scan. Two radiologists independently scored T1w VIBE Std/DL sequences both pre- and post-contrast and their calculated subtractions (SUBs) for image quality, sharpness, (motion)–artifacts, perceived signal-to-noise and diagnostic confidence with a Likert-scale from 1: Non-diagnostic to 5: Excellent. Lesion diameter was evaluated on the SUB for T1w VIBEStd/DL. All lesions were visually evaluated in T1w VIBEStd/DL pre- and post-contrast and their subtractions. Statistics included correlation analyses and paired t-tests. Results: Significantly higher Likert scale values were detected in the pre-contrast T1w VIBEDL compared to the T1w VIBEStd for image quality (each p < 0.001), image sharpness (p < 0.001), SNR (p < 0.001), and diagnostic confidence (p < 0.010). Significantly higher values for image quality (p < 0.001 in each case), image sharpness (p < 0.001), SNR (p < 0.001), and artifacts (p < 0.001) were detected in the post-contrast T1w VIBEDL and in the SUB. SUBDL provided superior diagnostic certainty compared to SUBStd in one reader (p = 0.083 or p = 0.004). Conclusions: Deep learning-enhanced T1w VIBEDL at 1.5T breast MRI offers superior image quality compared to T1w VIBEStd.
RATIONALE AND OBJECTIVES:Large Language Models (LLMs) offer a promising solution for extracting structured clinical information from free-text radiology reports. The Simplified Magnetic Resonance Index of Activity (sMARIA) is a validated scoring system used to quantify Crohn's disease (CD) activity based on Magnetic Resonance Enterography (MRE) findings. This study aims to evaluate the performance of two advanced LLMs in extracting key imaging features and computing sMARIA scores from free-text MRE reports. MATERIALS AND METHODS:This retrospective study included 117 anonymized free-text MRE reports from patients with confirmed CD. ChatGPT (GPT-4o) and DeepSeek (DeepSeek-R1) were prompted using a structured input designed to extract four key radiologic features relevant to sMARIA: bowel wall thickness, mural edema, perienteric fat stranding, and ulceration. LLM outputs were evaluated against radiologist annotations at both the segment and feature levels. Segment-level agreement was assessed using accuracy, mean absolute error (MAE) and Pearson correlation. Feature-level performance was evaluated using sensitivity, specificity, precision, and F1-score. Errors including confabulations were recorded descriptively. RESULTS:ChatGPT achieved a segment-level accuracy of 98.6%, MAE of 0.17, and Pearson correlation of 0.99. DeepSeek achieved 97.3% accuracy, MAE of 0.51, and correlation of 0.96. At the feature level, ChatGPT yielded an F1-score of 98.8% (precision 97.8%, sensitivity 99.9%), while DeepSeek achieved 97.9% (precision 96.0%, sensitivity 99.8%). CONCLUSIONS:LLMs demonstrate near-human accuracy in extracting structured information and computing sMARIA scores from free-text MRE reports. This enables automated assessment of CD activity without altering current reporting workflows, supporting longitudinal monitoring and large-scale research. Integration into clinical decision support systems may be feasible in the future, provided appropriate human oversight and validation are ensured.
RATIONALE AND OBJECTIVES:Photon Counting CT (PCCT) offers advanced imaging capabilities with potential for substantial radiation dose reduction; however, achieving this without compromising image quality remains a challenge due to increased noise at lower doses. This study aims to evaluate the effectiveness of a deep learning (DL)-based denoising algorithm in maintaining diagnostic image quality in whole-body PCCT imaging at reduced radiation levels, using real intraindividual cadaveric scans. MATERIALS AND METHODS:Twenty-four cadaveric human bodies underwent whole-body CT scans on a PCCT scanner (NAEOTOM Alpha, Siemens Healthineers) at four different dose levels (100%, 50%, 25%, and 10% mAs). Each scan was reconstructed using both QIR level 2 and a DL algorithm (ClariCT.AI, ClariPi Inc.), resulting in 192 datasets. Objective image quality was assessed by measuring CT value stability, image noise, and contrast-to-noise ratio (CNR) across consistent regions of interest (ROIs) in the liver parenchyma. Two radiologists independently evaluated subjective image quality based on overall image clarity, sharpness, and contrast. Inter-rater agreement was determined using Spearman's correlation coefficient, and statistical analysis included mixed-effects modeling to assess objective and subjective image quality. RESULTS:Objective analysis showed that the DL denoising algorithm did not significantly alter CT values (p ≥ 0.9975). Noise levels were consistently lower in denoised datasets compared to the Original (p < 0.0001). No significant differences were observed between the 25% mAs denoised and the 100% mAs original datasets in terms of noise and CNR (p ≥ 0.7870). Subjective analysis revealed strong inter-rater agreement (r ≥ 0.78), with the 50% mAs denoised datasets rated superior to the 100% mAs original datasets (p < 0.0001) and no significant differences detected between the 25% mAs denoised and 100% mAs original datasets (p ≥ 0.9436). CONCLUSION:The DL denoising algorithm maintains image quality in PCCT imaging while enabling up to a 75% reduction in radiation dose. This approach offers a promising method for reducing radiation exposure in clinical PCCT without compromising diagnostic quality.
Die Arbeitsgemeinschaft Gesundheitspolitische Verantwortung (AG GPV) innerhalb der Deutschen Röntgengesellschaft (DRG) ist mehr als ein Netzwerk: Sie ist eine Bewegung, die Radiolog:innen ermutigt, ihre Stimme im Gesundheitssystem einzubringen. Mit klaren Zielen und einer inspirierenden Vision bietet sie eine Plattform, auf der jede und jeder Einzelne aktiv zur Zukunft der Radiologie beitragen kann. In einer Zeit des demografischen Wandels, der Digitalisierung, der Ökonomisierung und des rasant zunehmenden Wissens spielt die Radiologie eine Schlüsselrolle. Engagieren Sie sich in unserer AG!
This study aimed to compare a conventional three-dimensional (3-D) magnetic resonance cholangiopancreatography (MRCP) sequence with a deep learning (DL)-accelerated MRCP sequence (hereafter, MRCPDL) regarding acquisition time and image quality. We conducted a prospective study of consecutive patients referred for MRCP between November 2023 and April 2024 at a single tertiary center. Each participant underwent 1.5T 3-D T2-weighted turbo spin echo MRCP using both a conventional sequence (threefold acceleration) and MRCPDL (eightfold acceleration). Three blinded readers independently evaluated image quality, including background signal suppression, bile and pancreatic duct visibility, artifact level, and diagnostic confidence on an ordinal four-point scale. Acquisition times were compared using a paired t-test. Image quality parameters were assessed with repeated measures ANOVA. Interreader agreement was analyzed using Fleiss' κ. Out of 419 consecutive patients, 30 participants were evaluated (mean age, 63 ± 15 years; 16 men, 14 women). The mean acquisition time was 10:30 ± 03:04 min for conventional MRCP and 3:57 ± 01:13 min for MRCPDL, P < 0.001. MRCPDL reduced acquisition time by 62.4