Background: MRI-based radiomics has shown promise in men with prostate cancer (PCa); however, successful clinical implementation is contingent upon on reproducible measurements. Purpose: We assessed the reproducibility of radiomics features extracted from bi-parametric prostate MRI (bpMRI) in prostate lesions and non-tumoral prostate tissue in men with PCa undergoing active surveillance (AS). Methods: This retrospective study included 47 men with biopsy-proven PCa undergoing AS (mean 68.9 ± 8.2 years, mean PSA density [PSAD] 0.08 ± 0.03 ng/mL/mL) who underwent two bpMRI approximately 12 months apart (range, 10-14 months; December 2018 to April 2020). The reproducibility of radiomics measurements was assessed using the same MRI platform (3T Skyra, Siemens Healthineers; inter-platform) (n = 37), different MRI vendors (Skyra, Siemens Healthineers; 3T Discovery MR750, GE Healthcare; inter-platform) (n = 10), and between observers (n = 10). Shape/1st-/2nd-order radiomics features were extracted from regions of interest on axial T2-weighted (T2-WI), diffusion-weighted imaging (DWI, b1600), and apparent diffusion coefficient (ADC) maps on prostate lesions, non-tumoral peripheral zones (PZs), and transition zones (TZs) using software. Reproducibility was evaluated by calculating the intraclass correlation coefficient (ICC) and coefficient of variation (CV). Associations of clinical variables and prostate volume were assessed. Results: PCa diagnoses included Gleason grade groups 1 (n = 46) and 2 (n = 1)]. Thirty-seven lesions (mean size 0.9 ± 0.4 cm) in 31 patients had PI-RADS v2.1 scores of 2 (n = 3)/3 (n = 12)/4 (n = 21)/5 (n = 1); 16 patients demonstrated diffuse PI-RADS 2 changes. Lesion radiomics features from T2-WI yielded a high proportion of good/moderate ICCs (intra-platform, 77.8%; inter-platform, 56.5%), whereas most DWI/ADC features yielded poor reproducibility. Similar results were observed for non-tumoral PZ/TZ. Intra-platform CVs were lowest for T2-WI lesion features (13.6%) and background PZ/TZ (<13.3%), while DWI/ADC exceeded 20%. Inter-platform CVs were lowest for lesions on T2-WI and were <18% for DWI/ADC; all background PZ/TZ CVs were < 16.4%. Inter-observer analyses showed good/moderate ICCs across all sequences and regions (57.4-92.6%). The distribution of ICC and CV values did not differ between intra- and inter-platform analyses (p > 0.05). Higher reproducibility (ICC > 0.5) was associated with larger prostate volume (intra-platform diagnostic odds ratio [DOR] = 2.58, 95% confidence interval [95%CI], 1.35-3.80, p = 0.01; inter-platform DOR = 3.48, 95%CI 1.79-5.17, p = 0.01) and older age (inter-platform DOR = 5.30, 95%CI 3.75-6.85, p < 0.01). Conclusions: Radiomics measurements from T2-WI demonstrated better intra-/inter-platform reproducibility than DWI/ADC for prostate lesions and non-tumoral tissue. Patient factors (larger prostate volumes and older age) influence radiomics stability. The optimization of diffusion-based radiomics features is needed to improve reproducibility given the essential role of DWI in prostate MRI.
Background Hepatocellular carcinoma (HCC) surveillance typically involves US, although gadoxetate-enhanced MRI may improve sensitivity. Identifying factors impacting image quality is essential, but comparative data between US and MRI are lacking. Purpose To compare image quality and its determinants in gadoxetate-enhanced MRI versus US in participants with cirrhosis who are undergoing HCC screening in a prospective bicenter study. Materials and Methods In a secondary analysis of a prospective bicenter North American HCC screening study (September 2020 to May 2023), participants with cirrhosis underwent contemporaneous gadoxetate-enhanced MRI and liver US. Two independent observers evaluated MRI scans for dynamic phase motion, hepatobiliary phase liver uptake, and diffusion-weighted imaging quality. Sums of these assessments were trichotomized into MRI scores (MR-A: no or minimal limitations; MR-B: moderate limitations; and MR-C: severe limitations). US quality was scored per the Liver Imaging Reporting and Data System (US-A: no or minimal limitations; US-B: moderate limitations; and US-C: severe limitations). Clinical factors and quality were assessed using univariable and/or multivariable analyses. Proportions of quality scores were compared (McNemar χ2 test). Results Among the 245 participants with cirrhosis (median age, 61 years; 133 men), MRI quality scores were classified as MR-A in 80.4% (197 of 245 participants), MR-B in 18.4% (45 of 245 participants), and MR-C in 1.2% (three of 245 participants), whereas available US visualization scores were classified as US-A in 24.2% (58 of 240 participants), US-B in 61.7% (148 of 240 participants), and US-C in 14.1% (34 of 240 participants). The proportion of examination scores with no or minimal limitations was higher for MRI than US in 240 participants with both set of scores (P < .001). Obesity (body mass index ≥ 30, calculated as weight in kilograms divided by height in meters squared) reduced quality for both modalities (MRI univariable odds ratio [OR], 4.20 [95% CI: 1.20, 22.6], P = .01; US OR, 2.50 [95% CI: 1.00, 6.00], P = .04), not confirmed at multivariable analysis. Child-Pugh score of B and/or C reduced MRI quality at univariable (OR, 2.60 [95% CI: 1.05, 6.30]; P = .03) and multivariable (adjusted OR, 3.95 [95% CI: 1.50, 10.42]; P = .006) analysis. Conclusion Most participants undergoing HCC screening had excellent MRI quality even when US was limited. Quality was adversely affected by obesity (both modalities) and Child-Pugh score (MRI only). ClinicalTrials.gov Identifier: NCT04539717 © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. Supplemental material is available for this article.
Emphysematous cholecystitis is a rare but severe variant of acute cholecystitis characterized by gas-forming organisms within the gallbladder wall or lumen. It progresses rapidly and carries substantial mortality, making early and accurate recognition essential. Although its pathogenesis involves gallbladder wall ischemia with superimposed infection by gas-producing bacteria-most commonly Clostridium species-the clinical presentation is often nonspecific, particularly in patients with diabetes mellitus or immunosuppression. Imaging therefore serves as the cornerstone of diagnosis. Abdominal radiographs may demonstrate intraluminal or intramural gas, while ultrasound can reveal echogenic foci with reverberation artifacts, though overlying bowel gas and diagnostic mimics may limit sensitivity. Computed tomography remains the most accurate modality, precisely delineating gas within the gallbladder wall, lumen, or adjacent tissues and facilitating urgent surgical or percutaneous intervention. Magnetic resonance imaging offers complementary soft tissue characterization when computed tomography is contraindicated. This review synthesizes traditional imaging findings and emerging diagnostic innovations by critically comparing modality-specific strengths, limitations, and pitfalls. Dual-energy and photon-counting computed tomography enhance tissue contrast and gas conspicuity, while artificial intelligence-assisted image analysis enables earlier detection and expedited triage in emergency settings. By integrating evolving technologies with established radiologic principles, this article provides a forward-looking framework for improving diagnostic precision and ultimately enhancing outcomes for patients with emphysematous cholecystitis.
Clinical practice guidelines recommend hepatocellular cancer (HCC) surveillance in patients with cirrhosis from any etiology and those with chronic hepatitis B virus (HBV) infection and additional risk factors. However, HCC incidence varies across groups. Several risk stratification models using clinical factors and/or biomarkers have been derived to facilitate tailored HCC surveillance. Although risk stratification models are used for patients with hepatitis B, few have been sufficiently validated in patients with cirrhosis. Indeed, many unanswered questions related to the development, validation, and impact evaluation of risk stratification models must be addressed before widespread implementation can be recommended. The National Cancer Institute's Translational Liver Cancer (TLC) Consortium was established to advance research focused on risk stratification and early detection of liver cancer. The TLC convened a multidisciplinary group, including clinicians, scientists, biostatisticians, and technology experts from the United States, Asia, and Europe, to provide a framework for the development, validation, and implementation of risk stratification models. The framework defines 4 phases of risk stratification model development and validation: phase 1-development and internal validation, phase 2-decision rule development, phase 3-external validation, and phase 4-impact evaluation. The group also defined a set of recommendations to improve the rigor of development and validation of HCC risk stratification strategies. This framework can inform best practices and highlight necessary steps for endorsement by practice guidelines and regulatory agencies, highlighting a path toward implementation in clinical practice.
Background Patients with solid renal masses (SRMs) are at risk of chronic kidney disease (CKD) after surgical resection without a reliable pre-operative predictor. Purpose To investigate whether pre-operative multiparametric MRI (mpMRI) can predict CKD development and progression to stage 3 CKD. Study Type Prospective. Population Forty-three participants (female = 13, mean age: 59 +/- 12 years) undergoing nephrectomy for SRM. Field Strength/Sequence 1.5 T, diffusion-weighted echo-planar imaging (DWI) using nine b-values (0-800 s/mm(2)), T-1-mapping using variable flip angle, multi-echo gradient-echo blood-oxygen-level-dependent (BOLD), and dynamic-contrast-enhanced MRI (DCE-MRI) using 3D T-1-weighted gradient-echo. Assessment A clinical CKD risk score was calculated from estimated glomerular filtration rate (eGFR), age, diabetes, and surgery (partial or radical nephrectomy). mpMRI parameters included cortical and medullary apparent diffusion coefficient (ADC), intravoxel incoherent motion (IVIM), tri-exponential diffusion (fast, medium, and slow), and spectral diffusion (vascular, tubule, and tissue) from DWI, native T-1 from T-1-mapping, R-2* from BOLD, and renal plasma flow and eGFR from DCE-MRI. Outcomes were a correlation with baseline eGFR, prediction of postoperative 12-month eGFR decline > 5 mL/min/1.73 m(2), and stage 3 CKD development (eGFR < 60 mL/min/1.73 m(2)). Statistical Tests Mann-Whitney U-test and Spearman's rank correlation coefficient (r). Diagnostic ability was determined by leave-one-out cross-validated logistic regression area-under-the-receiver-operator-curve (AUC) and diagnostic odds ratio (DOR) with p-value < 0.05 considered significant. Results Thirty of 43 (67%) participants had normal baseline renal function (eGFR >= 60 mL/min/1.73 m(2)). Twenty-nine participants completed 12-month follow-up: among 66% (19/29) who had baseline normal eGFR, 37% (7/19) developed stage 3 CKD. eGFR from DCE-MRI and tubule diffusion correlated with baseline eGFR ( = 0.43 and 0.33 respectively). Reduced vascular diffusion predicted eGFR decline (AUC = 0.75-0.83, DOR = 6.8-16.5). A larger contralateral ADC corticomedullary difference (AUC = 0.89; DOR = 22.5), and clinical CKD risk score (AUC = 0.81; DOR = 5.5) were the strongest predictors of CKD development. Data Conclusion Pre-operative mpMRI predicted post-nephrectomy CKD development. A larger corticomedullary difference in ADC may indicate reduced functional reserve.
Photon-counting CT (PCCT) is an emerging imaging modality that offers improved spatial resolution, soft tissue contrast, material decomposition, and dose efficiency over conventional CT with energy-integrating detectors (EID-CT). For patients with abdominal malignancies, who routinely undergo cross-sectional imaging for diagnosis, staging, response assessment, and surveillance, PCCT has the potential to enhance the detection and characterization of contrast-enhanced tumors and improve treatment response assessment. The objectives of this review are to provide a comprehensive overview of the technical principles of PCCT, optimization strategies, advantages in abdominal oncologic imaging, and current clinical applications geared towards clinical radiologists. We also address implementation challenges and discuss future directions for integrating PCCT into routine oncologic care. As PCCT utilization increases, its impact on abdominal cancer diagnosis and management is expected to grow, offering new opportunities for precision medicine and personalized imaging strategies. Not applicable.
Purpose To evaluate large language model (LLM)-based strategy performance for extraction and classification of incidental findings from whole-body (WB) imaging reports, particularly strategies incorporating Oncologically Relevant Findings Reporting and Data System (ONCO-RADS). Materials and Methods In this retrospective bicenter study, authors included all WB MRI reports from January 2016 to December 2023 at a referral center (internal dataset). Two observers extracted all incidental findings, and patient records were used to confirm final diagnoses. First, authors evaluated ONCO-RADS performance and the reproducibility of its incidental finding classifications by six radiologists. Then, authors evaluated the accuracy of three LLM-based strategies: (a) a fine-tuned DeBERTa/medical named entity recognition (NER) model; (b) zero-shot LLMs (ChatGPT-o1 [OpenAI], Gemini-2.5-Pro [Google]); and (c) reference-guided prompting of these LLMs using ONCO-RADS. Authors then expanded these strategies to an external dataset of 605 reports with multiple imaging techniques (405 WB MRI; 100 fluorodeoxyglucose PET/CT; and 100 chest-abdomen-pelvis CT acquisitions) from January 2022 to January 2025. Results The internal dataset included 823 patients (mean age, 63.7 years ± 11.7 [SD]; 457 male patients) with 1488 WB MRI reports. The average interobserver reproducibility of ONCO-RADS incidental finding classifications was excellent (Cohen κ, 0.87). The per-report accuracies of ONCO-RADS-guided LLMs (95.6% [151 of 158] and 86.7% [137 of 158] for ChatGPT-o1 and Gemini-2.5-Pro, respectively) were higher than those of the medical NER (69.0% [109 of 158]) and zero-shot LLMs (57.0% [90 of 158] and 70.9% [112 of 158] for ChatGPT-o1 and Gemini-2.5-Pro, respectively) (P < .001). In the external test set (mean age, 60.6 years ± 12.9; 330 male patients), the per-report accuracies of ONCO-RADS-guided ChatGPT-o1 (83.5% [505 of 605]) and Gemini-2.5-Pro (82.0% [496 of 605]) were higher than those of the models without ONCO-RADS prompting (63.1% [382 of 605] and 61.2% [370 of 605], respectively) and the medical NER (55.7% [337 of 605]) (P < .001). Conclusion Reference-guided prompting of the LLMs ChatGPT-o1 and Gemini-2.5-Pro with ONCO-RADS improved their performance in extracting and classifying incidental findings on WB imaging reports compared with zero-shot prompting and medical NER. Keywords: Large Language Models, Incidental Findings, Whole-Body MRI Supplemental material is available for this article. © RSNA, 2026.
Abstract Objective To assess the performance of a deep learning-based computer-aided detection (DL-CAD) algorithm for prostate lesion detection and classification on biparametric (bp)MRI. Materials and methods This retrospective, single-center study included men undergoing 3-T MRI of the prostate for suspected prostate cancer (PCa) between July and September of 2022. Using the radiology report as the reference standard, detection performance for high-risk lesions (defined as PI-RADS ≥ 3, 4, 5) by the DL-CAD was evaluated per-patient using sensitivity, specificity, PPV, NPV and AUC; and per-lesion using sensitivity and PPV. Kappa statistics was used to assess per-patient detection and per-lesion classification of PI-RADS ≥ 3 lesions. Clinical and imaging factors associated with discordance between DL-CAD and radiology reports were assessed using Mann–Whitney, Chi-square, and Fisher’s exact tests. Results 442 adult males (mean age 65 ± 9 years) were assessed. Per-patient sensitivity, specificity, PPV, and NPV for detection of PI-RADS ≥ 4 and 5 lesions were 65.3%/81.2%/62.7%/82.9% and 82.1%/93.8%/65.7%/97.3%, respectively. Per-patient performance for identifying PI-RADS ≥ 3/4/5 lesions was fair-to-excellent: AUC = 0.67 (0.62–0.71)/0.75 (0.71–0.80)/0.92 (0.89–0.96). For detection of PI-RADS ≥ 4 and 5, per-lesion sensitivity was 60.4% and 78.3%, while PPV was 55.0% and 60.3%. Per-patient agreement between DL-CAD and the reference increased with higher PI-RADS scores (kappa = 0.26 (0.18–0.35)/0.46 (0.37–0.55)/0.68 (0.59–0.78)). Agreement on classification of PI-RADS ≥ 3 lesions was moderate (kappa = 0.56 (0.45–0.68)). Conclusion A pre-trained DL-CAD showed good-to-excellent per-patient performance for the detection of PI-RADS ≥ 4 lesions and moderate performance of PI-RADS ≥ 3 lesion classification. Future prospective studies validating the DL algorithm with histopathologic correlation are warranted. Critical relevance statement A deep learning computer-aided detection (DL-CAD) algorithm showed good-to-excellent per-patient performance for detection of PI-RADS ≥ 4 lesions, moderate performance of PI-RADS ≥ 3 lesion classification and high negative predictive value, which can be applied in the clinic with knowledge of its limitations. Key Points Clinical validation of deep learning computer-aided detection (DL-CAD) models for the detection and classification of prostate lesions on MRI is urgently needed. A pre-trained DL-CAD algorithm showed fair-to-excellent per-patient performance for detection of prostate lesions on biparametric MRI, with moderate performance for PI-RADS ≥ 3 lesion classification. Identification of false negatives and false positives of prostate cancer detection DL-CAD algorithms is important for future improvement and clinical deployment. A DL-CAD-based prostate cancer detection algorithm with high NPV may reduce interpretation time. Graphical Abstract
Primary liver cancer, encompassing hepatocellular carcinoma and cholangiocarcinoma, represents a growing global health burden with significant morbidity and mortality. Imaging is central to the clinical management of these malignancies, informing surveillance, diagnosis, staging, and treatment response assessment. A comprehensive overview of current imaging modalities-including ultrasound, computed tomography, MRI, and PET/CT-and their respective roles across the spectrum of hepatic and biliary neoplasms is presented. Established diagnostic frameworks such as LI-RADS and staging systems including BCLC are reviewed, alongside emerging applications of radiomics and artificial intelligence in liver oncology.
Extraosseous (EO) involvement in multiple myeloma (MM) is a critical marker of systemic clonal escape, but its spatial and temporal dynamics are largely omitted from current risk models like R2-ISS. In this cohort of 543 patients with longitudinal follow-up between 2010 and 2025, whole-body (WB) imaging including WB-MRI ± FDG-PET/CT was used to characterize EO disease. EO status was modeled as a time-varying covariate in multivariable Cox proportional hazards models including clinical, biological, and genetic covariates to assess impact on overall survival (OS). EO disease was identified in 24.6% of cases (n = 134). We identified a distinct prognostic hierarchy for OS based on the timing and pattern of dissemination: compared to marrow-confined disease, prognosis was increasingly worse for primary paraskeletal (PS) lesions present at diagnosis (HR = 2.22), secondary PS (HR = 3.33), and reaching an ultra-high-risk for secondary extramedullary disease (EMD) (HR = 8.14). Patients with initial PS lesions faced a 2.6-fold increased risk of evolving into secondary EMD (p = 0.002). While lesion volume and anatomical localization of EO lacked prognostic value, secondary EMD, R2-ISS stage IV, number of lines of treatment and TP53 mutation at time of EO diagnosis independently predicted lower post-EO survival. Thus, EO involvement defines an aggressive phenotype in which prognosis is driven by the timing and dissemination pattern rather than EO tumor burden. WB imaging accurately identifies patients transitioning into these high-risk groups. Integrating timing and spatial EO features into future risk frameworks is essential for better prognostication, as the adverse impact of EO disease persists even after adjusting for R2-ISS and cytogenetic status.
Abstract Objectives Dynamic contrast-enhanced (DCE) magnetic resonance imaging (MRI) can assess kidney function, but artifacts and complex post-processing limit its use. We calculated estimated glomerular filtration rate (eGFR) and renal plasma flow (RPF) by combining a population-based arterial input function (pAIF) with a whole-kidney pharmacokinetic model (WKPM). We also compared DCE-MRI eGFR and RPF with serum eGFR and arterial spin labeling (ASL) derived RPF, respectively. Materials and methods In a prospective single-center study, 43 patients (30 M/13 F, 59.0 ± 11.8 y) with renal masses underwent multiparametric 1.5-T MRI, before and 3 months after nephrectomy (n = 15), including coronal, fat-saturated volumetric DCE-MRI (5-s temporal resolution) and background-suppressed pseudocontinuous ASL. DCE-MRI eGFR and RPF were measured by WKPM, incorporating individual-based arterial input function (iAIF) and population-based arterial input function (pAIF) as inputs. Pearson correlation, Bland-Altman analysis, and Mann-Whitney U statistics were used. Results Serum eGFR (mean 67.55 mL/min/1.73 m²) and DCE-MRI (mean eGFR pAIF 59.49, iAIF 63.60 mL/min/1.73 m²) were measured in 51 MRIs: correlation with serum eGFR was stronger for pAIF (r = 0.61, p < 0.001) than iAIF (r = 0.33, p = 0.018), with comparable Bland-Altman bias (-11.9% and -9.1%, respectively). RPF was measured by both DCE-MRI and ASL in 21 MRIs: mean RPF was 229.3 (ASL), 229.7 (pAIF), and 390.4 (iAIF) mL/min (p = 0.018). Correlation of pAIF RPF with ASL-derived RPF (r = 0.65, p < 0.001) was stronger than for iAIF RPF (r = 0.53, p = 0.014), with lower Bland-Altman bias (pAIF -1.0% versus iAIF 39.5%). Conclusion DCE-MRI using pAIF and WKPM provides simplified, robust single-kidney function estimates. Relevance statement This study proposes a simplified DCE-MRI post-processing method using a population-based arterial input function combined with a whole-kidney pharmacokinetic model. It avoids complex corticomedullary segmentation and minimizes aortic region-of-interest variability, and enables clinically feasible estimation of single-kidney function, supporting broader adoption of renal DCE-MRI in clinical practice. Key Points Population-based arterial input function reduces inter-observer variability and sensitivity to aortic region-of-interest placement artifacts. Whole-kidney modeling avoids complex segmentation of the cortex and medulla regions. DCE-MRI using population-AIF and whole-kidney modeling yields eGFR and RPF significantly correlated with serum and ASL references. Streamlined post-processing workflow supports broader routine clinical use of DCE-MRI. Graphical Abstract
Background Fontan-associated liver disease (FALD) is a consequence of Fontan circulation causing liver fibrosis, cirrhosis, portal hypertension, and potentially hepatocellular carcinoma (HCC), with imaging essential to evaluation. Objectives To describe the cross-sectional imaging features of hepatic morphology and focal liver lesions (FLLs) in patients with FALD. Methods This retrospective single-center study included patients post-Fontan procedure (10/2016-9/2022) who underwent CT or MRI and non-targeted liver biopsy for fibrosis staging. Two observers in consensus assessed CT/MRI for imaging findings of cirrhosis, portal hypertension, and FLL (>0.8cm) detection and characterization, including size, enhancement pattern, and MRI characteristics. A composite reference standard (pathology, multidisciplinary tumor board, imaging characteristics, or FLL stability) informed FLL diagnosis. Biopsies were staged using the Congestive Hepatic Fibrosis Score (CHFS) [range, 0(no fibrosis)-4 (cirrhosis)]. Associations between variables were analyzed using logistic regression and Fisher’s exact tests. Results Results from 41 patients [26M/15F, mean age=29.9y, MRI, n=24/CT, n=17] are presented. CHFS were 1(n=4)/2(n=19)/3(n=14)/4(n=4). Imaging signs of cirrhosis were common: liver surface nodularity (n=31) and volume redistribution (caudate lobe hypertrophy, n=35). Varices and ascites were observed in n=18 and n=16 patients. 62 FLL were identified in 15 patients (mean size=1.5±0.7cm). Diagnoses included benign-appearing enhancing lesions (n=52 lesions/10 patients), indeterminate (n=5 lesions/4 patients), HCC (n=4 lesions/3 patients), and sclerosing hemangioma (n=1). All HCC cases had CHFS=3. No association between laboratory, CHFS, and imaging findings of cirrhosis was found (p-values>0.12). Conclusions Imaging findings of cirrhosis are discordant with fibrosis stage in FALD. Enhancing FLLs, including HCC, are common and frequently observed in noncirrhotic liver.
Background Large language models (LLMs) can generate realistic synthetic medical images (deepfakes), which raise concerns about potential misuse. Purpose To assess the ability of radiologists and multimodal LLMs to distinguish ChatGPT-generated synthetic radiographs from authentic clinical images. Materials and Methods This retrospective diagnostic accuracy study conducted between April and August 2025 included 17 practicing radiologists from six countries with varying experience levels. In phase 1, the radiologists, blinded to the purpose of the study, assessed image quality and provided diagnoses for 154 radiographs from multiple anatomic regions (77 synthetic images generated using ChatGPT [GPT-4o; OpenAI] and 77 authentic images). In phase 2, after being informed of the study's purpose, the radiologists determined whether randomly presented radiographs were GPT-4o-generated or authentic. The same classification task was performed by four LLMs: GPT-4o, GPT-5 (OpenAI), Gemini 2.5 Pro (Google), and Llama 4 Maverick (Meta). In phase 3, an additional set of 110 chest radiographs (55 synthetic images generated using RoentGen and 55 authentic images) was analyzed to evaluate the performance of readers and LLMs in distinguishing synthetic versus authentic images. The McNemar test and t test were used for comparisons. Results Forty-one percent (seven of 17) of purpose-blinded radiologists spontaneously identified artificial intelligence-generated radiographs as being present in the dataset. After being informed that some radiographs were synthetic, there was no evidence of a difference in overall accuracy among all 17 radiologists in distinguishing synthetic images in the GPT-4o dataset (75% [95% CI: 68, 81]) versus in the RoentGen dataset (70% [95% CI: 62, 78]; P = .07). No tested LLM detected all synthetic radiographs in either dataset; however, GPT-4o-generated radiographs were more accurately differentiated from authentic ones by GPT-4o (accuracy, 85%) and GPT-5 (accuracy, 83%) compared with Llama 4 Maverick (accuracy, 59%) and Gemini 2.5 Pro (accuracy, 56%) (all P < .001). Common features of synthetic radiographs included bilateral symmetry, uniform grain or noise patterns, subtly unnatural soft-tissue textures, and overly smooth bone surfaces. Conclusion Synthetic radiographs (deepfakes) generated using an LLM were not easily distinguishable from authentic radiographs by either radiologists or LLMs. Training physicians and LLMs to recognize synthetic images is essential to mitigate risks. To support training, a curated deepfake dataset is available: https://noneedanick.github.io/DeepFakeXRay/. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by Bhayana and Krishna in this issue.
Background & Aims 18F-fluorodeoxyglucose positron emission tomography/computed tomography (18F-FDG PET/CT) is not routinely recommended for hepatocellular carcinoma (HCC) diagnosis or surveillance, and its prognostic role in patients with HCC remains unclear. This meta-analysis aims to evaluate the prognostic significance of 18F-FDG PET/CT in HCC. Methods A systematic literature search was conducted until August 2024 to identify studies evaluating the prognostic value of 18F-FDG PET/CT in patients with HCC. Studies were included if they reported an association between pre-treatment PET/CT and clinical outcomes, including overall survival (OS), recurrence-free survival (RFS), and progression-free survival (PFS). A random-effects model was employed with heterogeneity assessed using the I2 statistic. Results We identified 59 relevant studies (6 prospective and 53 retrospective) including 8,585 patients. Pre-treatment 18F-FDG PET/CT was positive in 30.8% of patients with HCC. Increased FDG activity was associated with poorer OS (hazard ratio [HR] 2.05, 95% CI 1.77–2.38, I2 = 71.6%), RFS (HR 3.18, 95% CI 2.08–4.86, I2 = 96.8%), and PFS (HR 1.64, 95% CI 1.36–1.98, I2 = 71.8%), respectively. FDG activity was significantly associated with OS for both curative (HR 2.10, 95% CI 1.58–2.80, I2 = 62%) and non-curative (HR 2.28, 95% CI 1.59–3.25, I2 = 75.3%) treatments. The prognostic significance of 18F-FDG PET/CT remained consistent across various PET/CT parameters, including SUVmax, tumor-to-liver ratio, and visual assessment of PET/CT positivity, and curative/non-curative therapeutic modalities, all demonstrating that higher FDG uptake is associated with poorer clinical outcomes. Conclusions Pre-treatment 18F-FDG PET/CT is positive in 30.8% of patients with HCC and is significantly associated with adverse clinical outcomes. These findings support its prognostic relevance in HCC, while further studies may help refine its clinical role and standardize 18F-FDG PET/CT-based risk stratification. PROSPERO ID CRD42025633333. Impact and implications This meta-analysis demonstrates that pre-treatment 18F-fluorodeoxyglucose positron emission tomography/computed tomography (18F-FDG PET/CT) has prognostic value in hepatocellular carcinoma (HCC), with higher FDG uptake associated with worse overall, recurrence-free, and progression-free survival. Although not routinely recommended for HCC staging, FDG PET/CT appears to identify a biologically more aggressive subgroup. These findings support its potential role as a complementary tool for risk stratification alongside conventional imaging and clinical markers. Incorporating PET/CT-derived metabolic information may help refine prognostication and guide treatment selection in both curative and non-curative settings. Prospective studies are needed to standardize its use in clinical decision-making.
Precise diagnosis of glioma remains challenging in two aspects: one is to accurately image the tumor margin before surgery, and the other is to distinguish pseudoprogression and true tumor progression after therapy. Here, we developed a non-metallic MRI contrast agent, composed of a serotonin-derived ionizable lipid nanoparticle (LNP) and an mRNA with its 3'UTR engineered for glioma-specific expression. To enhance the contrast signal of the glioma in MRI, the aquaporin-1 (AQP1) mRNA was delivered to encode a water channel on cell membranes. Furthermore, the engineered 3'UTR ensures that the mRNA is degraded in healthy tissues instead of glioma, making pseudoprogression easily detectable by MRI. The mRNA-based contrast agent, named "TARGET", provides an effective complement to gadolinium contrast agents in the diagnosis of brain tumors. This proof-of-concept study demonstrates that TARGET holds promise for substantially enhancing diagnostic accuracy of glioma and pseudoprogression in clinical translation.
Magnetic resonance elastography (MRE) measures liver stiffness for fibrosis staging, but its utility can be hindered by quality control (QC) challenges and measurement variability. The objective of the study was to fully automate liver MRE QC and liver stiffness measurement (LSM) using a deep learning (DL) method. In this retrospective, single center, IRB-approved human study, a curated dataset involved 897 MRE magnitude slices from 146 2D MRE scans [1.5 T and 3 T MRI, 2D Gradient Echo (GRE), and 2D Spin Echo-Echo Planar Imaging (SE-EPI)] of 69 patients (37 males, mean age 51.6 years). A SqueezeNet-based binary QC model was trained using combined and individual inputs of MRE magnitude slices and their 2D Fast-Fourier transforms to detect artifacts from patient motion, aliasing, and blurring. Three independent observers labeled MRE magnitude images as 0 (non-diagnostic quality) or 1 (diagnostic quality) to create a reference standard. A 2D U-Net segmentation model was trained on diagnostic slices with liver masks to support LSM. Intersection over union between the predicted segmentation and confidence masks identified measurable areas for LSM on elastograms. Cohen’s unweighted Kappa coefficient, mean LSM error (
PURPOSE:This systematic review and meta-analysis compared the diagnostic performance of MRI, [ 18 F]FDG-PET/CT, and [ 18 F]FDG-PET/MRI in detecting focal bone lesions and bone marrow infiltration in the initial staging of patients with multiple myeloma (MM) who underwent both MRI and [ 18 F]FDG-PET/CT or [ 18 F]FDG-PET/MRI studies. PATIENTS AND METHODS:A systematic literature search was conducted across PubMed, Embase, and Cochrane databases, including studies comparing the performance of MRI (WB-MRI or spine/pelvis MRI), [ 18 F]FDG-PET/CT, and/or [ 18 F]FDG-PET/MRI in the same patients for MM initial staging. Pooled sensitivities and concordance between imaging modalities were analyzed using R (package META-R). Heterogeneity and bias were assessed with the QUADAS-C tool. RESULTS:Twenty studies (published between 2007 and 2025) using the international MM diagnostic criteria as a reference standard met the inclusion criteria. Of these, 13 (n=742) compared per-patient sensitivity of [ 18 F]FDG-PET/CT and MRI, and 4 (n=224) compared MRI and [ 18 F]FDG-PET/MRI. Pooled sensitivities were 0.807 (95% CI: 0.74-0.86) for [ 18 F]FDG-PET/CT (with significant heterogeneity) versus 0.914 (95% CI: 0.88-0.94) for MRI (0.906 for spine/pelvis MRI and 0.920 for WB-MRI) ( P <0.001 for meta-regression analysis). Using contingency tables, 83% (599/721) of included patients had concordant [ 18 F]FDG-PET/CT and MRI results, while 14% (101/721) patients had negative [ 18 F]FDG-PET/CT and positive MRI with significant differences between the 2 techniques for the paired sample analysis ( P <0.001). The pooled sensitivity of the 4 studies including [ 18 F]FDG-PET/MRI was 0.944 (95% CI: 0.88-0.98). Consensus definitions for specificity in MM imaging should be standardized across studies. CONCLUSIONS:This systematic review and comparative meta-analysis demonstrates superior sensitivity of WB-MRI compared with [ 18 F]FDG-PET/CT for initial staging of MM patients. Future international guidelines might prioritize MRI and [ 18 F]FDG-PET/MRI for staging of MM patients. REGISTRATION:PROSPERO CRD42024564937.