
Background:Susceptibility-induced distortion in single-shot echo-planar imaging (ssEPI) prostate diffusion-weighted imaging (DWI) can degrade lesion conspicuity, apparent diffusion coefficient (ADC) estimation, and confidence in prostate MRI interpretation. Conventional correction strategies typically require additional acquisitions and may be limited in severe artifact settings. Purpose:To develop a physics-informed deep learning framework, distortion-guided restoration (DGR), for acquisition-free correction of ssEPI distortions in prostate DWI. Materials and methods:Distortion-guided restoration was proposed to learn the inverse of a physically simulated ssEPI distortion process. Clinical prostate DWI and T2-weighted (T2W) images from 408 3T MR examinations originating from 2 clinical datasets, acquired on Siemens 3T MAGNETOM Vida and Biograph mMR scanners, were used for paired distorted/undistorted training data across low-b DWI (b = 50 s/mm2), and scanner-provided ADC maps. The restoration network integrates a convolutional neural network (CNN)-based geometric correction module with a conditional diffusion refinement module, both guided by co-registered T2W images. Model selection and benchmarking were performed on a held-out synthetic test set (n = 34), against conventional algorithms using correction field mapping (FSL FUGUE and TOPUP). Model performance was evaluated in a retrospective study on 34 clinical scans with severe susceptibility artifacts drawn from the same clinical datasets, using quantitative image metrics and blinded radiologist scoring. Results:On synthetic data, the proposed DGR model achieved the highest peak signal-to-noise ratio (PSNR) and lowest normalized mean squared error (NMSE) across low-b DWI (0.089; 95% CI, 0.072-0.105) and ADC maps (0.062; 95% CI, 0.053-0.072), outperforming FSL TOPUP and FUGUE (all P < .001). In the 34 clinical cases affected by susceptibility artifacts, including 16 cases with histologically confirmed lesions (16/34, 47.1%), DGR substantially improved geometric fidelity (2.74 to 3.29), identified all histologically confirmed lesions (16/16, 100%), and enhanced overall image quality and diagnostic confidence (all P < .001). Conclusion:This proof-of-concept study suggests that a physics-informed hybrid CNN-diffusion framework offers a practical and acquisition-free solution for correcting severe prostate DWI distortions. It warrants further development for clinical utility and implementation.
Background:Cardiac magnetic resonance elastography (MRE) allows the non-invasive quantification of myocardial stiffness, but long-interval reproducibility data at 3.0 T single-frequency 140-Hz remain limited. Purpose:To evaluate the short- and long-interval reproducibility of 3.0 T cardiac MRE for left ventricular (LV) myocardial stiffness across healthy volunteers and patients with conditions associated with elevated myocardial stiffness. Materials and Methods:This prospective single-center study included 69 participants: 60 healthy volunteers and 9 with cardiovascular disease (6 with aortic stenosis and 3 with long-standing hypertension). All participants underwent baseline cardiac MRE (scanner: Signa Premier 3.0 T, GE Healthcare, driver: Resoundant), an immediate short-interval rescan (SIR) within 5 minutes, and a long-interval rescan (LIR) after approximately 1 month. Left ventricular myocardial stiffness and octahedral shear strain signal-to-noise ratio were measured independently by 2 readers. Reproducibility was assessed using Bland-Altman analysis, intraclass correlation coefficients (ICCs), within-subject coefficient of variation, standard error of measurement, and smallest detectable change at the 95% confidence (SDC95) level. Results:Baseline LV myocardial stiffness was higher in the disease group than in healthy volunteers, 11.73 ± 1.01 kPa and 7.80 ± 0.70 kPa (mean ± SD; P < .001), respectively. Inter-reader agreement was excellent across all timepoints, with ICCs of 0.970 at baseline (95% CI, 0.953-0.982), 0.966 for SIR (95% CI, 0.945-0.979), and 0.962 for LIR (95% CI, 0.939-0.976). Within-reader scan-rescan reproducibility was also excellent. For Reader 1, stiffness ICC was 0.979 for the SIR (95% CI, 0.966-0.987) and 0.957 for the LIR (95% CI, 0.932-0.973), with SDC95 values of 0.616 kPa (95% CI, 0.528-0.741) and 0.869 kPa (95% CI, 0.744-1.044), respectively. Reader 2 showed excellent reproducibility, with ICCs ≤0.943. Conclusions:3D single-frequency 140-Hz cardiac MRE at 3.0 T provides excellent short- and long-interval reproducibility for LV myocardial stiffness quantification. The SDC95 values may help distinguish true longitudinal stiffness changes from measurement variability.
Abstract Background Manual lesion measurements remain the standard for assessing oncologic treatment response, despite being time-consuming and prone to substantial inter-reader variability, which may lead to inconsistent response classification and subsequent variability in treatment decisions. Purpose To evaluate the impact of an artificial intelligence (AI) system for assisted lesion measurement on reading time and measurement consistency in follow-up CT examinations using the Response Evaluation Criteria in Solid Tumors (RECIST 1.1). Methods and Materials In this retrospective reader study, follow-up chest-abdomen-pelvis CT examinations from 212 oncology patients collected at two Dutch hospitals were assessed by 23 readers (15 radiologists, 8 residents) recruited from 11 international institutions under three conditions: unassisted, AI-assisted, and expert-assisted (using a prior radiologist’s unassisted measurement). To prevent bias related to the source of the measurements, readers were informed that all support was AI-derived. Primary outcomes were reading time to completion and inter-observer measurement variability. For each outcome, a Bayesian generalized linear mixed model was used to analyze the results. Results AI assistance significantly reduced per-patient reading time versus unassisted reading (-36.0 s; 95% CI: -53.0, -22.1). At the lesion level, it was associated with a small increase in variability relative to the expert-derived reference standard (1.32 mm; 95% CI: 0.83, 1.91). At the patient level, AI assistance did not meaningfully affect change in the sum of longest diameters (SLD; -0.38 mm; 95% CI: -1.57, 0.72), but increased RECIST outcome agreement by 7.7% (95% CI: 2.8, 12.7) compared with unassisted reading. Expert-assisted reading yielded even higher inter-reader agreement (13.3%; 95% CI: 8.6, 18.1). Conclusion AI assistance reduced reading time and improved patient-level RECIST agreement, despite a small increase in lesion-level measurement variability. These findings suggest that AI-assisted RECIST assessment may improve workflow and response classification consistency, while also providing a benchmark from expert-assisted reading for future AI development.
Background:Accurately estimating brain age can help identify deviations linked to neurodegenerative diseases, underscoring the need for robust models that accurately perform across heterogenous cohorts. Purpose:To develop an age prediction model that is interpretable and robust to demographic and technological variations in brain MRI. Materials and Methods:We propose a transformer-based brain age model that analyzes 3D T1-weighted MRI. Model performance was assessed using mean absolute error (MAE). Associations between brain age gap (BAG, ie, predicted minus chronological age) and chronological age were evaluated in cognitive normal (CN) participants. Clinical relevance was assessed by examining BAG differences across cognitive groups and correlations with Mini-Mental State Examination (MMSE) and Montreal Cognitive Assessment (MoCA). Results:We achieved an MAE 3.65 years on ADNI2 & 3 and OASIS3 test sets, and a high generalizability of MAE of 3.54 years on AIBL. In dementia, a notable increase in brain age gap (BAG) along with cognitive decline, with a mean of 0.15 years (95% CI: [-0.22, 0.51]) in CN, 2.55 years ([2.40, 2.70]) in mild cognitive impairment (MCI), and 6.12 years ([5.82, 6.43]) is noted. Negative correlation between BAG and cognitive scores was observed after adjustment for covariates, with r = -0.397 (P < 0.001) for MMSE and -0.393 (P < 0.001) for MoCA, where declining scores generally signify worsening cognitive performance. The saliency map highlighted white and deep gray matter structures as key regions influenced by brain aging. Conclusion:Our model effectively integrated multiview and volumetric information to achieve state-of-the-art brain age prediction, with improved generalizability, interpretability, and association with cognitive function.
Background:Radiology-risk communication affects multiple clinical specialties that use ionizing radiation, and many patients seek related information online. Prior expert evaluations found comparable performance between ChatGPT-generated and radiology-risk answers from official institutions, but patient perspectives have not been assessed. Purpose:To assess patients' perceptions of ChatGPT versus human-generated radiology-risk information. Methods and Materials:From December 2024 to March 2025, patients at 3 hospitals in the United States, Switzerland, and Lebanon were randomly assigned to 1 of 5 common radiology-risk questions. Participants, blinded to source, provided subjective ratings of both ChatGPT‑3.5 and human-generated institutional responses on 7-point Likert scales for satisfaction (primary outcome), comprehensibility, trust, and reassurance. Quantitative comparisons were performed with Inverse Normalizing Transformation, and free-text comments were analyzed using thematic coding. Results:A total of 328 patients participated (34% aged 18-39 years, 33% aged 40-59 years, 31% aged 60-79 years, 3% aged ≥80 years; 188 female). ChatGPT responses were rated significantly higher than human responses for satisfaction (0.70-point advantage; P < .001), comprehensibility (0.27 points; P < .01), trust (0.65 points; P < .001), and reassurance (0.51 points; P < .001). Findings converged with qualitative written comments (r = 0.91, P < .05), in which ChatGPT attracted 2.3× more positive comments while human-generated responses received 1.7× more negative comments. Conclusions:Unlike experts, patients preferred ChatGPT-generated responses to institutional materials for radiology risk questions. This divergence highlights the need for patient-centered communication and suggests that large language model-based styles, implemented with expert oversight, may improve the perceived clarity and trustworthiness of educational materials in medical specialties that use ionizing radiation.
Synovitis is the key inflammatory feature of rheumatoid arthritis (RA). Quantitative assessment of synovitis better correlates with patient outcomes than semi-quantitative assessment but it is time-consuming. To develop and validate an automated model for segmentation and quantification of wrist synovial tissue volume on post-contrast fat-suppressed T1-weighted MRI. Early rheumatoid arthritis patients (symptoms for ≤24 months) at a single center were recruited at baseline and were followed-up at year-1 and year-8. Post-contrast axial fat-suppressed T1-weighted images of the most symptomatic wrist were acquired at 3.0 T. One observer manually segmented consecutive synovitis areas on all MRI datasets. A framework, based on the convolutional neural network, nnU-Net, was trained and validated (5-fold cross-validation with image level splits) with 295 image datasets used for model training and validation. Rheumatoid arthritis MRI score (RAMRIS) was used to semi-quantitatively grade synovitis. Manually segmented synovial volume by a single reader was used as the reference standard. Forty-five external image datasets from two different imaging centers were used to test generalizable applicability. For automated synovitis segmentation, the overall Sørensen-Dice similarity coefficient (DSC) was 0.75 ± 0.11(mean ± sd) compared to manual segmentation. Higher DSC values were found in patients with moderate (0.80 ± 0.06) and severe (0.84 ± 0.05) degrees of synovitis. The model had a similar performance with external-acquired data (DSC value : 0.70 ± 0.20). Predicted and manually segmented synovitis volume measurements showed excellent agreement (Pearson correlation: r = 0.975, p < 0.001). A fully automated model quantified wrist synovial tissue volume with good agreement to manual reference and maintained performance on external data, supporting potential use in clinical studies and prospective evaluation in practice.
Background:Opportunistic screening of osteoporosis on CT has emerged as a cost-effective strategy for bone health assessment. However, vertebral attenuation varies substantially with CT tube voltage and spinal level, limiting standardization. Purpose:To model and standardize vertebral attenuation across tube voltages and spinal levels on low-dose chest CT for opportunistic osteoporosis screening. Materials and Methods:This retrospective study included 589 patients (336 women; mean age ± standard deviation, 65.9 ± 12.4 years) who underwent 2 noncontrast low-dose chest CT examinations and 1 dual-energy X-ray absorptiometry within 3 months. Vertebral attenuation was measured at T10-L2 on CT using a deep learning-based software (AVIEW SpineBH, Coreline Soft). Osteoporosis was diagnosed using dual-energy X-ray absorptiometry. A linear mixed-effects model regressed attenuation on the logarithm of tube voltage, incorporating fixed effects for spinal level (T10-L2) and patient-specific random effects. Model performance was evaluated by converting the L1-120 kVp reference to other tube-voltage-level combinations and comparing predicted with observed attenuation using Pearson correlation, mean bias, and 95% limits of agreement. Diagnostic applicability was assessed by comparing receiver operating characteristic-derived thresholds with model-predicted thresholds converted from the reference. Results:Vertebral attenuation decreased by approximately 6-11 Hounsfield units (HU) for every 10-kVp increase (β = -88.8 HU per log[kVp]; P < .001) and declined by approximately 4-14 HU per level from T10 to L2. Predicted values closely matched observed measurements (r = 0.89; mean bias, 2.2 HU; 95% limits of agreement, -54 to 59 HU). The model-predicted thresholds showed strong agreement with receiver operating characteristic-derived thresholds, differing by 1-10 HU across tube voltages and spinal levels. Conclusions:A log-linear mixed-effects model enables population-level standardization of vertebral attenuation across tube voltages and spinal levels, supporting the development of more consistent opportunistic CT-based osteoporosis screening methods on noncontrast chest CT.
Photon-counting detector CT (PCD-CT) directly detects individual photons and measures their energies, enabling the acquisition of multi-energy data from a single x-ray source with higher spatial resolution, thereby allowing visualization of smaller arteries and veins and other important imaging features not seen at conventional CT. It also substantially reduces calcium and stent blooming artifacts. Appropriate patient selection and task-based optimization of acquisition, reconstruction, and interpretation are critical to realize these benefits. Vascular imaging examinations for which early evidence shows that PCD-CT demonstrates higher diagnostic performance compared to conventional CT include coronary CT angiography (CTA) for evaluation of partially calcified plaques or stents, peripheral runoff CTA, detection of cerebrospinal fluid-venous fistulas, and head CTA for evaluation of small intracranial vessels and aneurysms. Current limitations include excessive image noise with the sharpest reconstruction kernels and challenges when adapting conventional CT protocols (eg, requiring sharper kernels, thinner slices, higher matrices, tailored scan modes) to achieve time-efficient image reconstruction and hanging protocols. New solutions and directions will include deep-learning-based image noise reduction, improved spectral separation, strategic triage of specific patient populations to PCD-CT systems, and subspecialty clinical practice recommendations for this new technology.
Background:Acute pancreatitis (AP) is a common gastrointestinal disease with a rising global incidence. While most cases are mild, severe AP (SAP) carries high mortality. Early and accurate severity prediction facilitates management optimization. However, existing clinical severity prediction models, such as Bedside Index of Severity in Acute Pancreatitis (BISAP) and Modified CT Severity Index (mCTSI), have modest accuracy and often rely on data unavailable at admission. Purpose:This study proposes a deep learning (DL) model to predict AP severity using abdominal contrast-enhanced CT scans acquired within 24 hours of admission. Materials and Methods:We collected 10 130 studies from 8335 patients across a multi-site U.S. health system and 3488 studies from public datasets. The DL model was trained in 2 stages: (1) self-supervised pretraining on 11 896 unlabeled studies and (2) fine-tuning on 550 labeled studies. Performance was evaluated against mCTSI on a hold-out internal test set (n = 100 patients) and against both mCTSI and BISAP on an external Hungarian AP registry (n = 518 patients). Results:On the internal test set, the model achieved areas under receiver operating curve (AUROCs) of 0.888 (95% CI: 0.800-0.960, sensitivity 0.50, specificity 0.92) for SAP and 0.888 (95% CI: 0.819-0.946, sensitivity 0.75, specificity 0.85) for mild AP (MAP), outperforming mCTSI (P = .002). External validation showed AUROCs of 0.887 (95% CI: 0.825-0.941, sensitivity 0.73, specificity 0.88) for SAP and 0.858 (95% CI: 0.826-0.888, sensitivity 0.38, specificity 0.98) for MAP, surpassing mCTSI (P = .024) and BISAP (P = .002). In retrospective triage analysis, the model correctly identified 50%-73% of patients who progressed to SAP and 40%-73% of those with MAP. Conclusion:The proposed DL model achieved performance comparable to or better than established prognostic tools and maintained robust external performance. These findings suggest that AI-assisted CT analysis may support early, automated risk stratification of AP.
Background:MR diffusion-weighted imaging (DWI), especially at high b-value, is a key acquisition to help identify clinically significant prostate cancer; however, it suffers from low signal-to-noise ratio (SNR), high noise floor, and susceptibility artifact. Purpose:To demonstrate the feasibility of improving DWI quality using a novel 50-channel pelvic coil in conjunction with a deep learning (DL)-based phase correction and a DL-denoising algorithm. Methods:In this prospective, single-center study, 24 consecutive men referred for prostate multiparametric MRI over 16 months were enrolled (age 47-79 years; mean, 68.1 years). Axial T2-weighted images and DWI were obtained using a prototype 50‑channel coil and standard clinical phased array (3 T Architect, GE HealthCare, USA). The DWI acquisitions were reconstructed with the vendor's deep learning denoising algorithm (ARDL). The same raw data were reconstructed offline using an investigational DL Phase Correction algorithm with ARDL (DLPC+ARDL). Two independent readers scored DWI and ADC series using 4 qualitative criteria. SNR and contrast-to-noise ratio (CNR) were measured on b = 1500 s/mm2 images. Combined reader scores were compared using the Wilcoxon matched‑pairs signed‑rank test, inter‑reader variability was assessed using Cohen's κ, and quantitative SNR/CNR values were compared using 2‑tailed paired t‑tests. Results:Twenty men were analyzable for qualitative and 18 for quantitative metrics (reported as mean ± SD). 50‑channel pelvic coil with DLPC+ARDL produced the highest SNR (99.70 ± 28.50) and CNR (91.68 ± 44.39), exceeding 50‑channel with ARDL alone (SNR 56.44 ± 28.50; CNR 51.11 ± 28.67), 30‑channel anterior array with ARDL (SNR 31.41 ± 13.18; CNR 26.26 ± 13.52), and DLPC + ARDL (SNR 49.4 ± 18.6; CNR 44.7 ± 18.3) (all P < .0001). Reader scores favored DLPC+ARDL in prostate border definition, peripheral/transition zone distinction, lesion conspicuity, and confidence of extraprostatic extension (all P values for DWI at b = 1500 s/mm2 < 0.0001; synthetic DWI at b = 2000 s/mm2: P = 7.5 × 10-7-0.01). Inter‑reader agreement was fair for acquired DWI (quadratic‑weighted κ = 0.34) and lower for synthetic DWI (κ = 0.19). Conclusion:DWI images acquired using the 50‑channel pelvic coil and reconstructed with the DLPC + ARDL pipeline yield the highest image quality compared to ARDL only pipeline and to all 30-channel coil imaging.
Background:Staging irreversible tissue injury in myocardial infarction (MI) enables risk assessment for post-MI major cardiovascular events. While cardiac MRI is the preferred modality for staging the severity of tissue injury, conventional scan protocols require long acquisition times with multiple breath-held and ECG-gated acquisitions, limiting its utilization. Purpose:To develop a free-breathing, whole-heart, non-ECG gated cardiac MRI for staging irreversible tissue injury in MI that can be completed in <20 minutes. Materials and Methods:A fast cardiac MRI (Biograph, Siemens Healthcare, 3 T) method based on a low-rank tensor framework was developed (reconstruction performed in MATLAB) and tested against the conventional approach using a pre-clinical canine model of reperfused MI (n = 15) with histological validation. Each subject underwent 2 exams that were randomized 2 days apart (day 6-8 post MI, respectively). Correlations between the proposed and conventional methods and left-ventricular ejection fraction (LVEF), MI size and transmurality, size of microvascular obstruction (MVO), and intramyocardial hemorrhage (IMH) volumes were assessed using linear regression and Bland-Altman analysis. Results:Twelve out of 15 subjects survived the initial reperfusion injury. The proposed method reduced acquisition time by >50%. The cardiac MRI evidence of tissue injury was confirmed on histopathology in all cases. The agreements between the proposed and conventional methods for LVEF, MI volume, persistent MVO volume and IMH volume were excellent; limits of agreement (LoA) were -2.1%-1.8%, -2.9 % -3.3%, -2.4%-4.1%, and -1.5%-1.5%, respectively. MI transmurality and early MVO showed good agreement; LoA were -6.8%-9.7% and -6.6%-8.2%, respectively. Conclusion:The proposed free-breathing, whole-heart, non-ECG gated cardiac MRI approach permits accurate determination of tissue injury in a canine model with >2-fold reduction in scan time. While the method remains to be tested in patients, it has the potential to facilitate efficient use of cardiac MRI for staging the severity of tissue injury in patients with reperfused MI.
Background:Pancreatic cancer (PC) is frequently missed on unenhanced CT examinations performed for unrelated clinical indications, where the pancreas is included incidentally and clinical suspicion is low. Purpose:To develop and validate a deep learning-based tool for PC diagnosis and risk stratification on unenhanced CT. Materials and Methods:This retrospective study included 3080 unenhanced CT studies of Taiwanese patients with PC, other pancreatic diseases and normal pancreas between 2004 and 2019 from a tertiary hospital, randomly divided into training, validation, and internal test sets. Unenhanced CT studies from United States institutions were used for external testing. A hybrid convolutional neural network-transformer model was trained for PC diagnosis and risk stratification. Performance was evaluated using sensitivity, specificity, and area under the curve (AUC), with comparisons to 2 radiologists by McNemar's test and exploratory decision curve analysis. Results:The internal dataset included 713 PCs (mean age, 64.6 ± 12.0 years; 384 men), 1661 normal pancreas and 706 other pancreatic diseases. In an exploratory comparison restricted to unenhanced CT (29 PCs, 31 controls), the sensitivity of computer-aided diagnosis (CAD) tool (89.7%, 72.6-97.8) seemed comparable with that of 1 radiologist (86.2%, 68.3-96.1) and higher than another (41.4%, 23.5-61.1); but wide confidence intervals and inter-radiologist variability warrant cautious interpretation. In the internal test set (142 PCs, 474 controls), sensitivity was 90.8% (84.9-95.0) and specificity 93.0% (90.4-95.2) (AUC: 0.98), with sensitivity comparable to radiologist reports based on enhanced and unenhanced CT (95.4%, 90.2-98.3; P = .21). In the external set (42 PCs, 22 controls), sensitivity was 76.2% (60.5-87.9) and specificity 86.4% (65.1-97.1) (AUC: 0.89). The tool stratified cases into 7 risk levels with likelihood ratios ranging from <0.01 to 173.46. Exploratory decision curve analysis suggested potential net benefit across threshold probabilities. Conclusion:This tool may assist in the opportunistic detection and risk stratification of PC on unenhanced CT.
Building on the commendable work by Thomas et al., this editorial proposes a simplified workflow for pre-treatment identification of patients needing macroaggregated-albumin-based lung shunt fraction prior to the 90Y treatment.
Screening whole-body MRI (WB-MRI) is gaining increasing attention as a tool for early disease detection, with growing adoption driven largely by consumer demand and direct-to-consumer private platforms. While WB-MRI has demonstrated utility in high-risk populations, its use in asymptomatic individuals remains controversial due to concerns about low diagnostic yield, false positives, overdiagnosis, and the lack of survival outcome data. Despite these limitations, the popularity of WB-MRI is expected to rise given the aging population and aggressive marketing by direct-to-consumer companies, underscoring the need for thoughtful and proactive engagement by radiologists. Radiologists have an obligation to ensure that scientific rigor, ethical oversight, and multidisciplinary collaboration guide the expansion of WB-MRI. This review outlines the current evidence and evolving landscape of screening WB-MRI, describes the development and implementation of a program within an academic radiology practice, and discusses the downstream implications and cost-effectiveness of WB-MRI screening. As this technology continues to expand beyond traditional indications, radiologists must play a leading role in defining best practices and ensuring that implementation remains evidence-based, transparent, and patient-centered.
Background:Patients access radiology reports without delay, which can cause anxiety and misunderstanding. While large language models (LLMs) can generate patient-friendly summaries (PS) to mitigate this, their potential to address literacy-based disparities remains unquantified. Purpose:To measure the effect of LLM-generated PS on the objective comprehension and subjective experiences of individuals reading lung cancer screening computed tomography reports, and to determine the differential impact with respect to self-rated English and health literacy. Materials and Methods:This cross-sectional survey (July 24-28, 2025) used a within-subjects design. Participants from the online research platform Prolific, self-enrolled from the general U.S. population, viewed 3 lung cancer screening reports (negative; negative with complex incidentals; suspicious for malignancy), first in their original format and then with an LLM-generated PS. Objective comprehension, subjective experiences (including anxiety, via a 5-point Likert scale), and hypothetical communication intent were assessed after each viewing. Univariate and multivariate analyses, including demographic subgroup comparisons, were performed. Results:A total of 1815 participants (mean age, 46 years ± 16 [SD]; 919 women) who completed the survey were evaluated. The addition of a PS improved objective comprehension (P < .001) and reduced anxiety (P < .001) for all scenarios. This effect was most pronounced for respondents with low self-rated English literacy, who had greater comprehension gain (P < .001) and anxiety reduction (P = .012) than those with high literacy. Individuals with low self-rated health literacy also experienced more anxiety reduction (P < .001). PS also increased the proportion of participants reporting that they would be willing to wait for a scheduled appointment to discuss their results (P < .001). Conclusion:LLM-generated PS of lung cancer screening reports increase comprehension and reduce anxiety, most notably among individuals with lower self-rated English and health literacy. If validated in a patient population, they represent a potential tool to improve communication.
Background:Cardiomegaly is a clinically significant incidental finding on chest computed tomography (CT) associated with heart failure, arrhythmias, and sudden cardiac death. Qualitative radiologist assessment is variable, and automated AI tools may enable objective opportunistic cardiac volumetry. Purpose:To evaluate whether AI-enabled total cardiac volume (TCVAI) derived from non-ECG-gated, non-contrast chest CT can identify cardiomegaly as defined by echocardiography. Materials and Methods:This retrospective study included 307 consecutive patients (median age, 67 years; 56% male) who underwent non-contrast chest CT at a single center on 7 scanner types (4 vendors) and clinically indicated echocardiography within 31 days. A commercially available AI tool (AI-Rad Companion, Siemens Healthineers) automatically quantified TCVAI, indexed to body surface area (TCVAI/BSA). Echocardiography reports were reviewed for chamber dilation and left ventricular hypertrophy (LVH), collectively defined as cardiomegaly. Associations between TCVAI/BSA and echocardiographic findings were assessed using correlation, ordinal regression, and receiver operating characteristic (ROC). Interscan repeatability was evaluated in 248 patients with 544 repeat CT examinations. Prespecified sex-specific thresholds were tested in a temporally independent validation cohort of 50 patients. Results:Median TCVAI was higher in patients with cardiomegaly than those without (1061.9 vs 798.4 mL; P < .001). TCVAI/BSA was associated with chamber dilation and LVH severity on univariate analysis and remained associated in multivariable ordinal models, except for right ventricular dilation. Discriminatory performance was fair to good, with area under the curve (AUC) 0.81 (95% CI, 0.75-0.87) in men and 0.77 (95% CI, 0.69-0.85) in women. Interscan repeatability was excellent (intraclass correlation coefficient [ICC]: 0.93). In independent validation, performance ranged from sensitivity 89.3%/specificity 27.3% at a high-sensitivity threshold to sensitivity 28.6%/specificity 100% at a high-specificity threshold. Conclusion:AI-derived cardiac volume from routine chest CT shows fair to good performance for identifying echocardiography-defined cardiomegaly with high measurement repeatability, supporting a potential role for automated cardiac volumetry as an objective, opportunistic biomarker.