
Research on AI in diagnostic radiology has focused on algorithm development and standalone performance, yet the human-AI interactions that ultimately determine the clinical value of AI are poorly understood. A narrative synthesis of the human-AI interaction and radiology AI literature highlighted three underrecognized determinants of successful human-AI collaboration in radiology. First, cognitive psychology dictates how automation bias, the framing of AI uncertainty, and AI-induced skill decay distort diagnostic reasoning. Second, the user interface (UI) and user experience (UX) of AI tools determine how the timing, salience, and documentation of AI findings shape radiologists' attention and reporting behavior; the influence of UI/UX is becoming even more critical as generative AI introduces novel applications and interaction paradigms. Third, algorithmic conformity leads radiologists to override their own judgment under medicolegal, transparency, and organizational pressures. For each domain, concrete research directions are proposed to bridge the gap between algorithmic capabilities and clinical utility.
Artificial intelligence (AI) is being deployed in radiology for image reconstruction and postprocessing to produce images that appear sharper, smoother, and more detailed. However, incorporating AI also introduces new failure modes and can exacerbate the disconnect between the perceived quality of an image and its diagnostic information content. Understanding the limitations of AI-enabled image reconstruction and postprocessing is essential for the safe and effective use of this technology. Therefore, the purpose of this report is to raise awareness of the limitations of AI-based image reconstruction and postprocessing to help users realize the benefits of the technology while minimizing risks. Accordingly, this report reviews approaches to image quality assessment, discusses the regulatory framework relevant to AI-enabled imaging devices, describes AI-specific failure modes, and outlines strategies to mitigate associated risks.
Purpose To evaluate the effect of an artificial intelligence (AI)-driven CT queue prioritization system on emergency department CT wait times using a discrete-event simulation. Materials and Methods Multimodal data from 313 966 emergency department visits (August 2020-August 2024) were retrospectively analyzed. A gradient boosting machine model was trained to predict the clinical actionability of CT studies at the time of order. Discrete-event simulations, calibrated to real-world operations, were conducted to compare a first-in, first-out policy against an AI-driven prioritization policy on the internal test set of 20 795 studies (March-August 2024). Wait times for actionable and nonactionable studies were compared between the first-in, first-out and AI-based dispatch policies, with bootstrap 95% CIs, and ceiling analyses using perfect predictions. Results Compared with the first-in, first-out policy, the AI-based dispatch policy reduced median wait times for actionable studies in the internal test set (7637 of 20 795 [36.73%]) by 10.75 minutes (95% CI: -12.40, -9.10) and 90th-percentile wait times by 43.36 minutes (95% CI: -50.58, -36.80). The proportion of actionable findings obtained within 1 hour increased from 48.33% (3691 of 7637) to 57.30% (4376 of 7637). For nonactionable studies, the median wait time was reduced by 5.80 minutes (95% CI: -7.00, -4.40), but the 90th-percentile wait time increased by 14.46 minutes (95% CI: 6.40, 23.20). The model captured 80%-87% of the maximum benefit achievable with perfect predictions. Conclusion AI-driven CT queue prioritization can help reduce the time to diagnosis for critical findings with minimal disruption to lower-acuity patients, using existing data streams and without the need for additional hardware. Keywords: Artificial Intelligence, CT, Emergency Department, Discrete-Event Simulation, Workflow, Queue Prioritization Supplemental material is available for this article. © RSNA, 2026 See also the editorial by Pfeiffer and Thalhammer in this issue.
Purpose To compare the performance of an artificial intelligence (AI) system with that of radiologists for estimating malignancy risk of indeterminate-size nodules (5-15 mm) at low-dose CT (LDCT) within a standardized and transparent evaluation framework. Materials and Methods Teams participating in the AI study had access to a public dataset of 555 malignant and 5608 benign nodules on 4069 baseline LDCT scans from the National Lung Screening Trial to develop AI systems. External testing was performed on 156 malignant and 312 benign size-matched nodules, all of indeterminate size, from 463 baseline scans collected from three large European lung cancer screening trials, and the best-performing AI system (based on area under the receiver operating characteristic curve [AUC]) was selected. An observer study was conducted in which radiologists assessed 300 randomly selected nodules (100 malignant, 200 benign) from the external test set. Radiologists categorized nodules as low, intermediate, or high risk, and the threshold of intermediate or greater risk (intermediate or high-risk) was used to define a positive test result. The selected AI system was compared with radiologists on this subset using the AUC. Results The selected AI system demonstrated superior performance to the 65 radiologists' mean (AUC, 0.78 [95% CI: 0.73, 0.84] vs 0.70 [95% CI: 0.65, 0.74]; P = .001). With use of the intermediate risk or greater threshold, the AI system correctly classified 12% more malignant nodules at matched specificity and yielded 20% fewer false-positive results at matched sensitivity. Conclusion The selected AI system was superior to radiologists in estimating malignancy risk of indeterminate lung nodules at LDCT. Keywords: CT, Thorax, Lung, Observer Performance, Screening, Supervised Learning, Lung Cancer Screening, Radiologists, Artificial Intelligence, Benchmarking, Pulmonary Nodule Malignancy Risk, Deep Learning Supplemental material is available for this article. © RSNA, 2026 See also commentary by Júdice de Mattos Farina and Szarf in this issue.
Purpose To develop and systematically evaluate an iterative training approach, termed the expert-guided annotation loop, for efficient reference standard segmentation generation, including assessment of two sample selection strategies and real-world clinical implementation. Materials and Methods This retrospective study included 10 datasets comprising 1941 CT and MRI scans from patients with autosomal dominant polycystic kidney disease, prostate cancer, uveal melanoma, thyroid eye disease, or non-small cell lung cancer. nnU-Net segmentation models were iteratively trained using an expert-guided annotation loop with random or active learning-based sample selection. In each iteration, additional samples were added to the training set, and model-generated presegmentations were corrected by expert radiologists to create reference standard annotations. Expert time required for manual segmentation versus presegmentations correction was measured. Model performance and efficiency were assessed using nonparametric tests, and cost savings were estimated for kidney and tumor segmentation using probabilistic sensitivity analysis. Feasibility of end-to-end no-code implementation was evaluated. Results Fifty-seven segmentation models were trained and evaluated. Final model mean Dice scores ranged from 0.67 to 0.97 for organ segmentation and from 0.64 to 0.69 for lung tumor segmentation across internal and external test sets. Maximum expert time savings were 90.3% for kidney and 48.2% for tumor segmentation (P < .001 and P = .003, respectively), corresponding to estimated per-examination cost savings of $14.30 (95% CI: 5.94, 26.87) and $5.63 (95% CI: -7.26, 26.09), respectively. No-code execution of the expert-guided annotation loop was feasible. Conclusion The expert-guided annotation loop reduced expert annotation time and enabled estimated cost savings while producing high-quality reference standard segmentations. The no-code workflow was implemented in a clinical environment. Keywords: Artificial Intelligence, CT, Human-in-the-Loop Machine Learning, Medical Image Segmentation, MRI Supplemental material is available for this article. © RSNA, 2026.
Purpose To investigate whether deep learning models trained on chest radiographs rely on radiographic exposure parameters as shortcut features and to quantify the resulting biases under controlled confounding and natural exposure regimes. Materials and Methods In this retrospective study, chest radiographs from MIMIC-CXR (January 2011-December 2016), the Medical Imaging and Data Resource Center (August 2020-May 2022), and EmoryCXR (September 2008-February 2023) were analyzed for pneumothorax detection, COVID-19 diagnosis, and race classification. Dataset-provided labels served as the reference standard. Three exposure parameters (ExposureTime, XRayTubeCurrent, and ExposureInuAs) were extracted from Digital Imaging and Communications in Medicine metadata. Models were trained under biased and balanced exposure label alignments and evaluated on matched and reversed distributions. A priori screening additionally identified high-risk exposure regimes. The area under the receiver operating characteristic curve (AUC) was compared using the DeLong test. Results A total of 727 604 chest radiographs in 240 681 patients (mean age ± SD, 60 years ± 17; 126 432 men, 114 128 women) were analyzed. For pneumothorax detection, AUC decreased from 0.94 (95% CI: 0.94, 0.95) to 0.56 (95% CI: 0.55, 0.58) on mismatched exposure distributions (ΔAUC, -0.38; P < .001). Similar declines were observed for COVID-19 (ΔAUC, -0.33; P < .001) and race classification (ΔAUC, -0.09; P < .001). The a priori exposure-regimen screening revealed high-risk regimes within the natural distribution that were associated with reduced model performance compared with typical exposures. Conclusion Deep learning models trained on chest radiographs may exploit exposure parameters as shortcut features; exposure-regimen audits may flag high-risk conditions before clinical deployment. Keywords: Computer Aided Diagnosis (CAD), Lung, Feature Detection, Diagnosis, Convolutional Neural Network (CNN), Chest Radiograph, Exposure Parameters, Deep Learning, Fairness, Shortcut Learning Supplemental material is available for this article. © RSNA, 2026 See also commentary by Le Guellec and Chassagnon in this issue.
Purpose To evaluate the pooled diagnostic accuracy of externally tested artificial intelligence (AI) models for malignancy classification of lung nodules at chest CT. Materials and Methods A systematic search of PubMed, Embase, Web of Science, Cumulative Index to Nursing and Allied Health Literature, and the Cochrane Library was performed in March 2023 and updated on January 12, 2025, to identify studies evaluating AI models for malignancy classification of lung nodules at chest CT using pathology and/or at least 2-year follow-up as reference standards. Risk of bias was assessed using the Quality Assessment of Diagnostic Accuracy Studies-2 tool, and pooled sensitivity and specificity were estimated using bivariate random-effects models. Results Twenty-one studies including 7454 nodules were analyzed, with lung cancer prevalence ranging from 5.7% (17 of 297) to 91.5% (214 of 234). All models were based on deep learning; 17 of the 21 studies (81%) involved Asian populations, 16 (76%) used nonscreening populations, 14 (67%) reported two-dimensional or three-dimensional convolutional neural network (CNN) architectures, and eight (38%) specified predefined malignancy thresholds. High risk of bias was identified in five studies for patient selection and in two for index testing. Pooled sensitivity was 88%, pooled specificity was 75%, positive likelihood ratio was 3.55, negative likelihood ratio was 0.16, area under the receiver operating characteristic curve was 0.89, and the diagnostic odds ratio was 22.4. Heterogeneity was high (I2 > 90%). Model architecture was associated with specificity, with higher values in studies reporting two-dimensional or three-dimensional CNNs compared with those without reported architecture (82%-83% vs 58%, P = .03; meta-regression P = .02); other subgroup analyses showed no evidence of differences. Conclusion Externally tested AI models demonstrated high sensitivity but moderate specificity for malignancy classification of lung nodules at chest CT, supporting a potential role in rule-out strategies. However, substantial heterogeneity, inconsistent reporting, and risk of bias limit interpretation. Keywords: Lung Cancer, Artificial Intelligence, Malignancy Classification, External Validation, Meta-Analysis Supplemental material is available for this article. © RSNA, 2026 See also commentary by Bressem and Kim in this issue.
Artificial intelligence (AI) has progressed from technical research to routine clinical use, reaching an inflection point where technological capabilities may exceed current regulatory and oversight frameworks. These systems are becoming more complex, progressing from narrow, task-specific algorithms to foundation models and early agentic prototypes. This progression has redistributed risk, responsibility, and clinical judgment, requiring radiologists and health care leaders to understand how policy choices affect patient safety and clinical innovation advancement. This special report provides a roadmap for aligning policy with clinical practice through a practical, lifecycle-based framework centered on patient safety. Translational bialignment is a concept that pairs regulatory science requirements (what AI systems should deliver to clinicians and patients) with implementation science capabilities (what institutions should provide for safe deployment of AI). This framework addresses the full AI lifecycle, from data stewardship and model development to validation, deployment, and monitoring, and articulates shared responsibilities for vendors, institutions, and clinicians grounded in trustworthy AI principles. The analysis focuses on U.S. regulatory frameworks, particularly Food and Drug Administration policies governing medical AI, with relevant highlights from international approaches. Concrete opportunities for radiologists to engage in policy formation, participate in oversight, and collaborate with industry and policymakers are provided to help shape a trustworthy and sustainable AI ecosystem. The alignment of policy, practice, and patient safety will enable medical AI to have a lasting impact on clinical care and public trust. The analysis and recommendations provided represent the authors' perspectives and do not necessarily reflect the official positions of the Radiological Society of North America. Keywords: Artificial Intelligence, Food and Drug Administration, Health Policy, Implementation Science, Large Language Models, Machine Learning, Patient Safety, Regulatory Science, Translational Bialignment © RSNA, 2026.
In the French national breast cancer screening program, second reading is performed only for mammograms interpreted as negative (Breast Imaging and Reporting Data System [BI-RADS] 1-2) at first reading, representing a unique screening workflow in Europe. This retrospective study assessed whether artificial intelligence (AI) could identify a subgroup of negative screening mammograms that could safely bypass second reading. A total of 55 589 screening mammograms from 42 419 women aged 50-74 years (January 2015-December 2019) whose mammograms were initially classified as BI-RADS 1-2 were analyzed. Second reading outcomes were compared with those of a commercial AI system using a predefined binary threshold (≥5). Among these examinations, 183 of 55 589 (0.33%) were recalled at second reading, yielding 12 cancers (positive predictive value, 6.6%; cancer detection rate, 0.22 per 1000 examinations). AI classified 42 606 of 55 589 (76.6%) examinations as low risk (≤4) and 12 983 of 55 589 (23.3%) as nonlow risk (≥5). One cancer was detected in the AI-low group (one of 55 589 [0.002%]) compared with 11 in the AI-nonlow group (11 of 55 589 [0.020%]; P < .001). Interval cancer rates were higher in the AI-nonlow group than in the AI-low group (2.16 vs 0.47 per 1000 examinations). These findings suggest that excluding AI-low examinations from second reading could reduce workload by approximately 77% while focusing radiologist review on higher-risk cases; prospective validation is needed. Keywords: Mammography, Breast, Computer Aided Diagnosis (CAD) Supplemental material is available for this article. © RSNA, 2026 See also commentary by Yao and Chae in this issue.
Purpose To develop a deep learning-based deformable registration method for dynamic contrast-enhanced (DCE) breast MRI that preserves tumor regions while maintaining global anatomic alignment during neoadjuvant chemotherapy (NAC) response assessment. Materials and Methods This retrospective study included internal and external cohorts of patients with breast cancer who were undergoing NAC. The internal cohort comprised patients who underwent DCE MRI from 2017 to 2020, and the external cohort was derived from the I-SPY2 trial. A conditional pyramid registration network integrating unsupervised keypoint detection with a volume-preserving mechanism was developed. Registration performance was evaluated using the Dice similarity coefficient (DSC), average landmark error, and tumor volume difference. A local-global biomarker derived from registered images was evaluated for predicting pathologic complete response (pCR) using the area under the receiver operating characteristic curve (AUC) and accuracy. Paired t tests were used for statistical comparisons. Results In 314 patients (all female; age, 50.6 years ± 12.0 [SD]) with 1630 scans in the internal cohort, the proposed method achieved a DSC of 0.95 ± 0.02, an average landmark error of 5.35 mm ± 3.46, and a tumor volume difference of 11.0% ± 10.7. In 100 patients (all female; age, 48.5 years ± 12.3) with 372 scans in the external cohort, the method achieved a DSC of 0.91 ± 0.09 and a tumor volume difference of 15.5% ± 13.8. Improvements in landmark distance and tumor preservation were statistically significant (P < .05) compared with most methods. For pCR prediction, incorporation of the proposed biomarker achieved an AUC of 0.81 ± 0.04 and an accuracy of 72.1% ± 5.0. Conclusion The proposed framework improved anatomic alignment while preserving tumor volume in longitudinal DCE breast MRI during NAC response assessment and enabled a registration-based biomarker for predicting treatment response. Keywords: MR Imaging, Image Postprocessing, Breast, Neural Networks, Radiomics, Prognosis, DCE Breast MRI, Deep Learning, Deformable Registration, Neoadjuvant Chemotherapy, Unsupervised Keypoint Detection Supplemental material is available for this article. © RSNA, 2026 See also the commentary by Zhang in this issue.
Purpose To develop a deep learning-enabled single breath-hold abbreviated MRI (DL-SBH-aMRI) protocol for hepatocellular carcinoma (HCC) diagnosis. Materials and Methods Patients at high risk for HCC from four institutions (January 2019-January 2025) were prospectively and retrospectively included. All patients underwent conventional complete MRI (cMRI) examinations, including precontrast T1-weighted imaging (pre-T1); postcontrast T1-weighted imaging in arterial, portal venous, and delayed phases; T2-weighted imaging (T2WI); diffusion-weighted imaging (DWI); and apparent diffusion coefficient mapping. Four generative models were trained to synthesize full MRI sequences (T2WI, DWI, apparent diffusion coefficient, arterial phase, portal venous phase, delayed phase) from pre-T1. The best-performing model was selected to generate synthetic sequences, which, combined with pre-T1 acquired from MRI, formed the DL-SBH-aMRI protocol. Image quality, perceptual realism, and lesion size measurement accuracy were evaluated for DL-SBH-aMRI versus cMRI; diagnostic performance at the patient and lesion levels was assessed using a reference standard based on histopathology and imaging findings. Results A total of 1008 patients were included (mean age ± SD, 56.8 years ± 11.8; 700 male patients). The diffusion-based generative model (Li-DiffNet) yielded the highest synthetic image quality and was selected as the backbone of DL-SBH-aMRI. DL-SBH-aMRI was noninferior to cMRI in subjective image quality scores (4.07-4.16 vs 4.18-4.19; P < .001). DL-SBH-aMRI demonstrated noninferior diagnostic performance (all P < .001) for HCC at the patient and lesion levels (sensitivity of 77.9%-88.7%; specificity of 91.6%-93.1%) compared with cMRI (sensitivity of 84.4%-92.5%; specificity of 94.1%-95.2%). Conclusion The DL-SBH-aMRI protocol may enable gadolinium-free, rapid MRI of the liver while preserving full-sequence diagnostic information for HCC diagnosis. Keywords: Hepatocellular Carcinoma, Deep Learning, Abbreviated MRI, Diagnostic Evaluation, Image Generation Supplemental material is available for this article. © The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. See also commentary by Wang and Gu in this issue.