Deep learning reconstruction (DLR) methods can generate false anatomical structures ("hallucinations") by interpreting noise or artifacts, especially at low radiation doses. This study assessed how dose reduction affects hallucinations in two DLR algorithms (DLR1 and DLR2) for high-resolution chest CT scans, specifically for interstitial lung disease (ILD). Full-dose projection data from a photon-counting detector CT scanner were simulated at 25% and 8% dose levels, with 8% approaching chest X-ray doses. Thirty image series (5 patients × 3 doses × 2 DLRs) were independently assessed by four thoracic radiologists under two references: routine reconstruction with a medium-sharp kernel (Qr56-IR3) and the sharpest quantitative kernel (Qr89-IR4). Images were rated on 5-point Likert scales for hallucination presence. DLR2 showed superior hallucination suppression versus DLR1 at 25% and 8% doses (p < 0.01), with comparable hallucination scores at full dose. Radiation dose reduction induces hallucinations in DLR images, with effects depending on algorithm. DLR2 consistently minimized hallucinations better than DLR1 at reduced doses.
Photon -counting CT (PCCT) is an emerging advanced CT technology that differs from conventional CT in its ability to directly convert incident x-ray photon energies into electrical signals. The detector design also permits substantial improvements in spatial resolution and radiation dose efficiency and allows for concurrent high -pitch and high -temporal -resolution multienergy imaging. This review summarizes (a) key differences in PCCT image acquisition and image reconstruction compared with conventional CT; (b) early evidence for the clinical benefit of PCCT for high -spatial -resolution diagnostic tasks in thoracic imaging, such as assessment of airway and parenchymal diseases, as well as benefits of high -pitch and multienergy scanning; (c) anticipated radiation dose reduction, depending on the diagnostic task, and increased utility for routine low -dose thoracic CT imaging; (d) adaptations for thoracic imaging in children; (e) potential for further quantitation of thoracic diseases; and (f) limitations and trade-offs. Moreover, important points for conducting and interpreting clinical studies examining the benefit of PCCT relative to conventional CT and integration of PCCT systems into multivendor, multispecialty radiology practices are discussed. (c) RSNA, 2024 Supplemental material is available for this article.
IntroductionVolume overload from mitral regurgitation can result in left ventricular systolic dysfunction. To prevent this, it is essential to operate before irreversible dysfunction occurs, but the optimal timing of intervention remains unclear. Current echocardiographic guidelines are based on 2D linear measurement thresholds only. We compared volumetric CT-based and 2D echocardiographic indices of LV size and function as predictors of post-operative systolic dysfunction following mitral repair.MethodsWe retrospectively identified patients with primary mitral valve regurgitation who underwent repair between 2005 and 2021. Several indices of LV size and function measured on preoperative cardiac CT were compared with 2D echocardiography in predicting post-operative LV systolic dysfunction (LVEFecho <50%). Area under the curve (AUC) was the primary metric of predictive performance.ResultsA total of 243 patients were included (mean age 57 ± 12 years; 65 females). The most effective CT-based predictors of post-operative LV systolic dysfunction were ejection fraction [LVEFCT; AUC 0.84 (95% CI: 0.77–0.92)] and LV end systolic volume indexed to body surface area [LVESViCT; AUC 0.88 (0.82–0.95)]. The best echocardiographic predictors were LVEFecho [AUC 0.70 (0.58–0.82)] and LVESDecho [AUC 0.79 (0.70–0.89)]. LVEFCT was a significantly better predictor of post-operative LV systolic dysfunction than LVEFecho (p = 0.02) and LVESViCT was a significantly better predictor than LVESDecho (p = 0.03). Ejection fraction measured by CT demonstrated significantly greater reproducibility than echocardiography.DiscussionCT-based volumetric measurements may be superior to established 2D echocardiographic parameters for predicting LV systolic dysfunction following mitral valve repair. Validation with prospective study is warranted.
Background Multiparametric MRI can help identify clinically significant prostate cancer (csPCa) (Gleason score ≥7) but is limited by reader experience and interobserver variability. In contrast, deep learning (DL) produces deterministic outputs. Purpose To develop a DL model to predict the presence of csPCa by using patient-level labels without information about tumor location and to compare its performance with that of radiologists. Materials and Methods Data from patients without known csPCa who underwent MRI from January 2017 to December 2019 at one of multiple sites of a single academic institution were retrospectively reviewed. A convolutional neural network was trained to predict csPCa from T2-weighted images, diffusion-weighted images, apparent diffusion coefficient maps, and T1-weighted contrast-enhanced images. The reference standard was pathologic diagnosis. Radiologist performance was evaluated as follows: Radiology reports were used for the internal test set, and four radiologists' PI-RADS ratings were used for the external (ProstateX) test set. The performance was compared using areas under the receiver operating characteristic curves (AUCs) and the DeLong test. Gradient-weighted class activation maps (Grad-CAMs) were used to show tumor localization. Results Among 5735 examinations in 5215 patients (mean age, 66 years ± 8 [SD]; all male), 1514 examinations (1454 patients) showed csPCa. In the internal test set (400 examinations), the AUC was 0.89 and 0.89 for the DL classifier and radiologists, respectively (P = .88). In the external test set (204 examinations), the AUC was 0.86 and 0.84 for the DL classifier and radiologists, respectively (P = .68). DL classifier plus radiologists had an AUC of 0.89 (P < .001). Grad-CAMs demonstrated activation over the csPCa lesion in 35 of 38 and 56 of 58 true-positive examinations in internal and external test sets, respectively. Conclusion The performance of a DL model was not different from that of radiologists in the detection of csPCa at MRI, and Grad-CAMs localized the tumor. © RSNA, 2024 Supplemental material is available for this article. See also the editorial by Johnson and Chandarana in this issue.
Purpose: Quantitative biomarkers from chest computed tomography (CT) can facilitate the incidental detection of important diseases. Atrial fibrillation (AFib) substantially increases the risk for comorbid conditions including stroke. This study investigated the relationship between AFib status and left atrial enlargement (LAE) on CT. Materials and Methods: A total of 500 consecutive patients who had undergone nongated chest CTs were included, and left atrium maximal axial cross-sectional area (LA-MACSA), left atrium anterior-posterior dimension (LA-AP), and vertebral body cross-sectional area (VB-Area) were measured. Height, weight, age, sex, and diagnosis of AFib were obtained from the medical record. Parametric statistical analyses and receiver operating characteristic curves were performed. Machine learning classifiers were run with clinical risk factors and LA measurements to predict patients with AFib. Results: Eighty-five patients with a diagnosis of AFib were identified. Mean LA-MACSA and LA-AP were significantly larger in patients with AFib than in patients without AFib (28.63 vs. 20.53 cm2, P<0.000001; 4.34 vs. 3.5 cm, P<0.000001, respectively), both with area under the curves (AUCs) of 0.73. Multivariable logistic regression analysis including age, sex, and VB-Area with LA-MACSA improved the AUC for predicting AFib (AUC=0.77). An LA-MACSA threshold of 30 cm2 demonstrated high specificity for AFib diagnosis at 92% and sensitivity of 48%, and LA-AP threshold at 4.5 cm demonstrated 90% specificity and 42% sensitivity. A Bayesian machine learning model using age, sex, height, body surface area, and LA-MACSA predicted AFib with an AUC of 0.743. Conclusions: LA-MACSA or LA-AP can be rapidly measured from routine chest CT, and when >30 cm2 and >4.5 cm, respectively, are specific indicators to predict patients at increased risk for AFib.
•Tuberculosis is a rare etiology of effusive-constrictive pericarditis.•It can present with systolic dysfunction, irrespective of myocardial involvement.•Anti-TB and steroid therapy are reasonable treatment options.
Objectives To evaluate quantitative computed tomography (QCT) features and QCT feature-based machine learning (ML) models in classifying interstitial lung diseases (ILDs). To compare QCT-ML and deep learning (DL) models’ performance. Methods We retrospectively identified 1085 patients with pathologically proven usual interstitial pneumonitis (UIP), nonspecific interstitial pneumonitis (NSIP), and chronic hypersensitivity pneumonitis (CHP) who underwent peri-biopsy chest CT. Kruskal-Wallis test evaluated QCT feature associations with each ILD. QCT features, patient demographics, and pulmonary function test (PFT) results trained eXtreme Gradient Boosting (training/validation set n = 911) yielding 3 models: M1 = QCT features only; M2 = M1 plus age and sex; M3 = M2 plus PFT results. A DL model was also developed. ML and DL model areas under the receiver operating characteristic curve (AUC) and 95% confidence intervals (CIs) were compared for multiclass (UIP vs. NSIP vs. CHP) and binary (UIP vs. non-UIP) classification performances. Results The majority (69/78 [88%]) of QCT features successfully differentiated the 3 ILDs (adjusted p ≤ 0.05). All QCT-ML models achieved higher AUC than the DL model (multiclass AUC micro-averages 0.910, 0.910, 0.925, and 0.798 and macro-averages 0.895, 0.893, 0.925, and 0.779 for M1, M2, M3, and DL respectively; binary AUC 0.880, 0.899, 0.898, and 0.869 for M1, M2, M3, and DL respectively). M3 demonstrated statistically significant better performance compared to M2 (∆AUC: 0.015, CI: [0.002, 0.029]) for multiclass prediction. Conclusions QCT features successfully differentiated pathologically proven UIP, NSIP, and CHP. While QCT-based ML models outperformed a DL model for classifying ILDs, further investigations are warranted to determine if QCT-ML, DL, or a combination will be superior in ILD classification. Key Points • Quantitative CT features successfully differentiated pathologically proven UIP, NSIP, and CHP. • Our quantitative CT-based machine learning models demonstrated high performance in classifying UIP, NSIP, and CHP histopathology, outperforming a deep learning model. • While our quantitative CT-based machine learning models performed better than a DL model, additional investigations are needed to determine whether either or a combination of both approaches delivers superior diagnostic performance.
BACKGROUND: Idiopathic pulmonary fibrosis (IPF) is a progressive, often fatal form of interstitial lung disease (ILD) characterized by the absence of a known cause and usual interstitial pneumonitis (UIP) pattern on chest CT imaging and/or histopathology. Distinguishing UIP/IPF from other ILD subtypes is essential given different treatments and prognosis. Lung biopsy is necessary when noninvasive data are insufficient to render a confident diagnosis. RESEARCH QUESTION: Can we improve noninvasive diagnosis of UIP be improved by predicting ILD histopathology from CT scans by using deep learning? STUDY DESIGN AND METHODS: This study retrospectively identified a cohort of 1,239 patients in a multicenter database with pathologically proven ILD who had chest CT imaging. Each case was assigned a label based on histopathologic diagnosis (UIP or non-UIP). A custom deep learning model was trained to predict class labels from CT images (training set, n = 894) and was evaluated on a 198-patient test set. Separately, two subspecialty-trained radiologists manually labeled each CT scan in the test set according to the 2018 American Thoracic Society IPF guidelines. The performance of the model in predicting histopathologic class was compared against radiologists' performance by using area under the receiveroperating characteristic curve as the primary metric. Deep learning model reproducibility was compared against intra-rater and inter-rater radiologist reproducibility. RESULTS: For the entire cohort, mean patient age was 62 +/- 12 years, and 605 patients were female (49%). Deep learning performance was superior to visual analysis in predicting histopathologic diagnosis (area under the receiver-operating characteristic curve, 0.87 vs 0.80, respectively; P < .05). Deep learning model reproducibility was significantly greater than radiologist inter-rater and intra-rater reproducibility (95% CI for difference in Krippendorff s alpha did not include zero). INTERPRETATION: Deep learning may be superior to visual assessment in predicting UIP/IPF histopathology from CT imaging and may serve as an alternative to invasive lung biopsy.
Imaging-based measurements form the basis of surgical decision making in patients with aortic aneurysm. Unfortunately, manual measurement suffer from suboptimal temporal reproducibility, which can lead to delayed or unnecessary intervention. We tested the hypothesis that deep learning could improve upon the temporal reproducibility of CT angiography-derived thoracic aortic measurements in the setting of imperfect ground-truth training data. To this end, we trained a standard deep learning segmentation model from which measurements of aortic volume and diameter could be extracted. First, three blinded cardiothoracic radiologists visually confirmed non-inferiority of deep learning segmentation maps with respect to manual segmentation on a 50-patient hold-out test cohort, demonstrating a slight preference for the deep learning method (p < 1e-5). Next, reproducibility was assessed by evaluating measured change (coefficient of reproducibility and standard deviation) in volume and diameter values extracted from segmentation maps in patients for whom multiple scans were available and whose aortas had been deemed stable over time by visual assessment (n = 57 patients, 206 scans). Deep learning temporal reproducibility was superior for measures of both volume (p < 0.008) and diameter (p < 1e-5) and reproducibility metrics compared favorably with previously reported values of manual inter-rater variability. Our work motivates future efforts to apply deep learning to aortic evaluation.
Background: There is intense interest and speculation in the application of artificial intelligence (AI) to radiology. The goals of this investigation were (1) to assess thoracic radiologists' perspectives on the role and expected impact of AI in radiology, and (2) to compare radiologists' perspectives with those of computer science (CS) experts working in the AI development. Methods: An online survey was developed and distributed to chest radiologists and CS experts at leading academic centers and societies, comparing their expectations of AI's influence on radiologists' jobs, job satisfaction, salary, and role in society. Results: A total of 95 radiologists and 45 computer scientists responded. Computer scientists reported having read more scientific journal articles on AI/machine learning in the past year than radiologists (mean [95% confidence interval]=17.1 [9.01-25.2] vs. 7.3 [4.7-9.9],P=0.0047). The impact of AI in radiology is expected to be high, with 57.8% and 73.3% of computer scientists and 31.6% and 61.1% of chest radiologists predicting radiologists' job will be dramatically different in 5 to 10 years, and 10 to 20 years, respectively. Although very few practitioners in both fields expect radiologists to become obsolete, with 0% expecting radiologist obsolescence in 5 years, in the long run, significantly more computer scientists (15.6%) predict radiologist obsolescence in 10 to 20 years, as compared with 3.2% of radiologists reporting the same (P=0.0128). Overall, both chest radiologists and computer scientists are optimistic about the future of AI in radiology, with large majorities expecting radiologists' job satisfaction to increase or stay the same (89.5% of radiologists vs. 86.7% of CS experts,P=0.7767), radiologists' salaries to increase or stay the same (83.2% of radiologists vs. 73.4% of CS experts,P=0.1827), and the role of radiologists in society to improve or stay the same (88.4% vs. 86.7%,P=0.7857). Conclusions: Thoracic radiologists and CS experts are generally positive on the impact of AI in radiology. However, a larger percentage, but still small minority, of computer scientists predict radiologist obsolescence in 10 to 20 years. As the future of AI in radiology unfolds, this study presents a historical timestamp of which group of experts' perceptions were closer to eventual reality.
PURPOSE:Echocardiography (echo) is widely used for right ventricular (RV) assessment. Current techniques for RV evaluation require additional imaging and manual analysis; machine learning (ML) approaches have the potential to provide efficient, fully automated quantification of RV function. METHODS:An automated ML model was developed to track the tricuspid annulus on echo using a convolutional neural network approach. The model was trained using 7791 image frames, and automated linear and circumferential indices quantifying annular displacement were generated. Automated indices were compared to an independent reference of cardiac magnetic resonance (CMR) defined RV dysfunction (RVEF < 50%). RESULTS:A total of 101 patients prospectively underwent echo and CMR: Fully automated annular tracking was uniformly successful; analyses entailed minimal processing time (<1 second for all) and no user editing. Findings demonstrate all automated annular shortening indices to be lower among patients with CMR-quantified RV dysfunction (all P < .001). Magnitude of ML annular displacement decreased stepwise in relation to population-based tertiles of TAPSE, with similar results when ML analyses were localized to the septal or lateral annulus (all P ≤ .001). Automated segmentation techniques provided good diagnostic performance (AUC 0.69-0.73) in relation to CMR reference and compared to conventional RV indices (TAPSE and S') with high negative predictive value (NPV 84%-87% vs 83%-88%). Reproducibility was higher for ML algorithm as compared to manual segmentation with zero inter- and intra-observer variability and ICC 1.0 (manual ICC: 0.87-0.91). CONCLUSIONS:This study provides an initial validation of a deep learning system for RV assessment using automated tracking of the tricuspid annulus.
Purpose: To test the performance of a deep learning (DL) model in predicting atrial fibrillation (AF) at routine nongated chest CT. Materials and Methods: A retrospective derivation cohort (mean age, 64 years; 51% female) consisting of 500 consecutive patients who underwent routine chest CT served as the training set for a DL model that was used to measure left atrial volume. The model was then used to measure atrial size for a separate 500-patient validation cohort (mean age, 61 years; 46% female), in which the AF status was determined by performing a chart review. The performance of automated atrial size as a predictor of AF was evaluated by using a receiver operating characteristic analysis. Results: There was good agreement between manual and model-generated segmentation maps by all measures of overlap and surface distance (mean Dice = 0.87, intersection over union = 0.77, Hausdorff distance = 4.36 mm, average symmetric surface distance = 0.96 mm), and agreement was slightly but significantly greater than that between human observers (mean Dice = 0.85 [automated] vs 0.84 [manual]; P =.004). Atrial volume was a good predictor of AF in the validation cohort (area under the receiver operating characteristic curve = 0.768) and was an independent predictor of AF, with an age-adjusted relative risk of 2.9. Conclusion: Left atrial volume is an independent predictor of the AF status as measured at routine nongated chest CT. Deep learning is a suitable tool for automated measurement.
Phase contrast (PC) cardiovascular magnetic resonance (CMR) is widely employed for flow quantification, but analysis typically requires time consuming manual segmentation which can require human correction. Advances in machine learning have markedly improved automated processing, but have yet to be applied to PC-CMR. This study tested a novel machine learning model for fully automated analysis of PC-CMR aortic flow. A machine learning model was designed to track aortic valve borders based on neural network approaches. The model was trained in a derivation cohort encompassing 150 patients who underwent clinical PC-CMR then compared to manual and commercially-available automated segmentation in a prospective validation cohort. Further validation testing was performed in an external cohort acquired from a different site/CMR vendor. Among 190 coronary artery disease patients prospectively undergoing CMR on commercial scanners (84% 1.5T, 16% 3T), machine learning segmentation was uniformly successful, requiring no human intervention: Segmentation time was < 0.01 min/case (1.2 min for entire dataset); manual segmentation required 3.96 ± 0.36 min/case (12.5 h for entire dataset). Correlations between machine learning and manual segmentation-derived flow approached unity (r = 0.99, p < 0.001). Machine learning yielded smaller absolute differences with manual segmentation than did commercial automation (1.85 ± 1.80 vs. 3.33 ± 3.18 mL, p < 0.01): Nearly all (98%) of cases differed by ≤5 mL between machine learning and manual methods. Among patients without advanced mitral regurgitation, machine learning correlated well (r = 0.63, p < 0.001) and yielded small differences with cine-CMR stroke volume (∆ 1.3 ± 17.7 mL, p = 0.36). Among advanced mitral regurgitation patients, machine learning yielded lower stroke volume than did volumetric cine-CMR (∆ 12.6 ± 20.9 mL, p = 0.005), further supporting validity of this method. Among the external validation cohort (n = 80) acquired using a different CMR vendor, the algorithm yielded equivalently small differences (∆ 1.39 ± 1.77 mL, p = 0.4) and high correlations (r = 0.99, p < 0.001) with manual segmentation, including similar results in 20 patients with bicuspid or stenotic aortic valve pathology (∆ 1.71 ± 2.25 mL, p = 0.25). Fully automated machine learning PC-CMR segmentation performs robustly for aortic flow quantification - yielding rapid segmentation, small differences with manual segmentation, and identification of differential forward/left ventricular volumetric stroke volume in context of concomitant mitral regurgitation. Findings support use of machine learning for analysis of large scale CMR datasets.
Throughout history, automation has brought about broad societal benefits at the expense of displaced workers. Although this can have negative consequences for certain industries, it is generally accepted that the widespread prosperity brought about by automation outweighs its temporary effects on employment. Medicine has remained largely insulated from job losses, but it is difficult to argue against the hypothetical benefits of automated physician labor, given that we are among the highest paid professionals and the fact that our services account for nearly 20% of the $3.3 trillion spent per year on health care in the United States alone [ 1 Centers for Disease Control and Prevention, National Center for Health StatisticsHealth expenditures. https://www.cdc.gov/nchs/fastats/health-expenditures.htmDate accessed: January 3, 2019 Google Scholar ]. Moreover, the resources necessary to support a workforce of physicians are a substantial barrier to health care access in the developing world, and automation could alleviate this.