Chest X-rays (CXRs) are the most commonly performed imaging investigation. In the UK, many centers experience reporting delays due to radiologist workforce shortages. Artificial intelligence (AI) tools capable of distinguishing "normal" from "abnormal" CXRs have emerged as a potential solution. If "normal" CXRs could be safely identified and reported without human input, a substantial portion of radiology workload could be reduced. This article examines the feasibility and implications of autonomous AI reporting of "normal" CXRs, using the United Kingdom as an example setting. Key issues include defining "normal," ensuring generalizability across populations, and managing the sensitivity-specificity trade-off. It also addresses legal and regulatory challenges, such as compliance with IR(ME)R and GDPR, and the lack of accountability frameworks for errors. Further considerations include the impact on radiologists practice, the need for robust post-market surveillance, and incorporation of patient perspectives. While the benefits are clear, adoption must be cautious, with strong governance, legal clarity, and rigorous clinical validation to ensure safe and sustainable use.
Background Various commercial artificial intelligence (AI) devices can identify lung cancer features on chest radiographs. Understanding their relative performance within the same patient population and health care setting is essential for making informed clinical deployment decisions. Purpose To determine and compare the diagnostic accuracy of multiple commercial AI devices for lung cancer detection on chest radiographs in a patient sample demonstrating a representative prevalence. Materials and Methods Consecutive chest radiographs requested from primary care for any indication in adult patients acquired at a single United Kingdom center between July 2020 and February 2021 were eligible for inclusion. Each radiograph was independently analyzed by each device. The multidisciplinary team decision served as the reference standard for diagnosis. Receiver operating characteristic analysis was performed for continuous scores, with the DeLong test used to compare the area under the receiver operating characteristic curve for the devices. Contingency tables of device classification results were used to calculate diagnostic accuracy metrics. The Cochran Q and McNemar tests were used to compare proportions of classification results between the devices. Fleiss κ was used to assess the agreement of classification results across the devices. Results A total of 5235 radiographs were obtained from 5235 patients (median age, 60 years; 53.4% female, 79.4% White, and 1.4% diagnosed with lung cancer with a visible tumor). Devices from seven manufacturers were tested. The area under the receiver operating characteristic curve varied (0.80-0.94; P < .05 for nine of 15 pairwise comparisons). The sensitivity (20.8%-77.8%), specificity (58.9%-98.4%), and positive predictive value (1.5%-28.4%) varied (P < .001 and P < .05 for 39 of 44 pairwise comparisons). The number of additional false-positive results for tumor detection compared with radiologist reporting ranged from 10 to 2039. Device classification results showed minimal agreement (κ = 0.24). Conclusion There was clinically and statistically significant variability in the diagnostic accuracy of commercial AI devices for lung cancer detection at chest radiography. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by Schaefer-Prokop and Schalekamp in this issue.
Background: There is clinical need to better quantify lung disease severity in pulmonary hypertension (PH), particularly in idiopathic pulmonary arterial hypertension (IPAH) and PH associated with lung disease (PH-LD). Purpose: To quantify fibrosis on CT pulmonary angiograms using an artificial intelligence (AI) model and to assess whether this approach can be used in combination with radiologic scoring to predict survival. Materials and Methods: This retrospective multicenter study included adult patients with IPAH or PH-LD who underwent incidental CT imaging between February 2007 and January 2019. Patients were divided into training and test cohorts based on the institution of imaging. The test cohort included imaging examinations performed in 37 external hospitals. Fibrosis was quantified using an established AI model and radiologically scored by radiologists. Multivariable Cox regression adjusted for age, sex, World Health Organization functional class, pulmonary vascular resistance, and diffusing capacity of the lungs for carbon monoxide was performed. The performance of predictive models with or without AI-quantified fibrosis was assessed using the concordance index (C index). Results: The training and test cohorts included 275 (median age, 68 years [IQR, 60-75 years]; 128 women) and 246 (median age, 65 years [IQR, 51-72 years]; 142 women) patients, respectively. Multivariable analysis showed that AI-quantified percentage of fibrosis was associated with an increased risk of patient mortality in the training cohort (hazard ratio, 1.01 [95% CI: 1.00, 1.02]; P = .04). This finding was validated in the external test cohort (C index, 0.76). The model combining AI-quantified fibrosis and radiologic scoring showed improved performance for predicting patient mortality compared with a model including radiologic scoring alone (C index, 0.67 vs 0.61; P < .001). Conclusion: Percentage of lung fibrosis quantified on CT pulmonary angiograms by an AI model was associated with increased risk of mortality and showed improved performance for predicting patient survival when used in combination with radiologic severity scoring compared with radiologic scoring alone.
Background: Computed tomography pulmonary angiography (CTPA) has been proposed to be diagnostic for pulmonary hypertension (PH) in multiple studies. However, the utility of the unenhanced CT measurements diagnosing PH has not been fully assessed. This study aimed to assess the diagnostic utility and reproducibility of cardiac and great vessel parameters on unenhanced computed tomography (CT) in suspected pulmonary hypertension (PH). Methods: In total, 42 patients with suspected PH who underwent unenhanced CT thorax and right heart catheterization (RHC) were included in the study. Three observers (a consultant radiologist, a specialist registrar in radiology, and a medical student) measured the parameters by using unenhanced CT. Diagnostic accuracy of the parameters was assessed by area under the receiver operating characteristic curve (AUC). Inter-observer variability between the consultant radiologist (primary observer) and the two secondary observers was determined by intra-class correlation analysis (ICC). Results: Overall, 35 patients were diagnosed with PH by RHC while 7 patients were not. Main pulmonary arterial (MPA) diameter was the strongest (AUC 0.79 to 0.87) and the most reproducible great vessel parameter. ICC comparing the MPA diameter measurement of the consultant radiologist to the specialist registrar’s and the medical student’s were 0.96 and 0.92, respectively. Right atrial area was the cardiac measurement with highest accuracy and reproducibility (AUC 0.76 to 0.79; ICC 0.980, 0.950) followed by tricuspid annulus diameter (AUC 0.76 to 0.79; ICC 0.790, 0.800). Conclusions: MPA diameter and right atrial areas showed high reproducibility. Diagnostic accuracies of these were within the range of acceptable to excellent, and might have clinical value. Tricuspid annular diameter was less reliable and less diagnostic and was therefore not a recommended diagnostic measurement.
BackgroundChronic pulmonary embolism (PE) may result in pulmonary hypertension (CTEPH). Automated CT pulmonary angiography (CTPA) interpretation using artificial intelligence (AI) tools has the potential for improving diagnostic accuracy, reducing delays to diagnosis and yielding novel information of clinical value in CTEPH. This systematic review aimed to identify and appraise existing studies presenting AI tools for CTPA in the context of chronic PE and CTEPH.MethodsMEDLINE and EMBASE databases were searched on 11 September 2023. Journal publications presenting AI tools for CTPA in patients with chronic PE or CTEPH were eligible for inclusion. Information about model design, training and testing was extracted. Study quality was assessed using compliance with the Checklist for Artificial Intelligence in Medical Imaging (CLAIM).ResultsFive studies were eligible for inclusion, all of which presented deep learning AI models to evaluate PE. First study evaluated the lung parenchymal changes in chronic PE and two studies used an AI model to classify PE, with none directly assessing the pulmonary arteries. In addition, a separate study developed a CNN tool to distinguish chronic PE using 2D maximum intensity projection reconstructions. While another study assessed a novel automated approach to quantify hypoperfusion to help in the severity assessment of CTEPH. While descriptions of model design and training were reliable, descriptions of the datasets used in training and testing were more inconsistent.ConclusionIn contrast to AI tools for evaluation of acute PE, there has been limited investigation of AI-based approaches to characterising chronic PE and CTEPH on CTPA. Existing studies are limited by inconsistent reporting of the data used to train and test their models. This systematic review highlights an area of potential expansion for the field of AI in medical image interpretation.There is limited knowledge of A systematic review of artificial intelligence tools for chronic pulmonary embolism in CT. This systematic review provides an assessment on research that examined deep learning algorithms in detecting CTEPH on CTPA images, the number of studies assessing the utility of deep learning on CTPA in CTEPH was unclear and should be highlighted.
Background High-resolution CT (HRCT) is central to the assessment of interstitial lung disease (ILD), and accurate classification of disease has important implications for patients. Evaluation of imaging features can be challenging, even for experienced thoracic radiologists. Previous work has provided equivocal evidence on the interpretation of HRCT features at ILD-related imaging. Purpose To perform a meta-analysis to assess the level of agreement among expert thoracic radiologists in interpreting ILD-related imaging. Materials and Methods A systematic literature search from January 2000 to October 2023 of the Ovid MEDLINE, Embase, and Cochrane Central Register of Controlled Trials databases was performed for articles reporting assessments of interobserver agreement between thoracic radiologists for evaluation of ILD findings, such as severity and progression of disease, presence of features such as honeycombing and ground-glass opacification, and classification based on the 2011 and 2018 American Thoracic Society/European Respiratory Society/Japanese Respiratory Society/Asociación Latinoamericana del Tórax (ATS/ERS/JRS/ALAT) guidelines for idiopathic pulmonary fibrosis (IPF). Meta-analysis was performed using a random-effects model to obtain pooled κ or intraclass correlation coefficient (ICC) values as measures of interobserver agreement. Results The final analysis included 13 studies consisting of 6943 images and 146 radiologists. In 10 studies assessing agreement of specific radiologic findings in ILD, the pooled κ value was 0.56 (95% CI: 0.43, 0.70). In eight studies, the assessed interobserver agreement of the ATS/ERS/JRS/ALAT diagnostic guidelines for IPF based on usual interstitial pneumonia (UIP) patterns, the pooled κ value was 0.61 (95% CI: 0.48, 0.74). One study reported a κ value of 0.87 for ILD progression. Seven studies assessing ILD severity could not be pooled; the individual κ values for ILD severity ranged from 0.64 to 0.90, and ICC values ranged from 0.63 to 0.96. Conclusion There was moderate agreement between thoracic radiologists when assessing ILD features and UIP pattern diagnosis but little evidence on agreement of disease severity, extent, or progression. Meta-analysis registry no. PROSPERO CRD42022361803 © RSNA, 2024 Supplemental material is available for this article. See also the editorial by Humbert in this issue.
OBJECTIVE:The introduction of Targeted Lung Health Checks (TLHC) to screen for lung cancer has highlighted that incidental findings are common and require management strategies. This study analyses retrospectively, incidentally detected breast lesions reported as part of the TLHC referred to the Breast Cancer clinicians.METHODS:All participants with incidental breast nodules referred to the Breast Cancer team in the first year of screening were reviewed.RESULTS:Fifty-two participants (48 female; 92.3%) were referred to the Breast Multidisciplinary Team Meeting for assessment of 43 breast nodules, 8 breast asymmetry/dense breasts, and 2 likely breast related metastatic disease. One participant declined breast team referral. For the 42 breast nodules investigated, the final diagnoses were 5 breast carcinomas, 10 normal breast tissue, and 27 benign nodules. One male patient was diagnosed with breast carcinoma. The 29 breast nodules classified as smooth and well defined were all benign. No malignancy was demonstrated in the group with asymmetric or dense breast tissue. Metastatic breast carcinoma was confirmed in two participants. Twenty-six out of thirty-seven (54%) females had prior breast screening mammograms precluding further investigation.CONCLUSION:Incidental breast nodules are common on THLC scans. Smooth, sharply defined breast nodules are likely to be benign but low-dose CT is poor at accurately assessing breast nodules. Agreed breast referral pathways prior to starting the Lung Cancer Screening programme are recommended. Access to screening mammograms can reduce referrals to the Breast clinic.ADVANCES IN KNOWLEDGE:Lessons learned from TLHC pilot studies can be useful to sites commencing national TLHC programme.
Objectives Early identification of lung cancer on chest radiographs improves patient outcomes. Artificial intelligence (AI) tools may increase diagnostic accuracy and streamline this pathway. This study evaluated the performance of commercially available AI-based software trained to identify cancerous lung nodules on chest radiographs.Design This retrospective study included primary care chest radiographs acquired in a UK centre. The software evaluated each radiograph independently and outputs were compared with two reference standards: (1) the radiologist report and (2) the diagnosis of cancer by multidisciplinary team decision. Failure analysis was performed by interrogating the software marker locations on radiographs.Participants 5722 consecutive chest radiographs were included from 5592 patients (median age 59 years, 53.8% women, 1.6% prevalence of cancer).Results Compared with radiologist reports for nodule detection, the software demonstrated sensitivity 54.5% (95% CI 44.2% to 64.4%), specificity 83.2% (82.2% to 84.1%), positive predictive value (PPV) 5.5% (4.6% to 6.6%) and negative predictive value (NPV) 99.0% (98.8% to 99.2%). Compared with cancer diagnosis, the software demonstrated sensitivity 60.9% (50.1% to 70.9%), specificity 83.3% (82.3% to 84.2%), PPV 5.6% (4.8% to 6.6%) and NPV 99.2% (99.0% to 99.4%). Normal or variant anatomy was misidentified as an abnormality in 69.9% of the 943 false positive cases.Conclusions The software demonstrated considerable underperformance in this real-world patient cohort. Failure analysis suggested a lack of generalisability in the training and testing datasets as a potential factor. The low PPV carries the risk of over-investigation and limits the translation of the software to clinical practice. Our findings highlight the importance of training and testing software in representative datasets, with broader implications for the implementation of AI tools in imaging.
Recent AI breakthroughs can automate quantification of imaging features. Distinguishing Group 1 and 3 PH is challenging in patients with 'mild' lung disease due to overlapping clinical characteristics. This multicentre study develops and deploys a CT AI model to quantify and establish the prognostic value of lung parenchymal changes. 521 consecutive patients between 2001-19 were included from the ASPIRE registry. A novel AI model which automatically classifies the lung parenchyma, and provides the percentage of normal, ground glass (GG), ground glass reticulation (GGR), honeycombing and emphysema was developed. Fibrosis severity was scored by specialist radiologists. Multivariate cox regression adjusted for age, sex, WHO function class, and DLCO. Findings were externally validated in 246 patients (33 centres, 37 scanners). All patterns were univariate predictors. GGR% (HR 1.02, p=0.015) and fibrosis% (HR 1.01, p=0.05) were continuous multivariate predictors. 2% GGR and 4% fibrosis corresponded to 20% 1-year mortality. In the external cohort, these thresholds were multivariate predictors (2% GGR HR 1.74, p=0.011 and 4% fibrosis HR 1.85, p=0.004). In 300 patients radiologically scored as having 'no' fibrosis, AI identified minor disease (1.2% GGR) which was prognostic (HR 1.03, p=0.006). Adding GGR to a predictive model of radiologically scored disease significantly improved the model (c-index 0.763 vs 0.742, p=0.038). This is the largest AI study in this domain. GGR% and fibrosis% are indepedent prognostic markers. AI is sensitive to minor disease, and provides additional predictive value used in combination with radiological reporting.
The patterns of idiopathic pulmonary fibrosis (IPF) lung disease that directly correspond to elevated hyperpolarised gas diffusion-weighted (DW) MRI metrics are currently unknown. This study aims to develop a spatial co-registration framework for a voxel-wise comparison of hyperpolarised gas DW-MRI and CALIPER quantitative CT patterns. Sixteen IPF patients underwent 3He DW-MRI and CT at baseline, and eleven patients had a 1-year follow-up DW-MRI. Six healthy volunteers underwent 129Xe DW-MRI at baseline only. Moreover, 3He DW-MRI was indirectly co-registered to CT via spatially aligned 3He ventilation and structural 1H MRI. A voxel-wise comparison of the overlapping 3He apparent diffusion coefficient (ADC) and mean acinar dimension (LmD) maps with CALIPER CT patterns was performed at baseline and after 1 year. The abnormal lung percentage classified with the LmD value, based on a healthy volunteer 129Xe LmD, and CALIPER was compared with a Bland–Altman analysis. The largest DW-MRI metrics were found in the regions classified as honeycombing, and longitudinal DW-MRI changes were observed in the baseline-classified reticular changes and ground-glass opacities regions. A mean bias of −15.3% (95% interval −56.8% to 26.2%) towards CALIPER was observed for the abnormal lung percentage. This suggests DW-MRI may detect microstructural changes in areas of the lung that are determined visibly and quantitatively normal by CT.
ObjectivesRight ventricle (RV) mass is an imaging biomarker of mean pulmonary artery pressure (MPAP) and pulmonary vascular resistance (PVR). Some methods of RV mass measurement on cardiac MRI (CMR) exclude RV trabeculation. This study assessed the reproducibility of measurement methods and evaluated whether the inclusion of trabeculation in RV mass affects diagnostic accuracy in suspected pulmonary hypertension (PH).Materials and methodsTwo populations were enrolled prospectively. (i) A total of 144 patients with suspected PH who underwent CMR followed by right heart catheterization (RHC). Total RV mass (including trabeculation) and compacted RV mass (excluding trabeculation) were measured on the end-diastolic CMR images using both semi-automated pixel-intensity-based thresholding and manual contouring techniques. (ii) A total of 15 healthy volunteers and 15 patients with known PH. Interobserver agreement and scan-scan reproducibility were evaluated for RV mass measurements using the semi-automated thresholding and manual contouring techniques.ResultsTotal RV mass correlated more strongly with MPAP and PVR (r = 0.59 and 0.63) than compacted RV mass (r = 0.25 and 0.38). Using a diagnostic threshold of MPAP ≥ 25 mmHg, ROC analysis showed better performance for total RV mass (AUC 0.77 and 0.81) compared to compacted RV mass (AUC 0.61 and 0.66) when both parameters were indexed for LV mass. Semi-automated thresholding was twice as fast as manual contouring (p < 0.001).ConclusionUsing a semi-automated thresholding technique, inclusion of trabecular mass and indexing RV mass for LV mass (ventricular mass index), improves the diagnostic accuracy of CMR measurements in suspected PH.
IntroductionSevere pulmonary hypertension (mean pulmonary artery pressure ≥35 mmHg) in chronic lung disease (PH-CLD) is associated with high mortality and morbidity. Data suggesting potential response to vasodilator therapy in patients with PH-CLD is emerging. The current diagnostic strategy utilises transthoracic Echocardiography (TTE), which can be technically challenging in some patients with advanced CLD. The aim of this study was to evaluate the diagnostic role of MRI models to diagnose severe PH in CLD.Methods167 patients with CLD referred for suspected PH who underwent baseline cardiac MRI, pulmonary function tests and right heart catheterisation were identified. In a derivation cohort (n = 67) a bi-logistic regression model was developed to identify severe PH and compared to a previously published multiparameter model (Whitfield model), which is based on interventricular septal angle, ventricular mass index and diastolic pulmonary artery area. The model was evaluated in a test cohort.ResultsThe CLD-PH MRI model [= (−13.104) + (13.059 * VMI)—(0.237 * PA RAC) + (0.083 * Systolic Septal Angle)], had high accuracy in the test cohort (area under the ROC curve (0.91) (p < 0.0001), sensitivity 92.3%, specificity 70.2%, PPV 77.4%, and NPV 89.2%. The Whitfield model also had high accuracy in the test cohort (area under the ROC curve (0.92) (p < 0.0001), sensitivity 80.8%, specificity 87.2%, PPV 87.5%, and NPV 80.4%.ConclusionThe CLD-PH MRI model and Whitfield model have high accuracy to detect severe PH in CLD, and have strong prognostic value.
Providing prognostic information is important when counseling patients and planning treatment strategies in chronic thromboembolic pulmonary hypertension (CTEPH). The aim of this study was to assess the prognostic value of gold standard imaging of cardiac structure and function using cardiac magnetic resonance imaging (CMR) in CTEPH. Consecutive treatment-naive patients with CTEPH who underwent right heart catheterization and CMR between 2011 and 2017 were identified from the ASPIRE (Assessing-the-Specturm-of-Pulmonary-hypertensIon-at-a-REferral-center) registry. CMR metrics were corrected for age and sex where appropriate. Univariate and multivariate regression models were generated to assess the prognostic ability of CMR metrics in CTEPH. Three hundred and seventy-five patients (mean+/-standard deviation: age 64+/-14 years, 49% female) were identified and 181 (48%) had pulmonary endarterectomy (PEA). For all patients with CTEPH, left-ventricular-stroke-volume-index-%predicted (LVSVI%predicted) (p = 0.040), left-atrial-volume-index (LAVI) (p = 0.030), the presence of comorbidities, incremental shuttle walking test distance (ISWD), mixed venous oxygen saturation and undergoing PEA were independent predictors of mortality at multivariate analysis. In patients undergoing PEA, LAVI (p < 0.010), ISWD and comorbidities and in patients not undergoing surgery, right-ventricular-ejection-fraction-%predicted (RVEF%pred) (p = 0.040), age and ISWD were independent predictors of mortality. CMR metrics reflecting cardiac function and left heart disease have prognostic value in CTEPH. In those undergoing PEA, LAVI predicts outcome whereas in patients not undergoing PEA RVEF%pred predicts outcome. This study highlights the prognostic value of imaging cardiac structure and function in CTEPH and the importance of considering left heart disease in patients considered for PEA.
Background Pulmonary hypertension (PH) in patients with chronic lung disease (CLD) predicts reduced functional status, clinical worsening and increased mortality, with patients with severe PH-CLD (≥35 mmHg) having a significantly worse prognosis than mild to moderate PH-CLD (21–34 mmHg). The aim of this cross-sectional study was to assess the association between computed tomography (CT)-derived quantitative pulmonary vessel volume, PH severity and disease aetiology in CLD. Methods Treatment-naïve patients with CLD who underwent CT pulmonary angiography, lung function testing and right heart catheterisation were identified from the ASPIRE registry between October 2012 and July 2018. Quantitative assessments of total pulmonary vessel and small pulmonary vessel volume were performed. Results 90 patients had PH-CLD including 44 associated with COPD/emphysema and 46 with interstitial lung disease (ILD). Patients with severe PH-CLD (n=40) had lower small pulmonary vessel volume compared to patients with mild to moderate PH-CLD (n=50). Patients with PH-ILD had significantly reduced small pulmonary blood vessel volume, compared to PH-COPD/emphysema. Higher mortality was identified in patients with lower small pulmonary vessel volume. Conclusion Patients with severe PH-CLD, regardless of aetiology, have lower small pulmonary vessel volume compared to patients with mild–moderate PH-CLD, and this is associated with a higher mortality. Whether pulmonary vessel changes quantified by CT are a marker of remodelling of the distal pulmonary vasculature requires further study.