Aims:Heart failure (HF) in non-ST-segment elevation acute coronary syndrome (NSTE-ACS) is associated with poor prognosis but often under-recognized. Coronary computed tomography angiography (CCTA), increasingly used in NSTE-ACS, contains cardiopulmonary features not routinely assessed for HF. We evaluated whether an artificial intelligence (AI) algorithm applied to CCTA could identify HF likelihood in NSTE-ACS. Methods and results:In this retrospective external validation study, the AI algorithm was applied without retraining or recalibration to CCTA scans from 1009 patients with NSTE-ACS in the VERDICT trial. Using a pre-specified threshold, patients were classified as low or high AI likelihood of HF. The primary outcome was HF during index hospitalization. The secondary outcome was post-discharge HF hospitalization among patients discharged alive without HF, with analyses adjusted for global registry of acute coronary events score >140 and severe coronary artery disease. Death was treated as a competing risk. Overall, 838 patients (83%) were classified as low AI likelihood and 171 (17%) as high. During index hospitalization, HF was diagnosed in 10 patients (1%) with low AI likelihood and 12 (7%) with high. Sensitivity was 55%, specificity 84%, positive predictive value 7%, and negative predictive value 99%. High AI likelihood was associated with increased risk of index HF (subdistribution hazard ratio, 5.39, 95% confidence interval (CI) 2.32-12.50). After discharge, HF hospitalization occurred in 25 patients (3%) with low AI likelihood and 14 (8%) with high. High AI likelihood remained associated with HF hospitalization (subdistribution hazard ratio 2.56, 95% CI 1.34-4.90). Conclusion:AI-based CCTA analysis identified a large low-risk subgroup and a smaller subgroup at increased HF risk, supporting further evaluation of opportunistic HF assessment from CCTA.
Abstract Objective Acute heart failure (AHF) is a common but underrecognized cause of dyspnea. Chest computed tomography (CT) can accurately assess pulmonary congestion, but radiologist reporting capacity may limit clinical utility. We hypothesized that an artificial intelligence (AI) model could automatically detect imaging signs of AHF and aimed to prospectively validate an AI model in an independent emergency department cohort, benchmarking its performance against radiologists and cardiologists. Materials and methods We prospectively validated a supervised machine-learning model in a single-center study of dyspneic patients undergoing low-dose, non-contrast chest CT and echocardiography. The primary analysis assessed diagnostic performance for CT-detected pulmonary congestion compatible with AHF, using radiologist-reported AHF as the reference and the area under the curve at receiver operating characteristic analysis (AUROC). Secondary analyses compared the AI model with blinded research radiologists and expert cardiologists. Results Of 234 patients (56% males), aged 74 ± 10 years (mean ± standard deviation), 61 (26%) had radiologist-reported AHF. The AI model achieved high diagnostic performance (AUROC 0.95 [95% confidence interval 0.93–0.98]), with 89% sensitivity [78–95] and 89% specificity [83–93]. At prespecified thresholds, rule-out maximized sensitivity (97% [89–100]) at the expense of specificity (74% [67–81]), whereas rule-in yielded high specificity (96% [92–98]) but lower sensitivity (66% [52–77]). In secondary analyses, the AI model achieved a median AUROC of 0.94 (range 0.91–0.96). Conclusion The AI model demonstrated high diagnostic performance for detecting AHF on chest CT in dyspneic patients. Integration into emergency workflows may support more consistent diagnosis, independent of clinician experience or time constraints. Relevance statement AI-based analysis of chest CT may enable earlier and more consistent detection of AHF, supporting timely triage and management, especially when specialist radiological expertise is limited or delayed. Key Points An AI model prospectively detected AHF on chest CT in dyspneic emergency department patients. In a prospective single-center cohort, AI achieved high diagnostic performance (AUROC 0.91–0.96), comparable to that of radiologists and cardiologists. AI-based chest CT interpretation may improve diagnostic consistency in the absence of standardized CT criteria for AHF. Graphical Abstract
INTRODUCTION:Acute appendicitis is a common surgical emergency, but accurate diagnosis remains challenging. Despite the 2020 World Society of Emergency Surgery Jerusalem guidelines recommending structured imaging pathways, including ultrasound and selective computed tomography, clinical assessment remains the primary diagnostic tool in Denmark. Contemporary practice emphasises rapid assessment, with imaging reserved for older or borderline cases. We hypothesised that reliance on clinical assessment alone contributes to a high number of negative appendectomies, particularly in young adults. METHODS:We conducted a single-centre retrospective cohort study of adults (> 18 years) undergoing surgery for suspected appendicitis at a tertiary university hospital in Denmark between January 2021 and December 2023. Data were extracted from electronic medical records and analysed in Stata 19.5. RESULTS:Among 613 patients, 522 had histologically confirmed appendicitis, yielding an overall negative appendectomy rate (NAR) of 14.9%. Patients without preoperative imaging (n = 279) had a NAR of 24.4%, compared with 6.9% among those who underwent preoperative imaging (n = 334). CONCLUSIONS:Reliance on clinical assessment alone results in a substantial number of unnecessary operations. Preoperative imaging significantly reduces NAR (p less-than 0.001), supporting broader adoption of guideline-based diagnostic strategies to improve diagnostic accuracy and optimise resource utilisation. FUNDING:None. TRIAL REGISTRATION:Not relevant.
Introduction: Chest CT scans are increasingly used in dyspneic patients where acute heart failure (AHF) is a key differential diagnosis. Interpretation remains challenging and radiology reports are frequently delayed due to a radiologist shortage, although flagging such information for emergency physicians would have therapeutic implication. Artificial intelligence (AI) can be a complementary tool to enhance the diagnostic precision. We aim to develop an explainable AI model to detect radiological signs of AHF in chest CT with an accuracy comparable to thoracic radiologists. Methods: A single-center, retrospective study during 2016-2021 at Copenhagen University Hospital - Bispebjerg and Frederiksberg, Denmark. A Boosted Trees model was trained to predict AHF based on measurements of segmented cardiac and pulmonary structures from acute thoracic CT scans. Diagnostic labels for training and testing were extracted from radiology reports. Structures were segmented with TotalSegmentator. Shapley Additive explanations values were used to explain the impact of each measurement on the final prediction. Results: Of the 4,672 subjects, 49 Conclusion: We developed an explainable AI model with strong discriminatory performance, comparable to thoracic radiologists. The AI model's stepwise, transparent predictions may support decision-making.
BACKGROUND:Dyspnea is a common cause of hospitalization, posing diagnostic challenges among older adult patients with multimorbid conditions. Chest computed tomography (CT) scans are increasingly used in patients with dyspnea and offer superior diagnostic accuracy over chest radiographs but face limited use due to a shortage of radiologists. OBJECTIVE:This study aims to develop and validate artificial intelligence (AI) algorithms to enable automatic analysis of acute CT scans and provide immediate feedback on the likelihood of pneumonia, pulmonary embolism, and cardiac decompensation. This protocol will focus on cardiac decompensation. METHODS:We designed a retrospective method development and validation study. This study has been approved by the Danish National Committee on Health Research Ethics (1575037). We extracted 4672 acute chest CT scans with corresponding radiological reports from the Copenhagen University Hospital-Bispebjerg and Frederiksberg, Denmark, from 2016 to 2021. The scans will be randomly split into training (2/3) and internal validation (1/3) sets. Development of the AI algorithm involves parameter tuning and feature selection using cross validation. Internal validation uses radiological reports as the ground truth, with algorithm-specific thresholds based on true positive and negative rates of 90% or greater for heart and lung diseases. The AI models will be validated in low-dose chest CT scans from consecutive patients admitted with acute dyspnea and in coronary CT angiography scans from patients with acute coronary syndrome. RESULTS:As of August 2025, CT data extraction has been completed. Algorithm development, including image segmentation and natural language processing, is ongoing. However, for pulmonary congestion, the algorithm development has been completed. Internal and external validation are planned, with overall validation expected to conclude in 2025 and the final results to be available in 2026. CONCLUSIONS:The results are expected to enhance clinical decision-making by providing immediate, AI-driven insights from CT scans, which will be beneficial for both clinicians and patients. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID):DERR1-10.2196/77030.
Mathematical morphology (MM) is an indispensable tool for post-processing. Several extensions of MM to categorical images, such as multi-class segmentations, have been proposed. However, none provide satisfactory definitions for morphology on probabilistic representations of categorical images. The categorical distribution is a natural choice for representing uncertainty about categorical images. Extending MM to categorical distributions is problematic because categories are inherently unordered. Without ranking categories, we cannot use the standard framework based on supremum and infimum. Ranking categories is impractical and problematic. Instead, we consider the probabilistic representation and operations that emphasize a single category. In this work, we review and compare previous approaches. We propose two approaches for morphology on categorical distributions: operating on Dirichlet distributions over the parameters of the distributions and operating directly on the distributions. We propose a "protected" variant of the latter and demonstrate the proposed approaches by fixing misclassifications and modeling annotator bias.
Screening for left ventricular systolic dysfunction (LVSD), defined as reduced left ventricular ejection fraction (LVEF), deserves renewed interest as the medical treatment for the prevention and progression of heart failure improves. We aimed to review the updated literature to outline the potential and caveats of using artificial intelligence-enabled electrocardiography (AIeECG) as an opportunistic screening tool for LVSD. We searched PubMed and Cochrane for variations of the terms "ECG," "Heart Failure," "systolic dysfunction," and "Artificial Intelligence" from January 2010 to April 2022 and selected studies that reported the diagnostic accuracy and confounders of using AIeECG to detect LVSD. Out of 40 articles, we identified 15 relevant studies; eleven retrospective cohorts, three prospective cohorts, and one case series. Although various LVEF thresholds were used, AIeECG detected LVSD with a median AUC of 0.90 (IQR from 0.85 to 0.95), a sensitivity of 83.3% (IQR from 73 to 86.9%) and a specificity of 87% (IQR from 84.5 to 90.9%). AIeECG algorithms succeeded across a wide range of sex, age, and comorbidity and seemed especially useful in non-cardiology settings and when combined with natriuretic peptide testing. Furthermore, a false-positive AIeECG indicated a future development of LVSD. No studies investigated the effect on treatment or patient outcomes. This systematic review corroborates the arrival of a new generic biomarker, AIeECG, to improve the detection of LVSD. AIeECG, in addition to natriuretic peptides and echocardiograms, will improve screening for LVSD, but prospective randomized implementation trials with added therapy are needed to show cost-effectiveness and clinical significance.
We report on the results of a small crowdsourcing experiment conducted at a workshop on machine learning for segmentation held at the Danish Bio Imaging network meeting 2020. During the workshop we asked participants to manually segment mitochondria in three 2D patches. The aim of the experiment was to illustrate that manual annotations should not be seen as the ground truth, but as a reference standard that is subject to substantial variation. In this note we show how the large variation we observed in the segmentations can be reduced by removing the annotators with worst pair-wise agreement. Having removed the annotators with worst performance, we illustrate that the remaining variance is semantically meaningful and can be exploited to obtain segmentations of cell boundary and cell interior.
Rapid advances in image processing capabilities have been seen across many domains, fostered by the application of machine learning algorithms to "big-data". However, within the realm of medical image analysis, advances have been curtailed, in part, due to the limited availability of large-scale, well-annotated datasets. One of the main reasons for this is the high cost often associated with producing large amounts of high-quality meta-data. Recently, there has been growing interest in the application of crowdsourcing for this purpose; a technique that has proven effective for creating large-scale datasets across a range of disciplines, from computer vision to astrophysics. Despite the growing popularity of this approach, there has not yet been a comprehensive literature review to provide guidance to researchers considering using crowdsourcing methodologies in their own medical imaging analysis. In this survey, we review studies applying crowdsourcing to the analysis of medical images, published prior to July 2018. We identify common approaches, challenges and considerations, providing guidance of utility to researchers adopting this approach. Finally, we discuss future opportunities for development within this emerging domain.
The recently introduced locally orderless tensor network (LoTeNet) for supervised image classification uses matrix product state (MPS) operations on grids of transformed image patches. The resulting patch representations are combined back together into the image space and aggregated hierarchically using multiple MPS blocks per layer to obtain the final decision rules. In this work, we propose a non-patch based modification to LoTeNet that performs one MPS operation per layer, instead of several patch-level operations. The spatial information in the input images to MPS blocks at each layer is squeezed into the feature dimension, similar to LoTeNet, to maximise retained spatial correlation between pixels when images are flattened into 1D vectors. The proposed multi-layered tensor network (MLTN) is capable of learning linear decision boundaries in high dimensional spaces in a multi-layered setting, which results in a reduction in the computation cost compared to LoTeNet without any degradation in performance.
Tensor networks are factorisations of high rank tensors into networks of lower rank tensors and have primarily been used to analyse quantum many-body problems. Tensor networks have seen a recent surge of interest in relation to supervised learning tasks with a focus on image classification. In this work, we improve upon the matrix product state (MPS) tensor networks that can operate on one-dimensional vectors to be useful for working with 2D and 3D medical images. We treat small image regions as orderless, squeeze their spatial information into feature dimensions and then perform MPS operations on these locally orderless regions. These local representations are then aggregated in a hierarchical manner to retain global structure. The proposed locally orderless tensor network (LoTeNet) is compared with relevant methods on three datasets. The architecture of LoTeNet is fixed in all experiments and we show it requires lesser computational resources to attain performance on par or superior to the compared methods.
Volumetric imaging is an essential diagnostic tool for medical practitioners. The use of popular techniques such as convolutional neural networks (CNN) for analysis of volumetric images is constrained by the availability of detailed (with local annotations) training data and GPU memory. In this paper, the volumetric image classification problem is posed as a multi-instance classification problem and a novel method is proposed to adaptively select positive instances from positive bags during the training phase. This method uses the extreme value theory to model the feature distribution of the images without a pathology and use it to identify positive instances of an imaged pathology. The experimental results, on three separate image classification tasks (i.e. classify retinal OCT images according to the presence or absence of fluid build-ups, emphysema detection in pulmonary 3D-CT images and detection of cancerous regions in 2D histopathology images) show that the proposed method produces classifiers that have similar performance to fully supervised methods and achieves the state of the art performance in all examined test cases.
Accurate assessment of pulmonary emphysema is crucial to assess disease severity and subtype, to monitor disease progression, and to predict lung cancer risk. However, visual assessment is time-consuming and subject to substantial inter-rater variability while standard densitometry approaches to quantify emphysema remain inferior to visual scoring. We explore if machine learning methods that learn from a large dataset of visually assessed CT scans can provide accurate estimates of emphysema extent and if methods that learn from emphysema extent scoring can outperform algorithms that learn only from emphysema presence scoring. Four Multiple Instance Learning classifiers, trained on emphysema presence labels, and five Learning with Label Proportions classifiers, trained on emphysema extent labels, are compared. Performance is evaluated on 600 low-dose CT scans from the Danish Lung Cancer Screening Trial and we find that learning from emphysema presence labels, which are much easier to obtain, gives equally good performance to learning from emphysema extent labels. The best performing Multiple Instance Learning and Learning with Label Proportions classifiers, achieve intra-class correlation coefficients around 0.90 and average overall agreement with raters of 78% and 79% compared to an inter-rater agreement of 83%.
Suppose one is faced with the challenge of tissue segmentation in MR images, without annotators at their center to provide labeled training data. One option is to go to another medical center for a trained classifier. Sadly, tissue classifiers do not generalize well across centers due to voxel intensity shifts caused by center-specific acquisition protocols. However, certain aspects of segmentations, such as spatial smoothness, remain relatively consistent and can be learned separately. Here we present a smoothness prior that is fit to segmentations produced at another medical center. This informative prior is presented to an unsupervised Bayesian model. The model clusters the voxel intensities, such that it produces segmentations that are similarly smooth to those of the other medical center. In addition, the unsupervised Bayesian model is extended to a semi-supervised variant, which needs no visual interpretation of clusters into tissues.
Emphysema is part of chronic obstructive pulmonary disease, a leading cause of mortality worldwide. Visual assessment of emphysema presence is useful for identifying subjects at risk and for research into disease development. We train a machine learning method to predict emphysema from visually assessed expert labels. We use a multiple instance learning approach to predict both scan-level and region-level emphysema presence. We evaluate performance on 600 low-dose CT scans from the Danish Lung Cancer Screening Study and achieve an AUC of 0.82 for scan-level prediction and AUCs between 0.76 and 0.88 for region-level prediction.
Supervised feature learning using convolutional neural networks (CNNs) can provide concise and disease relevant representations of medical images. However, training CNNs requires annotated image data. Annotating medical images can be a time-consuming task and even expert annotations are subject to substantial inter- and intra-rater variability. Assessing visual similarity of images instead of indicating specific pathologies or estimating disease severity could allow non-experts to participate, help uncover new patterns, and possibly reduce rater variability. We consider the task of assessing emphysema extent in chest CT scans. We derive visual similarity triplets from visually assessed emphysema extent and learn a low dimensional embedding using CNNs. We evaluate the networks on 973 images, and show that the CNNs can learn disease relevant feature representations from derived similarity triplets. To our knowledge this is the first medical image application where similarity triplets has been used to learn a feature representation that can be used for embedding unseen test images.
We propose an end-to-end deep learning method that learns to estimate emphysema extent from proportions of the diseased tissue. These proportions were visually estimated by experts using a standard grading system, in which grades correspond to intervals (label example: 1-5% of diseased tissue). The proposed architecture encodes the knowledge that the labels represent a volumetric proportion. A custom loss is designed to learn with intervals. Thus, during training, our network learns to segment the diseased tissue such that its proportions fit the ground truth intervals. Our architecture and loss combined improve the performance substantially (8% ICC) compared to a more conventional regression network. We outperform traditional lung densitometry and two recently published methods for emphysema quantification by a large margin (at least 7% AUC and 15% ICC), and achieve near-human-level performance. Moreover, our method generates emphysema segmentations that predict the spatial distribution of emphysema at human level.
Classification of emphysema patterns is believed to be useful for improved diagnosis and prognosis of chronic obstructive pulmonary disease. Emphysema patterns can be assessed visually on lung CT scans. Visual assessment is a complex and time-consuming task performed by experts, making it unsuitable for obtaining large amounts of labeled data. We investigate if visual assessment of emphysema can be framed as an image similarity task that does not require expert. Substituting untrained annotators for experts makes it possible to label data sets much faster and at a lower cost. We use crowd annotators to gather similarity triplets and use t-distributed stochastic triplet embedding to learn an embedding. The quality of the embedding is evaluated by predicting expert assessed emphysema patterns. We find that although performance varies due to low quality triplets and randomness in the embedding, we still achieve a median $$F_1$$ score of 0.58 for prediction of four patterns.
. Quantification of emphysema extent is important in diagnos-ing and monitoring patients with chronic obstructive pulmonary disease (COPD). Several studies have shown that emphysema quantification by supervised texture classification is more robust and accurate than tradi-tional densitometry. Current techniques require highly time consuming manual annotations of patches or use only weak labels indicating over-all disease status (e.g, COPD or healthy). We show how visual scoring of regional emphysema extent can be exploited in a learning with label proportions (LLP) framework to both predict presence of emphysema in smaller patches and estimate regional extent. We evaluate performance on 195 visually scored CT scans and achieve an intraclass correlation of 0.72 (0.65–0.78) between predicted region extent and expert raters. To our knowledge this is the first time that LLP methods have been applied to medical imaging data.
Tractography is the standard tool for automatic delineation of white matter tracts from diffusion weighted images. However, the output of tractography often requires post-processing to remove false positives and ensure a robust delineation of the studied tract, and this demands expert prior knowledge. Here we demonstrate how such prior knowledge, or indeed any prior spatial information, can be automatically incorporated into a shortest-path tractography approach to produce more robust results. We describe how such a prior can be automatically generated (learned) from a population, and we demonstrate that our framework also retains support for conventional interactive constraints such as waypoint regions. We apply our approach to the open access, high quality Human Connectome Project data, as well as a dataset acquired on a typical clinical scanner. Our results show that the use of a learned prior substantially increases the overlap of tractography output with a reference atlas on both populations, and this is confirmed by visual inspection. Furthermore, we demonstrate how a prior learned on the high quality dataset significantly increases the overlap with the reference for the more typical yet lower quality data acquired on a clinical scanner. We hope that such automatic incorporation of prior knowledge and the obviation of expert interactive tract delineation on every subject, will improve the feasibility of large clinical tractography studies.
Erik Dam合作论文数Nordic Bioscience ;Imaging Department ;Herlev Hovedgade 2071