The assessment of lumbar central canal stenosis (LCCS) is crucial for diagnosing and planning treatment for patients with low back pain and neurogenic pain. However, manual assessment methods are time-consuming, variable, and require axial MRIs. The aim of this study is to develop and validate an AI-based model that automatically classifies LCCS using sagittal T2-weighted MRIs. A pre-existing 3D AI algorithm was utilized to segment the spinal canal and intervertebral discs (IVDs), enabling quantitative measurements at each IVD level. Four musculoskeletal radiologists graded 683 IVD levels from 186 LCCS patients using the 4-class Lee grading system. A second consensus reading was conducted by readers 1 and 2, which, along with automatic measurements, formed the training dataset for a multiclass (grade 0–3) and binary (grade 0–1 vs. 2–3) random forest classifier with tenfold cross-validation. The multiclass model achieved a Cohen’s weighted kappa of 0.86 (95
Deep convolutional neural networks are widely used in medical image segmentation but require many labeled images for training. Annotating three-dimensional medical images is a time-consuming and costly process. To overcome this limitation, we propose a novel semi-supervised segmentation method that leverages mostly unlabeled images and a small set of labeled images in training. Our approach involves assessing prediction uncertainty to identify reliable predictions on unlabeled voxels from the teacher model. These voxels serve as pseudo-labels for training the student model. In voxels where the teacher model produces unreliable predictions, pseudo-labeling is carried out based on voxel-wise embedding correspondence using reference voxels from labeled images. We applied this method to automate hip bone segmentation in CT images, achieving notable results with just 4 CT scans. The proposed approach yielded a Hausdorff distance with 95th percentile (HD95) of 3.30 and IoU of 0.929, surpassing existing methods achieving HD95 (4.07) and IoU (0.927) at their best.
AI-assisted techniques for lesion registration and segmentation have the potential to make CT-based tumor follow-up assessment faster and less reader-dependent. However, empirical evidence on the advantages of AI-assisted volumetric segmentation for lymph node and soft tissue metastases in follow-up CT scans is lacking. The aim of this study was to assess the efficiency, quality, and inter-reader variability of an AI-assisted workflow for volumetric segmentation of lymph node and soft tissue metastases in follow-up CT scans. Three hypotheses were tested: (H1) Assessment time for follow-up lesion segmentation is reduced using an AI-assisted workflow. (H2) The quality of the AI-assisted segmentation is non-inferior to the quality of fully manual segmentation. (H3) The inter-reader variability of the resulting segmentations is reduced with AI assistance. The study retrospectively analyzed 126 lymph nodes and 135 soft tissue metastases from 55 patients with stage IV melanoma. Three radiologists from two institutions performed both AI-assisted and manual segmentation, and the results were statistically analyzed and compared to a manual segmentation reference standard. AI-assisted segmentation reduced user interaction time significantly by 33
Objectives Severity of degenerative scoliosis (DS) is assessed by measuring the Cobb angle on anteroposterior radiographs. However, MRI images are often available to study the degenerative spine. This retrospective study aims to develop and evaluate the reliability of a novel automatic method that measures coronal Cobb angles on lumbar MRI in DS patients. Materials and methods Vertebrae and intervertebral discs were automatically segmented using a 3D AI algorithm, trained on 447 lumbar MRI series. The segmentations were used to calculate all possible angles between the vertebral endplates, with the largest being the Cobb angle. The results were validated with 50 high-resolution sagittal lumbar MRI scans of DS patients, in which three experienced readers measured the Cobb angle. Reliability was determined using the intraclass correlation coefficient (ICC). Results The ICCs between the readers ranged from 0.90 (95% CI 0.83–0.94) to 0.93 (95% CI 0.88–0.96). The ICC between the maximum angle found by the algorithm and the average manually measured Cobb angles was 0.83 (95% CI 0.71–0.90). In 9 out of the 50 cases (18%), all readers agreed on both vertebral levels for Cobb angle measurement. When using the algorithm to extract the angles at the vertebral levels chosen by the readers, the ICCs ranged from 0.92 (95% CI 0.87–0.96) to 0.97 (95% CI 0.94–0.98). Conclusion The Cobb angle can be accurately measured on MRI using the newly developed algorithm in patients with DS. The readers failed to consistently choose the same vertebral level for Cobb angle measurement, whereas the automatic approach ensures the maximum angle is consistently measured. Clinical relevance statement Our AI-based algorithm offers reliable Cobb angle measurement on routine MRI for degenerative scoliosis patients, potentially reducing the reliance on conventional radiographs, ensuring consistent assessments, and therefore improving patient care. Key Points • While often available, MRI images are rarely utilized to determine the severity of degenerative scoliosis. • The presented MRI Cobb angle algorithm is more reliable than humans in patients with degenerative scoliosis. • Radiographic imaging for Cobb angle measurements is mitigated when lumbar MRI images are available. Graphical Abstract
In this study, we introduce a deep learning approach for segmenting kidney parenchyma and kidney abnormalities to support clinicians in identifying and quantifying renal abnormalities such as cysts, lesions, masses, metastases, and primary tumors. Our end-to-end segmentation method was trained on 215 contrast-enhanced thoracic-abdominal CT scans, with half of these scans containing one or more abnormalities. We began by implementing our own version of the original 3D U-Net network and incorporated four additional components: an end-to-end multi-resolution approach, a set of task-specific data augmentations, a modified loss function using top-$k$, and spatial dropout. Furthermore, we devised a tailored post-processing strategy. Ablation studies demonstrated that each of the four modifications enhanced kidney abnormality segmentation performance, while three out of four improved kidney parenchyma segmentation. Subsequently, we trained the nnUNet framework on our dataset. By ensembling the optimized 3D U-Net and the nnUNet with our specialized post-processing, we achieved marginally superior results. Our best-performing model attained Dice scores of 0.965 and 0.947 for segmenting kidney parenchyma in two test sets (20 scans without abnormalities and 30 with abnormalities), outperforming an independent human observer who scored 0.944 and 0.925, respectively. In segmenting kidney abnormalities within the 30 test scans containing them, the top-performing method achieved a Dice score of 0.585, while an independent second human observer reached a score of 0.664, suggesting potential for further improvement in computerized methods. All training data is available to the research community under a CC-BY 4.0 license on https://doi.org/10.5281/zenodo.8014289
PurposeLow back pain (LBP) is one of the most prevalent health condition worldwide and responsible for the most years lived with disability, yet the etiology is often unknown. Magnetic resonance imaging (MRI) is frequently used for treatment decision even though it is often inconclusive. There are many different image features that could relate to low back pain. Conversely, multiple etiologies do relate to spinal degeneration but do not actually cause the perceived pain. This narrative review provides an overview of all possible relevant features visible on MRI images and determines their relation to LBP.MethodsWe conducted a separate literature search per image feature. All included studies were scored using the GRADE guidelines. Based on the reported results per feature an evidence agreement (EA) score was provided, enabling us to compare the collected evidence of separate image features. The various relations between MRI features and their associated pain mechanisms were evaluated to provide a list of features that are related to LBP.ResultsAll searches combined generated a total of 4472 hits of which 31 articles were included. Features were divided into five different categories:'discogenic', 'neuropathic','osseous', 'facetogenic', and'paraspinal', and discussed separately.ConclusionOur research suggests that type I Modic changes, disc degeneration, endplate defects, disc herniation, spinal canal stenosis, nerve compression, and muscle fat infiltration have the highest probability to be related to LBP. These can be used to improve clinical decision-making for patients with LBP based on MRI.
Bone ranks as the third most frequent tissue affected by cancer metastases, following the lung and liver. Bone metastases are often painful and may result in pathological fracture, which is a major cause of morbidity and mortality in cancer patients. To quantify fracture risk, finite element (FE) analysis has shown to be a promising tool, but metastatic lesions are typically not specifically segmented and therefore their mechanical properties may not be represented adequately. Deep learning methods potentially provide the opportunity to automatically segment these lesions and change the mechanical properties more adequately. In this study, our primary focus was to gain insight into the performance of an automatic segmentation algorithm for femoral metastatic lesions using deep learning methods and the subsequent effects on FE outcomes. The aims were to determine the similarity between manual segmentation and automatic segmentation; the differences in predicted failure load between FE models with automatically segmented osteolytic and mixed lesions and the models with CT-based lesion values (the gold standard); and the effect on the BOne Strength (BOS) score (failure load adjusted for body weight) and subsequent fracture risk assessments. From two patient cohorts, a total number of 50 femurs with osteolytic and mixed metastatic lesions were included in this study. The femurs were segmented from CT images and transferred into FE meshes. The material behavior was implemented as non-linear isotropic. These FE models were considered as gold standard (Finite Element no Segmented Lesion: FE-no-SL), whereby the local calcium equivalent density of both femur and metastatic lesion was extracted from CT-values. Lesions in the femur were manually segmented by two biomechanical experts after which final lesion segmentation for each femur was obtained based on consensus of opinions between two observers. Subsequently, a self-configuring variant of the popular deep learning model U-Net known as nnU-Net was used to automatically segment metastatic lesions within the femur. For these models with segmented lesions (Finite Element with Segmented Lesion: FE-with-SL), the calcium equivalent density within the metastatic lesions was set to zero after being segmented by the neural network, simulating absence of load-bearing capacity of these lesions. The models (either with or without automatically segmented lesions) were loaded incrementally in axial direction until failure was simulated. Dice coefficient was used to evaluate the similarity of the manual and automatic segmentation. Mean calcium equivalent density values within the automatically segmented lesions were calculated. Failure loads and patterns were determined. Furthermore, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were calculated for both groups by comparing the predictions to the occurrence or absence of actual fracture within the patient cohorts. The automatic segmentation algorithm performed in a none-robust manner. Dice coefficients describing the similarity between consented manual and automatic segmentations were relatively low (mean 0.45 ± standard deviation 0.33, median 0.54). Failure load difference between the FE-no-SL and FE-with-SL groups varied from 0 % to 48 % (mean 6.6 %). Correlation analysis of failure loads between the two groups showed a strong relationship (R2 > 0.9). From the 50 cases, four cases showed clear deviations for which models with automatic lesion segmentation (FE-with-SL) showed considerably lower failure loads. In the whole database including osteolytic and mixed lesions, sensitivity and NPV remained the same, but specificity and PPV decreased from 94 % to 83 %, and from 78 % to 54 % respectively from FE-no-SL to FE-with-SL. This study indicates that the nnU-Net yielded none-robust outcomes in femoral lesion segmentation and that other segmentation algorithms should be considered. However, the difference in failure pattern and failure load between FE models with automatically segmented osteolytic and mixed lesions were relatively small in most cases with a few exceptions. On the other hand, the accuracy of fracture risk assessment using the BOS score was lower compared to the FE-no-SL. In conclusion, this study showed that automatic lesion segmentation is a none-solved issue and therefore, quantifying lesion characteristics and the subsequent effect on the fracture risk using deep learning will remain challenging.
Image registration is a fundamental medical image analysis task, and a wide variety of approaches have been proposed. However, only a few studies have comprehensively compared medical image registration approaches on a wide range of clinically relevant tasks. This limits the development of registration methods, the adoption of research advances into practice, and a fair benchmark across competing approaches. The Learn2Reg challenge addresses these limitations by providing a multi-task medical image registration data set for comprehensive characterisation of deformable registration algorithms. A continuous evaluation will be possible at https://learn2reg.grand-challenge.org. Learn2Reg covers a wide range of anatomies (brain, abdomen, and thorax), modalities (ultrasound, CT, MR), availability of annotations, as well as intra- and inter-patient registration evaluation. We established an easily accessible framework for training and validation of 3D registration methods, which enabled the compilation of results of over 65 individual method submissions from more than 20 unique teams. We used a complementary set of metrics, including robustness, accuracy, plausibility, and runtime, enabling unique insight into the current state-of-the-art of medical image registration. This paper describes datasets, tasks, evaluation methods and results of the challenge, as well as results of further analysis of transferability to new datasets, the importance of label supervision, and resulting bias. While no single approach worked best across all tasks, many methodological aspects could be identified that push the performance of medical image registration to new state-of-the-art performance. Furthermore, we demystified the common belief that conventional registration methods have to be much slower than deep-learning-based methods.
Transfer learning leverages pre-trained model features from a large dataset to save time and resources when training new models for various tasks, potentially enhancing performance. Due to the lack of large datasets in the medical imaging domain, transfer learning from one medical imaging model to other medical imaging models has not been widely explored. This study explores the use of transfer learning to improve the performance of deep convolutional neural networks for organ segmentation in medical imaging. A base segmentation model (3D U-Net) was trained on a large and sparsely annotated dataset; its weights were used for transfer learning on four new down-stream segmentation tasks for which a fully annotated dataset was available. We analyzed the training set size's influence to simulate scarce data. The results showed that transfer learning from the base model was beneficial when small datasets were available, providing significant performance improvements; where fine-tuning the base model is more beneficial than updating all the network weights with vanilla transfer learning. Transfer learning with fine-tuning increased the performance by up to 0.129 (+28\%) Dice score than experiments trained from scratch, and on average 23 experiments increased the performance by 0.029 Dice score in the new segmentation tasks. The study also showed that cross-modality transfer learning using CT scans was beneficial. The findings of this study demonstrate the potential of transfer learning to improve the efficiency of annotation and increase the accessibility of accurate organ segmentation in medical imaging, ultimately leading to improved patient care. We made the network definition and weights publicly available to benefit other users and researchers.
Segmentation of vertebrae and intervertebral discs (IVD) in MR images are important steps for automatic image analysis. This paper proposes an extension of an iterative vertebra segmentation method that relies on a 3D fully-convolutional neural network to segment the vertebrae one-by-one. We augment this approach with an additional segmentation step following each vertebra detection to also segment the IVD below each vertebra. To train and test the algorithm, we collected and annotated T2-weighted sagittal lumbar spine MR scans of 53 patients. The presented approach achieved a mean Dice score of 93 % ± 2 % for vertebra segmentation and 86 % ± 7 % for IVD segmentation. The method was able to cope with pathological abnormalities such as compression fractures, Schmorl’s nodes and collapsed IVDs. In comparison, a similar network trained for IVD segmentation without knowledge of the adjacent vertebra segmentation result did not detect all IVDs (89 %) and also achieved a lower Dice score of 83 % ± 9 %. These results indicate that combining IVD segmentation with vertebra segmentation in lumbar spine MR images can help to improve the detection and segmentation performance compared with separately segmenting these structures.
HomeRadiology: Artificial IntelligenceVol. 4, No. 2 PreviousNext CommentaryFree AccessAutomatic Brand Identification of Orthopedic Implants from Radiographs: Ready for the Next Step?Merel Huisman, Nikolas LessmannMerel Huisman, Nikolas LessmannAuthor AffiliationsFrom the Department of Radiology, University Medical Center Utrecht, Heidelberglaan 100, Utrecht 3508, the Netherlands (M.H.); and Department of Radiology and Nuclear Medicine, Radboud University Medical Center, Nijmegen, the Netherlands (N.L.).Address correspondence to M.H.Merel HuismanNikolas LessmannPublished Online:Mar 2 2022https://doi.org/10.1148/ryai.220008MoreSectionsPDF ToolsImage ViewerAdd to favoritesCiteTrack CitationsPermissionsReprints ShareShare onFacebookTwitterLinked In See article by Dutt and Mendonca et al in this issue.Merel Huisman, MD, PhD, is a radiologist with subspecialty interest in cardiothoracic radiology and musculoskeletal radiology. As an epidemiologist by training, her passion is the intersection between data science and clinical epidemiology. She is an EuSoMII board member and working group member of several national and international initiatives concerning standardization of artificial intelligence in health care. In 2021, she won an AuntMinnie Europe Award in the Rising Star category for her contribution to the field.Download as PowerPointOpen in Image Viewer Nikolas Lessmann, PhD, is an assistant professor at Radboud University Medical Center. He studied biomedical engineering at the University of Lübeck and obtained his doctorate from Utrecht University. In 2019, he joined the Diagnostic Image Analysis group, where his research interests focus on applications of machine learning and artificial intelligence in musculoskeletal image analysis.Download as PowerPointOpen in Image Viewer In an aging population, the need for revision surgery of orthopedic implants will become more prevalent. To adequately perform revision surgery, orthopedic surgeons must know the brand and model of the hardware in situ, for the simple reason that the appropriate instruments for removal have to be present in the operating room. If the available tools are incompatible with the hardware, delays and suboptimal procedures are the consequence, leading to higher costs, potentially higher complication rates, and considerable staff inconvenience. Usually, the surgeon relies on the original surgical notes to determine brand and model of the implanted hardware. A problem arises when the patient underwent surgery years ago, possibly in another country, and information on the implemented hardware is unavailable. In such cases, the surgeon has to rely on their experience or network to identify the implant preoperatively, which is often frustrating, time-consuming, and error prone. Radiologists do not typically recognize and/or report the brand and model of the hardware present on a radiograph. Even though this would be theoretically possible, it could be argued the advantages do not outweigh the disadvantages, as only in selected cases this would serve the orthopedic surgeon.Cervical spine hardware differs from hip and knee implants when it comes to the challenge of recognizing the brand and model in radiographs. Hip and knee implant registries have existed for many years now, and relatively few brands of hardware are used (1). Unlike prosthetic implants of the large joints, which can remain in situ for 15–20 years, spinal fixation hardware has a shorter life span; hence, developments are quicker and result in a wider variety of devices used. Therefore, a service that uses a diagnostic prediction model based on deep learning to quickly and reliably identify the implanted spinal hardware from radiographs in selected cases would be beneficial to patients and orthopedic surgeons alike.In this issue of Radiology: Artificial Intelligence, Dutt and Mendonca et al (2) pave the way for such a solution with a first step: the very successful proof of principle of a complete, replicable, and expandable pipeline for deep learning–based automatic detection and brand identification of cervical spinal hardware.In this single-center retrospective study (n = 984), Dutt and Mendonca et al describe and evaluate an artificial intelligence system that recognizes 10 of the most common cervical spine implants on anteroposterior or lateral view radiographs. Their system achieves an overall F1 score of 95% (95% CI: 92%, 98%) for both anterior and posterior hardware, indicating excellent model accuracy. An F1 score is a commonly used and robust accuracy measure and represents a combination of positive predictive value (precision) and sensitivity (recall) that is reliable regardless of the prevalence of the outcome. The corresponding area under the receiver operating characteristic curve is approaching 1, indicating near-perfect discrimination.The deep learning approach taken in this study does not offer many surprises at first sight. The authors relied on two popular neural network architectures, Efficient-Det and DenseNet, and assembled them into a straight-forward pipeline. However, with the excellent results that this approach achieves, the study adds to a growing body of evidence that orthopedic hardware identification might not be all too big of a challenge for modern neural networks (3). Several recent studies demonstrated that similarly standard approaches yield strong results for other types of hardware as well, such as knee (4) and hip (5,6) implants. In a study by Patel et al (7), such an approach was even more accurate in identifying implants than experienced orthopedic surgeons. These studies paint the picture of a technology mature enough to move beyond mere proof-of-principle studies.A current limitation that still needs to be addressed in general is the low number of implants that these systems can distinguish. Dutt and Mendonca et al built a system that can identify 10 cervical spine hardware models; other publications describe systems for other types of hardware that are equally restricted. Clearly, the practical value of automatic implant identification would be much greater if a single system could distinguish a broad range of implants, including various types of implants as well as older and rarer models that are especially problematic when planning revision surgery. Therefore, a pivotal question for further research will be how to efficiently scale these systems to a larger number of implants.For scaling up artificial intelligence models to larger datasets and more classes (ie, implant models), the approach of Dutt and Mendonca et al offers an attractive alternative to simply training a single prediction model with potentially hundreds of classes, corresponding to all kinds of implants. Their proposal is to divide the task into two serial tasks: a neural network specialized in detection (EfficientDet) first draws boxes around implants visible on the radiograph. It also labels the detected implants according to their general type, in this case, either anterior or posterior cervical spine hardware. Depending on which type of implant this gatekeeper network has identified, another neural network (DenseNet) specialized in that particular hardware type identifies the specific brand and model of the implant. With this strategy, the implant identification system becomes a collection of neural networks for different types of implants, and the addition of another hardware type becomes a matter of adding another network to this collection.Training these networks requires a set of images for which not only the type and model of the implant are known, but also in which the locations of the implants are annotated. These are used to train the gatekeeper network to detect implants and to confine the implant identification networks to relevant parts of the image. Rather than manually drawing boxes around the implants in the entire dataset, the article proposes a weakly supervised approach. Weak supervision is a form of supervised machine learning where the data are not hand annotated by an expert, but where less reliable and therefore weaker annotations are used instead. These can be, for instance, annotations crowdsourced from laypeople or annotations obtained with another machine learning model. The approach of Dutt and Mendonca et al requires annotation of a limited number of images, uses these to train the gatekeeper network that detects implants in the image, and uses that network to annotate the implant locations in the rest of the images. The study also gives an indication of how weak these annotations actually are: approximately 10% of the images were annotated, and in another 1900 randomly selected images, the automatically annotated implant location was reviewed, revealing mistakes in only 53 of 1900 images (3%).The study unfortunately does not give an indication of the minimum number of images that need to be hand annotated. This question stands at the beginning of most machine learning projects and ties in with the question of the minimum level of performance that the machine learning model needs to reach to become useful, be it for use in clinical practice or for weakly supervised learning. Finding reasonable estimates of these numbers will be important for making such approaches more attractive, making it easier for scientists to adopt the best approach.Another practical matter is collecting a large training set, consisting of radiographs of patients who received an implant of which brand and model can be determined reliably. Although Dutt and Mendonca et al manually reviewed surgical notes for almost 1000 patients, there is also promising research into automatic text analysis of surgical notes (8). Potentially, such methods could become a second source of weak labels and might facilitate assembly of much larger training sets.As with all single-center studies, external validation needs to be done on data that differ from the source population sampled in this study. We applaud the authors for including baseline characteristics of the patient population, including self-reported race and body mass index, which is useful for other institutions to make an estimation of how the results might translate to their local patient population. The authors provide free access to the final hardware localization and brand classification model at GitHub (a commonly used code repository for collaboration and version control) for other institutions to use and expand on.Dutt and Mendonca et al show impressive model performance in their cohort, indicating modern neural networks might have added value in clinical practice by reliably identifying medical devices. In a few years from now, one could think of a pay-per-use model for fast identification of any depicted medical device, perhaps in the form of a secured website or even a mobile app.Disclosures of Conflicts of Interest: M.H. Radiology: Artificial Intelligence trainee editorial board member. N.L. No relevant relationships.AcknowledgmentWe would like to thank J.J. Verlaan, MD, PhD (spine surgeon, University Medical Center Utrecht, Utrecht, the Netherlands) for his contributions.Authors declared no funding for this work.References1. Malchau H, Garellick G, Berry D, et al. Arthroplasty implant registries over the past five decades: Development, current, and future impact. J Orthop Res 2018;36(9):2319–2330. Crossref, Medline, Google Scholar2. Dutt R, Mendonca D, Phen M, et al. Automatic localization and brand detection of cervical spine hardware on radiographs using weakly supervised machine learning. Radiol Artif Intell 2022;4(2):e210099. Link, Google Scholar3. Ren M, Yi PH. Artificial intelligence in orthopedic implant model classification: a systematic review. Skeletal Radiol 2022;51(2):407–416. Crossref, Medline, Google Scholar4. Yi PH, Wei J, Kim TK, et al. Automated detection & classification of knee arthroplasty using deep learning. Knee 2020;27(2):535–542. Crossref, Medline, Google Scholar5. Karnuta JM, Haeberle HS, Luu BC, et al. Artificial Intelligence to Identify Arthroplasty Implants From Radiographs of the Hip. J Arthroplasty 2021;36(7S):S290–S294.e1. Crossref, Medline, Google Scholar6. Borjali A, Chen AF, Bedair HS, et al. Comparing the performance of a deep convolutional neural network with orthopedic surgeons on the identification of total hip prosthesis design from plain radiographs. Med Phys 2021;48(5):2327–2336. Crossref, Medline, Google Scholar7. Patel R, Thong EHE, Batta V, Bharath AA, Francis D, Howard J. Automated identification of orthopedic implants on radiographs using deep learning. Radiol Artif Intell 2021;3(4):e200183. Link, Google Scholar8. Sagheb E, Ramazanian T, Tafti AP, et al. Use of natural language processing algorithms to identify common data elements in operative notes for knee arthroplasty. J Arthroplasty 2021;36(3):922–926. Crossref, Medline, Google ScholarArticle HistoryReceived: Jan 13 2022Revision requested: Jan 19 2022Revision received: Jan 19 2022Accepted: Jan 25 2022Published online: Mar 02 2022 FiguresReferencesRelatedDetailsAccompanying This ArticleAutomatic Localization and Brand Detection of Cervical Spine Hardware on Radiographs Using Weakly Supervised Machine LearningJan 19 2022Radiology: Artificial IntelligenceRecommended Articles Automated Identification of Orthopedic Implants on Radiographs Using Deep LearningRadiology: Artificial Intelligence2021Volume: 3Issue: 4Deep Learning Improves Predictions of the Need for Total Knee ReplacementRadiology2020Volume: 296Issue: 3pp. 594-595Automatic Localization and Brand Detection of Cervical Spine Hardware on Radiographs Using Weakly Supervised Machine LearningRadiology: Artificial Intelligence2022Volume: 4Issue: 2Adventures and Misadventures in Plastic Surgery and Soft-Tissue ImplantsRadioGraphics2017Volume: 37Issue: 7pp. 2145-2163Convolutional Neural Networks for Automated Fracture Detection and Localization on Wrist RadiographsRadiology: Artificial Intelligence2019Volume: 1Issue: 1See More RSNA Education Exhibits Beyond "Prosthesis in Situ" - A Radiological Review of Normal Appearances and Complications of Orthopaedic ImplantsDigital Posters2019Elusive Complications in Hip Arthroplasty - Dare to Spot It: Imaging Features of Uncommon Postoperative ComplicationsDigital Posters2019Keep It Moving: Spondylosis and Posterior Spinal Motion Preserving SurgeriesDigital Posters2019 RSNA Case Collection Hip Polyethylene Liner DissociationRSNA Case Collection2021Silicone implant ruptureRSNA Case Collection2020Dysostosis Multiplex (Hurler's Syndrome)RSNA Case Collection2021 Vol. 4, No. 2 Metrics Downloaded 82 times Altmetric Score PDF download
Deep learning methods have demonstrated the ability to perform accurate coronary artery calcium (CAC) scoring. However, these methods require large and representative training data hampering applicability to diverse CT scans showing the heart and the coronary arteries. Training methods that accurately score CAC in cross-domain settings remains challenging. To address this, we present an unsupervised domain adaptation method that learns to perform CAC scoring in coronary CT angiography (CCTA) from non-contrast CT (NCCT). To address the domain shift between NCCT (source) domain and CCTA (target) domain, feature distributions are aligned between two domains using adversarial learning. A CAC scoring convolutional neural network is divided into a feature generator that maps input images to features in the latent space and a classifier that estimates predictions from the extracted features. For adversarial learning, a discriminator is used to distinguish the features between source and target domains. Hence, the feature generator aims to extract features with aligned distributions to fool the discriminator. The network is trained with adversarial loss as the objective function and a classification loss on the source domain as a constraint for adversarial learning. In the experiments, three data sets were used. The network is trained with 1,687 labeled chest NCCT scans from the National Lung Screening Trial. Furthermore, 200 labeled cardiac NCCT scans and 200 unlabeled CCTA scans were used to train the generator and the discriminator for unsupervised domain adaptation. Finally, a data set containing 313 manually labeled CCTA scans was used for testing. Directly applying the CAC scoring network trained on NCCT to CCTA led to a sensitivity of 0.41 and an average false positive volume 140 mm3/scan. The proposed method improved the sensitivity to 0.80 and reduced average false positive volume of 20 mm3/scan. The results indicate that the unsupervised domain adaptation approach enables automatic CAC scoring in contrast enhanced CT while learning from a large and diverse set of CT scans without contrast. This may allow for better utilization of existing annotated data sets and extend the applicability of automatic CAC scoring to contrast-enhanced CT scans without the need for additional manual annotations. The code is publicly available at https://github.com/qurAI-amsterdam/CACscoringUsingDomainAdaptation.
Purpose: Ensembles of convolutional neural networks (CNNs) often outperform a single CNN in medical image segmentation tasks, but inference is computationally more expensive and makes ensembles unattractive for some applications. We compared the performance of differently constructed ensembles with the performance of CNNs derived from these ensembles using knowledge distillation, a technique for reducing the footprint of large models such as ensembles. Approach: We investigated two different types of ensembles, namely, diverse ensembles of networks with three different architectures and two different loss-functions, and uniform ensembles of networks with the same architecture but initialized with different random seeds. For each ensemble, additionally, a single student network was trained to mimic the class probabilities predicted by the teacher model, the ensemble. We evaluated the performance of each network, the ensembles, and the corresponding distilled networks across three different publicly available datasets. These included chest computed tomography scans with four annotated organs of interest, brain magnetic resonance imaging (MRI) with six annotated brain structures, and cardiac cine-MRI with three annotated heart structures. Results: Both uniform and diverse ensembles obtained better results than any of the individual networks in the ensemble. Furthermore, applying knowledge distillation resulted in a single network that was smaller and faster without compromising performance compared with the ensemble it learned from. The distilled networks significantly outperformed the same network trained with reference segmentation instead of knowledge distillation. Conclusion: Knowledge distillation can compress segmentation ensembles of uniform or diverse composition into a single CNN while maintaining the performance of the ensemble.
BACKGROUND:A baseline computed tomography (CT) scan for lung cancer (LC) screening may reveal information indicating that certain LC screening participants can be screened less, and instead require dedicated early cardiac and respiratory clinical input. We aimed to develop and validate competing death (CD) risk models using CT information to identify participants with a low LC risk and a high CD risk.METHODS:Participant demographics and quantitative CT measures of LC, cardiovascular disease and chronic obstructive pulmonary disease were considered for deriving a logistic regression model for predicting 5-year CD risk using a sample from the National Lung Screening Trial (n=15 000). Multicentric Italian Lung Detection data were used to perform external validation (n=2287).RESULTS:Our final CD model outperformed an external pre-scan model (CD Risk Assessment Tool) in both the derivation (area under the curve (AUC) 0.744 (95% CI 0.727-0.761) and 0.677 (95% CI 0.658-0.695), respectively) and validation cohorts (AUC 0.744 (95% CI 0.652-0.835) and 0.725 (95% CI 0.633-0.816), respectively). By also taking LC incidence risk into consideration, we suggested a risk threshold where a subgroup (6258/23 096 (27%)) was identified with a number needed to screen to detect one LC of 216 (versus 23 in the remainder of the cohort) and ratio of 5.41 CDs per LC case (versus 0.88). The respective values in the validation cohort subgroup (774/2287 (34%)) were 129 (versus 29) and 1.67 (versus 0.43).CONCLUSIONS:Evaluating both LC and CD risks post-scan may improve the efficiency of LC screening and facilitate the initiation of multidisciplinary trajectories among certain participants.
Deep-learning-based registration methods emerged as a fast alternative to conventional registration methods. However, these methods often still cannot achieve the same performance as conventional registration methods because they are either limited to small deformation or they fail to handle a superposition of large and small deformations without producing implausible deformation fields with foldings inside. In this paper, we identify important strategies of conventional registration methods for lung registration and successfully developed the deep-learning counterpart. We employ a Gaussian-pyramid-based multilevel framework that can solve the image registration optimization in a coarse-to-fine fashion. Furthermore, we prevent foldings of the deformation field and restrict the determinant of the Jacobian to physiologically meaningful values by combining a volume change penalty with a curvature regularizer in the loss function. Keypoint correspondences are integrated to focus on the alignment of smaller structures. We perform an extensive evaluation to assess the accuracy, the robustness, the plausibility of the estimated deformation fields, and the transferability of our registration approach. We show that it achieves state-of-the-art results on the COPDGene dataset compared to conventional registration method with much shorter execution time. In our experiments on the DIRLab exhale to inhale lung registration, we demonstrate substantial improvements (TRE below $1.2$ mm) over other deep learning methods. Our algorithm is publicly available at https://grand-challenge.org/algorithms/deep-learning-based-ct-lung-registration/.
PURPOSE:To examine the prognostic value of location-specific arterial calcification quantities at lung screening low-dose CT for the prediction of cardiovascular disease (CVD) mortality. MATERIALS AND METHODS:This retrospective study included 5564 participants who underwent low-dose CT from the National Lung Screening Trial between August 2002 and April 2004, who were followed until December 2009. A deep learning network was trained to quantify six types of vascular calcification: thoracic aorta calcification (TAC); aortic and mitral valve calcification; and coronary artery calcification (CAC) of the left main, the left anterior descending, and the right coronary artery. TAC and CAC were determined in six evenly distributed slabs spatially aligned among chest CT images. CVD mortality prediction was performed with multivariable logistic regression using least absolute shrinkage and selection operator. The methods were compared with semiautomatic baseline prediction using self-reported participant characteristics, such as age, history of smoking, and history of illness. Statistical significance between the prediction models was tested using the nonparametric DeLong test. RESULTS:The prediction model was trained with data from 4451 participants (median age, 61 years; 37.9% women) and then tested on data from 1113 participants (median age, 61 years; 37.9% women). The prediction model using calcium scores achieved a C statistic of 0.74 (95% CI: 0.69, 0.79), and it outperformed the baseline model using only participant characteristics (C statistic, 0.69; P = .049). Best results were obtained when combining all variables (C statistic, 0.76; P < .001). CONCLUSION:Five-year CVD mortality prediction using automatically extracted image-based features is feasible at lung screening low-dose CT.© RSNA, 2021.
Plasma osteoprotegerin (OPG) and vascular smooth muscle cell (VSMC) derived extracellular vesicles (EVs) are important regulators in the process of vascular calcification (VC). In population studies, high levels of OPG are associated with events. In animal studies, however, high OPG levels result in reduction of VC. VSMC-derived EVs are assumed to be responsible for OPG transport and VC but this role has not been studied. For this, we investigated the association between OPG in plasma and circulating EVs with coronary artery calcium (CAC) as surrogate for VC in symptomatic patients. We retrospectively assessed 742 patients undergoing myocardial perfusion imaging (MPI). CAC scores were determined on the MPI-CT images using a previously developed automated algorithm. Levels of OPG were quantified in plasma and two EV-subpopulations (LDL and TEX), using an electrochemiluminescence immunoassay. Circulating levels of OPG were independently associated with CAC scores in plasma; OR 1.39 (95% CI 1.17-1.65), and both EV populations; EV-LDL; OR 1.51 (95% CI 1.27-1.80) and EV-TEX; OR 1.21 (95% CI 1.02-1.42). High levels of OPG in plasma were independently associated with CAC scores in this symptomatic patient cohort. High levels of EV-derived OPG showed the same positive association with CAC scores, suggesting that EV-derived OPG mirrors the same pathophysiological process as plasma OPG.
Importance Cardiovascular disease (CVD) is common in patients treated for breast cancer, especially in patients treated with systemic treatment and radiotherapy and in those with preexisting CVD risk factors. Coronary artery calcium (CAC), a strong independent CVD risk factor, can be automatically quantified on radiotherapy planning computed tomography (CT) scans and may help identify patients at increased CVD risk. Objective To evaluate the association of CAC with CVD and coronary artery disease (CAD) in patients with breast cancer. Design, Setting, and Participants In this multicenter cohort study of 15 915 patients with breast cancer receiving radiotherapy between 2005 and 2016 who were followed until December 31, 2018, age, calendar year, and treatment-adjusted Cox proportional hazard models were used to evaluate the association of CAC with CVD and CAD. Exposures Overall CAC scores were automatically extracted from planning CT scans using a deep learning algorithm. Patients were classified into Agatston risk categories (0, 1-10, 11-100, 101-399, >400 units). Main Outcomes and Measures Occurrence of fatal and nonfatal CVD and CAD were obtained from national registries. Results Of the 15 915 participants included in this study, the mean (SD) age at CT scan was 59.0 (11.2; range, 22-95) years, and 15 879 (99.8%) were women. Seventy percent (n = 11 179) had no CAC. Coronary artery calcium scores of 1 to 10, 11 to 100, 101 to 400, and greater than 400 were present in 10.0% (n = 1584), 11.5% (n = 1825), 5.2% (n = 830), and 3.1% (n = 497) respectively. After a median follow-up of 51.2 months, CVD risks increased from 5.2% in patients with no CAC to 28.2% in patients with CAC scores higher than 400. After adjustment, CVD risk increased with higher CAC score (hazard ratio [HR]CAC = 1-10 = 1.1; 95% CI, 0.9-1.4; HRCAC = 11-100 = 1.8; 95% CI, 1.5-2.1; HRCAC = 101-400 = 2.1; 95% CI, 1.7-2.6; and HRCAC>400 = 3.4; 95% CI, 2.8-4.2). Coronary artery calcium was particularly strongly associated with CAD (HRCAC>400 = 7.8; 95% CI, 5.5-11.2). The association between CAC and CVD was strongest in patients treated with anthracyclines (HRCAC>400 = 5.8; 95% CI, 3.0-11.4) and patients who received a radiation boost (HRCAC>400 = 6.1; 95% CI, 3.8-9.7). Conclusions and Relevance This cohort study found that coronary artery calcium on breast cancer radiotherapy planning CT scan results was associated with CVD, especially CAD. Automated CAC scoring on radiotherapy planning CT scans may be used as a fast and low-cost tool to identify patients with breast cancer at increased risk of CVD, allowing implementing CVD risk-mitigating strategies with the aim to reduce the risk of CVD burden after breast cancer. Trial Registration ClinicalTrials.gov Identifier: NCT03206333.
ObjectivesCombined assessment of cardiovascular disease (CVD), COPD and lung cancer may improve the effectiveness of lung cancer screening in smokers. The aims were to derive and assess risk models for predicting lung cancer incidence, CVD mortality and COPD mortality by combining quantitative computed tomography (CT) measures from each disease, and to quantify the added predictive benefit of self-reported patient characteristics given the availability of a CT scan.MethodsA survey model (patient characteristics only), CT model (CT information only) and final model (all variables) were derived for each outcome using parsimonious Cox regression on a sample from the National Lung Screening Trial (n=15 000). Validation was performed using Multicentric Italian Lung Detection data (n=2287). Time-dependent measures of model discrimination and calibration are reported.ResultsAge, mean lung density, emphysema score, bronchial wall thickness and aorta calcium volume are variables that contributed to all final models. Nodule features were crucial for lung cancer incidence predictions but did not contribute to CVD and COPD mortality prediction. In the derivation cohort, the lung cancer incidence CT model had a 5-year area under the receiver operating characteristic curve of 82.5% (95% CI 80.9–84.0%), significantly inferior to that of the final model (84.0%, 82.6–85.5%). However, the addition of patient characteristics did not improve the lung cancer incidence model performance in the validation cohort (CT model 80.1%, 74.2–86.0%; final model 79.9%, 73.9–85.8%). Similarly, the final CVD mortality model outperformed the other two models in the derivation cohort (survey model 74.9%, 72.7–77.1%; CT model 76.3%, 74.1–78.5%; final model 79.1%, 77.0–81.2%), but not the validation cohort (survey model 74.8%, 62.2–87.5%; CT model 72.1%, 61.1–83.2%; final model 72.2%, 60.4–84.0%). Combining patient characteristics and CT measures provided the largest increase in accuracy for the COPD mortality final model (92.3%, 90.1–94.5%) compared to either other model individually (survey model 87.5%, 84.3–90.6%; CT model 87.9%, 84.8–91.0%), but no external validation was performed due to a very low event frequency.ConclusionsCT measures of CVD and COPD provides small but reproducible improvements to nodule-based lung cancer risk prediction accuracy from 3 years onwards. Self-reported patient characteristics may not be of added predictive value when CT information is available.
Vertebral labelling and segmentation are two fundamental tasks in an automated spine processing pipeline. Reliable and accurate processing of spine images is expected to benefit clinical decision-support systems for diagnosis, surgery planning, and population-based analysis on spine and bone health. However, designing automated algorithms for spine processing is challenging predominantly due to considerable variations in anatomy and acquisition protocols and due to a severe shortage of publicly available data. Addressing these limitations, the Large Scale Vertebrae Segmentation Challenge (VerSe) was organised in conjunction with the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) in 2019 and 2020, with a call for algorithms towards labelling and segmentation of vertebrae. Two datasets containing a total of 374 multi-detector CT scans from 355 patients were prepared and 4505 vertebrae have individually been annotated at voxel-level by a human-machine hybrid algorithm (https://osf.io/nqjyw/, https://osf.io/t98fz/). A total of 25 algorithms were benchmarked on these datasets. In this work, we present the the results of this evaluation and further investigate the performance-variation at vertebra-level, scan-level, and at different fields-of-view. We also evaluate the generalisability of the approaches to an implicit domain shift in data by evaluating the top performing algorithms of one challenge iteration on data from the other iteration. The principal takeaway from VerSe: the performance of an algorithm in labelling and segmenting a spine scan hinges on its ability to correctly identify vertebrae in cases of rare anatomical variations. The content and code concerning VerSe can be accessed at: https://github.com/anjany/verse.