Purpose:The amount of fibroglandular tissue (FGT) and background parenchymal enhancement (BPE) are evaluated as part of the BI-RADS reporting standard for the diagnosis of breast cancer in magnetic resonance imaging (MRI). BPE is associated with breast cancer risk and affects the diagnostic accuracy of breast MRI readings but is subject to very high inter-reader variability. Approach:We developed a fully automatic observer-independent method to measure the amount of FGT and BPE based on breast, FGT, and lesion masks, which classifies each into four classes based on thresholds applied to those scores. This retrospective study involved 2840 patients from seven institutions for whom T1-weighted dynamic contrast-enhanced breast MRI sequences with and without fat suppression were acquired using 1.5 or 3T scanners. FGT and BPE assessments for model development and evaluation were performed according to the BI-RADS guidelines by radiologists from the respective institutions and by a radiographer if the assessments were missing. Results:The performance of the models was compared using four-class accuracy and a paired t-test ( α = 0.05 ). Our automatic classification of FGT and BPE in a multi-institutional setting achieved a four-class accuracy of 0.535 for FGT and 0.504 for BPE ( random = 0.25 ). The calibration of the classification thresholds for each site individually resulted in a significant improvement, reaching an accuracy of 0.669 ( p < 0.05 ) for FGT and 0.565 ( p < 0.05 ) for BPE. The area under curve of receiver operating characterstic (ROC-AUC) to classify FGT and BPE into low versus high is 0.883 for FGT and 0.757 for BPE when computed across all sites. The averaged per-site AUCs are 0.926 for FGT and 0.810 for BPE. Conclusion:FGT and BPE classification perform significantly better when classification thresholds are selected per clinical site.
Breast cancer is the most prevalent cancer among women globally, and magnetic resonance imaging (MRI) is increasingly being proposed as a screening tool in addition to its role as a diagnostic imaging modality. This creates a need for automated methods that analyze breast MR images. One of the tasks in BI-RADS-based breast MRI reading is to assess whether implants are present. We trained nnU-Net models using 5-fold crossvalidation for the task of breast implant segmentation and assessed the segmentation quality on a multi-centric dataset consisting of 941 MRI studies, 205 of which contained breast implants. Our models achieved a mean (standard deviation) Dice score of 0.96 (0.03) and an average symmetric surface distance of 0.95 (0.67) mm in cases with breast implants. Of the 736 cases without implants, 575 were correctly predicted to have no implants, which increased to 732 with simple volume filtering. Our findings suggest that breast implant detection and segmentation in MRI can be solved well by deep learning approaches, even with a limited amount of data.
Purpose To investigate whether selective removal of vascular structures can improve lesion visibility and interpretability in maximum intensity projection (MIP) images derived from dynamic contrast-enhanced MRI. Materials and Methods A retrospective analysis was conducted using breast MRI scans from the Duke-Breast-Cancer-MRI (Duke) dataset for model development. DeepVEST, a deep learning method for automatic vessel segmentation and removal, was developed. A reader study with five breast radiologists was conducted to evaluate the impact of vessel-removed MIPs on lesion assessment, using images from the Duke and Advanced-MRI-Breast-Lesions (AMBL) datasets. Segmentation performance was evaluated against manually annotated vessel segmentations using the Dice similarity coefficient (DSC), and perceived usefulness and quality were measured through reader study metrics. Interreader agreement on vessel removal quality and artifact impact was calculated using the Gwet agreement coefficient (AC1). Results DeepVEST achieved a DSC of 0.611 for vessel segmentation. In the reader study (150 assessments, 600 responses), vessels partially or fully obscured lesion margins in 91 of 150 assessments (60.7%). Vessel removal effectiveness was rated 3.820 ± 0.749 on a 5-point Likert scale. Artifacts were reported in 45 of 150 assessments (30.0%), with a low average severity score of 0.493 ± 0.900 (scale 0-5), indicating minimal image quality impact. Substantial interreader agreement was observed for vessel removal quality (AC1 = 0.703) and artifact evaluation (AC1 = 0.736). Conclusion DeepVEST enabled automatic vessel segmentation and removal in breast MRI MIP images, showing high effectiveness with minimal artifacts. Most readers found it helpful for lesion assessment in challenging cases. ©RSNA, 2026.
Enhance mammography quality to increase cancer detection by implementing continuous AI-driven feedback mechanisms, ensuring reliable, consistent, and high-quality screening by the ‘Perfect’, ‘Good’, ‘Moderate’, and ‘Inadequate’ (PGMI) criteria. To assess the impact of the AI software ‘b-boxTM’ on mammography quality, we conducted a comparative analysis of PGMI scores. We evaluated scores 50 days before (A) and after the software’s implementation in 2021 (B), along with assessments made in the first week of August 2022 (C1) and 2023 (C2), comparing them to evaluations conducted by two readers. Except for postsurgical patients, we included all diagnostic and screening mammograms from one tertiary hospital. A total of 4577 mammograms from 1220 women (mean age: 59, range: 21–94, standard deviation: 11.18) were included. 1728 images were obtained before (A) and 2330 images after the 2021 software implementation (B), along with 269 images in 2022 (C1) and 250 images in 2023 (C2). The results indicated a significant improvement in diagnostic image quality (p < 0.01). The percentage of ‘Perfect’ examinations rose from 22.34
OBJECTIVES:The aim of this study was to develop and validate a commercially available AI platform for the automatic determination of image quality in mammography and tomosynthesis considering a standardized set of features.MATERIALS AND METHODS:In this retrospective study, 11,733 mammograms and synthetic 2D reconstructions from tomosynthesis of 4200 patients from two institutions were analyzed by assessing the presence of seven features which impact image quality in regard to breast positioning. Deep learning was applied to train five dCNN models on features detecting the presence of anatomical landmarks and three dCNN models for localization features. The validity of models was assessed by the calculation of the mean squared error in a test dataset and was compared to the reading by experienced radiologists.RESULTS:Accuracies of the dCNN models ranged between 93.0% for the nipple visualization and 98.5% for the depiction of the pectoralis muscle in the CC view. Calculations based on regression models allow for precise measurements of distances and angles of breast positioning on mammograms and synthetic 2D reconstructions from tomosynthesis. All models showed almost perfect agreement compared to human reading with Cohen's kappa scores above 0.9.CONCLUSIONS:An AI-based quality assessment system using a dCNN allows for precise, consistent and observer-independent rating of digital mammography and synthetic 2D reconstructions from tomosynthesis. Automation and standardization of quality assessment enable real-time feedback to technicians and radiologists that shall reduce a number of inadequate examinations according to PGMI (Perfect, Good, Moderate, Inadequate) criteria, reduce a number of recalls and provide a dependable training platform for inexperienced technicians.
The heterogeneous group of B3 lesions in the breast harbors lesions with different malignant potential and progression risk. As several studies about B3 lesions have been published since the last Consensus in 2018, the 3rd International Consensus Conference discussed the six most relevant B3 lesions (atypical ductal hyperplasia (ADH), flat epithelial atypia (FEA), classical lobular neoplasia (LN), radial scar (RS), papillary lesions (PL) without atypia, and phyllodes tumors (PT)) and made recommendations for diagnostic and therapeutic approaches. Following a presentation of current data of each B3 lesion, the international and interdisciplinary panel of 33 specialists and key opinion leaders voted on the recommendations for further management after core-needle biopsy (CNB) and vacuum-assisted biopsy (VAB). In case of B3 lesion diagnosis on CNB, OE was recommended in ADH and PT, whereas in the other B3 lesions, vacuum-assisted excision was considered an equivalent alternative to OE. In ADH, most panelists (76%) recommended an open excision (OE) after diagnosis on VAB, whereas observation after a complete VAB-removal on imaging was accepted by 34%. In LN, the majority of the panel (90%) preferred observation following complete VAB-removal. Results were similar in RS (82%), PL (100%), and FEA (100%). In benign PT, a slim majority (55%) also recommended an observation after a complete VAB-removal. VAB with subsequent active surveillance can replace an open surgical intervention for most B3 lesions (RS, FEA, PL, PT, and LN). Compared to previous recommendations, there is an increasing trend to a de-escalating strategy in classical LN. Due to the higher risk of upgrade into malignancy, OE remains the preferred approach after the diagnosis of ADH.
INTRODUCTION:Contralateral axillary lymph node metastasis (CALNM) in breast cancer (BC) is considered a distant metastasis, marking stage 4cancer. Therefore, it is generally treated as an incurable disease. However, in clinical practice, staging and treatment remain controversial due to a paucity of data, and the St. Gallen 2021 consensus panel recommended a curative approach in patients with oligometastatic disease. Aberrant lymph node (LN) drainage following previous surgery or radiotherapy is common. Therefore, CALNM may be considered a regional event rather than systemic disease, and a re-sentinel procedure aided by lymphoscintigraphy permits adequate regional staging.CASE REPORT:Here, we report a 37-year-old patient with Lynch syndrome who presented with CALNM in an ipsilateral relapse of a moderately differentiated invasive ductal BC (ER 90%, PR 30%, HER2 negative, Ki-67 25%, microsatellite stable), 3 years after the initial diagnosis. Lymphoscintigraphy detected a positive sentinel LN in the contralateral axilla despite no sign of LN involvement or distant metastases on FDG PET/CT or MRI. The patient underwent bilateral mastectomy with sentinel node dissection, surgical reconstruction with histological confirmation of the CALNM, left axillary dissection, adjuvant chemotherapy, and anti-hormone therapy. In addition to her regular BC follow-up visits, the patient will undergo annual colonoscopy, gastroscopy, abdominal, and vaginal ultrasound screening. In January 2023, the patient was free of progression for 23 months after initiation of treatment for recurrent BC and CALNM.CONCLUSION:This case highlights the value of delayed lymphoscintigraphy and the contribution of sentinel procedure for local control in the setting of recurrent BC. Aberrant lymph node drainage following previous surgery may be the underlying cause of CALNM. We propose that CALNM without evidence of systemic metastasis should be considered a regional event in recurrent BC, and thus, a curative approach can be pursued. The next AJCC BC staging should clarify the role of CALNM in recurrent BC to allow for the development of specific treatment guidelines.
High breast density is a well-known risk factor for breast cancer. This study aimed to develop and adapt two (MLO, CC) deep convolutional neural networks (DCNN) for automatic breast density classification on synthetic 2D tomosynthesis reconstructions. In total, 4605 synthetic 2D images (1665 patients, age: 57 ± 37 years) were labeled according to the ACR (American College of Radiology) density (A-D). Two DCNNs with 11 convolutional layers and 3 fully connected layers each, were trained with 70
For AI-based classification tasks in computed tomography (CT), a reference standard for evaluating the clinical diagnostic accuracy of individual classes is essential. To enable the implementation of an AI tool in clinical practice, the raw data should be drawn from clinical routine data using state-of-the-art scanners, evaluated in a blinded manner and verified with a reference test. Three hundred and thirty-five consecutive CTs, performed between 1 January 2016 and 1 January 2021 with reported pleural effusion and pathology reports from thoracocentesis or biopsy within 7 days of the CT were retrospectively included. Two radiologists (4 and 10 PGY) blindly assessed the chest CTs for pleural CT features. If needed, consensus was achieved using an experienced radiologist’s opinion (29 PGY). In addition, diagnoses were extracted from written radiological reports. We analyzed these findings for a possible correlation with the following patient outcomes: mortality and median hospital stay. For AI prediction, we used an approach consisting of nnU-Net segmentation, PyRadiomics features and a random forest model. Specificity and sensitivity for CT-based detection of empyema (n = 81 of n = 335 patients) were 90.94 (95%-CI: 86.55–94.05) and 72.84 (95%-CI: 61.63–81.85%) in all effusions, with moderate to almost perfect interrater agreement for all pleural findings associated with empyema (Cohen’s kappa = 0.41–0.82). Highest accuracies were found for pleural enhancement or thickening with 87.02% and 81.49%, respectively. For empyema prediction, AI achieved a specificity and sensitivity of 74.41% (95% CI: 68.50–79.57) and 77.78% (95% CI: 66.91–85.96), respectively. Empyema was associated with a longer hospital stay (median = 20 versus 14 days), and findings consistent with pleural carcinomatosis impacted mortality.
Objectives: Rapid communication of CT exams positive for pulmonary embolism (PE) is crucial for timely initiation of anticoagulation and patient outcome. It is unknown if deep learning automated detection of PE on CT Pulmonary Angiograms (CTPA) in combination with worklist prioritization and an electronic notification system (ENS) can improve communication times and patient turnaround in the Emergency Department (ED). Methods: In 01/2019, an ENS allowing direct communication between radiology and ED was installed. Starting in 10/2019, CTPAs were processed by a deep learning (DL)-powered algorithm for detection of PE. CTPAs acquired between 04/2018 and 06/2020 (n = 1808) were analysed. To assess the impact of the ENS and the DL-algorithm, radiology report reading times (RRT), radiology report communication time (RCT), time to anticoagulation (TTA), and patient turnaround times (TAT) in the ED were compared for three consecutive time periods. Performance measures of the algorithm were calculated on a per exam level (sensitivity, specificity, PPV, NPV, F1score), with written reports and exam review as ground truth. Results: Sensitivity of the algorithm was 79.6 % (95 %CI:70.8-87.2%), specificity 95.0 % (95 %CI:92.0-97.1%), PPV 82.2 % (95 %CI:73.9-88.3), and NPV 94.1 % (95 %CI:91.4-96 %). There was no statistically significant reduction of any of the observed times (RRT, RCT, TTA, TAT). Conclusion: DL-assisted detection of PE in CTPAs and ENS-assisted communication of results to referring physicians technically work. However, the mere clinical introduction of these tools, even if they exhibit a good performance, is not sufficient to achieve significant effects on clinical performance measures.
Carcinoma of Unknown Primary presenting primarily as hepatic metastases encompasses a dismal subgroup of tumors with a median survival of 5.9 months. Adenocarcinoma is the most common histological subtype identified upon biopsy and the primary tumor remains undetectable in the majority of cases despite extensive workup. It is important to have a validated and standardized algorithm to follow these tumors to avoid unnecessary tests, as the wishes and health status of the patient represent the principal concerns. The purpose of this paper is to briefly review the current literature on carcinoma of unknown primary with hepatic metastases and propose a standardized diagnostic approach.
PurposeTo compare the segmentation and detection performance of a deep learning model trained on a database of human-labeled clinical stroke lesions on diffusion-weighted (DW) images to a model trained on the same database enhanced with synthetic stroke lesions.Materials and MethodsIn this institutional review board-approved study, a stroke database of 962 cases (mean patient age ± standard deviation, 65 years ± 17; 255 male patients; 449 scans with DW positive stroke lesions) and a normal database of 2027 patients (mean age, 38 years ± 24; 1088 female patients) were used. Brain volumes with synthetic stroke lesions on DW images were produced by warping the relative signal increase of real strokes to normal brain volumes. A generic three-dimensional (3D) U-Net was trained on four different databases to generate four different models: (a) 375 neuroradiologist-labeled clinical DW positive stroke cases (CDB); (b) 2000 synthetic cases (S2DB); (c) CDB plus 2000 synthetic cases (CS2DB); and (d) CDB plus 40 000 synthetic cases (CS40DB). The models were tested on 20% (n = 192) of the cases of the stroke database, which were excluded from the training set. Segmentation accuracy was characterized using Dice score and lesion volume of the stroke segmentation, and statistical significance was tested using a paired two-tailed Student t test. Detection sensitivity and specificity were compared with labeling done by three neuroradiologists.ResultsThe performance of the 3D U-Net model trained on the CS40DB (mean Dice score, 0.72) was better than models trained on the CS2DB (Dice score, 0.70; P < .001) or the CDB (Dice score, 0.65; P < .001). The deep learning model (CS40DB) was also more sensitive (91% [95% confidence interval {CI}: 89%, 93%]) than each of the three human readers (human reader 3, 84% [95% CI: 81%, 87%]; human reader 1, 78% [95% CI: 75%, 81%]; human reader 2, 79% [95% CI: 76%, 82%]), but was less specific (75% [95% CI: 72%, 78%]) than each of the three human readers (human reader 3, 96% [95% CI: 94%, 98%]; human reader 1, 92% [95% CI: 90%, 94%]; human reader 2, 89% [95% CI: 86%, 91%]).ConclusionDeep learning training for segmentation and detection of stroke lesions on DW images was significantly improved by enhancing the training set with synthetic lesions.Supplemental material is available for this article.© RSNA, 2020.
Poster: ECR 2019 / C-2207 / Towards tailored structured reporting: what do clinicians require in head and neck CTA reports by: N. Schmidt , L. Bonati, C. Glessgen, A. Jadczak, B. Stieltjes, K. Blackham; Basle/CH
Anti-angiogenic drugs cause a reduction in tumour density (Choi criteria) first and then in size [Response Evaluation Criteria In Solid Tumours (RECIST)]. The prognostic significance of changes in tumour density in metastatic renal cell carcinoma (mRCC) is unknown and was assessed in this study.