Breast US is an essential breast imaging tool that complements mammography and MRI. US is also often the primary imaging modality used to evaluate palpable breast masses and axillary lymph nodes and to guide percutaneous biopsy of breast masses and lymph nodes. Screening whole-breast US, with either handheld or automated technique, serves as a supplementary modality to screening mammography, particularly in women with dense breasts. Artificial intelligence (AI) has been adopted in US examinations to improve diagnostic accuracy and workflow. Analysis and quantification of background echotexture are emerging as a novel biomarker for breast cancer risk assessment. As US technology evolves and the scope of breast US widens, radiologists must understand the current and emerging US technology. They must also apply meticulous US scanning techniques to optimize image quality and ensure accurate diagnosis. This review provides a state-of-the-art summary of US technology and clinical applications as an adjuvant technique to mammography, MRI, and the clinical breast examination. The utility of breast US for screening, preoperative staging, and neoadjuvant treatment monitoring for breast cancer, breast intervention, and new techniques including AI, US tomography, optoacoustic imaging, and contrast-enhanced US will also be presented.
Background: Imaging-based breast cancer risk prediction models primarily use full-field digital mammography (FFDM). Although digital breast tomosynthesis (DBT) has become a predominant screening modality in the United States, its potential for long-term breast cancer risk prediction remains underexplored. Objective: The aim of this study was to develop and evaluate a deep learning model that uses longitudinal DBT examinations to predict long-term breast cancer risk. Methods: This retrospective study included 313,335 DBT examinations from 161,077 women (mean age, 58.5 ± 11.7 years) between January 2016 and August 2020 at a single health institution. A DBT-based risk prediction (DRP) model was developed to estimate 2- to 5-year breast cancer risk using longitudinal DBT examinations, patient age, and breast density. Model performance was compared with a single-timepoint DBT model, the Mirai model using same-day FFDM, and the Tyrer-Cuzick model using the AUC, time-dependent concordance index, and integrated Brier score. Results: In an independent test set (n = 34,570), the longitudinal DRP model achieved a 5-year AUC of 0.721 (95% CI, 0.698-0.744), improving on the single-timepoint DRP model (AUC, 0.707; 95% CI, 0.683-0.730; p < .001) and the Mirai model (AUC, 0.687; 95% CI, 0.663-0.710; p < .001). In a matched case-control cohort (n = 432), the DRP model achieved a 5-year AUC of 0.676 (95% CI, 0.626-0.726), compared with 0.563 (95% CI, 0.509-0.619; p < .001) for the Tyrer-Cuzick model. Among examinations of women with extremely dense breasts, the model classified 39.7% (746/1877) as average risk, with an observed 5-year cancer incidence of 0.8% (6/746). Among examinations of women with fatty breasts, the model classified 14.8% (386/2605) as high risk, with an observed 5-year cancer incidence of 2.6% (10/386). Conclusion: A deep learning model using longitudinal DBT examinations improved long-term breast cancer risk prediction compared with FFDM-based and clinical risk models. Clinical Impact: Longitudinal DBT-based risk prediction has the potential to inform dynamic risk assessment using screening images and to support future personalized screening strategies.
Purpose To develop and evaluate a deep learning model for automated quantification of breast arterial calcification (BAC) on screening mammography and to assess whether AI-derived BAC burden predicts major adverse cardiovascular events (MACE) in women. Methods In this retrospective study, 202,006 women who underwent screening mammography without history of MACE were included. A BAC segmentation model was trained on an expert-annotated dataset using a multi-task U-Net with a ResNet-18 encoder to detect and segment BAC. BAC burden was quantified as area (mm²) from model-generated masks using DICOM pixel spacing and categorized by tertiles into low, intermediate, and high. The PREVENT score and incident MACE were identified from electronic health records. Cox proportional hazards models were developed to evaluate AI-derived BAC burden and PREVENT score alone, and combined models for 5- and 10-year cardiovascular risk prediction. Results Among 202,006 women (mean age 54.8±11.7 years), 23.1% had AI-detected BAC, and 7,701 (3.8%) developed incident MACE during a median follow-up of 7.5 years. On the geographically held-out test set, the BAC model achieved an AUROC of 0.97, Dice score of 0.6678, and Pearson correlation of 0.961 between AI-derived and manually annotated BAC burden. BAC burden increased with age and was higher among women who developed MACE. Five-year MACE incidence increased across BAC categories from 1.5% in women without BAC to 6.9% in those with high BAC burden. BAC burden alone showed modest prediction of MACE, with 5-year and 10-year AUROCs of 0.661 and 0.650, respectively, while PREVENT achieved AUROCs of 0.781 and 0.771. Adding BAC to PREVENT produced minimal improvement in discrimination. Conclusion Deep learning-based BAC quantification from routine mammography is feasible, accurate, and associated with future cardiovascular risk. Although BAC added little to PREVENT for overall discrimination, it may serve as a scalable opportunistic imaging biomarker to identify women at elevated cardiovascular risk and support preventive care. ### Competing Interest Statement Dr. Reynolds has served as a consultant to Heartflow Inc and received in-kind donations for research from Abbott Vascular, Philips, SHL Biotelemetry, Siemens. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This retrospective study was approved by the NYU Grossman School of Medicine institutional review board and was compliant with the Health Insurance Portability and Accountability Act. The requirement for informed consent was waived. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data generated or analyzed during the study are included in the published paper.
MRI is the most effective method for screening high-risk breast cancer patients. While current exams rely on the qualitative evaluation of morphological features before and after contrast administration and less on contrast kinetic information, recent developments in fast acquisition methods aim to combine both. However, balancing spatial resolution, temporal resolution and scan time poses a considerable challenge in dynamic MRI. Here, we introduce a radial MRI reconstruction framework for Dynamic Contrast Enhanced (DCE) imaging, termed Enhanced Locally low-rank Imaging for Tissue contrast Enhancement (ELITE), to address these limitations. ELITE combines locally low-rank subspace modeling to capture spatially localized tissue dynamics with deep learning. We evaluate its effectiveness using the publicly available fastMRI breast initiative, demonstrating substantial improvements in CNR and noise reduction while enabling flexible temporal resolution down to 1 second. ELITE also shows benefits in neck and brain imaging, making it a viable alternative for other DCE-MRI applications.
Objective:To develop and evaluate a deep learning model for five-year breast cancer risk prediction from screening breast ultrasound (BUS) examinations. Methods:This retrospective study included 295,298 breast ultrasound examinations from 122,072 women imaged between 2012 and 2020. Patients were split into training, validation, and test sets; the test set included screening examinations only. BUS-Risk-Net aggregated image features using attention-based multiple instance learning and combined them with age and ultrasound-estimated breast density to predict 2- to 5-year risk. Performance was compared with the full Tyrer-Cuzick model in a matched case-control cohort and with a reduced Tyrer-Cuzick model in the held-out test set. Risk stratification was evaluated within BI-RADS density categories. Results:In the matched case-control cohort (n = 240 women), BUS-Risk-Net achieved a 5-year AUC of 0.632 (95% CI, 0.562-0.702), versus 0.514 for the full Tyrer-Cuzick model (95% CI, 0.440-0.588; p = 0.04). Among 19,548 examinations from 9,015 women eligible for 5-year evaluation in the test set, BUS-Risk-Net achieved an AUC of 0.679 (95% CI, 0.653-0.706), versus 0.594 for the reduced Tyrer-Cuzick model (95% CI, 0.564-0.623; P < .001). Observed 5-year cancer incidence increased across AI-defined risk tiers within each BI-RADS density category, ranging from 0.0% to 5.8% after AI stratification, compared with 2.1% to 3.6% across density categories alone. Discussion:Deep learning models applied to screening breast ultrasound could enable long-term breast cancer risk prediction and stratify risk beyond breast density alone. External and prospective validation is needed before clinical use.
Invasive lobular carcinoma (ILC) is the second-most common histologic subtype of breast cancer, constituting 5% to 15% of all breast cancers. It is characterized by an infiltrating growth pattern that may decrease detectability on mammography and US. The use of digital breast tomosynthesis (DBT) improves conspicuity of ILC, and sensitivity is 80% to 88% for ILC. Sensitivity of mammography is lower in dense breasts, and breast tomosynthesis has better sensitivity for ILC in dense breasts compared with digital mammography (DM). Screening US identifies additional ILCs even after DBT, with a supplemental cancer detection rate of 0 to 1.2 ILC per 1000 examinations. Thirteen percent of incremental cancers found by screening US are ILCs. Breast MRI has a sensitivity of 93% for ILC. Abbreviated breast MRI also has high sensitivity but may be limited due to delayed enhancement in ILC. Contrast-enhanced mammography has improved sensitivity for ILC compared with DM, with higher specificity than breast MRI. In summary, supplemental screening modalities increase detection of ILC, with MRI demonstrating the highest sensitivity.
Although digital breast tomosynthesis (DBT) improves diagnostic performance over full-field digital mammography (FFDM), false-positive recalls remain a concern in breast cancer screening. We developed a multi-modal artificial intelligence system integrating FFDM, synthetic mammography, and DBT to provide breast-level predictions and bounding-box localizations of suspicious findings. Our AI system, trained on approximately 500,000 mammography exams, achieved 0.945 AUROC on an internal test set. It demonstrated capacity to reduce recalls by 31.7 underscoring its potential to improve clinical workflows. External validation confirmed strong generalizability, reducing the gap to a perfect AUROC by 35.31 sites, the system reduced recall rates for low-risk cases. An improved version, trained on over 750,000 exams with additional labels, further reduced the gap by 18.86 underscore the importance of utilizing all available imaging modalities, demonstrate the potential for clinical impact, and indicate feasibility of further reduction of the test error with increased training set when using large-capacity neural networks.
This data curation work introduces the first large-scale dataset of radial k-space and DICOM data for breast DCE-MRI acquired in diagnostic breast MRI exams. Our dataset includes case-level labels indicating patient age, menopause status, lesion status (negative, benign, and malignant), and lesion type for each case. The public availability of this dataset and accompanying reconstruction code will support research and development of fast and quantitative breast image reconstruction and machine learning methods.
MRI is the most effective method for screening high-risk breast cancer patients. While current exams primarily rely on the qualitative evaluation of morphological features before and after contrast administration and less on contrast kinetic information, the latest developments in acquisition protocols aim to combine both. However, balancing between spatial and temporal resolution poses a significant challenge in dynamic MRI. Here, we propose a radial MRI reconstruction framework for Dynamic Contrast Enhanced (DCE) imaging, which offers a joint solution to existing spatial and temporal MRI limitations. It leverages a locally low-rank (LLR) subspace model to represent spatially localized dynamics based on tissue information. Our framework demonstrated substantial improvement in CNR, noise reduction and enables a flexible temporal resolution, ranging from a few seconds to 1-second, aided by a neural network, resulting in images with reduced undersampling penalties. Finally, our reconstruction framework also shows potential benefits for head and neck, and brain MRI applications, making it a viable alternative for a range of DCE-MRI exams.
OBJECTIVE:Radiology residents require timely, personalized feedback to develop accurate image analysis and reporting skills. Increasing clinical workload often limits attendings' ability to provide guidance. This study evaluates a HIPAA-compliant Generative Pretrained Transformer (GPT)-4o system that delivers automated feedback on breast imaging reports drafted by residents in real clinical settings. METHODS:We analyzed 5,000 resident-attending report pairs from routine practice at a multisite US health system. GPT-4o was prompted with clinical instructions to identify common errors and provide feedback. A reader study using 100 report pairs was conducted. Four attending radiologists and four residents independently reviewed each pair, determined whether predefined error types were present, and rated GPT-4o's feedback as helpful or not. Agreement between GPT and readers was assessed using percent match. Interreader reliability was measured with Krippendorff's α. Educational value was measured as the proportion of cases rated helpful. RESULTS:Three common error types were identified: (1) omission or addition of key findings, (2) incorrect use or omission of technical descriptors, and (3) final assessment inconsistent with findings. GPT-4o showed strong agreement with attending consensus: 90.5%, 78.3%, and 90.4% (Cohen's κ: 0.790, 0.550, and 0.615) across error types. Interreader reliability among all eight readers showed moderate to substantial variability (α = 0.767, 0.595, 0.567). When each reader was individually replaced with GPT-4o and interreader agreement among seven readers and GPT was recalculated, the effect was not statistically significant (Δ = -0.004 to 0.002, all P > .05). GPT's feedback was rated helpful in most cases: 89.8%, 83.0%, and 92.0%. DISCUSSION:ChatGPT-4o can reliably identify key educational errors. It may serve as a scalable tool to support radiology education.
Large language models (LLMs) and multi-modal large language models (MLLMs) represent the cutting-edge in artificial intelligence. This review provides a comprehensive overview of their capabilities and potential impact on radiology. Unlike most existing literature reviews focusing solely on LLMs, this work examines both LLMs and MLLMs, highlighting their potential to support radiology workflows such as report generation, image interpretation, EHR summarization, differential diagnosis generation, and patient education. By streamlining these tasks, LLMs and MLLMs could reduce radiologist workload, improve diagnostic accuracy, support interdisciplinary collaboration, and ultimately enhance patient care. We also discuss key limitations, such as the limited capacity of current MLLMs to interpret 3D medical images and to integrate information from both image and text data, as well as the lack of effective evaluation methods. Ongoing efforts to address these challenges are introduced.
Transformer-based detectors have shown success in computer vision tasks with natural images. These models, exemplified by the Deformable DETR, are optimized through complex engineering strategies tailored to the typical characteristics of natural scenes. However, medical imaging data presents unique challenges such as extremely large image sizes, fewer and smaller regions of interest, and object classes which can be differentiated only through subtle differences. This study evaluates the applicability of these transformer-based design choices when applied to a screening mammography dataset that represents these distinct medical imaging data characteristics. Our analysis reveals that common design choices from the natural image domain, such as complex encoder architectures, multi-scale feature fusion, query initialization, and iterative bounding box refinement, do not improve and sometimes even impair object detection performance in medical imaging. In contrast, simpler and shallower architectures often achieve equal or superior results. This finding suggests that the adaptation of transformer models for medical imaging data requires a reevaluation of standard practices, potentially leading to more efficient and specialized frameworks for medical diagnosis.
PURPOSE:To develop a digital reference object (DRO) toolkit to generate realistic breast DCE-MRI data for quantitative assessment of image reconstruction and data analysis methods. METHODS:A simulation framework in a form of DRO toolkit has been developed using the ultrafast and conventional breast DCE-MRI data of 53 women with malignant (n = 25) or benign (n = 28) lesions. We segmented five anatomical regions and performed pharmacokinetic analysis to determine the ranges of pharmacokinetic parameters for each segmented region. A database of the segmentations and their pharmacokinetic parameters is included in the DRO toolkit that can generate a large number of realistic breast DCE-MRI data. We provide two potential examples for our DRO toolkit: assessing the accuracy of an image reconstruction method using undersampled simulated radial k-space data and assessing the impact of the B 1 + $$ {\mathrm{B}}_1^{+} $$ field inhomogeneity on estimated parameters. RESULTS:The estimated pharmacokinetic parameters for each region showed agreement with previously reported values. For the assessment of the reconstruction method, it was found that the temporal regularization resulted in significant underestimation of estimated parameters by up to 57% and 10% with the weighting factor λ = 0.1 and 0.01, respectively. We also demonstrated that spatial discrepancy of v p $$ {v}_p $$ and PS $$ \mathrm{PS} $$ increase to about 33% and 51% without correction for B 1 + $$ {\mathrm{B}}_1^{+} $$ field. CONCLUSION:We have developed a DRO toolkit that includes realistic morphology of tumor lesions along with the expected pharmacokinetic parameter ranges. This simulation framework can generate many images for quantitative assessment of DCE-MRI reconstruction and analysis methods.
3D imaging enables accurate diagnosis by providing spatial information about organ anatomy. However, using 3D images to train AI models is computationally challenging because they consist of 10x or 100x more pixels than their 2D counterparts. To be trained with high-resolution 3D images, convolutional neural networks resort to downsampling them or projecting them to 2D. We propose an effective alternative, a neural network that enables efficient classification of full-resolution 3D medical images. Compared to off-the-shelf convolutional neural networks, our network, 3D Globally-Aware Multiple Instance Classifier (3D-GMIC), uses 77.98%-90.05% less GPU memory and 91.23%-96.02% less computation. While it is trained only with image-level labels, without segmentation labels, it explains its predictions by providing pixel-level saliency maps. On a dataset collected at NYU Langone Health, including 85,526 patients with full-field 2D mammography (FFDM), synthetic 2D mammography, and 3D mammography, 3D-GMIC achieves an AUC of 0.831 (95% CI: 0.769-0.887) in classifying breasts with malignant findings using 3D mammography. This is comparable to the performance of GMIC on FFDM (0.816, 95% CI: 0.737-0.878) and synthetic 2D (0.826, 95% CI: 0.754-0.884), which demonstrates that 3D-GMIC successfully classified large 3D images despite focusing computation on a smaller percentage of its input compared to GMIC. Therefore, 3D-GMIC identifies and utilizes extremely small regions of interest from 3D images consisting of hundreds of millions of pixels, dramatically reducing associated computational challenges. 3D-GMIC generalizes well to BCS-DBT, an external dataset from Duke University Hospital, achieving an AUC of 0.848 (95% CI: 0.798-0.896).
Full Field Digital Mammograms (FFDMs) and Digital Breast Tomosynthesis (DBT) are the two most widely used imaging modalities for breast cancer screening. Although DBT has increased cancer detection compared to FFDM, its widespread adoption in clinical practice has been slowed by increased interpretation times and a perceived decrease in the conspicuity of specific lesion types. Specifically, the non-inferiority of DBT for microcalcifications remains under debate. Due to concerns about the decrease in visual acuity, combined DBT-FFDM acquisitions remain popular, leading to overall increased exam times and radiation dosage. Enabling DBT to provide diagnostic information present in both FFDM and DBT would reduce reliance on FFDM, resulting in a reduction in both quantities. We propose a machine learning methodology that learns high-level representations leveraging the complementary diagnostic signal from both DBT and FFDM. Experiments on a large-scale data set validate our claims and show that our representations enable more accurate breast lesion detection than any DBT- or FFDM-based model.
Breast MRI has high sensitivity and negative predictive value, making it well suited to problem solving when other imaging modalities or physical examinations yield results that are inconclusive for the presence of breast cancer. Indications for problem-solving MRI include equivocal or uncertain imaging findings at mammography and/or US; suspicious nipple discharge or skin changes suspected to represent an abnormality when conventional imaging results are negative for cancer; lesions categorized as Breast Imaging Reporting and Data System 4, which are not amenable to biopsy; and discordant radiologic-pathologic findings after biopsy. MRI should not precede or replace careful diagnostic workup with mammography and US and should not be used when a biopsy can be safely performed. The role of MRI in characterizing calcifications is controversial, and management of calcifications should depend on their mammographic appearance because ductal carcinoma in situ may not appear enhancing on MR images. In addition, ductal carcinoma in situ detected solely with MRI is not associated with a higher likelihood of an upgrade to invasive cancer compared with ductal carcinoma in situ detected with other modalities. MRI for triage of high-risk lesions is a subject of ongoing investigation, with a possible future role for MRI in decreasing excisional biopsies. The accuracy of MRI is likely to increase with the use of advanced techniques such as deep learning, which will likely expand the indications for problem-solving MRI. ©RSNA, 2023 Quiz questions for this article are available in the supplemental material.
The use of digital breast tomosynthesis (DBT) in breast cancer screening has become widely accepted, facilitating increased cancer detection and lower recall rates compared with those achieved by using full-field digital mammography (DM). However, the use of DBT, as compared with DM, raises new challenges, including a larger number of acquired images and thus longer interpretation times. While most current artificial intelligence (AI) applications are developed for DM, there are multiple potential opportunities for AI to augment the benefits of DBT. During the diagnostic steps of lesion detection, characterization, and classification, AI algorithms may not only assist in the detection of indeterminate or suspicious findings but also aid in predicting the likelihood of malignancy for a particular lesion. During image acquisition and processing, AI algorithms may help reduce radiation dose and improve lesion conspicuity on synthetic two-dimensional DM images. The use of AI algorithms may also improve workflow efficiency and decrease the radiologist's interpretation time. There has been significant growth in research that applies AI to DBT, with several algorithms approved by the U.S. Food and Drug Administration for clinical implementation. Further development of AI models for DBT has the potential to lead to improved practice efficiency and ultimately improved patient health outcomes of breast cancer screening and diagnostic evaluation. See the invited commentary by Bahl in this issue. ©RSNA, 2022.
Pathology reports are considered the gold standard in medical research due to their comprehensive and accurate diagnostic information. Natural language processing (NLP) techniques have been developed to automate information extraction from pathology reports. However, existing studies suffer from two significant limitations. First, they typically frame their tasks as report classification, which restricts the granularity of extracted information. Second, they often fail to generalize to unseen reports due to variations in language, negation, and human error. To overcome these challenges, we propose a BERT (bidirectional encoder representations from transformers) named entity recognition (NER) system to extract key diagnostic elements from pathology reports. We also introduce four data augmentation methods to improve the robustness of our model. Trained and evaluated on 1438 annotated breast pathology reports, acquired from a large medical center in the United States, our BERT model trained with data augmentation achieves an entity F1-score of 0.916 on an internal test set, surpassing the BERT baseline (0.843). We further assessed the model's generalizability using an external validation dataset from the United Arab Emirates, where our model maintained satisfactory performance (F1-score 0.860). Our findings demonstrate that our NER systems can effectively extract fine-grained information from widely diverse medical reports, offering the potential for large-scale information extraction in a wide range of medical and AI research. We publish our code at https://github.com/nyukat/pathology_extraction.
Breast cancer screening, primarily conducted through mammography, is often supplemented with ultrasound for women with dense breast tissue. However, existing deep learning models analyze each modality independently, missing opportunities to integrate information across imaging modalities and time. In this study, we present Multi-modal Transformer (MMT), a neural network that utilizes mammography and ultrasound synergistically, to identify patients who currently have cancer and estimate the risk of future cancer for patients who are currently cancer-free. MMT aggregates multi-modal data through self-attention and tracks temporal tissue changes by comparing current exams to prior imaging. Trained on 1.3 million exams, MMT achieves an AUROC of 0.943 in detecting existing cancers, surpassing strong uni-modal baselines. For 5-year risk prediction, MMT attains an AUROC of 0.826, outperforming prior mammography-based risk models. Our research highlights the value of multi-modal and longitudinal imaging in cancer diagnosis and risk stratification.