As the use of artificial intelligence (AI) continues to grow in radiology, it has become clear that its real-world performance often differs from that demonstrated in premarket testing, underscoring the need for robust quality management (QM) programs at local institutions. For decades, a key mechanism to ensure QM in radiology practices has been ACR accreditation. However, no such program currently exists for AI in radiology. As leaders of the ACR Commissions on Quality and Safety and Informatics, we are dedicated to establishing ACR accreditation for radiology AI. In this article, we outline our plan for this effort. ACR accreditation is a peer-reviewed process that evaluates radiology practices according to ACR Practice Parameters and Technical Standards, which are consensus-based guidelines aimed at improving care quality and reducing variability. ACR Practice Parameters focus on clinical aspects like patient management, and Technical Standards address the performance of imaging and treatment equipment. To support the development of this accreditation program, the ACR Recognized Center for Healthcare-AI (ARCH-AI) program has been established as a precursor to formal accreditation. ARCH-AI participants attest to meeting minimum criteria in areas such as governance, model selection, acceptance testing, monitoring, and management of locally developed models. Insights gained from ARCH-AI will inform the development of the formal accreditation program, which will culminate in ACR Council approval, currently anticipated in spring 2027. The College remains committed to fostering dialogue among members and stakeholders to ensure AI fulfills its promise of enhancing patient care safely and effectively.
Purpose To evaluate the real-world performance of two FDA-approved artificial intelligence (AI)-based computer-aided triage and notification (CADt) detection devices and compare them with the manufacturer-reported performance testing in the instructions for use. Materials and methods Clinical performance of two FDA-cleared CADt large-vessel occlusion (LVO) devices was retrospectively evaluated at two separate stroke centers. Consecutive "code stroke" CT angiography examinations were included and assessed for patient demographics, scanner manufacturer, presence or absence of CADt result, CADt result, and LVO in the internal carotid artery (ICA), horizontal middle cerebral artery (MCA) segment (M1), Sylvian MCA segments after the bifurcation (M2), precommunicating part of cerebral artery, postcommunicating part of the cerebral artery, vertebral artery, basilar artery vessel segments. The original radiology report served as the reference standard, and a study radiologist extracted the above data elements from the imaging examination and radiology report. Results At hospital A, the CADt algorithm manufacturer reports assessment of intracranial ICA and MCA with sensitivity of 97% and specificity of 95.6%. Real-world performance of 704 cases included 79 in which no CADt result was available. Sensitivity and specificity in ICA and M1 segments were 85.3% and 91.9%. Sensitivity decreased to 68.5% when M2 segments were included and to 59.9% when all proximal vessel segments were included. At hospital B the CADt algorithm manufacturer reports sensitivity of 87.8% and specificity of 89.6%, without specifying the vessel segments. Real-world performance of 642 cases included 20 cases in which no CADt result was available. Sensitivity and specificity in ICA and M1 segments were 90.7% and 97.9%. Sensitivity decreased to 76.4% when M2 segments were included and to 59.4% when all proximal vessel segments are included. Discussion Real-world testing of two CADt LVO detection algorithms identified gaps in the detection and communication of potentially treatable LVOs when considering vessels beyond the intracranial ICA and M1 segments and in cases with absent and uninterpretable data.
OBJECTIVE:To demonstrate and test the capabilities of the ACR Connect and AI-LAB software platform by implementing multi-institutional artificial intelligence (AI) training and validation for breast density classification. METHODS:In this proof-of-concept study, six US-based hospitals installed Connect and AI-LAB. A breast density algorithm was trained and tested on retrospective mammograms. We recorded time to receive institutional review board approval, to install software locally, and to complete the testing and training. We calculated the performance of the breast density algorithm at each participating hospital and compared it to the performance of a holdout multi-institutional clinical trial testing dataset and a retrospective multi-institutional dataset. We calculated the performance of the locally fine-tuned models on the holdout test datasets. RESULTS:The median time to receive institutional review board approval was 66 days, and the median time to successfully install Connect and AI-LAB locally was 157 days. The median time to complete breast density algorithm testing and training was 216 days. The breast density algorithm performed worse at each hospital than on the holdout test dataset, suggesting poor generalizability of the base model. The fine-tuned models had mixed performance locally and performed poorly on the test dataset. DISCUSSION:In this study, we demonstrate the successful installation and implementation of Connect and AI-LAB software platforms at six facilities using a breast density algorithm. Our results suggest poor generalizability of an algorithm trained on a single dataset and algorithms fine-tuned at individual institutions, emphasizing the hypothetical importance of multi-institutional testing and training.
The correct interpretation of breast density is important in the assessment of breast cancer risk. AI has been shown capable of accurately predicting breast density, however, due to the differences in imaging characteristics across mammography systems, models built using data from one system do not generalize well to other systems. Though federated learning (FL) has emerged as a way to improve the generalizability of AI without the need to share data, the best way to preserve features from all training data during FL is an active area of research. To explore FL methodology, the breast density classification FL challenge was hosted in partnership with the American College of Radiology, Harvard Medical Schools’ Mass General Brigham, University of Colorado, NVIDIA, and the National Institutes of Health National Cancer Institute. Challenge participants were able to submit docker containers capable of implementing FL on three simulated medical facilities, each containing a unique large mammography dataset. The breast density FL challenge ran from June 15 to September 5, 2022, attracting seven finalists from around the world. The winning FL submission reached a linear kappa score of 0.653 on the challenge test data and 0.413 on an external testing dataset, scoring comparably to a model trained on the same data in a central location.
Multicenter clinical trials in radiology have long been a mainstay in our ability to translate clinical research to widespread clinical practice. They help ensure that promising results obtained by a single academic institution will generalize to the diversity of patients, imaging equipment and practice types present in the wider community. Multicenter studies reduce potential biases that may result from single institution results, promote health equity, and inform public policy. Outcomes data from multicenter trials have been used to support favorable recommendations from governmental agencies such as US Preventive Services Task Force and coverage and payment policy decisions from the Centers for Medicare and Medicaid Services (CMS) ( 1 Pisano ED Gatsonis C Hendrick E et al. Diagnostic performance of digital versus film mammography for breast-cancer screening. N Engl J Med. 2005; 353: 1773-1783 Crossref PubMed Scopus (1568) Google Scholar , 2 National Lung Screening Trial Research TeamReduced lung-cancer mortality with low-dose computed tomographic screening. N Engl J Med. 2011; 365: 395-409 Crossref PubMed Scopus (7067) Google Scholar ). The American College of Radiology (ACR) has a long history of supporting multicenter trials in diagnostic radiology and radiation oncology. These trials typically involve standardized protocol development, site management, the development of study data dictionaries and data collection and cleaning, independent interpretation of imaging findings and centralized support to collate, analyze and publish the results. Many of these trials have been funded by governmental agencies such the National Cancer Institute while others have been sponsored by industry and other institutions.
Purpose Medical imaging accounts for 85% of digital health's venture capital funding. As funding grows, it is expected that artificial intelligence (AI) products will increase commensurately. The study's objective is to project the number of new AI products given the statistical association between historical funding and FDA-approved AI products. Methods The study used data from the ACR Data Science Institute and for the number of FDA-approved AI products (2008-2022) and data from Rock Health for AI funding (2013-2022). Employing a 6-year lag between funding and product approved, we used linear regression to estimate the association between new products approved in a certain year, based on the lagged funding (ie, product-year funding). Using this statistical relationship, we forecasted the number of new FDA-approved products. Results The results show that there are 11.33 (95% confidence interval: 7.03-15.64) new AI products for every $1 billion in funding assuming a 6-year lag between funding and product approval. In 2022 there were 69 new FDA-approved products associated with $4.8 billion in funding. In 2035, product-year funding is projected to reach $30.8 billion, resulting in 350 new products that year. Conclusions FDA-approved AI products are expected to grow from 69 in 2022 to 350 in 2035 given the expected funding growth in the coming years. AI is likely to change the practice of diagnostic radiology as new products are developed and integrated into practice. As more AI products are integrated, it may incentivize increased investment for future AI products.
Abstract Objective To develop a free, vendor-neutral software suite, the American College of Radiology (ACR) Connect, which serves as a platform for democratizing artificial intelligence (AI) for all individuals and institutions. Materials and Methods Among its core capabilities, ACR Connect provides educational resources; tools for dataset annotation; model building and evaluation; and an interface for collaboration and federated learning across institutions without the need to move data off hospital premises. Results The AI-LAB application within ACR Connect allows users to investigate AI models using their own local data while maintaining data security. The software enables non-technical users to participate in the evaluation and training of AI models as part of a larger, collaborative network. Discussion Advancements in AI have transformed automated quantitative analysis for medical imaging. Despite the significant progress in research, AI is currently underutilized in current clinical workflows. The success of AI model development depends critically on the synergy between physicians who can drive clinical direction, data scientists who can design effective algorithms, and the availability of high-quality datasets. ACR Connect and AI-LAB provide a way to perform external validation as well as collaborative, distributed training. Conclusion In order to create a collaborative AI ecosystem across clinical and technical domains, the ACR developed a platform that enables non-technical users to participate in education and model development.
PURPOSE:The ACR Data Science Institute conducted its first annual survey of ACR members to understand how radiologists are using artificial intelligence (AI) in clinical practice and to provide a baseline for monitoring trends in AI use over time.METHODS:The ACR Data Science Institute sent a brief electronic survey to all ACR members via email. Invitees were asked for demographic information about their practice and if and how they were currently using AI as part of their clinical work. They were also asked to evaluate the performance of AI models in their practices and to assess future needs.RESULTS:Approximately 30% of radiologists are currently using AI as part of their practice. Large practices were more likely to use AI than smaller ones, and of those using AI in clinical practice, most were using AI to enhance interpretation, most commonly detection of intracranial hemorrhage, pulmonary emboli, and mammographic abnormalities. Of practices not currently using AI, 20% plan to purchase AI tools in the next 1 to 5 years.CONCLUSION:The survey results indicate a modest penetrance of AI in clinical practice. Information from the survey will help researchers and industry develop AI tools that will enhance radiological practice and improve quality and efficiency in patient care.
A core principle of ethical data sharing is maintaining the security and anonymity of the data, and care must be taken to ensure medical records and images cannot be reidentified to be traced back to patients or misconstrued as a breach in the trust between health care providers and patients. Once those principles have been observed, those seeking to share data must take the appropriate steps to curate the data in a way that organizes the clinically relevant information so as to be useful to the data sharing party, assesses the ensuing value of the data set and its annotations, and informs the data sharing contracts that will govern use of the data. Embarking on a data sharing partnership engenders a host of ethical, practical, technical, legal, and commercial challenges that require a thoughtful, considered approach. In 2019 the ACR convened a Data Sharing Workgroup to develop philosophies around best practices in the sharing of health information. This is Part 2 of a Report on the workgroup's efforts in exploring these issues.
The pace of regulatory clearance of artificial intelligence (AI) algorithms for radiology continues to accelerate, and numerous algorithms are becoming available for use in clinical practice. End users of AI in radiology should be aware that AI algorithms may not work as expected when used beyond the institutions in which they were trained, and model performance may degrade over time. In this article, we discuss why regulatory clearance alone may not be enough to ensure AI will be safe and effective in all radiological practices and review strategies available resources for evaluating before clinical use and monitoring performance of AI models to ensure efficacy and patient safety.
Radiology is at the forefront of the artificial intelligence transformation of health care across multiple areas, from patient selection to study acquisition to image interpretation. Needing large data sets to develop and train these algorithms, developers enter contractual data sharing agreements involving data derived from health records, usually with postacquisition curation and annotation. In 2019 the ACR convened a Data Sharing Workgroup to develop philosophies around best practices in the sharing of health information. The workgroup identified five broad domains of activity important to collaboration using patient data: privacy, informed consent, standardization of data elements, vendor contracts, and data valuation. This is Part 1 of a Report on the workgroup's efforts in exploring these issues.
ObjectiveWe developed deep learning algorithms to automatically assess BI-RADS breast density.MethodsUsing a large multi-institution patient cohort of 108,230 digital screening mammograms from the Digital Mammographic Imaging Screening Trial, we investigated the effect of data, model, and training parameters on overall model performance and provided crowdsourcing evaluation from the attendees of the ACR 2019 Annual Meeting.ResultsOur best-performing algorithm achieved good agreement with radiologists who were qualified interpreters of mammograms, with a four-class κ of 0.667. When training was performed with randomly sampled images from the data set versus sampling equal number of images from each density category, the model predictions were biased away from the low-prevalence categories such as extremely dense breasts. The net result was an increase in sensitivity and a decrease in specificity for predicting dense breasts for equal class compared with random sampling. We also found that the performance of the model degrades when we evaluate on digital mammography data formats that differ from the one that we trained on, emphasizing the importance of multi-institutional training sets. Lastly, we showed that crowdsourced annotations, including those from attendees who routinely read mammograms, had higher agreement with our algorithm than with the original interpreting radiologists.ConclusionWe demonstrated the possible parameters that can influence the performance of the model and how crowdsourcing can be used for evaluation. This study was performed in tandem with the development of the ACR AI-LAB, a platform for democratizing artificial intelligence.
Building robust deep learning-based models requires large quantities of diverse training data. In this study, we investigate the use of federated learning (FL) to build medical imaging classification models in a real-world collaborative setting. Seven clinical institutions from across the world joined this FL effort to train a model for breast density classification based on Breast Imaging, Reporting & Data System (BI-RADS). We show that despite substantial differences among the datasets from all sites (mammography system, class distribution, and data set size) and without centralizing data, we can successfully train AI models in federation. The results show that models trained using FL perform 6.3% on average better than their counterparts trained on an institute’s local data alone. Furthermore, we show a 45.8% relative improvement in the models’ generalizability when evaluated on the other participating sites’ testing data.
Purpose To determine in a large multicenter multireader setting the interreader reliability of Liver Imaging Reporting and Data System (LI-RADS) version 2014 categories, the major imaging features seen with computed tomography (CT) and magnetic resonance (MR) imaging, and the potential effect of reader demographics on agreement with a preselected nonconsecutive image set. Materials and Methods Institutional review board approval was obtained, and patient consent was waived for this retrospective study. Ten image sets, comprising 38-40 unique studies (equal number of CT and MR imaging studies, uniformly distributed LI-RADS categories), were randomly allocated to readers. Images were acquired in unenhanced and standard contrast material-enhanced phases, with observation diameter and growth data provided. Readers completed a demographic survey, assigned LI-RADS version 2014 categories, and assessed major features. Intraclass correlation coefficient (ICC) assessed with mixed-model regression analyses was the metric for interreader reliability of assigning categories and major features. Results A total of 113 readers evaluated 380 image sets. ICC of final LI-RADS category assignment was 0.67 (95% confidence interval [CI]: 0.61, 0.71) for CT and 0.73 (95% CI: 0.68, 0.77) for MR imaging. ICC was 0.87 (95% CI: 0.84, 0.90) for arterial phase hyperenhancement, 0.85 (95% CI: 0.81, 0.88) for washout appearance, and 0.84 (95% CI: 0.80, 0.87) for capsule appearance. ICC was not significantly affected by liver expertise, LI-RADS familiarity, or years of postresidency practice (ICC range, 0.69-0.70; ICC difference, 0.003-0.01 [95% CI: -0.003 to -0.01, 0.004-0.02]. ICC was borderline higher for private practice readers than for academic readers (ICC difference, 0.009; 95% CI: 0.000, 0.021). Conclusion ICC is good for final LI-RADS categorization and high for major feature characterization, with minimal reader demographic effect. Of note, our results using selected image sets from nonconsecutive examinations are not necessarily comparable with those of prior studies that used consecutive examination series. © RSNA, 2017.
Purpose To develop diagnostic reference levels (DRLs) and achievable doses (ADs) for the 10 most common adult computed tomographic (CT) examinations in the United States as a function of patient size by using the CT Dose Index Registry. Materials and Methods Data from the 10 most commonly performed adult CT head, neck, and body examinations from 583 facilities were analyzed. For head examinations, the lateral thickness was used as an indicator of patient size; for neck and body examinations, water-equivalent diameter was used. Data from 1 310 727 examinations (analyzed by using SAS 9.3) provided median values, as well as means and 25th and 75th (DRL) percentiles for volume CT dose index (CTDIvol), dose-length product (DLP), and size-specific dose estimate (SSDE). Applicable results were compared with DRLs from eight countries. Results More than 46% of the facilities were community hospitals; 13% were academic facilities. More than 48% were in metropolitan areas, 39% were suburban, and 13% were rural. More than 50% of the facilities performed fewer than 500 examinations per month. The abdomen and pelvis was the most frequently performed examination in the study (45%). For body examinations, DRLs (75th percentile) and ADs (median) for CTDIvol, SSDE, and DLP increased consistently with the patient's size (water-equivalent diameter). The relationships between patient size and DRLs and ADs were not as strong for head and neck examinations. These results agree well with the data from other countries. Conclusion DRLs and ADs as a function of patient size were developed for the 10 most common adult CT examinations performed in the United States. © RSNA, 2017.
PURPOSE To determine radiation dose indexes for computed tomography (CT) performed with renal colic protocols in the United States, including frequency of reduced-dose technique usage and any institutional-level factors associated with high or low dose indexes. MATERIALS AND METHODS The Dose Imaging Registry (DIR) collects deidentified CT data, including examination type and dose indexes, for CT performed at participating institutions; thus, the DIR portion of the study was exempt from institutional review board approval and was HIPAA compliant. CT dose indexes were examined at the institutional level for CT performed with a renal colic protocol at institutions that contributed at least 10 studies to the registry as of January 2013. Additionally, patients undergoing CT for renal colic at a single institution (with institutional review board approval and informed consent from prospective subjects and waiver of consent from retrospective subjects) were studied to examine individual renal colic CT dose index patterns and explore relationships between patient habitus, demographics, and dose indexes. Descriptive statistics were used to analyze dose indexes, and linear regression and Spearman correlations were used to examine relationships between dose indexes and institutional factors. RESULTS There were 49 903 renal colic protocol CT examinations conducted at 93 institutions between May 2011 and January 2013. Mean age ± standard deviation was 49 years ± 18, and 53.9% of patients were female. Institutions contributed a median of 268 (interquartile range, 77-699) CT studies. Overall mean institutional dose-length product (DLP) was 746 mGy ⋅ cm (effective dose, 11.2 mSv), with a range of 307-1497 mGy ⋅ cm (effective dose, 4.6-22.5 mSv) for mean DLPs. Only 2% of studies were conducted with a DLP of 200 mGy ⋅ cm or lower (a "reduced dose") (effective dose, 3 mSv), and only 10% of institutions kept DLP at 400 mGy ⋅ cm (effective dose, 6 mSv) or less in at least 50% of patients. CONCLUSION Reduced-dose renal protocol CT is used infrequently in the United States. Mean dose index is higher than reported previously, and institutional variation is substantial.
The potential risks associated with radiation exposure from medical imaging have received considerable attention in the media recently. In a desire to improve safety, efforts to reduce radiation exposure to appropriate levels are being made by organizations and facilities across the country and the world. But what is the “appropriate” level of radiation for a given examination? In fact, what is the national average level of radiation that is currently being administered by imaging facilities for a particular examination, for example, a CT scan of the head? The answer to this question is not known; but this is precisely the type of question for which the ACR Dose Index Registry (DIR) soon hopes to provide insight. The DIR is part of the ACR’s National Radiology Data Registry (NRDR), an information system that provides accurate and objective measures of practice processes and outcomes. Like other registries that are part of NRDR, the DIR will allow facilities to compare their CT dose indices to peer facilities and national values. The development of the registry has been a lengthy process requiring solutions to a variety of problems, including the standardization of procedure names and data elements, patient privacy, legal issues, technological issues, and vendor competition. Having overcome these hurdles, the ACR intends to launch the DIR in 2011.