Objective To describe the technical workflow enabling scalable automated artificial intelligence (AI) monitoring in the first national imaging AI registry, Assess-AI. Materials and Methods Large language model (LLM) prompts are developed to extract clinically relevant findings from radiology reports through collaboration between data scientists and subspecialty radiologists. Prompts are optimized using tuning cohorts of use case-specific radiology reports and LLMs available through AWS Bedrock. Such cohorts are used to evaluate prompt accuracy and consistency across 10 repeated runs. Report-AI result pairs submitted to the Assess-AI registry for actively monitored use cases are additionally used to further optimize corresponding prompts. Results Prompts were developed for nine use cases. In Stage 2 development cohorts, final-prompt agreement with hybrid report-derived reference standard labels was 0.985 for ICH and 0.997 for PE. Because these cohorts informed prompt refinement and label construction, they were not independent validation sets. Discussion The workflow demonstrates feasible report-finding extraction at scale; independent accuracy and clinical utility remain unestablished. Conclusion LLM-based extraction within Assess-AI enables scalable, report-anchored AI performance monitoring in radiology.
OBJECTIVE:To describe the design, infrastructure, and functionality of the ACR's Assess-AI registry, a national quality registry created to monitor the real-world performance of clinical imaging artificial intelligence (AI) models. METHODS:Assess-AI is a registry within the National Radiology Data Registry that enables participating facilities to submit de-identified AI output, radiology report text, and DICOM study metadata through the ACR Connect platform. Data are normalized and compared with surrogate labels extracted from radiology reports using large language model-based prompting pipelines. Concordance between AI outputs and extracted surrogate labels is computed centrally, with results delivered through interactive dashboards. Facilities may locally re-identify studies via the Forensics App for quality review. RESULTS:Assess-AI currently supports multiple imaging AI use cases including intracranial hemorrhage, pulmonary embolism, pneumothorax, large-vessel occlusion, bone age, and cervical spine fracture. Participating facilities can visualize data completeness, monitor longitudinal concordance trends, compare performance with registry benchmarks, and explore discordance by demographic or technical factors. Local review workflows allow detailed evaluation of discordant cases, supporting root-cause analysis and AI governance. DISCUSSION:Assess-AI provides a scalable, privacy-preserving framework for postdeployment performance monitoring of imaging AI models. By combining standardized data ingestion, large language model-based surrogate labeling, interactive analytics, and local adjudication, the registry addresses critical gaps in evaluating real-world AI performance. Ongoing expansion will incorporate additional modalities, model types, and risk-adjusted benchmarking to further enhance clinical utility.
BACKGROUND:Prekidney transplant evaluation routinely includes abdominal CT for presurgical vascular assessment. A wealth of body composition data are available from these CT examinations, but they remain an underused source of data, often missing from prognostication models, as these measurements require organ segmentation not routinely performed clinically by radiologists. We hypothesize that artificial intelligence facilitates accurate extraction of abdominal CT body composition data, allowing better prediction of outcomes. METHODS:We conducted a retrospective, single-center observational study of kidney transplant candidates wait-listed between January 1, 2007, and December 31, 2017, with available CT data. Validated deep learning models quantified body composition including fat, aortic calcification, bone density, and muscle mass. Logistic regression was used to compare body composition data to Expected Post-Transplant Survival Score (EPTS) as a predictor of 5-year wait-list mortality. RESULTS:In all, 899 patients were followed for a median 943 days (interquartile range 320-1,697). Of 899, 589 (65.5%) were men and 680 of 899 (75.6%) were White, non-Hispanic. Of 899, 167 patients (18.6%) died while on the waiting list. Myosteatosis (defined as the lowest tertile of muscle attenuation) and increased total aortic and abdominal calcification were associated with increased 5-year wait-list mortality. Logistic regression showed that imaging parameters performed similarly to EPTS at predicting 5-year wait-list mortality (area under receiver operating characteristic curve 0.70 [0.64-0.75] versus 0.67 [0.62-0.72], respectively), and combining body composition parameters with EPTS led to a slight improved survival prediction (area under receiver operating characteristic curve = 0.72, 95% confidence interval 0.66-0.76). CONCLUSIONS:Fully automated quantification of body composition in kidney transplant candidates is feasible. Myosteatosis and atherosclerosis are associated with 5-year wait-list mortality.
BACKGROUND:The PRECISE framework, defined as Patient-Focused Radiology Reports with Enhanced Clarity and Informative Summaries for Effective Communication, leverages GPT-4 to create patient-friendly summaries of radiology reports at a sixth-grade reading level. PURPOSE:The purpose of the study was to evaluate the effectiveness of the PRECISE framework in improving the readability, reliability, and understandability of radiology reports. We hypothesized that the PRECISE framework improves the readability and patient understanding of radiology reports compared to the original versions. MATERIALS AND METHODS:The PRECISE framework was assessed using 500 chest X-ray reports. Readability was evaluated using the Flesch Reading Ease, Gunning Fog Index, and Automated Readability Index. Reliability was gauged by clinical volunteers, while understandability was assessed by non-medical volunteers. Statistical analyses including t-tests, regression analyses, and Mann-Whitney U tests were conducted to determine the significance of the differences in readability scores between the original and PRECISE-generated reports. RESULTS:Readability scores significantly improved, with the mean Flesch Reading Ease score increasing from 38.28 to 80.82 (p-value < 0.001), the Gunning Fog Index decreasing from 13.04 to 6.99 (p-value < 0.001), and the ARI score improving from 13.33 to 5.86 (p-value < 0.001). Clinical volunteer assessments found 95 % of the summaries reliable, and non-medical volunteers rated 97 % of the PRECISE-generated summaries as fully understandable. CONCLUSION:The application of the PRECISE approach demonstrates promise in enhancing patient understanding and communication without adding significant burden to radiologists. With improved reliability and patient-friendly summaries, this approach holds promise for fostering patient engagement and understanding in healthcare decision-making. The PRECISE framework represents a pivotal step towards more inclusive and patient-centric care delivery.
Medical imaging is undergoing a transformation driven by the advent of new, highly effective, machine learning techniques paired with increases in computational capabilities (Cheng et al. 2021; Gilson et al. 2023; Almeida et al. 2024; Krishna et al. 2024). These advanced algorithms have the potential to improve disease detection, diagnosis, prognosis, and treatment outcomes. However, the complexity of machine learning models, the large amounts of curated and annotated data required by some methods, and the potential for bias and error make it challenging for individuals to safely and effectively leverage these methods (Lin et al. 2024; Guo et al. 2024; Xu et al. 2024; Linguraru et al. 2024; Wood et al. 2019). To address these challenges, the American Association of Physicists in Medicine (AAPM), American College of Radiology (ACR), Radiological Society of North America (RSNA), and Society for Imaging Informatics in Medicine (SIIM) have worked together to develop a syllabus detailing a recommended set of competencies for medical imaging professionals interacting with these systems. This guide is aimed at four different personas: users of AI systems, purchasers of AI systems, individuals who provide clinical expertise during the development of AI systems ("clinical collaborators"), and developers of AI systems.1 This is a syllabus, not a curriculum, and is intentional in this scope. Recognizing that individuals may benefit from different presentations of the same material, this work enumerates a series of relevant competencies but does not prescribe, nor offer, a method of instruction (Schuur, Rezazade Mehrizi, and Ranschaert 2021; Garin et al. 2023). By addressing the task-specific demands of each role, this guide will enable medical imaging professionals to utilize machine learning systems more safely and effectively, ultimately improving patient care and outcomes.
This retrospective diagnostic accuracy study compared radiologist-based qualitative assessments and radiomics-based analyses with an automated artificial intelligence (AI)–based volumetric approach for evaluating changes in kidney stone burden on follow-up CT examinations. With institutional review board approval, 157 patients (mean age, 61 ± 13 years; 99 men, 58 women) who underwent baseline and follow-up non-contrast abdomen–pelvis CT for kidney stone evaluation were included. The index test was an automated AI-based whole-kidney and stone segmentation radiomics prototype (Frontier, Siemens Healthineers), which segmented both kidneys and isolated stone volumes using a fixed threshold of 130 Hounsfield units, providing stone volume and maximum diameter per kidney. The reference standard was a threshold-defined volumetric assessment of stone burden change between baseline and follow-up CTs. The radiologist’s performance was assessed using (1) interpretations from clinical radiology reports and (2) an independent radiologist’s assessment of stone burden change (stable, increased, or decreased). Diagnostic accuracy was evaluated using multivariable logistic regression and receiver operating characteristic (ROC) analysis. Automated volumetric assessment identified stable (n = 44), increased (n = 109), and decreased (n = 108) stone burden across the evaluated kidneys. Qualitative assessments from radiology reports demonstrated weak diagnostic performance (AUC range, 0.55–0.62), similar to the independent radiologist (AUC range, 0.41–0.72) for differentiating changes in stone burden. A model incorporating higher-order radiomics features achieved an AUC of 0.71 for distinguishing increased versus decreased stone burdens compared with the baseline CT (p < 0.001), but did not outperform threshold-based volumetric assessment. The automated threshold-based volumetric quantification of kidney stone burdens provides higher diagnostic accuracy than qualitative radiologist assessments and radiomics-based analyses for identifying a stable, increased, or decreased stone burden on follow-up CT examinations.
Generative AI tools have proliferated across the market, garnered significant media attention, and increasingly found incorporation into the radiology practice setting. However, they raise a number of unanswered questions concerning governance and appropriate use. By their nature as general-purpose technologies, they strain the limits of existing FDA premarket review pathways to regulate them and introduce new sources of liability, privacy, and clinical risk. A multilayered governance approach is needed to balance innovation with safety. To address gaps in oversight, this piece establishes a trilaminar governance model for generative AI technologies. This treats federal regulations as a scaffold, upon which tiers of institutional guidelines and industry self-regulatory frameworks are added to create a comprehensive paradigm composed of interlocking parts. Doing so would provide radiologists with an effective risk management strategy for the future, foster continued technical development, and ultimately, promote patient care.
BACKGROUND:Hypertrophic cardiomyopathy (HCM) is associated with significant morbidity and mortality, including sudden cardiac death in the young. Its prevalence is estimated to be 1 in 500, although many people are undiagnosed. The ability to screen electrocardiograms for its presence could improve detection and enable earlier diagnosis. This study evaluated the accuracy of an artificial intelligence device (Viz HCM) in detecting HCM based on a 12-lead electrocardiogram. METHODS:The device was previously trained using deep learning and provides a binary outcome (HCM suspected or not suspected). This study included 293 HCM-positive and 2912 HCM-negative cases, which were selected from 3 hospitals based on chart review incorporating billing diagnostic codes, cardiac imaging, and electrocardiogram features. The device produced an output for 291 (99.3%) HCM-positive and 2905 (99.8%) HCM-negative cases. RESULTS:The device identified HCM with sensitivity of 68.4% (95% CI, 62.8-73.5%), specificity of 99.1% (95% CI, 98.7-99.4%), and area under the curve of 0.975 (95% CI, 0.965-0.982). With assumed population prevalence of 0.002 (1 in 500), the positive predictive value was 13.7% (95% CI, 10.1-19.9%) and the negative predictive value was 99.9% (95% CI, 99.9-99.9%). The device demonstrated broadly consistent performance across demographic and technical subgroups. CONCLUSIONS:The device identified HCM based on a 12-lead electrocardiogram with good performance. Coupled with clinical expertise, it has the potential to augment HCM detection and diagnosis.
We created and validated an open-access AI algorithm (AIc) for assessing image segmentation and patient centering in a multi-body-region, multi-center, and multi-scanner study. Our study included 825 head, chest, and abdomen-pelvis CT from 275 patients (153 females, 128 males; mean age 67 ± 14 years) scanned at five academic and community hospitals. CT images were processed with the AIc to determine vertical and horizontal centering at the skull base (head CT), carina (chest CT), and L2-L3 disc (abdomen CT). We manually measured the vertical and horizontal off-centering. We found strong correlations between AIc and manual estimate of off-centering in both the vertical (head, r = 0.93; chest, r = 0.94; abdomen, and r = 0.95) and horizontal directions (head CT, r = 0.85; chest, r = 0.85; abdomen, r = 0.8) and across age groups (r = 0.70–0.97), gender (r = 0.81–0.96), and multiple scanners from the five sites (r = 0.74–0.99). The AIc area under the receiver operating characteristic curve for centered and off-centered CT exams ranged from 0.72 (head) to 0.99 (chest). Therefore, our study showed that positron-emission tomography/CT (PET/CT) examinations commonly exhibit significant off-centering, particularly with vertical deviations often exceeding 30 mm and horizontal deviations between 10 and 30 mm. In addition, it demonstrated that our AI model can effectively assess both vertical and horizontal off-centering, although it performs better at estimating vertical off-centering.
The Biden Administration recently issued an executive order aimed at establishing a “whole-of-government” regulatory approach to AI products. The Executive Order recognizes the need for interagency coordination when developing frameworks for AI and encourages developers to follow new quality standards. This piece explores the executive order’s provisions and interprets their implications for the radiology space. Developing new radiologic AI tools would now require a more careful balancing act between workflow optimization and risk management, with a host of new required disclosures and mandated design components.
Background Intracranial hemorrhage is a critical finding on computed tomography (CT) of the head. This study compared the accuracy of an artificial intelligence (AI) model (Annalise Enterprise CTB Triage Trauma) to consensus neuroradiologist interpretations in detecting 4 hemorrhage subtypes: acute subdural/epidural hematoma, acute subarachnoid hemorrhage, intra‐axial hemorrhage, and intraventricular hemorrhage. Methods A retrospective stand‐alone performance assessment was conducted on data sets of cases of noncontrast CT of the head acquired between 2016 and 2022 at 5 hospitals in the United States for each hemorrhage subtype. The cases were obtained from patients aged ≥18 years. The positive cases were selected on the basis of the original clinical reports using natural language processing and manual confirmation. The negative cases were selected by taking the next negative case acquired from the same CT scanner after positive cases. Each case was interpreted independently by up to 3 neuroradiologists to establish consensus interpretations. Each case was then interpreted by the AI model for the presence of the relevant hemorrhage subtype. The neuroradiologists were provided with the entire CT study. The AI model separately received thin (≤1.5 mm) and thick (>1.5 and ≤5 mm) axial series as available. Results The 4 cohorts included 571 cases of acute subdural/epidural hematoma, 310 cases of acute subarachnoid hemorrhage, 926 cases of intra‐axial hemorrhage, and 199 cases of intraventricular hemorrhage. The AI model identified acute subdural/epidural hematoma with area under the curve of 0.973 (95% CI, 0.958–0.984) on thin series and 0.942 (95% CI, 0.921–0.959) on thick series; acute subarachnoid hemorrhage with area under the curve 0.993 (95% CI, 0.984–0.998) on thin series and 0.966 (95% CI, 0.945–0.983) on thick series; intraaxial hemorrhage with area under the curve of 0.969 (95% CI, 0.956–0.980) on thin series and 0.966 (95% CI, 0.953–0.976) on thick series; and intraventricular hemorrhage with area under the curve of 0.987 (95% CI, 0.969–0.997) on thin series and 0.983 (95% CI, 0.968–0.994) on thick series. Each finding had at least 1 operating point with sensitivity and specificity >80%. Conclusion The assessed AI model accurately identified intracranial hemorrhage subtypes in this CT data set. Its use could assist the clinical workflow, especially through enabling triage of abnormal CTs.
Purpose: We created an infrastructure for no code machine learning (NML) platform for non-programming physicians to create NML model. We tested the platform by creating an NML model for classifying radiographs for the presence and absence of clavicle fractures. Methods: Our IRB-approved retrospective study included 4135 clavicle radiographs from 2039 patients (mean age 52 +/- 20 years, F:M 1022:1017) from 13 hospitals. Each patient had two-view clavicle radiographs with axial and anterior -posterior projections. The positive radiographs had either displaced or non-displaced clavicle fractures. We configured the NML platform to automatically retrieve the eligible exams using the series' unique identification from the hospital virtual network archive via web access to DICOM Objects. The platform trained a model until the validation loss plateaus. Once the testing was complete, the platform provided the receiver operating characteristics curve and confusion matrix for estimating sensitivity, specificity, and accuracy. Results: The NML platform successfully retrieved 3917 radiographs (3917/4135, 94.7 %) and parsed them for creating a ML classifier with 2151 radiographs in the training, 100 radiographs for validation, and 1666 radiographs in testing datasets (772 radiographs with clavicle fracture, 894 without clavicle fracture). The network identified clavicle fracture with 90 % sensitivity, 87 % specificity, and 88 % accuracy with AUC of 0.95 (confidence interval 0.94 -0.96). Conclusion: A NML platform can help physicians create and test machine learning models from multicenter imaging datasets such as the one in our study for classifying radiographs based on the presence of clavicle fracture.
Importance Automatic generation of the impression section of radiology report can help make radiologists efficient and avoid reporting errors.Objective To evaluate the relationship, content, and accuracy of an Powerscribe Smart Impression (PSI) against the radiologists’ reported findings and impression (RDF).Design, Setting, and Participants The institutional review board approved retrospective study developed and trained an PSI algorithm (Nuance Communications, Inc.) with 9.8 million radiology reports from multiple sites to generate PSI based on information including the protocol name and the radiologists-dictated findings section of radiology reports. Three radiologists assessed 3879 radiology reports of multiple imaging modalities from 8 US imaging sites. For each report, we assessed if PSI can accurately reproduce the RDF in terms of the number of clinically significant findings and radiologists’ style of reporting while avoiding potential mismatch (with the findings section in terms of size, location, or laterality). Separately we recorded the word count for PSI and RDF. Data were analyzed with Pearson correlation and paired t-tests.Main Outcomes and Measures The data were ground truthed by three radiologists. Each radiologists recorded the frequency of the incidental/significant findings, any inconsistency between the RDF and PSI as well as the stylistic evaluation overall evaluation of PSI. Area under the curve (AUC), correlation coefficient, and the percentages were calculated.Results PSI reports were deemed either perfect (91.9%) or acceptable (7.68%) for stylistic concurrence with RDF. Both PSI (mismatched Haller’s Index) and RDF (mismatched nodule size) had one mismatch each. There was no difference between the word counts of PSI (mean 33±23 words/impression) and RDF (mean 35±24 words/impression) (p>0.1). Overall, there was an excellent correlation (r= 0.85) between PSI and RDF for the evolution of findings (negative vs. stable vs. new or increasing vs. resolved or decreasing findings). The PSI outputs (2%) requiring major changes pertained to reports with multiple impression items.Conclusion and Relevance In clinical settings of radiology exam interpretation, the Powerscribe Smart Impression assessed in our study can save interpretation time; a comprehensive findings section results in the best PSI output.### Competing Interest StatementThree coauthors (SA, RB, and SE) are employees of Nuance Communications. Two study coinvestigators (MKK and SRD) have received research grant funding for unrelated projects (Coreline Inc., Riverain Tech, Siemens Healthineers; Qure.AI, Lunit Inc., Vuno Inc.). There was no research grant, fund, or support provided for this study.### Funding StatementThis study did not receive any funding### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:Our retrospective study was approved by the institutional review board at Massachusetts General Brigham (IRB protocol number: 2020P003950) with a waiver of informed consent.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesN/A
Abstract Importance: The Agatston score is a measure of cardiovascular disease traditionally calculated on cardiac gated computed tomography (CT) of the chest. Cardiac gated CT is resource-intensive, can be hard to access, and involves extra radiation exposure. Artificial intelligence (AI) can be used to opportunistically calculate Agatston score on non-gated CTs performed for other indications. Objective: This study compared the accuracy of an AI model (Riverain Technologies ClearRead CT CAC) at calculating Agatston scores on non-gated CTs to both consensus radiologist interpretations on the same CTs and Agatston scores from paired cardiac gated CTs. Design: A retrospective standalone performance assessment was conducted on a dataset of non-contrast CT chest cases acquired between January 2022 and December 2023. Setting: The study was conducted at five hospitals in the United States. Participants: The cohort included non-gated CTs from 491 patients. It was enriched to ensure a representation of disease severity by selecting approximately two-thirds of patients using the originally reported Agatston score on a paired cardiac gated CT within the study timeframe. Main Outcome(s) and Measure(s): The study compared the agreement of Agatston categories (0, 1-99, 100-399 and ≥400) between the AI model and ground truth radiologists or original radiology reports using the quadratic weighted Kappa coefficient. Exposure(s): Each non-gated CT case was interpreted independently by three radiologists to establish consensus interpretations. Each CT was then interpreted by the AI model. The Agatston scores for paired cardiac gated CTs were obtained from original radiology reports. Results: The agreement between the AI model and ground truth radiologists was 0.959 (95% CI: 0.943-0.975). This result was broadly consistent across sex, age group, race, ethnicity and CT scanner manufacturer subgroups. The agreement between the AI model and paired cardiac gated CT was 0.906 (95% CI: 0.882-0.927). Conclusions and Relevance: The assessed AI model accurately calculated Agatston scores on non-gated CTs and produced similar scores to paired cardiac gated CTs. Its use could broaden screening for atherosclerotic cardiovascular disease, enabling opportunistic screening on CTs captured for other indications. ### Competing Interest Statement Source of Funding This study was funded by Riverain Technologies. Riverain Technologies was involved in the design and conduct of the study; preparation, review, and approval of the manuscript; and decision to submit the manuscript for publication. Riverain Technologies was not involved in the collection, management, analysis, and interpretation of the data. Disclosures Authors are employees of Mass General Brigham and/or Massachusetts General Hospital, which had received institutional funding from Riverain Technologies for the study. ### Funding Statement This study was funded by Riverain Technologies. Riverain Technologies was involved in the design and conduct of the study; preparation, review, and approval of the manuscript; and decision to submit the manuscript for publication. Riverain Technologies was not involved in the collection, management, analysis, and interpretation of the data. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved by the Mass General Brigham Institutional Review Board with a waiver of informed consent and ethical approval. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data is not publicly available as it contains Protected Health Information (PHI). We do not have IRB approval for public data-sharing.
Machine learning models can assist clinicians and researchers in many tasks within radiology such as diagnosis, triage, segmentation/measurement, and quality assurance. To better leverage machine learning we have developed a platform that allows users to label data and train models without requiring any programming knowledge. The technology stack consists of a TypeScript web application running on .NET for user interaction, Python, PyTorch, and MONAI for machine learning, DICOM WADO-RS to retrieve data from clinical systems, and Docker for model management. As a first trial of the system, researchers used it to train a model for clavicle fracture detection as part of an IRB-approved retrospective study. The researchers labeled 4,135 clavicle radiographs from 2,039 patients across 13 sites. The platform automatically split the data into training, validation, and test sets and trained a model until the validation loss plateaued. The system then returned a receiver operating characteristic curve, AUC, F1, and other metrics. The resulting model identifies clavicle fractures with 90% sensitivity, 87% specificity, and 88% accuracy with an AUC of 0.95. This model performance is equivalent to or better than similar models reported in the literature. More recently, our system was used to train a model to identify if ultrasound frames that contain personally identifiable information (PII). After validation, the model was used to help de-identify a large dataset that was to be used for research. This first-of-its-kind system streamlines model development and deployment and opens up an exciting new pathway for the use of AI within healthcare.### Competing Interest StatementMannudeep K. Kalra reports a relationship with Siemens Healthineers that includes: funding grants.### Funding StatementThis study did not receive any funding ### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:IRB of Mass General Brigham gave ethical approval for this work. Protocol #2023P000205.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesAll data produced in the present work are contained in the manuscript. Supporting spreadsheets are available upon reasonable request. The imaging data used for training will not be available.
Purpose: To assess the ability of the Annalise Enterprise CXR Triage Trauma (AnnaliseAI Pty Ltd, Sydney, NSW, Australia) artificial intelligence model to identify vertebral compression fractures on chest radiographs and its potential to address undiagnosed osteoporosis and its treatment. Materials and methods: This retrospective study used a consecutive cohort of 596 chest radiographs from four US hospitals between 2015 and 2021. Each radiograph included both frontal (anteroposterior or posteroanterior) and lateral projections. These radiographs were assessed for the presence of vertebral compression fracture in a consensus manner by up to three thoracic radiologists. The model then performed inference on the cases. A chart review was also performed for the presence of osteoporosis-related International Classification of Diseases, 10th revision diagnostic codes and medication use for the study period and an additional year of follow-up. Results: The model successfully completed inference on 595 cases (99.8%); these cases included 272 positive cases and 323 negative cases. The model performed with area under the receiver operating characteristic curve of 0.955 (95% confidence interval [CI]: 0.9390.968), sensitivity 89.3% (95% CI: 85.7%-92.7%) and specificity 89.2% (95% CI: 85.4%-92.3%). Out of the 236 true-positive cases (ie, correctly identified vertebral compression fractures by the model) with available chart information, only 86 (36.4%) had a diagnosis of vertebral compression fracture and 140 (59.3%) had a diagnosis of either osteoporosis or osteopenia; only 78 (33.1%) were receiving a disease-modifying medication for osteoporosis. Conclusion: The model identified vertebral compression fracture accurately with a sensitivity 89.3% (95% CI: 85.7%-92.7%) and specificity of 89.2% (95% CI: 85.4%-92.3%). Its automated use could help identify patients who have undiagnosed osteoporosis and who may benefit from taking disease-modifying medications.
PURPOSE:We compared the performance of generative artificial intelligence (AI) (Augmented Transformer Assisted Radiology Intelligence [ATARI, Microsoft Nuance, Microsoft Corporation, Redmond, Washington]) and natural language processing (NLP) tools for identifying laterality errors in radiology reports and images. METHODS:We used an NLP-based (mPower, Microsoft Nuance) tool to identify radiology reports flagged for laterality errors in its Quality Assurance Dashboard. The NLP model detects and highlights laterality mismatches in radiology reports. From an initial pool of 1,124 radiology reports flagged by the NLP for laterality errors, we selected and evaluated 898 reports that encompassed radiography, CT, MRI, and ultrasound modalities to ensure comprehensive coverage. A radiologist reviewed each radiology report to assess if the flagged laterality errors were present (reporting error-true-positive) or absent (NLP error-false-positive). Next, we applied ATARI to 237 radiology reports and images with consecutive NLP true-positive (118 reports) and false-positive (119 reports) laterality errors. We estimated accuracy of NLP and generative AI tools to identify overall and modality-wise laterality errors. RESULTS:Among the 898 NLP-flagged laterality errors, 64% (574 of 898) had NLP errors and 36% (324 of 898) were reporting errors. The text query ATARI feature correctly identified the absence of laterality mismatch (NLP false-positives) with a 97.4% accuracy (115 of 118 reports; 95% confidence interval [CI] = 96.5%-98.3%). Combined vision and text query resulted in 98.3% accuracy (116 of 118 reports or images; 95% CI = 97.6%-99.0%), and query alone had a 98.3% accuracy (116 of 118 images; 95% CI = 97.6%-99.0%). CONCLUSION:The generative AI-empowered ATARI prototype outperformed the assessed NLP tool for determining true and false laterality errors in radiology reports while enabling an image-based laterality determination. Underlying errors in ATARI text query in complex radiology reports emphasize the need for further improvement in the technology.
BACKGROUND AND PURPOSE:Mass effect and vasogenic edema are critical findings on CT of the head. This study compared the accuracy of an artificial intelligence model (Annalise Enterprise CTB) with consensus neuroradiologists' interpretations in detecting mass effect and vasogenic edema. MATERIALS AND METHODS:A retrospective stand-alone performance assessment was conducted on data sets of noncontrast CT head cases acquired between 2016 and 2022 for each finding. The cases were obtained from patients 18 years of age or older from 5 hospitals in the United States. The positive cases were selected consecutively on the basis of the original clinical reports using natural language processing and manual confirmation. The negative cases were selected by taking the next negative case acquired from the same CT scanner after positive cases. Each case was interpreted independently by up-to-three neuroradiologists to establish consensus interpretations. Each case was then interpreted by the artificial intelligence model for the presence of the relevant finding. The neuroradiologists were provided with the entire CT study. The artificial intelligence model separately received thin (≤1.5 mm) and/or thick (>1.5 and ≤5 mm) axial series. RESULTS:The 2 cohorts included 818 cases for mass effect and 310 cases for vasogenic edema. The artificial intelligence model identified mass effect with a sensitivity of 96.6% (95% CI, 94.9%-98.2%) and a specificity of 89.8% (95% CI, 84.7%-94.2%) for the thin series, and 95.3% (95% CI, 93.5%-96.8%) and 93.1% (95% CI, 89.1%-96.6%) for the thick series. It identified vasogenic edema with a sensitivity of 90.2% (95% CI, 82.0%-96.7%) and a specificity of 93.5% (95% CI, 88.9%-97.2%) for the thin series, and 90.0% (95% CI, 84.0%-96.0%) and 95.5% (95% CI, 92.5%-98.0%) for the thick series. The corresponding areas under the curve were at least 0.980. CONCLUSIONS:The assessed artificial intelligence model accurately identified mass effect and vasogenic edema in this CT data set. It could assist the clinical workflow by prioritizing interpretation of cases with abnormal findings, possibly benefiting patients through earlier identification and subsequent treatment.
The opportunistic use of radiological examinations for disease detection can potentially enable timely management. We assessed if an index created by an AI software to quantify chest radiography (CXR) findings associated with heart failure (HF) could distinguish between patients who would develop HF or not within a year of the examination. Our multicenter retrospective study included patients who underwent CXR without an HF diagnosis. We included 1117 patients (age 67.6 ± 13 years; m:f 487:630) that underwent CXR. A total of 413 patients had the CXR image taken within one year of their HF diagnosis. The rest (n = 704) were patients without an HF diagnosis after the examination date. All CXR images were processed with the model (qXR-HF, Qure.AI) to obtain information on cardiac silhouette, pleural effusion, and the index. We calculated the accuracy, sensitivity, specificity, and area under the curve (AUC) of the index to distinguish patients who developed HF within a year of the CXR and those who did not. We report an AUC of 0.798 (95%CI 0.77–0.82), accuracy of 0.73, sensitivity of 0.81, and specificity of 0.68 for the overall AI performance. AI AUCs by lead time to diagnosis (<3 months: 0.85; 4–6 months: 0.82; 7–9 months: 0.75; 10–12 months: 0.71), accuracy (0.68–0.72), and specificity (0.68) remained stable. Our results support the ongoing investigation efforts for opportunistic screening in radiology.