In recent years, there has been growing interest in leveraging imaging AI devices to reduce radiologist workload, particularly in screening scenarios. In a rule-out ("believe the negative" or BN) setting, patients deemed negative with high confidence by an AI device could bypass radiologist review, while in a rule-in ("believe the positive" or BP) setting, those identified by AI as highly suspicious would be autonomously recalled. This work proposes a theoretical approach to analyze rule-out, rule-in, and combination ROC curves given the marginals and correlations of the AI predictions and radiologist interpretations. An application to clinical mammography data shows that projected empirical radiologist performance under a rule-out or rule-in scenario is consistent with the theory.
Artificial intelligence (AI) is being deployed within radiology at a rapid pace. AI has proven an excellent tool for reconstructing and enhancing images that appear sharper, smoother, and more detailed, can be acquired more quickly, and allowing clinicians to review them more rapidly. However, incorporation of AI also introduces new failure modes and can exacerbate the disconnect between perceived quality of an image and information content of that image. Understanding the limitations of AI-enabled image reconstruction and enhancement is critical for safe and effective use of the technology. Hence, the purpose of this communication is to bring awareness to limitations when AI is used to reconstruct or enhance a radiological image, with the goal of enabling users to reap benefits of the technology while minimizing risks.
Multiple diagnostic tests are frequently used to determine the presence of a disease condition in patients. In this paper, we use bivariate copulas to examine the properties of receiver operating characteristic (ROC) curves formed when two correlated diagnostic tests are used together to rule-out ("believe the negative") and rule-in ("believe the positive") patients for disease. We use this theory to analyze three mammography data sets where AI devices are applied to reduce radiologists' workload or improve diagnostic performance. Our analysis shows with generality that increasing the radiologist-AI correlation for diseased cases enhances the area under the ROC curve (AUC) of a radiologist-AI rule-out curve, whereas decreasing correlation for non-diseased cases has a similar effect. The opposite trends hold for rule-in scenarios. Applications to clinical mammography data show that projected empirical radiologist performance under a rule-out or rule-in scenario is consistent with the theory.
Artificial intelligence (AI) is being deployed in radiology for image reconstruction and postprocessing to produce images that appear sharper, smoother, and more detailed. However, incorporating AI also introduces new failure modes and can exacerbate the disconnect between the perceived quality of an image and its diagnostic information content. Understanding the limitations of AI-enabled image reconstruction and postprocessing is essential for the safe and effective use of this technology. Therefore, the purpose of this report is to raise awareness of the limitations of AI-based image reconstruction and postprocessing to help users realize the benefits of the technology while minimizing risks. Accordingly, this report reviews approaches to image quality assessment, discusses the regulatory framework relevant to AI-enabled imaging devices, describes AI-specific failure modes, and outlines strategies to mitigate associated risks.
BackgroundAn artificial intelligence (AI)-enabled rule-out device may autonomously remove patient images unlikely to have cancer from radiologist review. Many published studies evaluate this type of device by retrospectively applying the AI to large datasets and use sensitivity and specificity as the performance metrics. However, these metrics have fundamental shortcomings because sensitivity will always be negatively affected in retrospective studies of rule-out applications of AI.MethodWe reviewed 2 performance metrics to compare the screening performance between the radiologist-with-rule-out-device and radiologist-without-device workflows: positive/negative predictive values (PPV/NPV) and expected utility (EU). We applied both methods to a recent study that reported improved performance in the radiologist-with-device workflow using a retrospective US dataset. We then applied the EU method to a European study based on the reported recall and cancer detection rates at different AI thresholds to compare the potential utility among different thresholds.ResultsFor the US study, neither PPV/NPV nor EU can demonstrate significant improvement for any of the algorithm thresholds reported. For the study using European data, we found that EU is lower as AI rules out more patients including false-negative cases and reduces the overall screening performance.ConclusionsDue to the nature of the retrospective simulated study design, sensitivity and specificity can be ambiguous in evaluating a rule-out device. We showed that using PPV/NPV or EU can resolve the ambiguity. The EU method can be applied with only recall rates and cancer detection rates, which is convenient as ground truth is often unavailable for nonrecalled patients in screening mammography.HighlightsSensitivity and specificity can be ambiguous metrics for evaluating a rule-out device in a retrospective setting. PPV and NPV can resolve the ambiguity but require the ground truth for all patients. Based on utility theory, expected utility (EU) is a potential metric that helps demonstrate improvement in screening performance due to a rule-out device using large retrospective datasets.We applied EU to a recent study that used a large retrospective mammography screening dataset from the United States. That study reported an improvement in specificity and decrease in sensitivity when using their AI as a rule-out device retrospectively. In terms of EU, we cannot conclude a significant improvement when the AI is used as a rule-out device.We applied the method to a European study that reported only recall rates and cancer detection rates. Since there is no established EU baseline value in European mammography screening workflow, we estimated the EU baseline using data from previous literature. We cannot conclude a significant improvement when the AI is used as a rule-out device for the European study.In this work, we investigated the use of EU to evaluate rule-out devices using large retrospective datasets. This metric, used with retrospective clinical data, could be used as supporting evidence for rule-out devices.
Objective: To quantify the impact of workflow parameters on time-savings in report turnaround time (TAT) due to an AI-triage device that prioritized pulmonary embolism (PE) in chest CT pulmonary angiography (CTPA) exams. Methods: This retrospective study analyzed 11252 adult CTPA exams conducted for suspected PE at a single tertiary academic medical center. Data was divided into two periods: pre-AI and post-AI. For PE-positive exams, TAT - defined as the duration from patient scan completion to the first preliminary report completion - was compared between the two periods. Time-savings were reported separately for work-hour and off-hour cohorts. To characterize radiologist workflow, 527234 records were retrieved from the PACS and workflow parameters such as exam inter-arrival time and radiologist read-time extracted. These parameters were input into a computational model to predict time-savings following deployment of an AI-triage device and to study the impact of workflow parameters. Results: The pre-AI dataset included 4694 chest CTPA exams with 13.3
Objective: To quantify the impact of workflow parameters on time savings in report turnaround time due to an AI triage device that prioritized pulmonary embolism (PE) in chest CT pulmonary angiography (CTPA) examinations. Methods: This retrospective study analyzed 11,252 adult CTPA examinations conducted for suspected PE at a single tertiary academic medical center. Data was divided into two periods: pre-artificial intelligence (AI) and post-AI. For PE-positive examinations, turnaround time (TAT)- defined as the duration from patient scan completion to the first preliminary report completion-was compared between the two periods. Time savings were reported separately for work-hour and off-hour cohorts. To characterize radiologist workflow, 527,234 records were retrieved from the PACS and workflow parameters such as examination interarrival time and radiologist read time extracted. These parameters were input into a computational model to predict time savings after deployment of an AI triage device and to study the impact of workflow parameters. Results: The pre-AI dataset included 4,694 chest CTPA examinations with 13.3% being PE-positive. The post-AI dataset comprised 6,558 examinations with 16.2% being PE-positive. The mean TAT for pre-AI and post-AI during work hours are 68.9 (95% confidence interval 55.0-82.8) and 46.7 (38.1-55.2) min, respectively, and those during off-hours are 44.8 (33.7-55.9) and 42.0 (33.6-50.3) min. Clinically observed time savings during work hours (22.2 [95% confidence interval: 5.85-38.6] min) were significant (P = .004), while off-hour (2.82 [- 11.1 to 16.7] min) were not (P = .345). Observed time savings aligned with model predictions (29.6 [95% range: 23.2-38.1] min for work hours; 2.10 [1.76, 2.58] min for off-hours). Discussion: Consideration and quantification of the clinical workflow contributes to the accurate assessment of the expected time savings in report TAT after deployment of an AI triage device.
The deployment of multiple AI-triage devices in radiology departments has grown rapidly, yet the cumulative impact on patient wait-times across different disease conditions remains poorly understood. This research develops a comprehensive mathematical and simulation framework to quantify wait-time trade-offs when multiple AI-triage devices operate simultaneously in a clinical workflow. We created multi-QuCAD, a software tool that models complex multi-AI, multi-disease scenarios using queueing theory principles, incorporating realistic clinical parameters including disease prevalence rates, radiologist reading times, and AI performance characteristics from FDA-cleared devices. The framework was verified through four experimental scenarios ranging from simple two-disease workflows to complex nine-disease systems, comparing preemptive versus non-preemptive scheduling disciplines and priority versus hierarchical triage protocols. Analysis of brain imaging workflows demonstrated that while AI-triage devices significantly reduce wait-times for target conditions, they can substantially delay diagnosis of non-targeted, yet urgent conditions. The study revealed that hierarchical protocol generally provides more wait-time savings for the highest-priority conditions compared to the priority protocol, though at the expense of more delays to lower-priority patients with other time-sensitive conditions. The quantitative framework presented provides essential insights for orchestrating multi-AI deployments to maximize overall patient time-saving benefits while minimizing unintended delay for other important patient populations.
We investigated the use of equivalent relative utility (ERU) to evaluate the effectiveness of artificial intelligence (AI)-based rule-out algorithms designed to autonomously remove non-cancer patient images from radiologist review. Two evaluation metrics are explored: positive/negative predictive values and ERU. We applied both methods to a recent US study that concluded an improved specificity by retrospectively applying their AI algorithm to analyze a large mammography dataset. The ERU values are also calculated given the recall and cancer detection rates from a European mammography screening study. Without large prospective studies, ERU may provide insights in the effectiveness of a rule-out algorithm.
Artificial intelligence (AI) is playing an increasingly important role in medicine. We examine the current trends in the evolution of this role and attempt to understand how it may develop over the next two hundred years, as the very nature of human existence may be altered. We are concerned with weaknesses in the AI development and deployment processes, in particular with lapses in the use of AI evaluation methodology and in the tendency of both humans and AI systems to accept unquestioningly the conclusions reached by such systems. We posit the need for a “humble” AI aware of its own limitations and of the limitations of its ambit. Without serious attention to these aspects of AI, it may come to represent an existential threat to our species. We reference examples from popular culture to illustrate these concerns.
In the past decade, artificial intelligence (AI) algorithms have made promising impacts in many areas of healthcare. One application is AI-enabled prioritization software known as computer-aided triage and notification (CADt). This type of software as a medical device is intended to prioritize reviews of radiological images with time-sensitive findings, thus shortening the waiting time for patients with these findings. While many CADt devices have been deployed into clinical workflows and have been shown to improve patient treatment and clinical outcomes, quantitative methods to evaluate the wait-time-savings from their deployment are not yet available. In this paper, we apply queueing theory methods to evaluate the wait-time-savings of a CADt by calculating the average waiting time per patient image without and with a CADt device being deployed. We study two workflow models with one or multiple radiologists (servers) for a range of AI diagnostic performances, radiologist’s reading rates, and patient image (customer) arrival rates. To evaluate the time-saving performance of a CADt, we use the difference in the mean waiting time between the diseased patient images in the with-CADt scenario and that in the without-CADt scenario as our performance metric. As part of this effort, we have developed and also share a software tool to simulate the radiology workflow around medical image interpretation, to verify theoretical results, and to provide confidence intervals for the performance metric we defined. We show quantitatively that a CADt triage device is more effective in a busy, short-staffed reading setting, which is consistent with our clinical intuition and simulation results. Although this work is motivated by the need for evaluating CADt devices, the evaluation methodology presented in this paper can be applied to assess the time-saving performance of other types of algorithms that prioritize a subset of customers based on binary outputs.
One of the main goals in the use of model observers is to serve as an accurate predictor of human observer performance. However, there have been relatively few studies that evaluate model observers in untrained conditions. In this work we evaluate the generalization properties of commonly used model observers using results from a psychophysical study investigating the effect of apodization in low-dose CT from a set of three lesion discrimination tasks. The study involved a total of 24 experimental conditions across three factors (task, system resolution, and apodization). This data allows us to explore the effects of different training regimes on predictive accuracy.We evaluate the Pre-Whitening Matched Filter (PWMF), “Eye-Filtered” Non-Pre-Whitening (NPWE) and Sparse-Channelized Difference-of-Gaussian (SDOG) models for predictive performance, and we compare various training and testing regimens. These include “training” by using reported values from the literature, training and testing on the same set of experimental conditions, and training and testing on different sets of experimental conditions. Of this latter category, we use both leave-one-condition-out for training and testing as well as a leave-one-factor-out strategy, where all conditions with a given factor level are withheld for testing. Our approach may be considered a fixed-reader approach, since we use all available readers for both training and testing.Our results show that training models improves predictive accuracy in these tasks, with predictive errors dropping by a factor of two or more in absolute deviation. However, the fitted models are not fully capturing the effects apodization and other factors in these tasks.
BACKGROUND This study reports the results of a set of discrimination experiments using simulated images that represent the appearance of subtle lesions in low-dose computed tomography (CT) of the lungs. Noise in these images has a characteristic ramp-spectrum before apodization by noise control filters. We consider three specific diagnostic features that determine whether a lesion is considered malignant or benign, two system-resolution levels, and four apodization levels for a total of 24 experimental conditions. PURPOSE The goal of the investigation is to better understand how well human observers perform subtle discrimination tasks like these, and the mechanisms of that performance. We use a forced-choice psychophysical paradigm to estimate observer efficiency and classification images. These measures quantify how effectively subjects can read the images, and how they use images to perform discrimination tasks across the different imaging conditions. MATERIALS AND METHODS The simulated CT images used as stimuli in the psychophysical experiments are generated from high-resolution objects passed through a modulation transfer function (MTF) before down-sampling to the image-pixel grid. Acquisition noise is then added with a ramp noise-power spectrum (NPS), with subsequent smoothing through apodization filters. The features considered are lesion size, indistinct lesion boundary, and a nonuniform lesion interior. System resolution is implemented by an MTF with resolution (10% max.) of 0.47 or 0.58 cyc/mm. Apodization is implemented by a Shepp-Logan filter (Sinc profile) with various cutoffs. Six medically naïve subjects participated in the psychophysical studies, entailing training and testing components for each condition. Training consisted of staircase procedures to find the 80% correct threshold for each subject, and testing involved 2000 psychophysical trials at the threshold value for each subject. Human-observer performance is compared to the Ideal Observer to generate estimates of task efficiency. The significance of imaging factors is assessed using ANOVA. Classification images are used to estimate the linear template weights used by subjects to perform these tasks. Classification-image spectra are used to analyze subject weights in the spatial-frequency domain. RESULTS Overall, average observer efficiency is relatively low in these experiments (10%-40%) relative to detection and localization studies reported previously. We find significant effects for feature type and apodization level on observer efficiency. Somewhat surprisingly, system resolution is not a significant factor. Efficiency effects of the different features appear to be well explained by the profile of the linear templates in the classification images. Increasingly strong apodization is found to both increase the classification-image weights and to increase the mean-frequency of the classification-image spectra. A secondary analysis of "Unapodized" classification images shows that this is largely due to observers undoing (inverting) the effects of apodization filters. CONCLUSIONS These studies demonstrate that human observers can be relatively inefficient at feature-discrimination tasks in ramp-spectrum noise. Observers appear to be adapting to frequency suppression implemented in apodization filters, but there are residual effects that are not explained by spatial weighting patterns. The studies also suggest that the mechanisms for improving performance through the application of noise-control filters may require further investigation.
In the past decade, Artificial Intelligence (AI) algorithms have made promising impacts to transform healthcare in all aspects. One application is to triage patients' radiological medical images based on the algorithm's binary outputs. Such AI-based prioritization software is known as computer-aided triage and notification (CADt). Their main benefit is to speed up radiological review of images with time-sensitive findings. However, as CADt devices become more common in clinical workflows, there is still a lack of quantitative methods to evaluate a device's effectiveness in saving patients' waiting times. In this paper, we present a mathematical framework based on queueing theory to calculate the average waiting time per patient image before and after a CADt device is used. We study four workflow models with multiple radiologists (servers) and priority classes for a range of AI diagnostic performance, radiologist's reading rates, and patient image (customer) arrival rates. Due to model complexity, an approximation method known as the Recursive Dimensionality Reduction technique is applied. We define a performance metric to measure the device's time-saving effectiveness. A software tool is developed to simulate clinical workflow of image review/interpretation, to verify theoretical results, and to provide confidence intervals of the performance metric we defined. It is shown quantitatively that a triage device is more effective in a busy, short-staffed setting, which is consistent with our clinical intuition and simulation results. Although this work is motivated by the need for evaluating CADt devices, the framework we present in this paper can be applied to any algorithm that prioritizes customers based on its binary outputs.
In the last 5 to 10 years, there has been an enormous increase in the interest and use of network models in imaging. These are being considered for numerous imaging applications, including denoising, decision support, learned-feature selection, and many others. Network models "learn" solutions to imaging problems from labelled training data and an elaborate training regime. When a successful model is developed, it represents a computational algorithm for performing some task of interest. But it also encodes a solution to an imaging problem that may be intractable by conventional analytical means. Network models are therefore of interest for how they formulate a solution to a problem of interest. This work focuses on that process. We present two case studies in the analysis of neural networks. The first consists of a denoising network for digital breast tomosynthesis (DBT) images developed using a complex anatomical simulation of breast tissues and realistic x-ray transport physics. The second looks at a lesion detection network, also for DBT images, based on the same anatomical simulation model. For the denoising network, we find that it is very well represented by a linear operation that is effectively a Gaussian convolution kernel. The detection filter appears to be locally linear, but the filter profile appears to depend on what stimulus is used to probe the network. There does not appear to be any clear structure in quadratic components from reverse correlation. Overall, this study shows how regression and reverse-correlation techniques can be used to analyze network models.
Medical image interpretation is central to detecting, diagnosing, and staging cancer and many other disorders. At a time when medical imaging is being transformed by digital technologies and artificial intelligence, understanding the basic perceptual and cognitive processes underlying medical image interpretation is vital for increasing diagnosticians’ accuracy and performance, improving patient outcomes, and reducing diagnostician burnout. Medical image perception remains substantially understudied. In September 2019, the National Cancer Institute convened a multidisciplinary panel of radiologists and pathologists together with researchers working in medical image perception and adjacent fields of cognition and perception for the “Cognition and Medical Image Perception Think Tank.” The Think Tank’s key objectives were to identify critical unsolved problems related to visual perception in pathology and radiology from the perspective of diagnosticians, discuss how these clinically relevant questions could be addressed through cognitive and perception research, identify barriers and solutions for transdisciplinary collaborations, define ways to elevate the profile of cognition and perception research within the medical image community, determine the greatest needs to advance medical image perception, and outline future goals and strategies to evaluate progress. The Think Tank emphasized diagnosticians’ perspectives as the crucial starting point for medical image perception research, with diagnosticians describing their interpretation process and identifying perceptual and cognitive problems that arise. This article reports the deliberations of the Think Tank participants to address these objectives and highlight opportunities to expand research on medical image perception.
A Computer-Aided Triage and Notification (CADt) device uses artificial intelligence (AI) to prioritize radiological medical images and speed up reviews of diseased cases in time-sensitive conditions such as stroke, intercranial hemorrhage, and pneumothorax. However, questions remain on the quantitative assessment of the clinical effectiveness of CADt devices for speeding the review of patient images with time-sensitive conditions. This work presents an analytical method based on queueing theory to quantify the wait-time-savings and to study the impacts of CADt in various clinical settings. Theoretical results are consistent with clinical intuition and are verified by Monte Carlo simulations.
Welcome and Introduction to SPIE Medical Imaging conference 11599: Image Perception, Observer Performance, and Technology Assessment
A recent study reported on an in-silico imaging trial that evaluated the performance of digital breast tomosynthesis (DBT) as a replacement for full-field digital mammography (FFDM) for breast cancer screening. In this in-silico trial, the whole imaging chain was simulated, including the breast phantom generation, the x-ray transport process, and computational readers for image interpretation. We focus on the design and performance characteristics of the computational reader in the above-mentioned trial. Location-known lesion (spiculated mass and clustered microcalcifications) detection tasks were used to evaluate the imaging system performance. The computational readers were designed based on the mechanism of a channelized Hotelling observer (CHO), and the reader models were selected to trend human performance. Parameters were tuned to ensure stable lesion detectability. A convolutional CHO that can adapt a round channel function to irregular lesion shapes was compared with the original CHO and was found to be suitable for detecting clustered microcalcifications but was less optimal in detecting spiculated masses. A three-dimensional CHO that operated on the multiple slices was compared with a two-dimensional (2-D) CHO that operated on three versions of 2-D slabs converted from the multiple slices and was found to be optimal in detecting lesions in DBT. Multireader multicase reader output analysis was used to analyze the performance difference between FFDM and DBT for various breast and lesion types. The results showed that DBT was more beneficial in detecting masses than detecting clustered microcalcifications compared with FFDM, consistent with the finding in a clinical imaging trial. Statistical uncertainty smaller than 0.01 standard error for the estimated performance differences was achieved with a dataset containing approximately 3000 breast phantoms. The computational reader design methodology presented provides evidence that model observers can be useful in-silico tools for supporting the performance comparison of breast imaging systems.