
Digital Breast Tomosynthesis (DBT) increases breast cancer detection rates but produces a significantly greater number of images for screeners to read compared to traditional two-dimensional (2-D) mammograms. Putting screeners at risk of fatigue and therefore error in detecting cancers. The aim of this study was to explore if screeners showed differences in subjective fatigue, blink metrics and diagnostic accuracy during a DBT reading session with and without breaks. Prospective study including 45 participants from 6 different hospital sites around England between December 2020 to April 2022. Non-intrusive, screen mounted eye tracking cameras (60Hz sampling rate) were set up in the participant’s natural reading environment. Forty DBT cases were read in a random order (47.5% malignant, 12.5% benign, 40% normal). Each breast was rated as normal or benign (return to screen) or indeterminate, suspicious or highly suspicious (recall). Twenty-one participants had a break at approximately 40 minutes into the session. Participants without a break showed a significantly greater difference in subjective fatigue before and after the reporting session (44% vs 33%, p=0.037). Furthermore, those without breaks exhibited significantly greater blinks per minute (15.75 vs 13.25, p<0.001) and blink duration (milliseconds) (296 vs 286, p<0.001). There was no significant difference in overall accuracy between the cohorts (p=0.921). Blink metrics have the potential to be used in identifying early onset of fatigue during reading sessions.
The final stage in the medical imaging diagnostic system is the radiologist’s interpretation of the images, though research on the factors influencing performance in digital breast tomosynthesis (DBT) is inconclusive. This study seeks to understand the performance of radiologists in reading DBT images and the parameters impacting observer performance in three different countries. The study used a DBT mammogram test to compare the performance of radiologists from Australia, China and Iran in reading thirty-five DBT cases. A range of performance metrics including specificity, sensitivity, lesion sensitivity, ROC AUC and JAFROC FOM were generated for each radiologist upon the conclusion of the test set. The radiologists also provided demographic information relating to their experience in reading digital mammograms and DBT. Each country had a greater percentage of radiologists that have completed a breast imaging fellowship compared to those that have not. Australia had a greater percentage of radiologists that have completed training in DBT reading (Australia=88.2%), while China and Iran had a smaller percentage of radiologists that have not completed training in DBT reading (China=37%, Iran=40%). Significant differences were identified between the three countries in specificity (p=.001), lesion sensitivity (p=.016), ROC (p<.001) and JAFROC (p<.001). Australia had the highest mean value for all performance metrics, while China had the lowest mean value for all performance metrics. Australian radiologists have a moderate positive correlation between lesion sensitivity and the number of years reading DBT images (r=.513, p=.042). Iranian radiologists who read more than 20 DBT cases per week obtained significantly higher performance in lesion sensitivity 73.3% vs. 51.8%; p=.032) than the ones who read less than 20 DBT cases per week.
Deep convolutional neural networks (CNNs) have demonstrated high accuracy in a wide range of computer vision applications, including medical and biological imaging. Many CNNs are fully supervised learning algorithms, and their performance is directly associated with the quality of the training data labels, which are human-defined. In this work, we investigate the fidelity of human-defined truth for cell detection, segmentation, and classification tasks in multiplex microscopy images. We compare manual annotations from human readers on three tasks. Readers were asked to (1) segment all cells in single-channel fluorescence images of a pannuclear stain (DAPI), (2) segment cells in two-channel fluorescence images (CD20/DAPI), only identifying cells with both nuclear signal (DAPI) and signal from a cell surface marker (CD20), and (3) segment two separate cell classes in three-channel fluorescence images (CD3/CD4/DAPI). In this third task, readers were asked to identify cells that had nuclear signal and were CD3+/CD4- and CD3+/CD4+. By comparing these manual segmentations within and between readers, we demonstrate that human readers show the least variability in single-channel DAPI segmentation (p<<0.05, F test for equal variance). We also compared the agreement of human readers with one another to the agreement of an object-detection network, Yolov5, on cell detection in DAPI images. All pairwise comparisons of human readers with other human readers yielded an average F1-score of 0.83±0.14, and comparisons of Yolov5 with human readers yielded an average F1-score of 0.84±0.12 (p=0.26, Welch’s T test). We therefore demonstrate that out of the provided tasks, DAPI detection provides the highest fidelity ground truth, and were unable to show a difference between Yolov5 and human readers in this task.
Computer-aided diagnostic schemes have been developed with the primary aim of supporting diagnostic accuracy in the investigation of mammography images. A robust mammograms database is a priority requirement to assist in testing the effectiveness of techniques associated with these schemes. However, such datasets–with information on radiological and biopsy reports, different types of findings and good quality images-are difficult to be available, mainly due to restrictions of different radiology centers and hospitals or limited to the research team involved. Aiming to overcome these issues, we are developing an alternative based on images of a structured breast phantom, previously developed in our group. Although having been validated in terms of their physical characteristics, the investigation on the images produced from this phantom exposure on a digital mammography system is an important task with regard to their appearance compared to actual breasts images. For this purpose, a software was developed for managing comparative tests between images from actual breasts and from this phantom, considering internal regions of interest. The intention is to evaluate how much the simulated and real images are confused according to the human visual perception of different observers, seeking to validate the use of this breast phantom for structuring a new mammography images database aimed at evaluation tests of CAD schemes in mammography. Results were promising showing a variation of 50-60% in the rate of correct answers among observers, indicating a reasonable level of confusion between the two types of images.
Previous studies reported that the cancer subtypes radiologists struggling to detect successfully varied across countries in mammography interpretation. However, little is known whether such variation is also in radiologists’ perception of local cancer-free areas. This study compared the cancer-free areas incorrectly flagged as cancer by radiologists from two populations in reading dense screening mammograms. We collected reading data from 20 Chinese and 16 Australian radiologists who previously evaluated 60 dense screening cases. For each cohort, findings from all readers were pooled together, and the local cancer-free areas classified as cancer were identified. Particularly the areas misclassified by readers from both cohorts were recognized and displayed on the mammograms as overlaps. For each overlap, we counted the error rate, the proportion of readers failing to distinguish between normality and abnormality, as a measure of the actual difficulty level for each reader cohort. Afterward, the Spearman correlation was performed to explore whether the calculated cohort-specific difficulty levels were correlated. A similar analysis was conducted on two geographically-distant groups within China. Results showed that between Chinese and Australian radiologists, the correlation was only found in the cancer-free views of cancer cases (r=0.902, p=0.004). However, between the two groups within China, we found strong correlations in both cancer-containing (r=0.833, p=0.333) and cancer-free views (r=0.955, p=0.022) of cancer cases, despite an insignificant correlation in normal cases. In conclusion, radiologists from different populations display various error-making patterns in reading dense screening mammograms, while those with similar demographic characteristics share the diagnosis to a certain degree.
Multiple imaging modalities are commonly jointly used for investigating biological phenomena or diagnosis purposes. In this study, we propose a deep-learning-based cross-modality imaging technique that utilizes one imaging modality to computationally predict another. A novel neural network architecture, featuring recurrent multi-stage refinement controlled by gated activation, was developed for this purpose. To demonstrate the effectiveness of the proposed method, we conducted experiments on predicting organelle fluorescence images from stimulated Raman scattering (SRS) imaging. The results of the experiments indicate that our method outperforms the current state-of-the-art techniques across multiple datasets, in terms of both accuracy and efficiency. The neural network architecture was able to produce high-quality predictions with clear boundaries and high prediction accuracy through the multi-stage refinement process. The proposed method presents a versatile framework that addresses the limitations of current deep-learning-enabled cross-modality image prediction techniques and has potential applications in the field of medical and biological imaging.
Tools for computer-aided diagnosis based on deep learning have become increasingly important in the medical field. Such tools can be useful, but require effective communication of their decision-making process in order to safely and meaningfully guide clinical decisions. Inherently interpretable models provide an explanation for each decision that matches their internal decision-making process. We present a user interface that incorporates the Interpretable AI Algorithm for Breast Lesions (IAIA-BL) model, which interpretably predicts both mass margin and malignancy for breast lesions. The user interface displays the most relevant aspects of the model’s explanation including the predicted margin value, the AI confidence in the prediction, and the two most highly activated prototypes for each case. In addition, this user interface includes full-field and cropped images of the region of interest, as well as a questionnaire suitable for a reader study. Our preliminary results indicate that the model increases the readers’ confidence and accuracy in their decisions on margin and malignancy.
Cerebral vascular abnormalities, such as unruptured intracranial aneurysms (UIAs), can result in poor clinical outcomes or death. A statistical atlas of the cerebral vasculature based on healthy controls can be used to identify abnormalities in patients. This atlas was registered to 19 patients with an UIA and 18 healthy controls. Z-score values were computed for the vessels in the Circle of Willis, indicating if and where the vessel radius deviates from the atlas. A color-coded Z-score map was made, to aid radiologists in identifying cerebral vascular abnormalities. Results showed that patients with UIAs have statistically significantly higher Z-score values than healthy controls. In addition, in 17/19 patients the Z-score values at the location of the UIA were statistically significantly higher than elsewhere in the Circle of Willis. This indicates that this technique can both be used to identify patients-at-risk with an abnormal cerebral vasculature, as well as guide visual search for UIAs with 89 % detection sensitivity.
Model observers are one of the most common tools to evaluate signal detectability on medical images. There is a large literature of model observers applied to simple signals such as disks or Gaussian blobs. However, these signals do not represent realistic scenarios where lesion shapes are irregular and lesion location is unknown. For instance, in breast imaging, masses can be elongated, spiculated or bumpy. We study how different model observers perform on irregularly shaped signals including the Non-Pre-Whitening (NPW) and two variations of the Channelized Hotelling Observer (CHO), one with symmetrical Laguerre-Gauss channels and one with adaptive Laguerre-Gauss channels. The novel adaptive Laguerre-Gauss channels are built taking into account the shape of the lesion by modifying the distance matrix used to build channels that adopt the shape of the signal. We embedded four different signal shapes in 2D realistic backgrounds in a digital mammography image generated using the VICTRE in silico breast imaging system and studied the models' performance for detection and search. Our results show that the adaptive channels perform better than symmetrical channels with the most irregular signal (a star and an elongated mass).
Medical imaging systems are commonly assessed and optimized by the use of objective measures of image quality (IQ). The performance of the ideal observer (IO) acting on imaging measurements has long been advocated as a figure-of-merit to guide the optimization of imaging systems. For computed imaging systems, the performance of the IO acting on imaging measurements also sets an upper bound on task-performance that no image reconstruction method can transcend. As such, estimation of IO performance can provide valuable guidance when designing under-sampled data-acquisition techniques by enabling the identification of designs that will not permit the reconstruction of diagnostically inappropriate images for a specified task - no matter how advanced the reconstruction method is or how plausible the reconstructed images appear. The need for such analysis is urgent because of the substantial increase of medical device submissions on deep learning-based image reconstruction methods and the fact that they may produce clean images disguising the potential loss of diagnostic information when data is aggressively under-sampled. Recently, convolutional neural network (CNN) approximated IOs (CNN-IOs) was investigated for estimating the performance of data space IOs to establish task-based performance bounds for image reconstruction, under an X-ray computed tomographic (CT) context. In this work, the application of such data space CNN-IO analysis to multi-coil magnetic resonance imaging (MRI) systems has been explored. This study utilized stylized multi-coil sensitivity encoding (SENSE) MRI systems and deep-generated stochastic brain models to demonstrate the approach. Signal-known-statistically and background-known-statistically (SKS/BKS) binary signal detection tasks were selected to study the impact of different acceleration factors on the data space IO performance.
MIDRC was created to facilitate machine learning research for tasks including early detection, diagnosis, prognosis, and assessment of treatment response related to the COVID-19 pandemic and beyond. The purpose of the Technology Development Project (TDP) 3c is to create resources to assist researchers in evaluating the performance of their machine learning algorithms. An interactive decision tree has been developed, organized by the type of task that the machine learning algorithm is being trained to perform. The user can select information such as: (a) the type of task, (b) the nature of the reference standard, and (c) the type of the algorithm output. Based on the user responses, they can obtain recommendations regarding appropriate performance evaluation approaches and metrics, including literature references, short video tutorials, and links to available software. Five tasks have been identified for the decision tree: (a) classification, (b) detection/localization, (c) segmentation, (d) time-to-event analysis, and (e) estimation. As an example, the classification branch of the decision tree includes binary and multi-class classification tasks and provides suggestions for methods and metrics as well as software recommendations, and literature references for situations where the algorithm produces either binary or non-binary (e.g., continuous) output and for reference standards with negligible or non-negligible variability and unreliability. The decision tree has been made publicly available on the MIDRC website to assist researchers in conducting task-specific performance evaluations, including classification, detection/localization, segmentation, estimation, and time-to-event tasks.
Purpose: Blinded independent central review is recommended by the US FDA for registration of oncology trials as it provides bias-free image assessment and avoidance of potential unblinding of patient data. Double read with adjudication is a highly advocated review model used in such trials. Disagreement between readers is natural and inevitable. Radiological disagreement rates or Adjudication Rate (AR) of 30–65% are reported by several papers since 1959 for different oncologic indications. The aim of the study is to develop and use an algorithm to identify reader pair with predicted high AR and investigate if overall study AR can be kept constant or improved by assigning less cases to a reader pair with high AR. Methods: A retrospective analysis was performed of 285 subjects with 3351 post-baseline timepoints reviewed by board-certified radiologist reviewers using Response Evaluation Criteria in Solid Tumors (RECIST) 1.1 criteria in a BICR set up. The reader adjudication rate was calculated and analyzed throughout the duration of review. The distribution of cases per reader and distribution of cases per reader pair was calculated and overall study AR, and each reader pair AR were calculated. Data Analysis Methods: Data was prepared and analyzed using linear regression with MS Excel and R programming script (R version 4.1.2 (2021-11-01) -- "Bird Hippie" augmented by RStudio 2022.07.0+548 "Spotted Wakerobin". Results: Using the data at completion of 50% reads, predicted AR per reader pair was found to correlate well for five out of six reader pairs on the study. This predicted AR can then be used to assign more cases to reader pair with low AR and less cases to reader pair with high AR to keep study level variability low. Conclusions: Predictive modeling of AR using linear regression can provide an insight into variability of each reader pair, which in turn determines the AR of that reader pair and collectively determines study AR. However, it is still not clear to what level of prioritization in case assignment to specific pairs can be considered acceptable and not artificial.
Chest x-ray radiography (CXR) is widely used in screening and detecting lung diseases. However, reading CXR images is often difficult resulting in diagnostic errors and inter-reader variability. To address this clinical challenge, a Multi-task, Optimal-recommendation, and Max-predictive Classification and Segmentation (MOM-ClaSeg) system is developed to detect and delineate different abnormal regions of interest (ROIs) on CXR, make multiple recommendations of abnormalities sorted by the generated probability scores, and automatically generate diagnostic report. MOM-ClaSeg consists of convolutional neural networks to generate a detection, finer-grained segmentation and prediction score for each ROI based on augmented MaskRCNN framework, and multi-layer perception neural networks to fuse results to generate the optimal recommendation for each detected ROI based on decision fusion framework. Total of 310,333 adult CXR containing 67,071 normal and 243,262 abnormal images depicting 307,415 confirmed ROIs of 65 different abnormalities were assembled as to train MOM-ClaSeg. An independent 22,642 CXR was assembled to test MOMClaSeg. Radiologists detected 6,646 ROIs that depict 43 different types of abnormalities on 4,068 CXR images. Comparing with radiologists’ detection results, MOM-ClaSeg system detected 6,009 true-positive ROIs and 6,379 false-positive ROIs, which represents 90.3% sensitivity and 0.28 false-positive ROIs per image. For the eight common diseases, the computed areas under ROC curves ranged from 0.880 to 0.988. Additionally, 70.4% of MOM-ClaSeg system-detected abnormalities along with system-generated diagnostic reports were directly accepted by radiologists. This study presents the first AI-based multi-task prediction system to detect different abnormalities and generate diagnostic reports to assist radiologists accurately and/or efficiently detecting lung diseases.
Artificial intelligence (AI)-based methods are showing substantial promise in segmenting oncologic positron emission tomography (PET) images. For clinical translation of these methods, assessing their performance on clinically relevant tasks is important. However, these methods are typically evaluated using metrics that may not correlate with the task performance. One such widely used metric is the Dice score, a figure of merit that measures the spatial overlap between the estimated segmentation and a reference standard (e.g., manual segmentation). In this work, we investigated whether evaluating AI-based segmentation methods using Dice scores yields a similar interpretation as evaluation on the clinical tasks of quantifying metabolic tumor volume (MTV) and total lesion glycolysis (TLG) of primary tumor from PET images of patients with non-small cell lung cancer. The investigation was conducted via a retrospective analysis with the ECOG-ACRIN 6668/RTOG 0235 multi-center clinical trial data. Specifically, we evaluated different structures of a commonly used AI-based segmentation method using both Dice scores and the accuracy in quantifying MTV/TLG. Our results show that evaluation using Dice scores can lead to findings that are inconsistent with evaluation using the task-based figure of merit. Thus, our study motivates the need for objective task-based evaluation of AI-based segmentation methods for quantitative PET.
Deep-learning (DL)-based methods have shown significant promise in denoising myocardial perfusion SPECT images acquired at low dose. For clinical application of these methods, evaluation on clinical tasks is crucial. Typically, these methods are designed to minimize some fidelity-based criterion between the predicted denoised image and some reference normal-dose image. However, while promising, studies have shown that these methods may have limited impact on the performance of clinical tasks in SPECT. To address this issue, we use concepts from the literature on model observers and our understanding of the human visual system to propose a DL-based denoising approach designed to preserve observer-related information for detection tasks. The proposed method was objectively evaluated on the task of detecting perfusion defect in myocardial perfusion SPECT images using a retrospective study with anonymized clinical data. Our results demonstrate that the proposed method yields improved performance on this detection task compared to using low-dose images. The results show that by preserving task-specific information, DL may provide a mechanism to improve observer performance in low-dose myocardial perfusion SPECT.
Digital Pathology (DP) reporting workstations permit eye tracking experiments which can aid our understanding of reading strategies and medical errors in pathology. However, eye tracking with DP slides is complex due to the nature of the slide viewing process: slide panning and zooming. Eye tracking technology records gaze coordinates to a screen surface, but these coordinates do not account for the ever-changing on-screen content (due to slide navigation), and therefore it is essential to track pathologists’ slide navigation to determine where on the slide the pathologist has viewed and what features were fixated. Additionally, visualising the resulting eye tracking data proves challenging due to the zooming component. Other eye tracking studies in DP have accounted for slide navigation by employing custom slide viewers that output slide movements as a data stream with the eye tracking data which are co-registered for analysis. We are conducting a DP eye tracking study using a commercial slide viewer which has been adopted at selected UK hospital sites, but slide movement data cannot be outputted as a data stream in this context. Therefore, we’re developing a software platform using computer vision techniques that can be applied to the recorded screen capture of the DP workstation which is synchronised with the eye tracking data. The developed algorithm could be adapted for use with other commercial slide viewers for future studies. Here, we explore how studies have addressed these issues and we discuss our approach.
Deep learning algorithms for detection and segmentation have been shown to be vulnerable to single-pixel attacks. These attacks can lead to catastrophic failure of the deep learning algorithm. In the case of biomedical imaging, this can result in significant damage to clinical outcomes. While single-pixel attacks have been studied within the field of digital pathology, they have yet to be studied within the realm of radiology, in particular with volumetric U-Net or V-Net architectures. In this work, we demonstrated that using gradcam++, we could identify vulnerable voxels for the single-voxel attacks that were slightly negative in value towards the boundary of kidney segmentations that lead to the significant distortion of the output kidney classification. Figure 1 demonstrates the graphical abstract for this work.
Modern generative models, such as generative adversarial networks (GANs), hold tremendous promise for several applications in medical imaging that include unconditional medical image synthesis, image translation, and optimization of imaging systems. However, the extent to which a GAN learns image statistics that are relevant to a diagnostic task is unknown. In this work, canonical stochastic image models (SIMs) that simulate realistic mammographic textures are employed to evaluate GAN-based SIMs with respect to detection, detection-localization, and detection-estimation tasks. It is shown that the specific GAN architecture considered has higher propensity to generate statistics that confound the observers performing the three considered tasks. This work highlights the need for continued development of objective metrics for evaluating GANs.
Purpose: Variability in observer performance BICR is common but not well understood and various measures like AR, AAR, RDI help quantify it which leads to multiple complex data points. Network analysis uses mathematically based algorithms to characterize the components of a network of entities and identifying, visualizing, and analysing their relationships. In a network, variables are represented by nodes, the relationships represented by edges between these nodes. The visualization technique involves mapping relationships among entities based on the symmetry or asymmetry of data. Maps from data points generated during double read adjudication study can provide the performance of each reader pair primarily based on AAR. Methods: Adjudication data from four oncology clinical trials with 2163 subjects, 16937 post-baseline responses was analyzed. Performance metrics included number of cases, adjudication rate, adjudication agreement rate for each read and reader pair. The data were aggregated and prepared for network analysis in Python-a high-level, cross-platform, and open-sourced programming language released under a GPL-compatible license. Python Software Foundation (PSF), a non-profit organization, holds the copyright. Url-https://www.python.org Version 3.9.0 Results: This graphic visualization provides simplistic organization of a complicated data analysis and supports the quality monitoring process of independent reviews. The tool provides a snapshot of the review performance of all the readers in the trial allowing the study team to investigate and intervene in a timely manner with the intent of supporting robust and accurate data analysis. Conclusions: Network analysis plots for reader performance metrics in BICR provide excellent visual mapping to interpret multiple critical metrics in a single plot which would otherwise require multiple plots and tables. Timely review of these plots during the trial can help demonstrate the effectiveness of interventions as well.
One of the main goals in the use of model observers is to serve as an accurate predictor of human observer performance. However, there have been relatively few studies that evaluate model observers in untrained conditions. In this work we evaluate the generalization properties of commonly used model observers using results from a psychophysical study investigating the effect of apodization in low-dose CT from a set of three lesion discrimination tasks. The study involved a total of 24 experimental conditions across three factors (task, system resolution, and apodization). This data allows us to explore the effects of different training regimes on predictive accuracy.We evaluate the Pre-Whitening Matched Filter (PWMF), “Eye-Filtered” Non-Pre-Whitening (NPWE) and Sparse-Channelized Difference-of-Gaussian (SDOG) models for predictive performance, and we compare various training and testing regimens. These include “training” by using reported values from the literature, training and testing on the same set of experimental conditions, and training and testing on different sets of experimental conditions. Of this latter category, we use both leave-one-condition-out for training and testing as well as a leave-one-factor-out strategy, where all conditions with a given factor level are withheld for testing. Our approach may be considered a fixed-reader approach, since we use all available readers for both training and testing.Our results show that training models improves predictive accuracy in these tasks, with predictive errors dropping by a factor of two or more in absolute deviation. However, the fitted models are not fully capturing the effects apodization and other factors in these tasks.