Background Incorrect artificial intelligence (AI) suggestions can lead to automation bias; however, their impact on medical image interpretation is underresearched. Purpose To assess how incorrect AI suggestions influence the diagnostic accuracy, read times, and visual search behavior of readers interpreting screening mammograms. Materials and Methods In this retrospective multireader paired study conducted between September 2024 and February 2025, 10 National Health Service Breast Screening Programme mammography readers evaluated a test set of two-view mammography screening examinations. The test set included true-positive (TP), false-negative (FN), false-positive (FP), and true-negative (TN) AI suggestions, verified by 3 years of follow-up or histopathologic analysis. In round 1, readers interpreted cases without AI. In round 2, conducted 6 weeks later, a commercially available AI tool was used as decision support, displaying prompts with a region score of 10 or higher (scale, 0-100). Eye-tracking cameras recorded readers' fixations-maintained gaze-over specific image areas. Wilcoxon signed rank tests were used for paired comparisons between rounds, and Kruskal-Wallis tests compared cases with different AI outcomes. Results The test set (n = 60) included cases with 26 TP, 14 FN, 14 FP, and six TN AI suggestions. Median reader sensitivity was lower for cases with FN AI suggestions when reading cases with AI (39%) compared with unassisted reading (71%; P = .002). Reader specificity was higher for cases with FP AI suggestions (39% vs 21%; P = .004). A greater number of visible (TP and FP) AI prompts led to longer median read times, from 25 seconds (zero prompts) to 34 seconds (four or more prompts) (P = .001). Readers fixated less when reviewing cancer cases that AI failed to detect (FN suggestions) compared with unassisted reading (0.44 vs 0.47 fixations per second; P = .03). Shorter fixation durations were observed when readers interpreted cases with FP AI suggestions compared with unassisted reading (0.54 vs 0.56 second; P = .001). Conclusion Incorrect AI suggestions influenced both reader accuracy and visual search behaviors during mammography interpretation. The greatest negative impact was observed with FN AI suggestions; therefore, AI thresholds should be calibrated accordingly. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by Clauser in this issue. See also the editorial by Abbasi and Giess in this issue.
Objectives: The interpretation of digital breast tomosynthesis (DBT) screening examinations is a complex task for an already overstretched workforce which has the potential to increase pressure on readers leading to fatigue and patient safety issues. Studies in non-medical and medical settings have suggested that changes in blink characteristics can reflect fatigue. The purpose of this study is to investigate the use of blink characteristics as an objective marker of fatigue in readers interpreting DBT breast screening examinations. Methods: Twenty-six DBT readers involved in the UK PROSPECTS trial interpreted a test set of 40 DBT cases while being observed by an eye tracking device from November 2019 to February 2021. Raw data from the eye tracker were collected and automated processing software was used to produce eye blinking characteristics data which were analysed using multiple linear regression statistical models. Results: Of the 26 DBT readers recruited, eye tracking data from 23 participants were analysed due to missing data rendering 3 participants’ data uninterpretable. The mean reading time per DBT case was 2.81 min. There was a statistically significant increase in blinking duration of 0.38 ms/case as the reading session progressed (p < 0.0001). This was the result of a significant decrease in the number of ultra-short blinks lasting ≤50 ms (p = 0.0005) and a significant increase in longer blinks lasting 51–100 ms (p = 0.008). Conclusion Changes in blinking characteristics could serve as objective measures of reader fatigue and may prove useful in the development of DBT reading protocols. Advances in knowledge: Blink characteristics can be used as an objective measure of fatigue; however there is limited evidence of their use in radiological settings. Our study suggests that changes in blink duration and frequency could be used to monitor fatigue in DBT reading sessions.
Objectives Digital breast tomosynthesis (DBT) can improve diagnostic accuracy compared to 2D mammography, but DBT reporting is time-consuming and potentially more fatiguing. Changes in diagnostic accuracy and subjective and objective fatigue were evaluated over a DBT reporting session, and the impact of taking a reporting break was assessed. Materials and methods Forty-five National Health Service (NHS) mammography readers from 6 hospitals read a cancer-enriched set of 40 DBT cases whilst eye tracked in this prospective cohort study, from December 2020 to April 2022. Eye-blink metrics were assessed as objective fatigue measures. Twenty-one readers had a reporting break, 24 did not. Subjective fatigue questionnaires were completed before and after the session. Diagnostic accuracy and subjective and objective fatigue measures were compared between the cohorts using parametric and non-parametric significance testing. Results Readers had on average 10 years post-training breast screening experience and took just under 2 h (105.8 min) to report all cases. Readers without a break reported greater levels of subjective fatigue (44% vs. 33%, p = 0.04), which related to greater objective fatigue: an increased average blink duration (296 ms vs. 286 ms, p < 0.001) and a reduced eye-opening velocity (76 mm/s vs. 82 mm/s, p < 0.001). Objective fatigue increased as the trial progressed for the no break cohort only ( p s < 0.001). No difference was identified in diagnostic accuracy between the groups (accuracy: 87% vs. 87%, p = 0.92). Conclusions Implementing a break during a 2-h DBT reporting session resulted in lower levels of subjective and objective fatigue. Breaks did not impact diagnostic accuracy, which may be related to the extensive experience of the readers. Clinical relevance statement DBT is being incorporated into many mammography screening programmes. Recognising that reporting breaks are required when reading large volumes of DBT studies ensures this can be factored in when setting up reading sessions. Trial registration Clinical trials registration number: NCT03733106 Key Points • Use of digital breast tomosynthesis (DBT) in breast screening programmes can cause significant reader fatigue. • The effectiveness of incorporating reading breaks into DBT reporting sessions, to reduce mammography reader fatigue, was investigated using eye tracking. • Integrating breaks into DBT reporting sessions reduced reader fatigue; however, diagnostic accuracy was unaffected.
Digital Pathology (DP) reporting workstations permit eye tracking experiments which can aid our understanding of reading strategies and medical errors in pathology. However, eye tracking with DP slides is complex due to the nature of the slide viewing process: slide panning and zooming. Eye tracking technology records gaze coordinates to a screen surface, but these coordinates do not account for the ever-changing on-screen content (due to slide navigation), and therefore it is essential to track pathologists’ slide navigation to determine where on the slide the pathologist has viewed and what features were fixated. Additionally, visualising the resulting eye tracking data proves challenging due to the zooming component. Other eye tracking studies in DP have accounted for slide navigation by employing custom slide viewers that output slide movements as a data stream with the eye tracking data which are co-registered for analysis. We are conducting a DP eye tracking study using a commercial slide viewer which has been adopted at selected UK hospital sites, but slide movement data cannot be outputted as a data stream in this context. Therefore, we’re developing a software platform using computer vision techniques that can be applied to the recorded screen capture of the DP workstation which is synchronised with the eye tracking data. The developed algorithm could be adapted for use with other commercial slide viewers for future studies. Here, we explore how studies have addressed these issues and we discuss our approach.
Purpose: The introduction of whole slide imaging and digital pathology has enabled greater scrutiny of visual search behaviors among pathologists. We aim to investigate zooming and panning behaviors, external markers of visual processing capabilities, and the changes with experience. Approaches: Twenty digitized breast core needle biopsy histopathology slides were obtained from the circulating slides from the main digital pathology trial (IRAS number: 258799). These were presented to five pathologists with varying experience (1.5 to 40 years) whose examinations were recorded. Data of visual fixations were collected using eye-tracking cameras, and the magnification data and zooming behaviors were extracted in an objective fashion by an automated algorithm. The relationship between experience and metrics was analyzed using mixed-effects regression analyses. Results: There was a significant association between experience and both reading times ( p < 0.001 ) and a number of fixations ( p < 0.001 ), with these relationships being inversely proportional. The greater experience was also associated with greater diagnostic accuracy ( p = 0.033 ). We found that experience was significantly associated with greater use of magnification changes ( p < 0.001 ). Conversely, less experience showed a near significant association with the increased proportion of time spent panning ( p = 0.070 ). Conclusions: Fewer fixations needed to reach a diagnosis and quicker reading times are indicative of greater cognitive and visual processing capabilities with greater experience. These cognitive capabilities may be a prerequisite for the more frequent zooming changes that are more prevalent with increasing experience.
Purpose Digital breast tomosynthesis (DBT) exhibits increased sensitivity and specificity compared to 2D mammography (DM), but DBT images are complex and interpretation takes longer. Clinicians may fatigue or hit a cognitive limit sooner when reading DBT, potentially reducing diagnostic accuracy. Eye blink behaviour was investigated to explore fatigue and cognitive load. Methods Screeners (N=47) from five UK breast screening centres were eye tracked as they read 40 DBT cases (15 normal, 6 benign and 19 malignant), from November 2019-July 2021. Differences in diagnostic accuracy and blink behaviour were analysed over the course of the reading session. Blink rates and case durations were investigated by case malignancy and outcome using T-tests and ANOVAs (α=0.05). Results Blink rates were higher on malignant cases than on normal cases (p=0.004), and blink rates were higher for cases with true positive outcomes than for cases with true negative outcomes (p=0.013). Participants spent less time on malignant cases than normal or benign cases (ps=<0.0001), whilst spending more time on cases with a false positive outcome than on cases with a true negative or true positive outcome (ps<0.0001). No significant difference in blink rate or diagnostic performance by time through reporting session. Conclusion Differences in blink rate and time on case are associated with case malignancy and outcome, potentially reflecting varying cognitive demand and interpretation strategies. Further investigation into blinking during medical image interpretation may identify robust signals of cognition and fatigue that could be used for education and training purposes, whilst indicating optimal screening session duration.
Currently in the UK, a national trial to test the effect of a transition from traditional Full Field Digital Mammography (FFDM) to Digital Breast Tomosynthesis (DBT) is being conducted. DBT, having a higher sensitivity and specificity as compared to FFDM alone, could be a better modality in national breast cancer screening. However, its incorporation in the incredibly busy and detailed UK screening program is difficult. Reading times in DBT have been shown to be longer and strenuous (Connor et al, 2012). Therefore, much research needs to be completed to develop recommendations for its efficiency. One key factor in DBT reading is the progression of fatigue, as both a cause and effect of prolonged reading times. We aimed to develop a program to process real time raw eye tracking data to identify a change in fatigue-state through blink detection. Our focus was on analysing the whole data set and defining blinks through observed events. Two real time signals which the eye tracker generates, namely the left and right 'Eyelid Opening' value, were considered. Through assessment of these signals, blinks of varying duration were identified. Additional parameters such as recorded frame sequences and time stamps were added to the processing to delineate the exact occurrence of these blinks during the reading process. We aim to analyse past and future large DBT eye tracked files, with our processing software, to identify the point of fatigue onset in a DBT reading session.
The UK national screening program for breast cancer currently uses Full Field Digital Mammography (FFDM). Various studies have shown that DBT has a higher sensitivity and specificity in identifying early breast cancer apart from benign pathologies, even in very dense breasts. This potentially makes DBT a better screening modality to detect early breast cancer, as well as minimize false positive recall rates. However, DBT has multiple image slices and thereby makes reading cases inherently a longer and potentially more visually fatiguing task. Our previous studies (Dong et al, 2017 and 2018) have demonstrated the impact of institutional training on reading techniques in DBT. The reading technique itself appears to have an effect on total reading time. In other follow-on studies we have employed eye tracking which gives rise to complex data sets, including parameters such as eyelid opening and pupil diameter measures, which can then be employed to gauge blinks and fatigue onset. Findings from this work have guided changes in our blink identification techniques and we have now developed semi-automated programmed processes which can analyze the large data set and provide a more accurate assessment of fatigue and vigilance parameters through blink detection. Here, we have considered ‘eyelid opening’ parameters of both the left and the right eye separately. Having such a separated approach allowed us to tease out particular aspects of blinking. Similar to Schleicher et al (2008), we found there to be ultra-short blinks (30-50 milli seconds), short blinks (51- 100 msecs), long blinks (101-500 msecs) and also microsleeps (>500 msecs). We argue that the changes observed in the frequencies of these blinks can be used as a measure of vigilance and fatigue during DBT reading.
Radiology: Volume 284: Number 2—August 2017 n radiology.rsna.org 413 1 From the Centre for Medical Imaging, University College London, 3rd Floor East, 250 Euston Rd, London NW1 2PG, England (A.A.P., S.A.T., S.H.); Health and Medical Sciences Group, University of Cumbria, Lancaster, England (P.P.); Department of Primary Care Health Sciences, University of Oxford, Oxford, England (G.S., T.F.); Institute of Applied Health Sciences, University of Birmingham, Birmingham, England (S.M.). Received September 12, 2016; revision requested November 8 and received December 6; accepted January 4, 2017; final version accepted January 12. Address correspondence to A.A.P. (e-mail: andrew.plumb@ ucl.ac.uk).
The 90% CI of mean AUC differences between the exploratory sample and the validation sample corresponding to severe AKI. Dotted line are the equivalence margin with Δ = 20% (wider interval) and Δ = 10% (narrower interval)
[This corrects the article DOI: 10.1186/s41512-016-0001-y.].
INTRODUCTION:To design, implement and evaluate the effect of an educational intervention on student radiographer attitudes across their educational tenure.METHODS:In the first phase, an educational intervention that involved didactic lectures, reflective exercises and simulation suits, aimed at improving student radiographer attitudes towards the older person, was designed and implemented. Kogan's attitudes towards older people (KoP) scale was administrated at five test points; pre-intervention; post-intervention; 6 months post intervention; 12 months post intervention and 24 months post intervention. At the final test point these quantitative data was supplemented with qualitative data for triangulation of the findings.RESULTS:Students held positive attitudes towards older people pre intervention, these increased significantly post intervention (p = 0.01). However, this increase in positive scores was not noted at 6 months and 12-months post intervention. At 24-months post intervention, although there was a slight increase in positive attitudes when compared to the 6 and 12 month scores, this increase was not found to be significant (p = 0.178) CONCLUSION: The results post-intervention suggested that an educational intervention can have a significant impact on student radiographer's attitudes towards older people. However, the qualitative data suggests that experiences on initial clinical placement can be detrimental to attitudinal scores, particularly if the intervention does not include Dementia care strategies.
Purpose To investigate the effect of increasing navigation speed on the visual search and decision making during polyp identification for computed tomography (CT) colonography Materials and Methods Institutional review board permission was obtained to use deidentified CT colonography data for this prospective reader study. After obtaining informed consent from the readers, 12 CT colonography fly-through examinations that depicted eight polyps were presented at four different fixed navigation speeds to 23 radiologists. Speeds ranged from 1 cm/sec to 4.5 cm/sec. Gaze position was tracked by using an infrared eye tracker, and readers indicated that they saw a polyp by clicking a mouse. Patterns of searching and decision making by speed were investigated graphically and by multilevel modeling. Results Readers identified polyps correctly in 56 of 77 (72.7%) of viewings at the slowest speed but in only 137 of 225 (60.9%) of viewings at the fastest speed (P = .004). They also identified fewer false-positive features at faster speeds (42 of 115; 36.5%) of videos at slowest speed, 89 of 345 (25.8%) at fastest, P = .02). Gaze location was highly concentrated toward the central quarter of the screen area at faster speeds (mean gaze points at slowest speed vs fastest speed, 86% vs 97%, respectively). Conclusion Faster navigation speed at endoluminal CT colonography led to progressive restriction of visual search patterns. Greater speed also reduced both true-positive and false-positive colorectal polyp identification. © RSNA, 2017 Online supplemental material is available for this article.
Objective: The aim of this study was to explore reader gaze, performance, and preference during interpretation of cranial computed tomography (cCT) in stack mode at two different sizes. Background: Digital display of medical images allows for the manipulation of many imaging factors, like image size, by the radiologists, yet it is often not known what display parameters better suit human perception. Method: Twenty-one radiologists provided informed consent to be eye tracked while reading 20 cCT cases. Half of these cases were presented at a size of 14 × 14 cm (512 × 512 pixels), half at 28 × 28 cm (1,024 × 1,024 pixels). Visual search, performance, and preference for the two image sizes were assessed. Results: When reading small images, significantly fewer, but longer, fixations were observed, and these fixations covered significantly more slices. Time to first fixation of true positive findings was faster in small images, but dwell time on true findings was longer. Readers made more false positive decisions in small images, but no overall difference in either jackknife alternative free-response receiver operating characteristic or reading time was found. Conclusion: Overall performance is not affected by image size. However, small-stack-mode cCT images may better support the use of motion perception and acquiring an overview, whereas large-stack-mode cCT images seem better suited for detailed analyses. Application: Subjective and eye-tracking data suggest that image size influences how images are searched and that different search strategies might be beneficial under different circumstances.
In this paper a new approach is proposed to track the perceptual behaviour of radiologists when they examine mammographic images displayed on large dual clinical monitors. Zooming and panning are inevitably performed by the radiologist to examine such large images by using the DICOM viewing software. Such image manipulating movements on the target displays makes eye tracking techniques difficult to perform and also the size of the dual clinical monitors makes existing eye tracking techniques generally inadequate. Hence a method using the Smart Eye Pro eye tracker and optical character recognition techniques was designed to relate the recorded radiologists’ eye gaze behaviour on the monitors to the actual zoomed and panned medical image areas. This then allows clinical studies involving radiologists interacting with these mammographic images to be successfully carried out.
OBJECTIVE:To assess the effect of expected abnormality prevalence on visual search and decision-making in CT colonography (CTC). METHODS:13 radiologists interpreted endoluminal CTC fly-throughs of the same group of 10 patient cases, 3 times each. Abnormality prevalence was fixed (50%), but readers were told, before viewing each group, that prevalence was either 20%, 50% or 80% in the population from which cases were drawn. Infrared visual search recording was used. Readers indicated seeing a polyp by clicking a mouse. Multilevel modelling quantified the effect of expected prevalence on outcomes. RESULTS:Differences between expected prevalence were not statistically significant for time to first pursuit of the polyp (median 0.5 s, each prevalence), pursuit rate when no polyp was on screen (median 2.7 s(-1), each prevalence) or number of mouse clicks [mean 0.75/video (20% prevalence), 0.93 (50%), 0.97 (80%)]. There was weak evidence of increased tendency to look outside the central screen area at 80% prevalence and reduction in positive polyp identifications at 20% prevalence. CONCLUSION:This study did not find a large effect of prevalence information on most visual search metrics or polyp identification in CTC. Further research is required to quantify effects at lower prevalence and in relation to secondary outcome measures. ADVANCES IN KNOWLEDGE:Prevalence effects in evaluating CTC have not previously been assessed. In this study, providing expected prevalence information did not have a large effect on diagnostic decisions or patterns of visual search.