Medical image interpretation is central to detecting, diagnosing, and staging cancer and many other disorders. At a time when medical imaging is being transformed by digital technologies and artificial intelligence, understanding the basic perceptual and cognitive processes underlying medical image interpretation is vital for increasing diagnosticians’ accuracy and performance, improving patient outcomes, and reducing diagnostician burnout. Medical image perception remains substantially understudied. In September 2019, the National Cancer Institute convened a multidisciplinary panel of radiologists and pathologists together with researchers working in medical image perception and adjacent fields of cognition and perception for the “Cognition and Medical Image Perception Think Tank.” The Think Tank’s key objectives were to identify critical unsolved problems related to visual perception in pathology and radiology from the perspective of diagnosticians, discuss how these clinically relevant questions could be addressed through cognitive and perception research, identify barriers and solutions for transdisciplinary collaborations, define ways to elevate the profile of cognition and perception research within the medical image community, determine the greatest needs to advance medical image perception, and outline future goals and strategies to evaluate progress. The Think Tank emphasized diagnosticians’ perspectives as the crucial starting point for medical image perception research, with diagnosticians describing their interpretation process and identifying perceptual and cognitive problems that arise. This article reports the deliberations of the Think Tank participants to address these objectives and highlight opportunities to expand research on medical image perception.
Cancer‐related cognitive impairments (CRCI) are frequently reported among cancer survivors, and attention is the most frequently assessed cognitive domain in CRCI. However, there is no consensus as to whether attention is impaired. We suggest that a major reason for this lack of agreement is a lack of construct validity for neuropsychological attention tests. We propose to assess the construct validity of neuropsychological attention tests with respect to experimental paradigms from cognitive psychology.
We investigated whether standardized neuropsychological tests and experimental cognitive paradigms measure the same cognitive faculties. Specifically, do neuropsychological tests commonly used to assess attention measure the same construct as attention paradigms used in cognitive psychology and neuroscience? We built on the "general attention factor", comprising several widely used experimental paradigms (Huang et al., 2012). Participants (n = 636) completed an on-line battery (TestMyBrain.org) of six experimental tests [Multiple Object Tracking, Flanker Interference, Visual Working Memory, Approximate Number Sense, Spatial Configuration Visual Search, and Gradual Onset Continuous Performance Task (Grad CPT)] and eight neuropsychological tests [Trail Making Test versions A & B (TMT-A, TMT-B), Digit Symbol Coding, Forward and Backward Digit Span, Letter Cancellation, Spatial Span, and Arithmetic]. Exploratory factor analysis in a subset of 357 participants identified a five-factor structure: (1) attentional capacity (Multiple Object Tracking, Visual Working Memory, Digit Symbol Coding, Spatial Span), (2) search (Visual Search, TMT-A, TMT-B, Letter Cancellation); (3) Digit Span; (4) Arithmetic; and (5) Sustained Attention (GradCPT). Confirmatory analysis in 279 held-out participants showed that this model fit better than competing models. A hierarchical model where a general cognitive factor was imposed above the five specific factors fit as well as the model without the general factor. We conclude that Digit Span and Arithmetic tests should not be classified as attention tests. Digit Symbol Coding and Spatial Span tap attentional capacity, while TMT-A, TMT-B, and Letter Cancellation tap search (or attention-shifting) ability. These five tests can be classified as attention tests.
Interoception refers to the representation of the internal states of an organism, and includes the processes by which it senses, interprets, integrates, and regulates signals from within itself. This review presents a unified research framework and attempts to offer definitions for key terms to describe the processes involved in interoception. We elaborate on these definitions through illustrative research findings, and provide brief overviews of central aspects of interoception, including the anatomy and function of neural and non-neural pathways, diseases and disorders, manipulations and interventions, and predictive modeling. We conclude with discussions about major research gaps and challenges.
Objectives: Rotating shift work is associated with adverse outcomes due to circadian misalignment, sleep curtailment, work-family conflicts, and other factors. We tested a bright light countermeasure to enhance circadian adaptation on a counterclockwise rotation schedule. Methods: Twenty-nine adults (aged 20–40 years; 15 women) participated in a 4-week laboratory simulation with weekly counterclockwise transitions from day, to night, to evening, to day shifts. Each week consisted of five 8-hour workdays including psychomotor vigilance tests, two days off, designated 8-hour sleep episodes every day, and an assessment of circadian melatonin secretion. Participants were randomized to a treatment group (N=14), receiving intermittent bright light during work designed to facilitate circadian adaptation, or a control group (N=15) working in indoor light. Adaptation was measured by how much of the melatonin secretion episode overlapped with scheduled sleep timing. Results: On the last night shift, there was a greater overlap between melatonin secretion and scheduled sleep time in the treatment group [mean 4.90, standard deviation (SD) 2.8 hours] compared to the control group (2.62, SD 2.8 hours; P=0.002), with night shift adaptation strongly influenced by baseline melatonin timing (r2= -0.71, P=0.01). While the control group exhibited cognitive deficits on the last night shift, the treatment group’s cognitive deficits on the last night and evening shifts were minimized. Conclusions: In this laboratory setting, intermittent bright light during work hours enhanced adaptation to night work and subsequent readaptation to evening and day work. Light regimens scheduled to shift circadian timing should be tested in actual shift workers on counterclockwise schedules as a workplace intervention.
In an era of unprecedented technological advances across the eld of medical imaging, it is critically important to understand the cognitive and perceptual functioning of the “human in the loop.” These are the clinicians who make the nal decisions on detection, diagnosis, and treatment. However, research on the cognitive and perceptual processes underlying clinician performance is chronically understudied.In September 2019, the National Cancer Institute (NCI) convened the “Cognition and Medical Image Perception Think Tank” in order to advance research on medical image perception 1 and cognition. The Think Tank brought radiologists and pathologists together with researchers working in medical image perception and adjacent elds of cognition and perception, along with representatives from interested federal government agencies (including NCI, National Institute of Biomedical Imaging and Bioengineering …
We used the Multi-Item Localisation (MILO) task to examine search through two sequences. In Sequential blocks of trials, six letters and six digits were touched in order. In Mixed blocks, participants alternated between letters and digits. These conditions mimic the A and B variants of the Trail Making Test (TMT). In both block types, targets either vanished or remained visible after being touched. There were two key findings. First, in Mixed blocks, reaction times exhibited a saw-tooth pattern, suggesting search for successive pairs of targets. Second, reaction time patterns for vanish and remain conditions were identical in Sequential blocks—indicating that participants could ignore past targets—but diverged in Mixed blocks. This suggests a breakdown of inhibitory tagging. These findings may help explain the elevated completion times observed in TMT-B, relative to TMT-A.
Radiologists can identify the gist of a medical image (abnormal vs. normal) better than chance in static 2D images, after presentations of half a second or less (Evans et al., 2013; 2016). The gist of 2D real-world scenes is carried by low spatial frequency channels, which convey the structural layout of scenes (Schyns & Oliva, 1994). In contrast the gist signal in 2D mammography is carried by high spatial frequency channels (Evans et al. 2016). Standard practice in radiology is moving to 3D modalities, where each case consists of a series of images that are assembled into a virtual stack. Radiologists can extract gist from movies of these stacks (Trevino et al., 2019; Wu & Wolfe, 2019). We do not know which channels carry the gist from 3D stacks. We tested 51 radiologists with prostate mpMRI experience on 56 cases, each comprising a stack of 26 T2-weighted prostate mpMRI images. Lesions (Gleason scores 6-9) were present in 50% of cases. A trial consisted of a movie of a single case presented at 48 ms/slice. After each case, participants localized the cancerous lesion on a prostate sector map, then indicated whether a cancerous lesion was presented, and gave a confidence rating. Radiologists were divided equally into three groups who viewed either unfiltered images, low-pass (< 2 cycles/°) filtered images, or high-pass (> 6 cycles/°) filtered images. Unfiltered detection and localization performance were higher than chance (d’ = 0.28; localization = 31%). Radiologists performed at chance when detecting lesions in high-pass filtered images (d’ = 0.16), and were significantly lower than chance for low-pass filtered images (d’ = -0.23). Our data indicate that gist perception from 3D prostate MRI relies on spatial frequency channels between 2 and 6 cycles/°. These findings emphasize that scene gist is highly dependent on task and context.
Abstract. Radiologists can identify whether a radiograph is abnormal or normal at above chance levels in breast and lung images presented for half a second or less. This early perceptual processing has only been demonstrated in static two-dimensional images (e.g., mammograms). Can radiologists rapidly extract the “gestalt” from more complex imaging modalities? For example, prostate multiparametric magnetic resonance imaging (mpMRI) displays a series of images as a virtual stack and comprises multiple imaging sequences: anatomical information from the T2-weighted (T2W) sequence, functional information from diffusion-weighted imaging, and apparent diffusion coefficient sequences. We first tested rapid perceptual processing in static T2W images then among the two functional sequences. Finally, we examined whether this rapid radiological perception could be observed using T2W multislice imaging. Readers with experience in prostate mpMRI could detect and localize lesions in all sequences after viewing a 500-ms static image. Experienced prostate readers could also detect and localize lesions when viewing multislice image stacks presented as brief movies, with image slices presented at either 48, 96, or 144 ms. The ability to quickly extract the perceptual gestalt may be a general property of expert perception, even in complex imaging modalities.
This article introduces a mobile app version of the Multi-Item Localization (MILO) task. The MILO task was designed to explore the temporal context of search through a sequence and has proven useful in both basic and applied research settings. Here, we describe the basic features of the app and how it can be obtained, installed, and modified. We also provide example data files and present two new sets of empirical data to verify that previous findings concerning prospective planning and retrospective memory (i.e., inhibitory tagging) are reproducible with the app. We conclude by discussing ongoing studies and future modifications that illustrate the flexibility and potential of the MILO Mobile app.
Do neuropsychological tests commonly used to assess attention in clinical populations measure the same construct as experimental attention tests? We followed up a factor analysis by Huang et al. (2012), who proposed that many visual cognition attention paradigms load on a “general attention factor”, a. Adult participants (N = 488) completed a comprehensive 90-minute on-line battery (TestMyBrain.org). Five visual cognition paradigms (Multiple Object Tracking (MOT), Flanker Interference, Visual Change Detection (VCD), Approximate Number Sense (ANS), L/T Visual Search Task) were selected to match the general attention factor. We included the Gradual Onset Continuous Performance Task (Grad CPT), hypothesizing that some neuropsychological tests might be measuring sustained rather than selective attention. Neuropsychological tests, selected according to popularity in the domain of cancer-related cognitive impairments (Horowitz et al. 2019), comprised Trail Making Test versions A & B (TMT), Digit Symbol Substitution (DSS), Forward and Backward Digit Span, Letter Cancellation, Spatial Span, and Arithmetic. We obtained a four-factor solution: (1) GradCPT, MOT, VCD, and ANS, along with Spatial Span and DSS; (2) Digit Span Forward and Backward; (3) TMT A & B, Letter Cancellation, and Visual Search; (4) Arithmetic. Flanker Interference did not load on any factor. These results do not fully replicate a, as Visual Search and Flanker Interference were not related to other attentional paradigms. Of neuropsychological measures, Spatial Span and DSS were related to the main attention factor, while those with a search component (e.g., TMT) were related to Visual Search. These results help us to understand the structure of our visual attention paradigms, and to connect visual cognition to neuropsychology. We recommend that clinical studies should be cautious about attributing attention deficits; Digit Span, for example, should not be characterized as an attention measure. Visual search may be distinct from other attentional paradigms.
This symposium aims to show how visual search works in children, adults and older age, in realistic settings and environments. We will review what we know about visual search in real and virtual scenes, and its applications to solving global human challenges. Insights of brain processes underlying visual search during life will also be shown. The final objective is to better understand visual search as a whole in the lifespan, and in the real world; and to demonstrate how science can be transferred to society improving human lives, involving children, as well as younger and older adults.
Radiologists can identify the gist of a radiograph (i.e., abnormal vs. normal) better than chance in breast, lung, and prostate images presented for half a second. However, this rapid perceptual gist processing has only been demonstrated in static two-dimensional images. Standard practice in radiology is moving to three-dimensional (3D) “volumetric” modalities. In volumetric imaging, such as multiparametric MRI (mpMRI), used in prostate screening, a single case consists of a series of image slices through the body that are assembled into a virtual stack. Radiologists can acquire a 3D representation of organ structures by scrolling through stacks. Can radiologists extract perceptual gist from this more complex imaging modality? We tested 14 radiologists with prostate mpMRI experience on 56 cases, each comprising a stack of 26 T2-weighted prostate mpMRI slices. Lesions (Gleason scores 6–9) were present in 50% of cases. A trial consisted of a single movie of the stack. After each case, participants localized the cancerous lesion on a prostate sector map, then indicated whether a cancerous lesion was presented, and gave a confidence rating. Presentation duration was varied between groups. Radiologists were divided into three groups who viewed cases presented at either 48 ms/slice (20.8 Hz, n = 5), 96 ms/slice (10.4 Hz, n = 5), or 144 ms/slice (6.9 Hz, n = 4). Performance declined as slice duration increased (d’ [95% CI]: 48 ms = 0.77 [−.08 - 1.6]; 96 ms = 0.71 [0.17 - 1.24]; 144 ms = 0.47 [0.25 - 0.69]), though gist perception was not statistically significant for the 48 ms group. Localization accuracy (chance ~= 0.08) was 0.40, 0.47, and 0.48, respectively. Our data indicate that radiologists do develop gist perception for 3D modalities. Furthermore, slower presentation rates did not improve performance; there may be an optimal framerate for processing this type of 3D information.
Humans can extract the gist of a visual scene in a fraction of a second, categorizing it as, say, indoor or outdoor, open or closed. Expert radiologists can do this for complex medical images, such as those generated by prostate multiparametric magnetic resonance imaging (mpMRI). MpMRI combines anatomical information from T2-weighted (T2W) sequences, and functional sequences such as conventional diffusion-weighted imaging (DWI) and the apparent diffusion coefficient (ADC). Standard workstation formats present these imaging modalities side-by-side. The goal of this study was to study the nature of mpMRI gist in different modalities. Which modality generates the strongest gist? Are anatomical or functional sequences more useful? Do these modalities provide independent gist information? We tested three groups of five radiologists with prostate mpMRI experience. Each group was shown 100 images from a single modality (T2W, DWI, or ADC). The same cases were used across groups to allow comparison across modalities. Lesions (Gleason scores 6–9) were present in 50% of the images. Images were taken from the base, mid, or apex regions of the prostate. Stimuli were presented for 500 ms, followed by a sector map of the prostate. Participants were first asked to localize the lesion on the sector map (even if they did not see a lesion), then indicate whether or not they thought a lesion was present, and then provided a confidence rating. All three groups detected lesions better than chance [d’ mean(sd): T2W 0.83(0.51); DWI 0.80 (0.29); ADC 1.16(0.31)]. These results suggest that both anatomical and functional information contribute to mpMRI gist. Furthermore, there was little consistency from modality to modality as to which cases produced the best performance, indicating that each modality contributes unique information.
A large body of evidence indicates that cancer survivors who have undergone chemotherapy have cognitive impairments. Substantial disagreement exists regarding which cognitive domains are impaired in this population. We suggest that is in part due to inconsistency in how neuropsychological tests are assigned to cognitive domains. The purpose of this paper is to critically analyze the meta-analytic literature on cancer-related cognitive impairments (CRCI) to quantify this inconsistency. We identified all neuropsychological tests reported in seven meta-analyses of the CRCI literature. Although effect sizes were generally negative (indicating impairment), every domain was declared to be impaired in at least one meta-analysis and unimpaired in at least one other meta-analysis. We plotted summary effect sizes from all the meta-analyses and quantified disagreement by computing the observed and ideal distributions of the one-way χ2 statistic. The actual χ2 distributions were noticeably more peaked and shifted to the left than the ideal distributions, indicating substantial disagreement among the meta-analyses in how neuropsychological tests were categorized to domains. A better understanding of the profile of impairments in CRCI is essential for developing effective remediation methods. To accomplish this goal, the research field needs to promote better agreement on how to measure specific cognitive functions.
In visual search tasks, observers can guide their attention towards items in the visual field that share features with the target item. In this series of studies, we examined the time course of guidance toward a subset of items that have the same color as the target item. Landolt Cs were placed on 16 colored disks. Fifteen distractor Cs had gaps facing up or down while one target C had a gap facing left or right. Observers searched for the target C and reported which side contained the gap as quickly as possible. In the absence of other information, observers must search at random through the Cs. However, during the trial, the disks changed colors. Twelve disks were now of one color and four disks were of another color. Observers knew that the target C would always be in the smaller color set. The experimental question was how quickly observers could guide their attention to the smaller color set. Results indicate that observers could not make instantaneous use of color information to guide the search, even when they knew which two colors would be appearing on every trial. In each study, it took participants 200–300 ms to fully utilize the color information once presented. Control studies replicated the finding with more saturated colors and with colored C stimuli (rather than Cs on colored disks). We conclude that segregation of a display by color for the purposes of guidance takes 200–300 ms to fully develop.
Cancer-related cognitive impairment (CRCI) is a widespread problem for the increasing population of cancer survivors. Our understanding of the nature, causes, and prevalence of CRCI is hampered by a reliance on clinical neuropsychological methods originally designed to detect focal lesions. Future progress will require collaboration between neuroscience and clinical neuropsychology.