Your understanding of what you see now surely influences what you will look at next. Yet this simple concept has only recently begun to be systematically studied and elaborated within theoretical frameworks. The Scene Perception & Event Comprehension Theory (SPECT) distinguishes between front-end and back-end processes that occur while viewers perceive and comprehend dynamic real-world events. Front-end processes occur during each eye fixation (information extraction, attentional selection) and back-end processes occur in memory (the current event model, prior knowledge, and executive processes). We begin with a selective review of the scene perception literature on bottom-up and top-down effects on attentional selection in scenes, and highlight unanswered questions regarding the impact of the viewer's event model–their understanding of what is happening now. Then, we outline the SPECT theoretical framework, and review empirical evidence about how the viewer's current event model influences attentional selection. This influence is contrasted with those of visual saliency (e.g., color, brightness, motion, etc.) and task-driven control (i.e., goal setting, attentional control, inhibition). From this review, we specify a hierarchy of factors affecting attentional selection, in the order of task-driven control, visual saliency, and event models. We then propose several mechanisms by which the viewer’s event model influences attentional selection, and propose a systematic approach to investigating how that happens while watching dynamic scenes. (217 words)
People spontaneously segment an observed everyday activity into discrete, meaningful events, but segmentation can be modified by task goals. Asking young adults to attend to event segmentation while watching movies of everyday actions improved their memory up to 1 month later (Flores et al., 2017). Does attending to event segmentation improve memory across the lifespan? Participants between the ages of 20 and 79 watched movies of actors performing everyday activities while intentionally encoding them for a recall and a recognition memory test 1 week (Experiment 1) or 1 month (Experiment 2) later. In addition to intentionally encoding the movies, half of the participants segmented the movies into fine-grained events. Young adults who segmented recalled more words in their recall responses than those who intentionally encoded 1 week and 1 month later. Middle-aged adults benefited from the intervention after a 1-week delay but not after a 1-month delay. Older adults over the age of 70 did not benefit from attending to segmentation. Of those who segmented, young and older adults showed similar agreement about the locations of event boundaries. Together, the results suggest that older adults are less able, compared to young adults, to maintain or retrieve well-encoded event memories after a delay. In addition, individual differences in segmentation agreement predicted memory up to 1 month later, regardless of age. These results suggest a practical and easy-to-implement intervention for improving recall of everyday events in young and middle-aged adults that is ineffective in older adults. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
People remember events as a series of interconnected actions organized in time. The ability to perceive, represent and remember events is essential to understanding the world, adapting to changing circumstances and surviving. Extensive research has been conducted to understand memory representations of simple, static 'events' that often lack real-world contextual meaning. Although this work lays the foundation of current theories of cognition, it is disconnected from how memory functions in the real world. In this Review, we discuss how events are perceived from a continuous stream of experience, and how perceived events organize memory for real-world experiences. Further, we discuss how cognitive ageing, mental health conditions and neurodegenerative diseases impact event memory. We provide a cohesive overview of event memory for naturalistic stimuli and suggest future directions for the field.
Scene Perception and Event Comprehension Theory (SPECT) posits that understanding picture stories depends upon a coordination of two processes: (1) integrating new information into the current event model that is coherent with it (i.e., mapping) and (2) segmenting experiences into distinct event models (i.e., shifting). In two experiments, we investigated competing hypotheses regarding how viewers coordinate the mapping process of bridging inference generation and the shifting process of event segmentation by manipulating the presence/absence of Bridging Action pictures (i.e., creating coherence gaps) in wordless picture stories. The Computational Effort Hypothesis says that experiencing a coherence gap prompts event segmentation and the additional computational effort to generate bridging inferences. Thus, it predicted a positive relationship between event segmentation and explanations when Bridging Actions were absent. Alternatively, the Coherence Gap Resolution Hypothesis says that experiencing a coherence gap prompt generating a bridging inference to close the gap, which obviates segmentation. Thus, it predicted a negative relationship between event segmentation and the production of explanations. Replicating prior work, viewers were more likely to segment and generate explanations when Bridging Action pictures were absent than when they were present. Crucially, the relationship between explanations and segmentation was negative when Bridging Action pictures were absent, consistent with the Coherence Gap Resolution Hypothesis. Unexpectedly, the relationship was positive when Bridging Actions were present. The results are consistent with SPECT’s assumption that mapping and shifting processes are coordinated, but how they are coordinated depends upon the experience of a coherence gap.
Spatial memory is important for supporting the successful completion of everyday activities and is a particularly vulnerable domain in late life. Grouping items together in memory, or chunking, can improve spatial memory performance. In memory for desktop scale spaces and well-learned large-scale environments, error patterns suggest that information is chunked in memory. However, the chunking mechanisms involved in learning new large-scale, navigable environments are poorly understood. In five experiments, two of which included young and older adult samples, participants watched movies depicting routes through building-sized environments while attempting to remember the locations of cued objects. We tested memory for the cued objects with virtual pointing, distance estimation, and map drawing tasks after participants viewed each route. Patterns of error failed to show consistent evidence of chunking in spatial memory across all experiments. One possibility is that chunking in spatial memory relies on visual perceptual grouping mechanisms that are not in play during encoding of large-scale spaces encountered through extended route experiences that do not afford concurrent viewing of target locations.
People spontaneously segment continuous ongoing actions into sequences of events. Prior research found that gaze similarity and pupil dilation increase at event boundaries and that older adults segment more idiosyncratically than do young adults. We used eye tracking to explore age-related differences in gaze similarity (i.e., the extent to which individuals look at the same places at the same time as others) and pupil dilation at event boundaries. Older and young adults watched naturalistic videos of actors performing everyday activities while we tracked their eye movements. Afterward, they segmented the videos into subevents. Replicating prior work, we found that pupil size and gaze similarity increased at event boundaries. Thus, there were fewer individual differences in eye position at boundaries. We also found that young adults had higher gaze similarity than older adults throughout an entire video and at event boundaries. This study is the first to show that age-related differences in how people parse continuous everyday activities into events may be partially explained by individual differences in gaze patterns. Those who segment less normatively may do so because they fixate less normative regions. Results have implications for future interventions designed to improve encoding in older adults. (PsycInfo Database Record (c) 2024 APA, all rights reserved).
We segment what we read into meaningful events, each separated by a discrete boundary. How does event segmentation during encoding relate to the structure of story information in long-term memory? To evaluate this question, participants read stories of fictional historical events and then engaged in a post-reading verb arrangement task. In this task, participants saw verbs from each of the events placed randomly on a computer screen, and then they arranged the verbs into groups onscreen based on their understanding of the story. Participants who successfully comprehended the story placed verbs from the same event closer to each other than verbs from different events, even after controlling for orthographic, text-based, semantic, and situational overlap between verbs. Thus, how people structure story information into separate events during online comprehension is associated with how that information is stored in memory. Specifically, story information within an event is bound together in memory more so than information between events.
Memory for personally experienced events (episodic memory) declines with age; however, memory for facts about the world and the steps involved in completing common actions (semantic memory) remains relatively stable with age. In a series of recent studies, we examined whether older adults could leverage their semantic knowledge to facilitate the acquisition of new episodic memories. Previous studies provided mixed results as to whether top-down processing, such as relevant knowledge, affects encoding, but these studies used young adult participants. Across multiple experiments, we found that semantic knowledge influences encoding processes, such as the ability to parse ongoing activity into discrete events, a process called event segmentation. Here, we review previous literature, and our own recent studies, that examined age-related changes in event segmentation and memory for everyday events, and how semantic knowledge affects these processes. We relate our current findings to previous results, and then we discuss the implications for theories of event cognition and open questions for the field.
Rapid scene categorization is typically argued to be purely feed-forward. Yet, when navigating in our environment, we usually see predictable sequences of scene categories (e.g., offices followed by hallways, parking lots followed by sidewalks, etc.). Previous work showed that scenes were both easier to recognize, and to discriminate from phase-randomized noise, when shown in ecologically valid, predictable sequences than in randomized sequences (Smith & Loschky, 2019). But, when in scene processing do sequential predictions facilitate scene categorization? We examined this question using EEG. Participants saw scenes in either spatiotemporally coherent sequences (first-person viewpoint of navigating, from, say, an office to a classroom) or their randomized versions. Participants saw 288 scene RSVP sequences (each 1 target and 9 primes), while we recorded their event-related potentials (ERPs). Participants had to categorize one randomly selected target image on each trial, in an 8-AFC task. We found reduced ERP amplitudes for targets in coherent sequences roughly 140 milliseconds after image onset--when ERPs typically first index rapid scene categorization--and during the N300 and N400 components, suggesting both reduced identification costs and semantic integration costs in coherent sequences. Interestingly, such ERP amplitude reductions were predicted by low-level visual similarity between sequential prime-target pairs, suggesting that visual similarity might explain the reduced processing costs in coherent sequences. To test this hypothesis, in Experiment 2, we reran Experiment 1 behaviorally, but replaced the targets with noise images, and asked participants to predict the categories of the missing scenes. Target scenes were more predictable in coherent sequences. Importantly, the correlations of Experiment 2 image predictability with Experiment 1 ERP amplitudes (from 140-450 ms) were greater than for image similarity with ERP amplitudes. Thus, contrary to purely feed-forward accounts of rapid scene category recognition, both predictions for an upcoming scene category and visual similarity between successive scenes facilitate scene “gist”.
While semantic and episodic memory may be distinct memory systems, their interdependence is substantial. For instance, decades of work have shown that semantic knowledge facilitates episodic memory. Here, we aim to clarify this interactive relationship by determining whether semantic knowledge facilitates the acquisition of new episodic memories, in part, by influencing an encoding mechanism, event segmentation. In the current study, we evaluated the extent to which semantic knowledge shapes how people segment ongoing activity and how such knowledge-related benefits in segmentation affect episodic memory performance. To investigate these effects, we combined data across three studies that had young and older adults segment and remember videos of everyday activities that were either familiar or unfamiliar to their age group. We found age-related differences in event-segmentation ability and memory performance, but only when older adults lacked semantic knowledge. Most importantly, when they had access to relevant semantic knowledge, older adults segmented and remembered information similar to young adults. Our findings indicate that older adults can use semantic knowledge to effectively encode and retrieve everyday information. These effects suggest that future interventions can leverage older adults' intact semantic knowledge to attenuate age-related deficits in event segmentation and episodic long-term memory.
How does viewers' knowledge guide their attention while they watch everyday events, how does it affect their memory, and does it change with age? Older adults have diminished episodic memory for everyday events, but intact semantic knowledge. Indeed, research suggests that older adults may rely on their semantic memory to offset impairments in episodic memory, and when relevant knowledge is lacking, older adults' memory can suffer. Yet, the mechanism by which prior knowledge guides attentional selection when watching dynamic activity is unclear. To address this, we studied the influence of knowledge on attention and memory for everyday events in young and older adults by tracking their eyes while they watched videos. The videos depicted activities that older adults perform more frequently than young adults (balancing a checkbook, planting flowers) or activities that young adults perform more frequently than older adults (installing a printer, setting up a video game). Participants completed free recall, recognition, and order memory tests after each video. We found age-related memory deficits when older adults had little knowledge of the activities, but memory did not differ between age groups when older adults had relevant knowledge and experience with the activities. Critically, results showed that knowledge influenced where viewers fixated when watching the videos. Older adults fixated less goal-relevant information compared to young adults when watching young adult activities, but they fixated goal-relevant information similarly to young adults, when watching more older adult activities. Finally, results showed that fixating goal-relevant information predicted free recall of the everyday activities for both age groups. Thus, older adults may use relevant knowledge to more effectively infer the goals of actors, which guides their attention to goal-relevant actions, thus improving their episodic memory for everyday activities.
How do the limits of peripheral vision versus the effects of attentional selection influence change blindness? We investigated this in four Experiments. We first characterizing the difficulty of numerous change blindness demonstrations using the standard Flicker paradigm in Experiment 1. Experiment 2 evaluated viewers’ ability to peripherally discriminate change pairs of varying difficulty when changes were fully attended. We first showed participants a change. Participants then fixated predetermined locations, 1.25-10 degrees eccentricity from each change, and performed a peripheral ABX discrimination task for versions A and B. Discrimination performance declined with increasing eccentricity for changes “difficult” to find in the Flicker paradigm, but “easy” changes showed little effect of eccentricity. Experiment 3 simulated the effects of crowding in peripheral vision using the Texture Tiling Model (TTM). We created “mongrel” texture versions of each image pair, with simulated “foveation” at each of the fixation locations and eccentricities of Experiment 2. Participants’ discriminated whether each mongrel was generated from image version A or B. Performance decreased with simulated eccentricity, with a steeper slope for “difficult” change pairs than “easy” pairs, suggesting that peripheral information loss partially explains change blindness. In Experiment 4, we manipulated attention to the change by manipulating participants’ knowledge of the change before the ABX peripheral task. The pre-cued condition (same as Experiment 2) replicated the finding that the changes easiest to peripherally discriminate were those easiest to find in the Flicker paradigm. However, un-cued “difficult” changes were uniformly poorly discriminated at all eccentricities, and “easy” changes were only above-chance at 1.25 deg eccentricity. This suggests that change detection involves two stages: 1) attending to the change; 2) peripherally or foveally discriminating the two versions. Thus, both attentional selection and the limitations of peripheral vision influence change detection. The latter can be roughly approximated by the TTM.
What guides eye movements while viewing visual narratives? More specifically, do comprehension processes influence attentional selection when reading wordless picture stories? According to the Scene Perception & Event Comprehension Theory (SPECT) there are front-end processes, such as attentional selection, that occur during single eye fixations, and back-end processes, such as building an event model, that occur in working memory and long-term memory. Here we have investigated how attentional selection may be influenced by event models while people view visual narratives. Prior research has shown that as more situational changes occur in a visual narrative (e.g., space, time, characters, goals and sub-goals), viewers are more likely to perceive an event boundary (i.e., beginning of an event) (Magliano et al., 2011). Other research has shown that viewing times increase at event boundaries (Hard, Recchia & Tversky, 2011; Smith, Newberry & Bailey, 2019). In an eye-tracking study using the “Boy, Dog, Frog” picture stories, we replicated the findings that spatiotemporal and character changes produced higher event segmentation. We also replicated the findings that viewing time was longer at event boundaries. We then extended those findings to eye movements. We asked whether the longer viewing times at event boundaries, when more event indices changed, were due to longer fixations (i.e., increased processing load), or more fixations (i.e., more search for information). Longer viewing times were strongly associated with more fixations, not longer fixations, supporting the search hypothesis. Both viewing times and the number of fixations were found to be significantly predicted by spatiotemporal changes, character changes, and the beginning of superordinate goals. Further analyses will assess whether these additional fixations are simply allocated to the new and salient content in an image, or if they are also specifically directed towards the contents necessary to make inferences about the characters’ goals in the narrative.
We deconstruct continuous streams of action into smaller, meaningful events. Research has shown that the ability to segment continuous activity into such events and remember their contents declines with age; however, knowledge improves with age. We investigated how young and older adults use knowledge to more efficiently encode and later remember information from everyday events by having participants view a series of self-paced slideshows depicting everyday activities. For some activities, older adults produce more normative scripts than do young adults (older adult activities) and for other activities, young adults produce more normative scripts than do older adults (young adult activities). Overall, participants viewed event boundaries longer than within events (i.e., the event boundary advantage) replicating prior research (e.g., Hard, Recchia, & Tversky, 2011). Importantly, older adults demonstrated the boundary advantage for the older adult activities but not the young adult activities, and they also had better recognition memory for the older adult activities than the young adult activities. We also found that the magnitude of a participant's boundary advantage was associated with better memory, but only for the less knowledgeable activities. Results indicate that older adults use their intact knowledge to better encode and remember everyday activities, but that knowledge and event segmentation may have independent influences on event memory.
Past research suggests that recognizing scene gist, a viewer's holistic semantic representation of a scene acquired within a single eye fixation, involves purely feed-forward mechanisms. We investigated whether expectations can influence scene categorization. To do this, we embedded target scenes in more ecologically valid, first-person-viewpoint image sequences, along spatiotemporally connected routes (e.g., an office to a parking lot). We manipulated the sequences' spatiotemporal coherence by presenting them either coherently or in random order. Participants identified the category of one target scene in a 10-scene-image rapid serial visual presentation. Categorization accuracy was greater for targets in coherent sequences. Accuracy was also greater for targets with more visually similar primes. In Experiment 2, we investigated whether targets in coherent sequences were more predictable and whether predictable images were identified more accurately in Experiment 1 after accounting for the effect of prime-to-target visual similarity. To do this, we removed targets and had participants predict the category of the missing scene. Images were more accurately predicted in coherent sequences, and both image predictability and prime-to-target visual similarity independently contributed to performance in Experiment 1. To test whether prediction-based facilitation effects were solely due to response bias, participants performed a two-alternative forced-choice task in which they indicated whether the target was an intact or a phase-randomized scene. Critically, predictability of the target category was irrelevant to this task. Nevertheless, results showed that sensitivity, but not response bias, was greater for targets in coherent sequences. Predictions made prior to viewing a scene facilitate scene-gist recognition.
Past research has argued that scene gist, a holistic semantic representation of a scene acquired within a single fixation, is extracted using purely feed-forward mechanisms. As such, scene gist recognition studies have presented scenes from multiple categories in randomized sequences. We tested whether rapid scene categorization could be facilitated by priming from sequential expectations. We created more ecologically valid, first-person viewpoint, image sequences, along spatiotemporally connected routes (e.g., an office to a parking lot). Participants identified target scenes at the end of rapid serial visual presentations. Critically, we manipulated whether targets were in coherent or randomized sequences. Target categorization was more accurate in coherent sequences than in randomized sequences. Furthermore, categorization was more accurate for a target following one or more images within the same category than following a switch between categories. Accuracy was also higher following two primes from the same category than following only one between scene categories (e.g., multiple office primes facilitated recognition of a hallway). Likewise, accuracy was higher for targets more visually similar to their immediately preceding primes. This suggests that prime-to-target visual similarity may explain the coherent sequence advantage. We tested this hypothesis in a second experiment, which was identical except that target images were removed from the sequences, and participants were asked to predict the scene category of the missing target. Missing images in coherent sequences were more accurately predicted than missing images in randomized sequences, and more predictable images were identified more accurately in Experiment 1. Importantly, partial correlations revealed that image predictability and prime-to-target visual similarity independently contributed to rapid scene gist categorization accuracy.
Scene Perception and Event Comprehension Theory (SPECT) is an integrative theoretical framework for understanding both visual narrative perception and comprehension. It incorporates theories from the domains of scene perception, event perception, and narrative comprehension. SPECT assumes that the back-end processes involved in the construction of the event model are general cognitive mechanisms, which extend beyond the language-based modality to include visual narratives. This assumption has been supported by the data from our studies, which show evidence consistent with three stages of building the current event model: laying the foundation, mapping, and shifting. After having laid the foundation, according to SPECT, the viewer continues to map new incoming information to the current event model. A unique contribution of SPECT is that it allows us to test novel hypotheses about the interactions between front-end and back-end mechanisms, which are assumed to be bidirectional.
What role do scene category expectations play in scene gist recognition? Research has shown viewers accurately identify the gist of briefly flashed scenes presented in randomized sequences suggesting it involves purely feed-forward mechanisms. We investigated if sequential expectations for scenes could influence their gist recognition. We created spatial narrative sequences of images linked along spatio-temporal routes from starting points to destinations (e.g., office, hallway, stairwell, sidewalk, parking lot). 10 scene images from each narrative were presented in an RSVP sequence, with 9 of the 10 images given 300 ms of processing time. The target image was presented for 24 ms simultaneously with an attentional alerting tone and followed by a 48 ms 1/f noise mask. Following presentation of the target and mask, the participant selected the target category from an 8-AFC array of scene category labels. Temporal position of target images within each 10 image sequence was equated and counter-balanced so participants could not guess when the target image would appear. To reduce predictability, we showed 1-3 image subsequences from each category within each narrative so targets were preceded by 0, 1, or 2 images of the same semantic category (e.g., 0, 1, or 2 offices preceded a target office). Scenes were presented in both coherent and randomized sequences to test two competing hypotheses. The "Narrative coherence" hypothesis predicted accuracy would be higher for images in coherent narrative sequences as expectations prime to-be-presented representations. Alternatively, the "Feed-forward" hypothesis predicted accuracy would not differ between coherent and randomized sequences. We found that images presented in coherent sequences were identified more accurately than targets within randomized sequences. Target images preceded by sequential exposures of the same scene category were identified more accurately than targets that were not. Further research will identify whether the facilitation is due to increased sensitivity or bias. Meeting abstract presented at VSS 2018
Observers can categorize a novel scene within the first 100 ms of its onset. Researchers have suggested this is accomplished through a rapid feed-forward sweep of neural activation. Consequently, researchers have focused on examining the minimal perceptual information diagnostic of a scene's semantic category, rather than investigating the role that top-down processes play in rapid scene categorization. Thus, most scene gist studies present scenes from multiple categories in randomized sequences.Conversely, in this experiment, we tested whether scene gist recognition is facilitated by sequential expectations.To do this, we created more ecologically valid spatio-temporal "narrative" sequences of images along spatially connected routes from starting points to destinations (e.g., office, hallway, stairwell, sidewalk, parking lot). Scene images were presented one-at-a-time, and based on pilot-testing, briefly flashed (24 ms) and masked (48 ms), followed by selecting the scene category from an 8-AFC array. To reduce predictability of subsequent images, we included subsequences of randomly 1-4 scenes from each category (e.g., 1-4 office images followed by 1-4 hallways, etc.). Critically, scenes were shown in either coherent or randomized narrative sequences to test two competing alternative hypotheses: 1) "Narrative coherence": accuracy is higher for images in coherent narrative sequences because scene category expectations prime representations for to-be-presented scenes; 2) "Feed-forward": accuracy does not differ between coherent and randomized image sequences because it is a purely feed-forward process. Results: images presented in randomized sequences were categorized just as accurately as images presented in coherent sequences, consistent with the "Feed-forward" hypothesis. Furthermore, accuracy did not increase as a function of the number of sequential exposures to a single scene category (e.g., 1-4 office images in a row), suggesting that the mechanisms responsible for perceiving a scene's meaning are so rapid that prior exposures do not support extracting the gist of subsequent scenes. Meeting abstract presented at VSS 2017